diff --git a/.claude/skills/ai-app-rigs/SKILL.md b/.claude/skills/ai-app-rigs/SKILL.md new file mode 100644 index 0000000..0c7ecbd --- /dev/null +++ b/.claude/skills/ai-app-rigs/SKILL.md @@ -0,0 +1,244 @@ +--- +name: ai-app-rigs +description: ai-app's test rigs, harness scripts and reference measurements - ui-sandbox.sh, debug-transcript.sh, transcript-bench.sh, stream-bench.sh, trace-draw.sh, the /usage fixture vocabulary, the fake CLI, the rule that no UI-driving script may tap a coordinate, how to test llama.cpp and ssh on this machine, how importing behaves, and the scroll/stream/explorer numbers not worth re-measuring. Read before running or writing a benchmark, driving the app's UI from a script, exercising the session lifecycle, testing a llama or remote session, or touching the import screen. +--- + +# ai-app: rigs, harnesses and measurements + +Moved out of `AGENTS.md` on 2026-09-04 so it is read when it is relevant +rather than sent with every request in this repo -- it was 12 KB of the 35 KB +that file cost on every one. Unchanged in the move, and still the only copy. + +## The rigs + +Each exists because something was invisible without it. + +- **`app/ui-sandbox.sh`** — a second `ai-server` with its own `$HOME`, config + and data directory, holding eight invented Claude Code transcripts and a + `claude` that is two lines of shell. **That isolation is the point**: the + import screen lists whatever is in `~/.claude/projects`, which in this VM is + real agent transcripts, so exercising *delete* against the ordinary server + deletes somebody's conversation and exercising *import* starts a real + `--resume` on the owner's account. + Its port and root derive from the checkout's name, so two checkouts' + sandboxes cannot reach each other, and its token is generated once into + `~/.config/ai-app/sandbox-token` and carried across restarts along with any + the enrolment flow appended — so the emulator app is enrolled **once** (the + start banner prints the command) and stays enrolled. It shares the real TLS + certificates, because the installed APK pins that CA. + Driving verbs, so none of this is re-derived per session: + `./ui-sandbox.sh spawn [title]` (an echo session, prints its id), + `./ui-sandbox.sh send SID text|@file`, and + `./ui-sandbox.sh api /path [curl args]`. + `./ui-sandbox.sh keep` restarts the server without wiping the sessions and + enrolment already there — for when the fixture under test was expensive to + build; plain `start` wipes them, which is right for the list-screen + fixtures and wrong for that. + It passes `--delay` by default, and `AI_SANDBOX_BIG_MB` puts one large + transcript among the small ones while `AI_SANDBOX_SPAWN_DELAY` makes the + fake CLI slow to start. Both exist because operations that finish in + milliseconds have states on the way that nothing can observe, and an + unobservable state is one where broken and working look identical. + It also builds a fixture tree at the sandbox home's `~/files` for the + explorer, holding the states otherwise only reachable by finding a real + machine in one: an empty directory, a name with a tab and one with an + apostrophe, a binary file, one over `FILE_LIMIT`, one `chmod 000`, a + symlink to a directory and a broken one, a source file per language, and + the three sizes the limits were measured against (`edit-32k.rs`, + `edit-128k.rs`, `big-source.rs`). Point a session at it with + `./ui-sandbox.sh api /sessions//cwd -X POST -H 'content-type: application/json' -d '{"cwd":"~/files"}'`. + The explorer's 409 is produced by editing the file on the machine + (`printf … > file`) between pressing the pencil and pressing save. +- **`app/debug-transcript.sh`** — a real conversation on the emulator. The + echo driver is the right rig for most things and the wrong one for anything + whose cost scales with what was actually written: a real reply is longer, + is real markdown, and carries tool calls whose input and output are + kilobytes. Two faults were invisible until a real transcript was loaded — a + page of history landing mid-fling threw the reader back to the newest end, + and parsing one real reply took 51ms against 4.6ms for a synthetic one. + `-b` takes the biggest conversation on the machine rather than the newest, + which is what a scrolling test wants; `--stop` takes it down. + It copies the transcript into `/tmp` and gives the server a `HOME` of its + own, so the import can only see the copy — importing spawns `claude + --resume`, and against the real file that is a second CLI writing to a + conversation somebody may still be in. **A transcript never goes in this + repository**: they hold whatever was said, read and written in that + session, and `~/repos` is shared with the host besides. +- **`/usage` in an echo session puts up an invented meter**, which is how the + rate-limit screens' states are reached without spending quota: `/usage 42`, + `/usage 95 20` (minutes left), `/usage 42 never` (the between-blocks window + with no reset time), `/usage 42 unreadable`, `/usage notloggedin`, + `/usage unreachable`, `/usage failed`, `/usage off`. The vocabulary is + `usage::Fixture`'s, since those are its states. With none set an echo + session meters nothing, which is the ordinary case and draws no bar. +- **A fake CLI exercises the process lifecycle without a token.** Point a + `claude_cli` provider's `command` at a two-line script — `#!/bin/sh` and + `cat > /dev/null` — and it behaves the way the lifecycle code cares about: + it holds the fifo open, records a real pid, writes nothing, and dies on a + signal. So adopt, stop, restart and start are all drivable without a real + `--resume` and without spending a turn on somebody's account. Reach for + this when what is under test is *whether a process is running*, and for + `debug-transcript.sh` when it is *what the transcript draws*. +- **`app/transcript-bench.sh`** is the standard scroll measurement: it opens + the first session (or `-k` keeps the current screen), scrolls a fixed + gesture loop, and prints the app's render report — the same one the in-app + copy button produces, whose `on screen:` line names what the viewport was + holding. Compare two runs with the same gestures; the emulator's absolute + frame times transfer nothing, the report's accounting does. Run it either + side of any change under `Markdown*.kt`, `Transcript*.kt` or + `SessionScreen.kt`'s list, and put the report in the commit. The numbers + that move first are the worst `record: one block`, the reparse mean while + streaming, and the draw phase's accounting line. +- **`app/stream-bench.sh [-k] FILE`** is that measurement for a reply still + arriving. It taps "Jump to latest" so the list is pinned to the newest end, + resets the report, sends FILE, waits for the transcript to stop growing, + and prints. Both of those are corrections to a first version that measured + nothing: a transcript parked further back never redraws while a reply + streams into it, and a session is idle at *both* ends of a turn, so polling + for idle answers before the turn has started. +- **`app/trace-draw.sh`** names what a scrolling frame spends inside the + framework, from `atrace` text output with no trace processor needed. It is + how the cost of a layout node per link was attributed to the framework + rather than guessed at. + +### Driving the UI + +**No script that drives this app's UI presses a coordinate.** Every control +is found by the name it already carries for assistive technology — +`ui-trace record --do "tap 'Session settings'"` — which resolves the label +against the screen at the moment of the gesture and fails the whole run when +it is not there. `app/bench-lib.sh` is what the bench scripts share for it. A +coordinate is a position measured once by hand, and anything that moves the +control makes the tap land on whatever now sits there — the bench then +reports a number that was never measured, which reads exactly like a result. +Both bench scripts pressed the render report at `tap 723 205` until that +button moved into the session settings dialog on 2026-09-03. The check that +none has crept back: + + grep -n "tap [0-9]" app/*.sh + +Swipes are still coordinates, deliberately: a gesture across a scrolling area +is a distance rather than a control. + +**Two traps in the emulator bench loop**, each of which cost a run. +`adb shell pm clear` removes the enrolment and the notification permission +along with the saved anchors, so the next run measures a permission dialog — +re-enrol with the command `ui-sandbox.sh` prints, and +`pm grant … POST_NOTIFICATIONS`. And a saved scroll anchor is per session id, +so the only way two builds start a scroll from the same place is a *fresh +session for each*. + +**The emulator is `~/repos/emulator-tools`' business, not this repo's.** +`emu up` creates and boots the AVD named after this checkout — whatever `emu +name` prints, never a name typed out here, since this file is the same in +every clone. `run-android.sh` is that plus a build and an install. The `adb` +on `PATH` after sourcing `android-env.sh` is that repo's wrapper, which fills +in `-s` from the same rule. Gradle does not go through it, so a Gradle init +script from `emulator-tools` runs `emu check` before `installDebug`, +`uninstallDebug` and `connectedAndroidTest` and fails rather than fanning out +to every attached device; when it refuses, say which device you mean at the +moment you use it — `ANDROID_SERIAL=$(emu serial) ./gradlew …`. + +### Testing llama.cpp and ssh here + +**Both are set up here as of 2026-09-04** and need nothing typed. The +prebuilt CPU llama.cpp lives outside the repo at `~/.local/opt/llama.cpp` +(the 15 MB `ubuntu-x64` release asset) and is symlinked as +`/usr/local/bin/llama-server`, which is what makes **discovery find it over +ssh**: `~/.local/bin` is not on the PATH a non-interactive ssh session gets. +It resolves its own libraries through `$ORIGIN`, so no `LD_LIBRARY_PATH` is +needed. One model is downloaded — `unsloth/Qwen3-0.6B-GGUF/Qwen3-0.6B-Q8_0.gguf`, +639 MB under `~/.local/share/ai-app/models` — and answers at usable speed on +this VM's 8 cores. **Do not test with a 2-bit quant**: the +IQ2_XXS of that model produces fluent nonsense, which reads exactly like a +broken driver — `llama-cli` produces the same from the file directly, which +is how to tell the two apart in a hurry. + +There is no second machine, so **ssh this VM to itself**. That is set up +too: the key is `~/.config/ai-app/ssh-self` (its public half is in +`~/.ssh/authorized_keys`, labelled removable), and the real config carries a +setup called **"this vm over ssh"** — `bob@127.0.0.1` with that +`identityFile` plus +`options: ["StrictHostKeyChecking=no", "UserKnownHostsFile=/tmp/ai-app-known-hosts"]` +so it touches nothing real — offering `claude-cli` and `llama-cpp`. It is the +whole rig for "does a remote llama session work", since the far machine is +this one and the model file is the same file. For a throwaway setup of your +own, point a provider's `command` at something harmless like `/bin/echo` +rather than at `claude`: the transport is what is under test, the process +exiting immediately is the signal, and it costs no tokens. The remote login +shell here is **fish**; the +remote script and `ssh.rs`'s POSIX quoting happen to mean the same thing in +both, but that is luck rather than design, and a shell that is neither is the +thing to suspect first if a remote spawn ever mangles an argument. + +## Importing + +The import list reports each session's **size as well as its line count**, +because the two disagree in the way that matters: these transcripts embed +screenshots as base64, so one line can be a megabyte. On this machine a 69 MB +session has 3,427 lines and a 44 MB one has 6,792 — nothing about a line +count tells you what continuing a session will cost. Shown, not warned about; +importing a large session is a choice somebody is entitled to make. + +**Never import a Claude Code session that is open in a terminal.** The app +refuses it — see PLAN.md for the incident that made that a refusal rather +than a warning. + +**One Claude Code session id can name two files, and the listing offers it +once.** Resuming from a different working directory makes the CLI write a +second transcript with the same id under that directory's project folder — an +ordinary state of a machine, not corruption. Everything downstream addresses +a session by id, and the phone keyed its list on it, so two rows sharing one +**closed the app** on a Compose duplicate-key throw. `parse_listing` keeps +the copy with the most lines, because the other is usually a few-hundred-byte +stub and is often the *newer* of the two, so recency is the wrong key. +Deleting removes every copy rather than the first, or the row came back after +a delete that reported success. The phone's half is `uniqueItems`, which +every list keyed on a server-chosen id goes through: a repeat there must +never be able to close the app, whatever produced it. + +**Deleting a session offers to take the machine's own transcript with it** — +`DELETE /sessions/{id}?deleteForeign=true`, behind a switch in the +confirmation, and only where the driver keeps a record of its own +(`keepsOwnTranscript`, which today means Claude Code). Off by default, +because leaving that copy is what makes an ordinary delete recoverable — and +the dialog's paragraph is rewritten when it is on rather than appended to, +since the sentence promising the conversation "should still be there to +import again" is exactly the one the switch makes false. The server deletes +the machine's copy *first*, so a machine it cannot reach leaves the session +where it was instead of half-deleted. + +## Measurements worth not re-taking + +- **What the transcript screen costs to scroll.** Taken 2026-08-30 on the GPU + emulator against a real imported transcript with the server at + `--delay 120`. Settled and flinging fast, both into fresh history and back + through rows already drawn: **5.2–5.9% janky frames, 99th percentile + 29–32ms, 0–2 slow UI-thread frames.** The stock Settings app on the same + device is 3.3% and 38ms, so this is at the platform floor. The number that + is *not* at the floor is the first few seconds after opening a session, + where every row on the way is being composed for the first time; that is + inherent to a lazy list and it is why a measurement taken before the screen + settles reads three times worse. **Settle first, then reset `gfxinfo`.** +- **The reset path is not reachable by reopening a session.** Measured + 2026-09-04 against a session streaming at 20 events a second: reopening one + with an anchor 1,800 events back connects **87–119 events behind**, well + under `CATCH_UP_LIMIT`'s 200, because the restore is two requests — the + opening page, then one span covering the whole distance. To exercise the + reset at all you have to lower `CATCH_UP_LIMIT` in a throwaway build; at 5 + the app takes the reset on a live connection, clears, refills and carries + on without reconnecting. +- **The session screen's stream survives backgrounding here** — 20 seconds at + the launcher while 415 events were produced brought no reconnect at all, + which is not what the comment above that loop expects, and is most likely + this emulator being headless rather than the phone's behaviour. +- **Reopening a cached session costs one request for one event** (the probe), + and scrolling the whole conversation back costs nothing more; a cold open + of the same 500-event session is two pages, 100 events. Measured + 2026-09-04 on the emulator against the sandbox. +- **Reading is cheap and editing is not.** The viewer handles a 1 MiB, + 28,000-line file because it draws one row per line; the editor is one + `BasicTextField`, which costs two seconds a frame at 128 kB and stops the + app at 1 MiB, so `EDIT_LIMIT` caps it at 32 kB with the reason said on + screen. If you make the editor faster, that number is what to move. + EXPLORER.md's "What the measurements said" has the rest. diff --git a/DECISIONS.md b/DECISIONS.md new file mode 100644 index 0000000..65ff286 --- /dev/null +++ b/DECISIONS.md @@ -0,0 +1,44 @@ +# Decisions awaiting review + +Choices made while working autonomously, for Bryan to keep or change. Each +says what was picked and why; the detail is in the design doc it names. +Delete an entry once it has been looked at. + +## Subagent views (2026-09-05, `SUBAGENTS.md`) + +Made on my own judgement, limited blast radius: + +1. **A subagent is a transcript, not a session.** It has no process, + controls or settings; it is addressed as `/sessions/{id}/subagents/{sub}` + and stored under the session's directory, so deleting the session takes + it. Alternative rejected: registering it as a session of its own, which + would give it a card in the main list and a driver that can do nothing. +2. **Read-only view is the session screen minus its controls**, rather than + a second, simpler transcript screen. Keeps paging, caching, selection + and rendering in one place. Cost: a `readOnly` mode threaded through + `SessionScreen`. +3. **The list only carries a count.** Each session row says how many + subagents it has; their titles and statuses are fetched when the card is + expanded. Keeps `GET /sessions` from reading every subagent transcript. + Consequence: an expanded card's statuses refresh with the list, not live. +4. **Expanded/collapsed is remembered per session on the phone**, not on + the server. Collapsed by default, per the transcript convention that new + things arrive collapsed. +5. **Subagents of imported sessions are not shown.** The import path still + skips `isSidechain` records; the CLI's own `subagents/agent-*.jsonl` files + are not read. Only subagents run while this backend was watching exist. +6. **Echo grows `/subagent [n]`** as the test rig, so nothing here needs a + paid turn to exercise. + +Deferred, because they reach further than this feature: + +- **Live status on the list.** Whether the session list should follow a + stream at all (it refreshes on demand today) decides whether subagent + status can ever be live there. Not changed. +- **Nested subagents.** A subagent's own Task calls are shown as tool calls + in its transcript and are not given transcripts of their own. Supporting + that is the same mechanism one level down, but the UI would need nested + expanders. + +- **The subagent status row says "context unknown".** Nothing measures a + subagent's context; the row could leave it out rather than admit it. diff --git a/SUBAGENTS.md b/SUBAGENTS.md new file mode 100644 index 0000000..0d8a1c7 --- /dev/null +++ b/SUBAGENTS.md @@ -0,0 +1,141 @@ +# Subagents + +A session's subagents -- the helpers a Claude Code session starts through its +Task tool -- each get a transcript of their own, listed under the session's +card and readable in the same transcript view the session has. Designed +2026-09-05; the decisions Bryan has not yet reviewed are in `DECISIONS.md`. + +## What a subagent is here + +**A subagent is a second transcript owned by a session, in the same event +model, with no process and no controls.** It is not a session: it cannot be +messaged, stopped or started, and it has no setup, model or usage of its +own. Everything it shares with a session -- the transcript file format, the +paging routes, the SSE stream, the phone's cache and rendering -- is reused +by addressing, not by copying. + +The CLI reports a subagent's messages on the parent's own stream-json +output, each carrying `parent_tool_use_id` = the id of the Task `tool_use` +that started it. Before this the translator dropped those lines +(`subagent_events_are_not_duplicated_into_the_transcript`); now it routes +them to that subagent's own translator and transcript. The parent's +transcript still shows only the Task call itself. + +## Storage + +Under the session directory: + +``` +/subagents//meta.json {title, created} +/subagents//transcript.jsonl same SeqEvent lines as the session's +``` + +The id is the Task tool_use id (`toolu_…`), which is unique, stable across a +backend restart, and already the key everything on the parent side uses. +Only ids matching `[A-Za-z0-9_-]+` are ever created or looked up, since the +id becomes a path. + +The transcript's sequence numbers are its own, starting at 1. `Transcript`, +`read_window`, `catch_up` and `read_after` work on it unchanged. + +Its path out: deleting the session deletes its directory, subagents included. +There is no separate delete. + +## Lifecycle, as events in the subagent's transcript + +1. Created on the first child line for an unseen parent id (or, when the + parent Task call was seen, at that call). First lines written: + `Status Running`, then `UserMessage { text: }` when + the prompt is known -- it genuinely is the subagent's first user turn. +2. Every child line is translated by that subagent's own `Translator` + (one per subagent: tool ids are unique but streaming deltas are by + content-block index, and parallel subagents interleave). +3. **The parent's `tool_result` never finishes a subagent.** The Task tool + runs in the background by default: the `tool_result` -- "Async agent + launched..." -- arrives the moment it *starts*, while the subagent goes + on working for however long its own turn takes, sometimes minutes. What + ends it is its own turn ending: the raw API's `message_delta` on its + stream carrying `stop_reason: "end_turn"` (a `stop_reason` of `tool_use` + is the model about to call one, not an end), or a `result` line for its + own turn if a future CLI version ever sends one. Either maps to + `Status Exited`; the subagent's vocabulary has no `Idle`, so the + equivalent event `dispatch` produces for an ordinary session is dropped + rather than written. A shipped version of this finished on the + `tool_result` instead, which read a running background agent as + "finished" with its transcript truncated at the moment it launched. +4. **A child line for a subagent that already finished reopens it** + (`Status Running`) rather than being dropped: a background Task can be + sent another message long after its first turn ended, and that is + exactly what a further line for it means. Same transcript, same child + `Translator`, just picking back up. +5. When the parent session's process exits (`Status Exited` on the + session), every subagent still `Running` gets `Status Exited` too: its + process was the parent's. + +A subagent that was mid-flight when the backend restarted keeps working: +the registry reopens the existing transcript on the next child line, and +the file continues its sequence -- the same reopening #4 describes, whether +what closed it was a restart or its own `end_turn`. If its turn ended while +the backend was down nothing recorded that until the next line arrives, so +its last status stays `Running`, which the list reports as **unknown** +rather than as running (see the wire shape) until then. + +Title: the Task call's `description` input, then ` ()` when +one is given; falling back to the tool's name when the child arrives before +(or without) the parent call being seen. + +## Server layout + +- `session/subagent.rs` -- the registry: `Subagents` (per session, in + `Shared`), `Subagent` (its `Transcript` behind a mutex plus a + `broadcast::Sender`), `record(id, event)`, `start(id, title, + prompt)`, `finish(id)`, `reopen(id)`, `finish_all()`, `list()` from disk. Drivers get an + `Arc` beside their `EventSink`; llama ignores it. +- `session/claude/translate.rs` -- routes child lines by parent id, holds + one child `Translator` per subagent, remembers pending Task calls' + description/prompt/subagent_type. +- `session/echo.rs` -- `/subagent [n]`: the test rig. Starts *n* (default 1) + subagents at once, each named "helper k". Each writes the prompt as its + user message, streams a few words of text, runs one `Bash` tool call, then + finishes about three seconds after starting, and the parent's Task calls + end when their subagent does. Three seconds so the running state can be + seen on the phone. +- `routes.rs` -- three routes, in the doc table. + +## Wire shape + +``` +GET /sessions/{id} SessionInfo gains `subagents: N` (count, 0 when none) +GET /sessions same field on each row +GET /sessions/{id}/subagents [{id, title, status, created, lastActivity}], oldest first +GET /sessions/{id}/subagents/{sub}/transcript exactly the session transcript's query and answer +GET /sessions/{id}/subagents/{sub}/events?after=N exactly the session events stream +``` + +`status` is the transcript's last `Status` event, serialised like a session's +(`running`, `exited`), except that a subagent whose session is not itself +running cannot be running: the list answers `unknown` for that one. The +phone words these as *running*, *finished* and *unknown* on the subcard. + +The count on `SessionInfo` is a directory listing, so the list stays cheap. +The per-subagent status is only read when the list route is asked for. + +## Phone + +- `SessionSummary.subagents: Int`. A card with a non-zero count ends in an + expander row -- a full-width `Chevron(Pointing.Down)` row that flips to + `Pointing.Up` -- collapsed by default. Expanding fetches + `/sessions/{id}/subagents` and draws one `OutlinedCard` per subagent, + indented inside the session card, the way dev-updater draws a project's + components: title, then the status word and a relative time. The + expansion state is per session id and survives a refresh of the list. +- Tapping a subcard opens `Screen.Subagent`, which is `SessionScreen` in + **read-only** form: the same transcript, paging, cache, selection, + images and status row, with the composer, the process button, the model + picker, the files button, the settings cog and the usage bar left out. + The header shows the subagent's title with the session's title beneath + it. Back returns to the list. +- Addressing: `fetchTranscript`, `EventStream`, `TranscriptSource` and the + cache take a transcript address rather than a session id -- + `sessions/{id}` or `sessions/{id}/subagents/{sub}` -- so the cache nests a + subagent's copy under its session's and the same code serves both. diff --git a/app/androidApp/src/main/kotlin/com/example/aiapp/Api.kt b/app/androidApp/src/main/kotlin/com/example/aiapp/Api.kt index f1b1466..832f89b 100644 --- a/app/androidApp/src/main/kotlin/com/example/aiapp/Api.kt +++ b/app/androidApp/src/main/kotlin/com/example/aiapp/Api.kt @@ -142,6 +142,20 @@ data class SessionSummary( val keepsOwnTranscript: Boolean, /** How much the session asks before acting; null when it was never set. */ val permissionMode: String?, + /** + * How hard the model thinks, or null for the CLI's own default. + * + * Null is a level somebody can choose, not only one to start in -- see [EFFORT_LEVELS]. It is + * reported rather than assumed for the same reason [permissionMode] is. + */ + val effort: String?, + /** + * Whether a thinking level does anything here -- a Claude CLI session, not a llama or echo one. + * + * Asked of the server rather than worked out from the provider's name, because this is a + * property of the driver's *kind* and the phone only has the name. + */ + val takesEffort: Boolean, /** * Whether this continues a session the machine already had, which changes what deleting means. */ @@ -153,6 +167,24 @@ data class SessionSummary( * itself from a default is one you can turn off while believing you are reading it. */ val notify: Boolean, + /** + * Whether this session sends itself a message once its account's usage limit lifts, and what + * that message says. + * + * The message is what the server would actually send, with its own default already filled in, + * so the field shows the words rather than an empty box standing for them. + */ + val autoResume: Boolean, + val autoResumeMessage: String, + /** + * When the server next intends to check whether the limit has lifted, in epoch seconds, or null + * when nothing is waiting. + * + * A time to *ask*, not a time to resume: the server checks the meter at that moment and waits + * again if the limit is still on. Worded that way wherever it is shown, because a promise this + * app cannot keep is worse than no time at all. + */ + val resumeAt: Double?, /** * The directory the session works in, or null where it was never given one. * @@ -178,8 +210,27 @@ data class SessionSummary( * server because that is where a provider's kind is known. */ val maxImageEdge: Int?, + /** + * Which of `GET /usage`'s snapshots is about this session, and null where nothing meters it. + * + * The rate-limit bar answers a question about an *account*, and what decides which account -- + * if any -- is the provider this session runs, not the machine it runs on. Pairing by machine + * alone drew the Claude CLI's five-hour window under every echo session on a machine that also + * has the CLI: a quota that session cannot spend and could never run down. Decided by the + * server for the same reason [maxImageEdge] is -- it is a fact about the provider's kind, and + * this app has only its name. + */ + val usageProvider: String?, val status: String, val lastActivity: Double, + /** + * How many subagents this session has, however their own status now reads. + * + * A directory listing on the server rather than a status read per subagent, so the list stays + * cheap; the per-subagent state is only fetched when the card is expanded. Zero on a server + * that predates subagents, so this app still opens against one. + */ + val subagents: Int, ) private fun parseSession(session: JSONObject) = @@ -192,14 +243,24 @@ private fun parseSession(session: JSONObject) = title = session.getString("title"), model = session.optString("model").ifEmpty { null }, permissionMode = session.optString("permissionMode").ifEmpty { null }, + effort = session.optString("effort").ifEmpty { null }, + takesEffort = session.optBoolean("takesEffort", false), imported = session.optBoolean("imported", false), notify = session.optBoolean("notify", true), + autoResume = session.optBoolean("autoResume", false), + // The server sends its own default rather than nothing, so an empty answer means an older + // server -- and this app's word for it is the same word. + autoResumeMessage = + session.optString("autoResumeMessage").ifEmpty { DEFAULT_RESUME_MESSAGE }, + resumeAt = if (session.has("resumeAt")) session.getDouble("resumeAt") else null, cwd = session.optString("cwd").ifEmpty { null }, contextTokens = if (session.has("contextTokens")) session.getLong("contextTokens") else null, maxImageEdge = session.optInt("maxImageEdge", 0).takeIf { it > 0 }, + usageProvider = session.optString("usageProvider").ifEmpty { null }, status = session.getString("status"), lastActivity = session.getDouble("lastActivity"), + subagents = session.optInt("subagents", 0), ) fun fetchSessions(settings: ServerSettings): List = @@ -215,6 +276,35 @@ fun fetchSessions(settings: ServerSettings): List = fun fetchSession(settings: ServerSettings, sessionId: String): SessionSummary = requestFromServer(settings, "/sessions/$sessionId") { parseSession(it.jsonObject()) } +/** + * One row of `GET /sessions/{id}/subagents`, oldest first. + * + * A subagent is a second transcript owned by a session -- no process, no controls of its own -- so + * this carries only what a card needs to draw and to open it; see SUBAGENTS.md. [status] is + * "running", "exited" or "unknown": a subagent whose session is not itself running cannot be + * running, and the list says so rather than reporting a state that cannot hold. + */ +data class SubagentSummary( + val id: String, + val title: String, + val status: String, + val created: Double, + val lastActivity: Double, +) + +fun fetchSubagents(settings: ServerSettings, sessionId: String): List = + requestFromServer(settings, "/sessions/$sessionId/subagents") { + it.jsonObjects { row -> + SubagentSummary( + id = row.getString("id"), + title = row.getString("title"), + status = row.getString("status"), + created = row.getDouble("created"), + lastActivity = row.getDouble("lastActivity"), + ) + } + } + // What the server offers, so the spawn screen has no hardcoded lists: a setup added to the server's // config.ron appears here with no app rebuild. // @@ -375,6 +465,12 @@ data class SshDetails( * Where files attached from here land on that machine; null for the session's own directory. */ val attachmentsDir: String? = null, + /** + * Where that machine keeps its GGUF models; null for the same place the backend keeps its own + * (`~/.local/share/ai-app/models`, read on that machine). A llama.cpp session serves the file + * from the machine it runs on, so this is where its models are looked for and listed. + */ + val modelsDir: String? = null, ) private fun SshDetails.toJson() = @@ -382,6 +478,7 @@ private fun SshDetails.toJson() = if (port != null) put("port", port) if (!identityFile.isNullOrBlank()) put("identityFile", identityFile) if (!attachmentsDir.isNullOrBlank()) put("attachmentsDir", attachmentsDir) + if (!modelsDir.isNullOrBlank()) put("modelsDir", modelsDir) } /** What a machine turns out to have, without saving anything. */ @@ -450,6 +547,8 @@ fun spawnSession( model: String? = null, cwd: String? = null, permissionMode: String? = null, + /** Null for whatever the server's default is; see [fetchDefaultEffort]. */ + effort: String? = null, params: Map = emptyMap(), /** Continue this Claude Code session instead of starting an empty one. */ import: String? = null, @@ -467,6 +566,7 @@ fun spawnSession( if (!model.isNullOrBlank()) put("model", model) if (!cwd.isNullOrBlank()) put("cwd", cwd) if (!permissionMode.isNullOrBlank()) put("permissionMode", permissionMode) + if (!effort.isNullOrBlank()) put("effort", effort) if (!import.isNullOrBlank()) put("import", import) if (params.isNotEmpty()) { put("params", JSONObject(params.toMap())) @@ -906,7 +1006,7 @@ fun startImport( */ fun fetchTranscript( settings: ServerSettings, - sessionId: String, + address: TranscriptAddress, before: Long? = null, limit: Int = 80, // Count [limit] in rows, not events, joining a reply's streamed deltas into one -- so a page of @@ -925,7 +1025,7 @@ fun fetchTranscript( if (coalesce) append("&coalesce=true") if (after != null) append("&after=").append(after) } - return requestFromServer(settings, "/sessions/$sessionId/transcript$query") { connection -> + return requestFromServer(settings, "/${address.urlPath}/transcript$query") { connection -> val body = JSONArray(connection.inputStream.bufferedReader().readText()) // The text as well as the event: the transcript cache stores the one and the fold needs the // other, and they have to be the same line. @@ -973,6 +1073,56 @@ fun setSessionModel(settings: ServerSettings, sessionId: String, model: String) */ val PERMISSION_MODES = listOf("manual", "acceptEdits", "auto", "bypassPermissions", "plan") +/** + * What a new session's thinking level is when nothing chose one, or null for the CLI's own. + * + * Held by the server rather than by this phone, because a second device would otherwise spawn + * sessions at a level the first one's owner never picked. + */ +fun fetchDefaultEffort(settings: ServerSettings): String? = + requestFromServer(settings, "/defaults") { + it.jsonObject().optString("effort").ifEmpty { null } + } + +/** Sets what new sessions start at. Nothing already running changes. */ +fun setDefaultEffort(settings: ServerSettings, level: String?) { + requestFromServer( + settings, + "/defaults", + method = "POST", + jsonBody = JSONObject().put("effort", level ?: JSONObject.NULL).toString(), + ) {} +} + +/** + * How hard the model thinks, as `claude --effort` takes them, cheapest first. + * + * Not offered alongside the model and the permission mode on the session's own bar, because it does + * not behave like them: the CLI has a control request for those two and none for this (checked + * against 2.1.258), so a level is settled when the process is launched. Changing it therefore stops + * the process, which is what the working directory beside it in this dialog does, and why it is + * here rather than on a bar whose other controls take effect mid-turn. + */ +val EFFORT_LEVELS = listOf("low", "medium", "high", "xhigh", "max") + +/** What the picker shows, and sends as null, for a session that has chosen no level. */ +const val DEFAULT_EFFORT = "default" + +/** + * Records how hard a session thinks and **stops its process**, since the level is read when the + * process is launched. The next message, or Start, runs one that has it. + * + * [level] is null for the CLI's own default. + */ +fun setSessionEffort(settings: ServerSettings, sessionId: String, level: String?) { + requestFromServer( + settings, + "/sessions/$sessionId/effort", + method = "POST", + jsonBody = JSONObject().put("effort", level ?: JSONObject.NULL).toString(), + ) {} +} + /** Switches how much a running session asks before acting, also in place. */ fun setSessionPermissionMode(settings: ServerSettings, sessionId: String, mode: String) { requestFromServer( @@ -984,6 +1134,36 @@ fun setSessionPermissionMode(settings: ServerSettings, sessionId: String, mode: } /** Turns this session's notifications on or off. Stored on the backend -- see `SessionConfig`. */ +/** + * What an auto-resume says when nothing else was typed. Mirrors the server's own default, so a + * cleared field shows the word that would actually be sent instead of going blank. + */ +const val DEFAULT_RESUME_MESSAGE = "continue" + +/** + * Turns auto-resume on or off and sets what it would say, in one request because they are one + * decision -- see the server's `/sessions/{id}/auto-resume`. + */ +fun setSessionAutoResume( + settings: ServerSettings, + sessionId: String, + autoResume: Boolean, + message: String?, +) { + requestFromServer( + settings, + "/sessions/$sessionId/auto-resume", + method = "POST", + jsonBody = + JSONObject() + .put("autoResume", autoResume) + // Empty means the server's default rather than a session poked with nothing to + // read, which is the same rule the server applies to the field. + .put("message", message?.trim()?.ifEmpty { null } ?: JSONObject.NULL) + .toString(), + ) {} +} + fun setSessionNotify(settings: ServerSettings, sessionId: String, notify: Boolean) { requestFromServer( settings, @@ -1074,6 +1254,26 @@ private fun parseDownload(o: JSONObject) = error = if (o.has("error")) o.getString("error") else null, ) +/** + * The models on one machine, which is the list a llama.cpp session there can choose from. + * + * Not [fetchModels], which is what the *backend* has downloaded. A session serves its model from + * the machine it runs on, so for a machine reached over ssh those are two different lists -- and + * offering the backend's would name files that are not there, turning a choice that cannot work + * into a session that fails when it tries to load one. + */ +fun fetchSetupModels(settings: ServerSettings, setupId: String): List = + requestFromServer(settings, "/setups/${setupId.urlEncoded()}/models") { connection -> + JSONArray(connection.inputStream.bufferedReader().readText()).mapObjects { m -> + LocalModel( + key = m.getString("key"), + repo = m.getString("repo"), + file = m.getString("file"), + bytes = m.getLong("bytes"), + ) + } + } + fun fetchModels(settings: ServerSettings): Models = requestFromServer(settings, "/models") { connection -> val body = JSONObject(connection.inputStream.bufferedReader().readText()) diff --git a/app/androidApp/src/main/kotlin/com/example/aiapp/AppRoot.kt b/app/androidApp/src/main/kotlin/com/example/aiapp/AppRoot.kt index c1b000a..c8efd06 100644 --- a/app/androidApp/src/main/kotlin/com/example/aiapp/AppRoot.kt +++ b/app/androidApp/src/main/kotlin/com/example/aiapp/AppRoot.kt @@ -1,7 +1,9 @@ package com.example.aiapp import androidx.activity.compose.BackHandler +import androidx.compose.foundation.background import androidx.compose.foundation.layout.Box +import androidx.compose.foundation.layout.fillMaxSize import androidx.compose.foundation.layout.imePadding import androidx.compose.foundation.layout.padding import androidx.compose.material3.AlertDialog @@ -34,7 +36,23 @@ import kotlinx.coroutines.withContext * session, spawning one, and settings. */ private sealed class Screen { - data object Main : Screen() + /** + * The session list, with a subagent's own transcript over it when [subagent] is set. + * + * A layer on this screen rather than a screen of its own, for the same reason [Session.files] + * is: [SessionListScreen] owns which cards are expanded and what each expansion fetched, kept + * in `remember`, and a subagent is opened from a card's expander. As a sibling `Screen` it was + * disposed and recreated on every return, which lost that state -- an expanded card collapsed + * itself the moment its own subagent's view was closed. + */ + data class Main(val subagent: SubagentTarget? = null) : Screen() + + /** + * One subagent's own transcript, read-only. See [SessionScreen]'s `subagent` parameter and + * SUBAGENTS.md's "Phone". Closing it returns to [Main] under it, not to [Session]: a subagent + * is opened from the session list's card rather than from inside the session it belongs to. + */ + data class SubagentTarget(val summary: SessionSummary, val subagent: SubagentSummary) /** * One session, with the file explorer over it when [files] is set. @@ -81,7 +99,7 @@ fun AppRoot( val context = LocalContext.current val scope = rememberCoroutineScope() var settings by remember(settingsVersion) { mutableStateOf(loadServerSettings(context)) } - var screen by remember { mutableStateOf(Screen.Main) } + var screen by remember { mutableStateOf(Screen.Main()) } // A notification tap this could not follow, and why. Null both before one is asked for and // after one succeeds, since success is a screen rather than a message. var failedOpen by remember { mutableStateOf(null) } @@ -96,7 +114,7 @@ fun AppRoot( share = shareRequest // A session already open takes it. Otherwise the list is where the choice is made, // whatever screen was showing: Spawn and Settings have nowhere to put a file. - if (screen !is Screen.Session) screen = Screen.Main + if (screen !is Screen.Session) screen = Screen.Main() } } @@ -123,7 +141,7 @@ fun AppRoot( existing = null, onSaved = { saved -> settings = saved - screen = Screen.Main + screen = Screen.Main() }, onBack = null, ) @@ -136,7 +154,7 @@ fun AppRoot( // shows, so it always refetches. val goToMain = { reloadToken++ - screen = Screen.Main + screen = Screen.Main() } if (screen !is Screen.Main) { BackHandler(onBack = goToMain) @@ -185,6 +203,9 @@ fun AppRoot( reloadToken = reloadToken, share = share, onOpen = { screen = Screen.Session(it) }, + onOpenSubagent = { summary, subagent -> + screen = here.copy(subagent = Screen.SubagentTarget(summary, subagent)) + }, onSpawn = { screen = Screen.Spawn }, onImported = { imported -> reloadToken++ @@ -192,6 +213,27 @@ fun AppRoot( }, onSettings = { screen = Screen.Settings }, ) + // Its own back handler is registered after MainScreen's, so it is the one the + // platform asks first while a subagent is open -- the same rule the files + // explorer's handler follows over its session, below. + here.subagent?.let { target -> + BackHandler { screen = here.copy(subagent = null) } + // Its own opaque background: this screen was always the sole content under + // the theme's own Surface before, so it never had to paint one -- stacked over + // the list here, the space between its own cards let the list underneath show + // through without this. The same fix FilesScreen needed over its session. + Box(Modifier.fillMaxSize().background(MaterialTheme.colorScheme.background)) { + key(target.summary.id, target.subagent.id) { + SessionScreen( + settings = current, + summary = target.summary, + onBack = { screen = here.copy(subagent = null) }, + onFiles = {}, + subagent = target.subagent, + ) + } + } + } } is Screen.Session -> // Keyed on the id, because a different session is a different screen rather than this diff --git a/app/androidApp/src/main/kotlin/com/example/aiapp/Dividers.kt b/app/androidApp/src/main/kotlin/com/example/aiapp/Dividers.kt index 01a13ab..57600e1 100644 --- a/app/androidApp/src/main/kotlin/com/example/aiapp/Dividers.kt +++ b/app/androidApp/src/main/kotlin/com/example/aiapp/Dividers.kt @@ -12,6 +12,10 @@ import androidx.compose.ui.Alignment import androidx.compose.ui.Modifier import androidx.compose.ui.graphics.Color import androidx.compose.ui.unit.dp +import java.time.Instant +import java.time.ZoneId +import java.time.format.DateTimeFormatter +import java.time.format.FormatStyle /** * A line across the transcript saying what left the session's context. @@ -47,3 +51,40 @@ fun TranscriptDivider(text: String, color: Color, modifier: Modifier = Modifier) fun ClearedRow(modifier: Modifier = Modifier) { TranscriptDivider("Context cleared", clearedColor, modifier) } + +/** + * The mark running out of quota leaves. + * + * The same red the usage bar takes when a window is spent, because it is the same fact in a second + * place: colour by consequence, so "there is nothing left to spend" is learned once. + * + * A time rather than a countdown. The row is folded once and never re-measured, so a span would go + * stale on screen the moment it was drawn; and this is when the *account* said it would reset, + * which is not a promise about when the session picks back up. A limit the session was told no + * reset time for says nothing about one -- that state has its own words rather than a plausible + * number. + */ +@Composable +fun LimitRow(item: TranscriptItem.LimitNote, modifier: Modifier = Modifier) { + TranscriptDivider(limitSummary(item.resetsAt, ZoneId.systemDefault()), overLimitColor, modifier) +} + +/** + * What the row says. Split out so the wording is testable without a screen, since the two states it + * has to keep apart -- a reset time that arrived and one that never did -- are exactly the pair + * that reads the same when it goes wrong. + * + * [zone] is a parameter rather than read here so a test says the same thing wherever it runs. + */ +fun limitSummary(resetsAt: Double?, zone: ZoneId): String { + val at = resetsAt?.let { + try { + DateTimeFormatter.ofLocalizedTime(FormatStyle.SHORT) + .withZone(zone) + .format(Instant.ofEpochSecond(it.toLong())) + } catch (_: Exception) { + null + } + } + return if (at == null) "Usage limit reached" else "Usage limit reached • resets $at" +} diff --git a/app/androidApp/src/main/kotlin/com/example/aiapp/EventStream.kt b/app/androidApp/src/main/kotlin/com/example/aiapp/EventStream.kt index 748b3fc..6f30cb2 100644 --- a/app/androidApp/src/main/kotlin/com/example/aiapp/EventStream.kt +++ b/app/androidApp/src/main/kotlin/com/example/aiapp/EventStream.kt @@ -14,7 +14,7 @@ private const val RESET_EVENT = "reset" * mean. [close] from any thread ends it, and the caller owns reconnecting -- with the last seq it * saw as the new cursor. */ -class EventStream(settings: ServerSettings, private val sessionId: String) { +class EventStream(settings: ServerSettings, private val address: TranscriptAddress) { private val stream = Sse(settings) fun close() = stream.close() @@ -35,7 +35,7 @@ class EventStream(settings: ServerSettings, private val sessionId: String) { // one and the screen folds the other, and they have to be the same line. onEvent: (raw: String, event: SeqEvent) -> Unit, ) { - stream.run("/sessions/$sessionId/events?after=$after", onOpen) { name, data -> + stream.run("/${address.urlPath}/events?after=$after", onOpen) { name, data -> // A named frame carries no payload and a data frame has no name. if (name == RESET_EVENT) onReset() else if (data.isNotEmpty()) onEvent(data, parseSeqEvent(data)) diff --git a/app/androidApp/src/main/kotlin/com/example/aiapp/Events.kt b/app/androidApp/src/main/kotlin/com/example/aiapp/Events.kt index f266f70..8580867 100644 --- a/app/androidApp/src/main/kotlin/com/example/aiapp/Events.kt +++ b/app/androidApp/src/main/kotlin/com/example/aiapp/Events.kt @@ -157,6 +157,18 @@ sealed class SessionEvent { */ data object Cleared : SessionEvent() + /** + * The session stopped because its account's usage limit was reached. + * + * Its own event rather than an [Error] carrying the CLI's sentence, because it is a state + * rather than something that went wrong -- and because the raw sentence is `Claude AI usage + * limit reached|1788546972`, which is not readable by the person it is shown to. + * + * [resetsAt] is epoch seconds and null where the session was told nothing. Only the server acts + * on it; what this draws it as is a time, not a countdown, because nothing here re-measures it. + */ + data class LimitReached(val resetsAt: Double?) : SessionEvent() + data class Error(val message: String) : SessionEvent() /** @@ -261,6 +273,10 @@ fun parseSeqEvent(json: String): SeqEvent { trigger = body.optString("trigger").ifEmpty { null }, ) "cleared" -> SessionEvent.Cleared + "limitReached" -> + SessionEvent.LimitReached( + if (body.has("resetsAt")) body.getDouble("resetsAt") else null + ) "error" -> SessionEvent.Error(body.getString("message")) else -> SessionEvent.Unknown(type) } diff --git a/app/androidApp/src/main/kotlin/com/example/aiapp/MainScreen.kt b/app/androidApp/src/main/kotlin/com/example/aiapp/MainScreen.kt index f18ef22..4252ac4 100644 --- a/app/androidApp/src/main/kotlin/com/example/aiapp/MainScreen.kt +++ b/app/androidApp/src/main/kotlin/com/example/aiapp/MainScreen.kt @@ -48,6 +48,8 @@ fun MainScreen( /** What another app shared in and no session has taken yet; see [ShareRequest]. */ share: ShareRequest? = null, onOpen: (SessionSummary) -> Unit, + /** Opens one session's subagent, from the expander under its card. */ + onOpenSubagent: (SessionSummary, SubagentSummary) -> Unit, onSpawn: () -> Unit, onImported: (SessionSummary) -> Unit, onSettings: () -> Unit, @@ -139,6 +141,7 @@ fun MainScreen( settings = settings, reloadToken = token, onOpen = onOpen, + onOpenSubagent = onOpenSubagent, onSpawn = onSpawn, ) MainTab.Import -> diff --git a/app/androidApp/src/main/kotlin/com/example/aiapp/SessionListScreen.kt b/app/androidApp/src/main/kotlin/com/example/aiapp/SessionListScreen.kt index 1c3cc2b..27faf3f 100644 --- a/app/androidApp/src/main/kotlin/com/example/aiapp/SessionListScreen.kt +++ b/app/androidApp/src/main/kotlin/com/example/aiapp/SessionListScreen.kt @@ -1,7 +1,9 @@ package com.example.aiapp import androidx.compose.foundation.ExperimentalFoundationApi +import androidx.compose.foundation.clickable import androidx.compose.foundation.combinedClickable +import androidx.compose.foundation.layout.Arrangement import androidx.compose.foundation.layout.Box import androidx.compose.foundation.layout.Column import androidx.compose.foundation.layout.Row @@ -9,6 +11,7 @@ import androidx.compose.foundation.layout.Spacer import androidx.compose.foundation.layout.fillMaxSize import androidx.compose.foundation.layout.fillMaxWidth import androidx.compose.foundation.layout.height +import androidx.compose.foundation.layout.heightIn import androidx.compose.foundation.layout.padding import androidx.compose.foundation.layout.width import androidx.compose.foundation.lazy.LazyColumn @@ -17,6 +20,7 @@ import androidx.compose.material3.Card import androidx.compose.material3.CircularProgressIndicator import androidx.compose.material3.FloatingActionButton import androidx.compose.material3.MaterialTheme +import androidx.compose.material3.OutlinedCard import androidx.compose.material3.Switch import androidx.compose.material3.Text import androidx.compose.material3.TextButton @@ -30,6 +34,8 @@ import androidx.compose.runtime.setValue import androidx.compose.ui.Alignment import androidx.compose.ui.Modifier import androidx.compose.ui.platform.LocalContext +import androidx.compose.ui.semantics.contentDescription +import androidx.compose.ui.semantics.semantics import androidx.compose.ui.unit.dp import kotlinx.coroutines.Dispatchers import kotlinx.coroutines.launch @@ -47,12 +53,40 @@ fun SessionListScreen( settings: ServerSettings, reloadToken: Int, onOpen: (SessionSummary) -> Unit, + /** Opens one session's subagent, from the expander under its card. */ + onOpenSubagent: (SessionSummary, SubagentSummary) -> Unit, onSpawn: () -> Unit, ) { val scope = rememberCoroutineScope() var listState by remember { mutableStateOf>>(LoadState.Loading) } var confirmingDelete by remember { mutableStateOf(null) } + // Which session cards are expanded to show their subagents, and what each expansion fetched. + // Ids rather than a flag on the row for the same reason `deleting` is: the rows are rebuilt + // from + // whatever the server last said, and this belongs to the reader's own choice, which survives a + // refresh. + var expandedSessions by remember { mutableStateOf(setOf()) } + var subagentLoads by remember { + mutableStateOf(mapOf>>()) + } + + fun loadSubagents(sessionId: String) { + subagentLoads = subagentLoads + (sessionId to LoadState.Loading) + scope.launch { + subagentLoads = + subagentLoads + + (sessionId to + try { + LoadState.Loaded( + withContext(Dispatchers.IO) { fetchSubagents(settings, sessionId) } + ) + } catch (e: ApiException) { + LoadState.failed(e) + }) + } + } + // Failures that belong to one session rather than to the list, keyed by its id and shown on its // own card. The two scopes are decided by whether the server answered: it answered and refused, // so this says nothing about the other rows. @@ -84,6 +118,14 @@ fun SessionListScreen( withContext(Dispatchers.IO) { transcriptCache.retainOnly(loaded.value.map { it.id }.toSet()) } + // A session gone from this answer cannot still be expanded, and an expanded one + // that is still here asks again -- its subagents may have changed since the + // last + // fetch. + val ids = loaded.value.map { it.id }.toSet() + expandedSessions = expandedSessions intersect ids + subagentLoads = subagentLoads.filterKeys { it in ids } + expandedSessions.forEach(::loadSubagents) loaded } catch (e: ApiException) { LoadState.failed(e) @@ -127,6 +169,17 @@ fun SessionListScreen( deleting = session.id in deleting, onOpen = { onOpen(session) }, onLongPress = { confirmingDelete = session }, + expanded = session.id in expandedSessions, + subagents = subagentLoads[session.id], + onToggleSubagents = { + if (session.id in expandedSessions) { + expandedSessions = expandedSessions - session.id + } else { + expandedSessions = expandedSessions + session.id + loadSubagents(session.id) + } + }, + onOpenSubagent = { subagent -> onOpenSubagent(session, subagent) }, ) Spacer(Modifier.height(12.dp)) } @@ -225,7 +278,7 @@ fun SessionListScreen( deleteSession(settings, session.id, alsoDeleteForeign) // After it succeeded, not before: a refused delete leaves the // session exactly as it was, and its transcript with it. - transcriptCache.session(session.id).purge() + transcriptCache.session(TranscriptAddress(session.id)).purge() } // Only this row, and only what changed. Refetching the list instead // put every other session back through loading and handed the @@ -276,6 +329,12 @@ private fun SessionCard( deleting: Boolean, onOpen: () -> Unit, onLongPress: () -> Unit, + /** Whether the expander below is open. Collapsed by default; see [SessionListScreen]. */ + expanded: Boolean, + /** What the expander's own fetch answered, or null before it has been asked. */ + subagents: LoadState>?, + onToggleSubagents: () -> Unit, + onOpenSubagent: (SubagentSummary) -> Unit, ) { BusyItem(label = if (deleting) "deleting" else null) { Card( @@ -332,11 +391,105 @@ private fun SessionCard( color = MaterialTheme.colorScheme.error, ) } + // Nothing at all for a card with no subagents: a disabled expander here would be + // noise on every ordinary session's card. Its own row at the bottom rather than + // beside the title or the machine line, so opening it never displaces text that was + // already on screen -- see UI_RULES on a control not displacing the text beside it. + if (session.subagents > 0) { + Spacer(Modifier.height(8.dp)) + // The platform's minimum touch height, not the chevron's own ten or so dp: + // at the chevron's height a tap meant for it landed on the first subcard + // beneath and opened a subagent instead. + Row( + horizontalArrangement = Arrangement.Center, + verticalAlignment = Alignment.CenterVertically, + modifier = + Modifier.fillMaxWidth() + .heightIn(min = 48.dp) + .clickable(enabled = !deleting, onClick = onToggleSubagents) + .semantics { + contentDescription = + if (expanded) "Collapse subagents" else "Expand subagents" + }, + ) { + Chevron(if (expanded) Pointing.Up else Pointing.Down) + } + if (expanded) { + Spacer(Modifier.height(4.dp)) + Column(verticalArrangement = Arrangement.spacedBy(8.dp)) { + when (subagents) { + null, + is LoadState.Loading -> + CircularProgressIndicator( + modifier = Modifier.width(20.dp).height(20.dp), + strokeWidth = 2.dp, + ) + is LoadState.Error -> + // Said here rather than left silent: a fetch that failed and an + // expander that simply found nothing must not look the same -- + // see UI_RULES on designing the unknown state first. + Text( + subagents.message, + style = MaterialTheme.typography.bodySmall, + color = MaterialTheme.colorScheme.error, + ) + is LoadState.Loaded -> + subagents.value.forEach { subagent -> + SubagentCard( + subagent, + onClick = { onOpenSubagent(subagent) }, + ) + } + } + } + } + } } } } } +/** + * One subagent, indented inside its session's card -- the way dev-updater draws a project's + * components (`ComponentCard`, `UpdaterScreen.kt`): an outlined card, not the session card's own + * filled one, so the nesting reads as one step rather than as another session. + */ +@Composable +private fun SubagentCard(subagent: SubagentSummary, onClick: () -> Unit) { + OutlinedCard(Modifier.fillMaxWidth().clickable(onClick = onClick)) { + Column(Modifier.padding(horizontal = 12.dp, vertical = 8.dp)) { + Text(subagent.title, style = MaterialTheme.typography.titleSmall) + Spacer(Modifier.height(2.dp)) + Row(modifier = Modifier.fillMaxWidth()) { + Text( + subagentStatusLabel(subagent.status), + style = MaterialTheme.typography.bodySmall, + color = MaterialTheme.colorScheme.onSurfaceVariant, + modifier = Modifier.weight(1f), + ) + Text( + relativeTime(subagent.lastActivity), + style = MaterialTheme.typography.bodySmall, + color = MaterialTheme.colorScheme.onSurfaceVariant, + ) + } + } + } +} + +/** + * The subcard's word for a subagent's status -- see SUBAGENTS.md's "Wire shape". Its own function + * rather than a branch inside [StatusText], because a subagent's three states are not that + * composable's five: "exited" reads as "finished" here, since its process was always its parent's + * and never something of its own to have merely stopped. + */ +private fun subagentStatusLabel(status: String) = + when (status) { + "running" -> "running" + "exited" -> "finished" + else -> "unknown" + } + @Composable fun StatusText(status: String) { val (label, color) = diff --git a/app/androidApp/src/main/kotlin/com/example/aiapp/SessionScreen.kt b/app/androidApp/src/main/kotlin/com/example/aiapp/SessionScreen.kt index 1d674cf..5d96007 100644 --- a/app/androidApp/src/main/kotlin/com/example/aiapp/SessionScreen.kt +++ b/app/androidApp/src/main/kotlin/com/example/aiapp/SessionScreen.kt @@ -216,16 +216,33 @@ fun SessionScreen( share: ShareRequest? = null, /** Said once [share] has been attached here, so it is not attached again. */ onShareTaken: () -> Unit = {}, + /** + * Draws this screen read-only, on a subagent's own transcript instead of the session's. + * + * A subagent has no process and no controls of its own -- see SUBAGENTS.md's "Phone" -- so + * every gate below keyed on this switches off the composer, the files button, the settings cog, + * the usage bar and notifications, while everything that draws a transcript (paging, cache, + * selection, images, the status row, stream reconnects) is reused unchanged, pointed at + * [address] instead of the session's own. + */ + subagent: SubagentSummary? = null, ) { DebugStats.count("session screen recomposed") + val isSubagent = subagent != null + val address = TranscriptAddress(summary.id, subagent?.id) val scope = rememberCoroutineScope() val topEdgeHeld = remember { TopEdgeHold() } var items by remember { mutableStateOf(listOf()) } - var status by remember { mutableStateOf(summary.status) } + var status by remember { mutableStateOf(subagent?.status ?: summary.status) } // Seeded from the row this screen was opened from, so a conversation already under way says how // much it is holding before any turn happens here. Null is "nobody has measured it", which is a // different answer from an empty context and is drawn differently. - var contextTokens by remember(summary.id) { mutableStateOf(summary.contextTokens) } + // + // A subagent has no context measurement of its own, so it always starts unmeasured rather than + // borrowing the parent session's figure -- see UI_RULES on not showing an inferred value as one + // that was measured. + var contextTokens by + remember(address) { mutableStateOf(if (isSubagent) null else summary.contextTokens) } // When the current compaction started. The moment comes off the `compacting` status event // itself -- the server timestamps every transcript line -- rather than off this device noticing // one, which is what makes it survive leaving the session and reopening it. @@ -241,7 +258,13 @@ fun SessionScreen( val context = LocalContext.current // Seeded from what was left in the box last time and written back on every keystroke, so // leaving the screen does not throw away a half-typed message. See `Drafts.kt`. - var input by remember(summary.id) { mutableStateOf(atEnd(loadDraft(context, summary.id))) } + // + // A subagent has no box to type into, so it never touches a draft at all -- not this session's, + // which is what reading one keyed only by `summary.id` would do here. + var input by + remember(summary.id) { + mutableStateOf(if (isSubagent) atEnd("") else atEnd(loadDraft(context, summary.id))) + } // A model the reader has chosen and not yet confirmed. See [ModelSwitchWarning]: switching // makes the session re-read the whole conversation. var pendingModel by remember { mutableStateOf(null) } @@ -294,27 +317,26 @@ fun SessionScreen( // Reload throws away what it was reading from. val cache = remember(settings) { TranscriptCache(cacheRoot(context, settings)) } val source = - remember(summary.id, epoch) { - TranscriptSource(settings, summary.id, cache.session(summary.id)) - } + remember(address, epoch) { TranscriptSource(settings, address, cache.session(address)) } // Whether the cached tail has been shown to still be the server's own line. Nothing is resumed // from a cached cursor until it has, and a probe that could not be made leaves this false for // the stream loop to try again. - var probePassed by remember(summary.id, epoch) { mutableStateOf(false) } + var probePassed by remember(address, epoch) { mutableStateOf(false) } // Whether the opening effect is still settling that question. It draws the cached rows and // lifts [ready] before the answer arrives, which is the point of the cache -- so the stream // below waits for this rather than for `ready`, or it asks the same question twice. - var probing by remember(summary.id, epoch) { mutableStateOf(true) } + var probing by remember(address, epoch) { mutableStateOf(true) } // The oldest sequence number loaded, and whether there is more behind it. Paging backwards is // what keeps opening a long session cheap. var oldestSeq by remember { mutableLongStateOf(0L) } - // Where this session was last being read, from this device's own store. Read once, because the - // answer stops being interesting the moment the list is on screen. - val savedAnchor = remember(summary.id, epoch) { loadScrollAnchor(context, summary.id) } + // Where this transcript was last being read, from this device's own store, keyed by the address + // rather than the session id so a subagent's saved position cannot collide with its session's. + // Read once, because the answer stops being interesting the moment the list is on screen. + val savedAnchor = remember(address, epoch) { loadScrollAnchor(context, address.cachePath) } // Whether the saved position is still being put back. Nothing is drawn while it is: opening at // the newest end and then travelling to the anchor is exactly the journey a reader must never // see. - var restoring by remember(summary.id, epoch) { mutableStateOf(savedAnchor != null) } + var restoring by remember(address, epoch) { mutableStateOf(savedAnchor != null) } // Messages the server has taken and the session has not read yet, by the id that will resolve // them. From the event stream rather than from what this screen sent, so they survive leaving // the session -- and a message sent from another device is drawn waiting on this one too. @@ -327,11 +349,11 @@ fun SessionScreen( var loadingHistory by remember { mutableStateOf(false) } var ready by remember { mutableStateOf(false) } // Replies parsed ahead of the rows that draw them; see [ParsedReplies]. - val replies = remember(summary.id) { ParsedReplies() } - // Keyed like everything else describing one session's transcript. `rememberLazyListState` saves - // through `rememberSaveable`, and this screen restores by its own anchor instead -- two - // restores would fight over the first frame. - val listState = remember(summary.id) { LazyListState() } + val replies = remember(address) { ParsedReplies() } + // Keyed like everything else describing one transcript. `rememberLazyListState` saves through + // `rememberSaveable`, and this screen restores by its own anchor instead -- two restores would + // fight over the first frame. + val listState = remember(address) { LazyListState() } // Whether the newest message is on screen right now. The list is reversed, so the newest end is // the scrolling start: nothing behind you is exactly being at the bottom. Asked of the scroll // state rather than of item indices, because a zero-height first item makes an index ambiguous. @@ -637,7 +659,7 @@ fun SessionScreen( // ended and carries live events only. The window comes from this phone's own copy when there is // one, and then costs a single request to check that the server's transcript is still the one // it came from. See TRANSCRIPT_CACHE.md. - LaunchedEffect(summary.id, epoch) { + LaunchedEffect(address, epoch) { /** * One opening window onto the screen, whichever side it came from. * @@ -667,11 +689,16 @@ fun SessionScreen( // A replay is as old as the last visit; the row this screen was opened from was // fetched moments ago. So the transcript comes from the cache and everything that // is not the transcript comes from the summary -- otherwise a session that finished - // an hour ago opens saying "working" until the stream connects. - status = summary.status - model = summary.model - permissionMode = summary.permissionMode ?: "auto" - if (summary.status != "compacting") compactingSince = null + // an hour ago opens saying "working" until the stream connects. A subagent's status + // comes from its own summary, never the parent session's: they are two different + // things running or not, and the parent's model and permission mode do not apply to + // it at all. + status = subagent?.status ?: summary.status + if (!isSubagent) { + model = summary.model + permissionMode = summary.permissionMode ?: "auto" + } + if (status != "compacting") compactingSince = null // Nothing to put back, so these rows are the screen and the probe can return under // them. A restore still has history to fetch and is gated below. if (savedAnchor == null) ready = true @@ -798,7 +825,7 @@ fun SessionScreen( // at the top on their return. Switching apps is a choice somebody made, not a fault to report. // Stopping the stream deliberately makes the drop a close rather than an error, and resuming // reconnects from the same cursor. - LaunchedEffect(summary.id, ready, epoch, lifecycleOwner) { + LaunchedEffect(address, ready, epoch, lifecycleOwner) { if (!ready) return@LaunchedEffect // The opening effect draws cached rows and lifts `ready` *before* it has checked that the // cursor under them is still the server's, so `ready` is no longer the whole gate. Without @@ -868,17 +895,22 @@ fun SessionScreen( // The screen going away entirely, which the lifecycle scope above does not cover: a composable // can leave the composition while the activity stays started. Keyed on the epoch as well, so // Reload's replacement source is the one a later disposal closes. - DisposableEffect(summary.id, epoch) { onDispose { source.close() } } + DisposableEffect(address, epoch) { onDispose { source.close() } } // Nothing gets announced about the session somebody is reading; see NotificationService. // RESUMED rather than STARTED because "looking at it" means the foreground. - LaunchedEffect(summary.id, lifecycleOwner) { - lifecycleOwner.repeatOnLifecycle(Lifecycle.State.RESUMED) { - NotificationService.showing(context, summary.id) - try { - awaitCancellation() - } finally { - NotificationService.stoppedShowing(summary.id) + // + // Not for a subagent: it has no notifications of its own, and it is not the session this would + // otherwise mark as being read. + if (!isSubagent) { + LaunchedEffect(summary.id, lifecycleOwner) { + lifecycleOwner.repeatOnLifecycle(Lifecycle.State.RESUMED) { + NotificationService.showing(context, summary.id) + try { + awaitCancellation() + } finally { + NotificationService.stoppedShowing(summary.id) + } } } } @@ -924,7 +956,7 @@ fun SessionScreen( val (index, offset, awayFromNewest) = settled saveScrollAnchor( context, - summary.id, + address.cachePath, // Nothing to restore at the newest end, which is where a session with no anchor // opens anyway. One *before* the index, because item zero is the "below" slot. if (!awayFromNewest) null @@ -947,7 +979,7 @@ fun SessionScreen( // // There is no correction beside this one. Following the newest message is not an effect: the // list is reversed, so an arriving message extends the end the viewport is pinned to. - val unitSizes = remember(summary.id) { HashMap() } + val unitSizes = remember(address) { HashMap() } LaunchedEffect(listState, moreHistory) { snapshotFlow { listState.layoutInfo } .collect { info -> @@ -983,21 +1015,25 @@ fun SessionScreen( } } - LaunchedEffect(summary.setupName, summary.provider) { - offeredModels = - try { - withContext(Dispatchers.IO) { - fetchSetups(settings) - .firstOrNull { it.name == summary.setupName } - ?.providers - ?.firstOrNull { it.name == summary.provider } - ?.models - .orEmpty() + // Only for the model picker, which a subagent does not have. + if (!isSubagent) { + LaunchedEffect(summary.setupName, summary.provider) { + offeredModels = + try { + withContext(Dispatchers.IO) { + fetchSetups(settings) + .firstOrNull { it.name == summary.setupName } + ?.providers + ?.firstOrNull { it.name == summary.provider } + ?.models + .orEmpty() + } + } catch (_: Exception) { + // Not worth reporting: the picker simply has nothing to offer, which is + // visible. + emptyList() } - } catch (_: Exception) { - // Not worth reporting: the picker simply has nothing to offer, which is visible. - emptyList() - } + } } /** @@ -1148,8 +1184,9 @@ fun SessionScreen( } // One poll for the machines' limits, read by everything on this screen that reports them. - val usageFeed = rememberUsageFeed(settings) - val usage = usageFeed.forSetup(summary.setup) + // Nothing meters a subagent -- it has no account of its own -- so it never starts this poll. + val usageFeed = if (isSubagent) null else rememberUsageFeed(settings) + val usage = usageFeed?.forSession(summary) ?: SessionUsage.NotMetered RecordFrames() var usageOpen by remember { mutableStateOf(false) } var settingsOpen by remember { mutableStateOf(false) } @@ -1256,20 +1293,35 @@ fun SessionScreen( // A ring's worth, which is what the arrow already keeps on its other three sides. Spacer(Modifier.width(GLYPH_BUTTON_MARGIN)) Column(Modifier.weight(1f)) { - Text(title, style = MaterialTheme.typography.titleMedium) - // Machine first, then what runs on it -- the same order and the same wording - // everywhere this pair appears, so it reads as one fact rather than two - // sentences with different grammar. - // - // No model. The picker in the footer already shows what this session is set to, - // and showing it twice means two things to keep in step -- they disagreed for a - // moment on every model change. - Text( - "${summary.setupName} · ${summary.provider}", - style = MaterialTheme.typography.bodySmall, - color = MaterialTheme.colorScheme.onSurfaceVariant, - ) + // A subagent's own title, with the session's beneath it in a smaller style -- + // the header says whose conversation this is as well as what it is. Otherwise + // just the session's title, as before. + if (subagent != null) { + Text(subagent.title, style = MaterialTheme.typography.titleMedium) + Text( + title, + style = MaterialTheme.typography.bodySmall, + color = MaterialTheme.colorScheme.onSurfaceVariant, + ) + } else { + Text(title, style = MaterialTheme.typography.titleMedium) + // Machine first, then what runs on it -- the same order and the same + // wording everywhere this pair appears, so it reads as one fact rather than + // two sentences with different grammar. + // + // No model. The picker in the footer already shows what this session is set + // to, and showing it twice means two things to keep in step -- they + // disagreed for a moment on every model change. + Text( + "${summary.setupName} · ${summary.provider}", + style = MaterialTheme.typography.bodySmall, + color = MaterialTheme.colorScheme.onSurfaceVariant, + ) + } } + // None of this is a subagent's: it has no files of its own to browse, no settings, + // and nothing meters it -- see SUBAGENTS.md's "Phone". + // // Beside the provider it reports on, which is the line directly to its left. Its // real home is this provider's settings, which do not exist yet. A session on a // provider with no such service gets an honest "unavailable" rather than a hidden @@ -1284,42 +1336,47 @@ fun SessionScreen( // Usage, files, settings -- widest scope first, narrowing to the right, so the cog // stays at the end where every other screen keeps it. Asked for in this order by // Iris on 2026-09-03. - Row { - GlyphButton( - USAGE_GLYPH, - "Usage", - { usageOpen = true }, - colour = usageGlyphColour(usage), - ) - // The machine's files, which is where the answer to "what did it actually - // change" is. It opens *over* this screen rather than replacing it. - GlyphButton( - FOLDER_GLYPH, - "Files", - onClick = { - onFiles( - FilesTarget( - setup = summary.setup, - setupName = summary.setupName, - // Where this session works, and the machine's own home when it - // was never given a directory -- resolved there rather than - // guessed at here, since this app does not know that home. - start = summary.cwd?.takeIf { it.isNotBlank() } ?: "~", + if (!isSubagent) { + Row { + GlyphButton( + USAGE_GLYPH, + "Usage", + { usageOpen = true }, + colour = usageGlyphColour(usage), + ) + // The machine's files, which is where the answer to "what did it actually + // change" is. It opens *over* this screen rather than replacing it. + GlyphButton( + FOLDER_GLYPH, + "Files", + onClick = { + onFiles( + FilesTarget( + setup = summary.setup, + setupName = summary.setupName, + // Where this session works, and the machine's own home when + // it was never given a directory -- resolved there rather + // than guessed at here, since this app does not know that + // home. + start = summary.cwd?.takeIf { it.isNotBlank() } ?: "~", + ) ) - ) - }, - ) - // What it opens is about this session, so it sits at the end of the session's - // own row. A cog and not a word because there will be more, and a bar of words - // has nowhere to put it. - GlyphButton(SETTINGS_GLYPH, "Session settings", { settingsOpen = true }) + }, + ) + // What it opens is about this session, so it sits at the end of the + // session's own row. A cog and not a word because there will be more, and a + // bar of words has nowhere to put it. + GlyphButton(SETTINGS_GLYPH, "Session settings", { settingsOpen = true }) + } } } // Under the header, above everything the session itself says: it is a fact about the // machine rather than a turn in the conversation, and it is the number that decides - // whether to keep going. - SessionUsageBar(usage) + // whether to keep going. Nothing meters a subagent. + if (!isSubagent) { + SessionUsageBar(usage) + } (streamError ?: actionError)?.let { message -> Text( @@ -1565,6 +1622,7 @@ fun SessionScreen( is TranscriptItem.ClearedNote -> ClearedRow() is TranscriptItem.CompactedNote -> CompactedRow(item) + is TranscriptItem.LimitNote -> LimitRow(item) // Never reached: a peer message is flattened into // its own units. Here because a `when` over the // item kinds has to stay exhaustive. @@ -1663,185 +1721,208 @@ fun SessionScreen( ) } + // Kept for a subagent -- see SUBAGENTS.md's "Phone" -- with the wording that turns + // "exited" into "finished" for one, since it has no process to leave running or stop. SessionStatusRow( status = status, compactingFor = compactingFor, contextTokens = contextTokens, + subagent = isSubagent, ) - // Between the transcript and the box: above what is being typed, so the list does not - // cover the thing the command is about, and below everything that explains it. - CommandSuggestions( - // Nothing to suggest about a suggestion that was just taken. `/compact` is a whole - // command *and* a prefix of itself, so picking it left the list standing there with - // the one row already chosen. Held by what was picked rather than by a flag, so - // typing anything else brings the list back without a second thing to reset. - commands = if (input.text == picked) emptyList() else suggestedCommands(input.text), - onPick = { command -> - // At the end of what was inserted, which is where the reader carries on typing: - // a command with an argument is put in the box half-written, and a cursor left - // at the front makes the next keystroke the first character of "/rename". - input = atEnd(command.typed()) - picked = command.typed() - }, - ) - - // Always enabled -- a send while the session is running becomes a steering message - // injected at the next tool boundary, which is the point of the whole app. - // - // The field gets a row of its own, above the buttons: sharing one put the full width - // behind three controls, so the thing being typed into was the narrowest on the row. - Column(Modifier.fillMaxWidth().padding(8.dp)) { - // Directly above the box they will be sent from, so what is attached is visible - // rather than counted: the "+2" on the button below said how many and never which. - PendingAttachments( - settings = settings, - sessionId = summary.id, - refs = pendingAttachments, - onRemove = { pendingAttachments = pendingAttachments - it }, - ) - OutlinedTextField( - value = input, - onValueChange = { - input = it - saveDraft(context, summary.id, it.text) + // Everything from here down is the composer: a subagent cannot be messaged, so none of + // it applies -- see SUBAGENTS.md's "Phone". + if (!isSubagent) { + // Between the transcript and the box: above what is being typed, so the list does + // not cover the thing the command is about, and below everything that explains it. + CommandSuggestions( + // Nothing to suggest about a suggestion that was just taken. `/compact` is a + // whole command *and* a prefix of itself, so picking it left the list standing + // there with the one row already chosen. Held by what was picked rather than by + // a flag, so typing anything else brings the list back without a second thing + // to + // reset. + commands = + if (input.text == picked) emptyList() else suggestedCommands(input.text), + onPick = { command -> + // At the end of what was inserted, which is where the reader carries on + // typing: a command with an argument is put in the box half-written, and a + // cursor left at the front makes the next keystroke the first character of + // "/rename". + input = atEnd(command.typed()) + picked = command.typed() }, - modifier = Modifier.fillMaxWidth(), - // No longer "(+image)": the images are on screen above this, and a placeholder - // saying so said it in words beside the thing itself. - placeholder = { Text("Message") }, - maxLines = 4, ) - Row( - verticalAlignment = Alignment.CenterVertically, - modifier = Modifier.fillMaxWidth(), - ) { - // Photo or file, asked here rather than by two buttons: the row is full, and - // attaching is one action whichever picker answers it. - var attaching by remember { mutableStateOf(false) } - Box { - // Just "+". The count it used to carry was standing in for showing them. - BubbleButton(onClick = { attaching = true }) { Text("+") } - DropdownMenu( - expanded = attaching, - onDismissRequest = { attaching = false }, - // See PickerButton: without this the menu opens a status bar's height - // away from the button in an edge-to-edge activity. - properties = PopupProperties(clippingEnabled = false), - shape = BubbleMenuShape, - ) { - DropdownMenuItem( - text = { Text("Photo") }, - onClick = { - attaching = false - pickImage.launch( - PickVisualMediaRequest( - ActivityResultContracts.PickVisualMedia.ImageOnly - ) - ) - }, - ) - DropdownMenuItem( - text = { Text("File") }, - onClick = { - attaching = false - pickFile.launch(arrayOf("*/*")) - }, - ) - } - } - // The settings share what is left after the actions have taken what they need. - // A Row hands out intrinsic widths in order and clips whatever runs past the - // edge, so with these laid out first the arrival of Stop pushed Send off the - // screen entirely -- the app's central control, gone at the moment it is most - // in use. + + // Always enabled -- a send while the session is running becomes a steering message + // injected at the next tool boundary, which is the point of the whole app. + // + // The field gets a row of its own, above the buttons: sharing one put the full + // width + // behind three controls, so the thing being typed into was the narrowest on the + // row. + Column(Modifier.fillMaxWidth().padding(8.dp)) { + // Directly above the box they will be sent from, so what is attached is visible + // rather than counted: the "+2" on the button below said how many and never + // which. + PendingAttachments( + settings = settings, + sessionId = summary.id, + refs = pendingAttachments, + onRemove = { pendingAttachments = pendingAttachments - it }, + ) + OutlinedTextField( + value = input, + onValueChange = { + input = it + saveDraft(context, summary.id, it.text) + }, + modifier = Modifier.fillMaxWidth(), + // No longer "(+image)": the images are on screen above this, and a + // placeholder saying so said it in words beside the thing itself. + placeholder = { Text("Message") }, + maxLines = 4, + ) Row( verticalAlignment = Alignment.CenterVertically, - modifier = Modifier.weight(1f), + modifier = Modifier.fillMaxWidth(), ) { - if (offeredModels.isNotEmpty()) { + // Photo or file, asked here rather than by two buttons: the row is full, + // and + // attaching is one action whichever picker answers it. + var attaching by remember { mutableStateOf(false) } + Box { + // Just "+". The count it used to carry was standing in for showing + // them. + BubbleButton(onClick = { attaching = true }) { Text("+") } + DropdownMenu( + expanded = attaching, + onDismissRequest = { attaching = false }, + // See PickerButton: without this the menu opens a status bar's + // height away from the button in an edge-to-edge activity. + properties = PopupProperties(clippingEnabled = false), + shape = BubbleMenuShape, + ) { + DropdownMenuItem( + text = { Text("Photo") }, + onClick = { + attaching = false + pickImage.launch( + PickVisualMediaRequest( + ActivityResultContracts.PickVisualMedia.ImageOnly + ) + ) + }, + ) + DropdownMenuItem( + text = { Text("File") }, + onClick = { + attaching = false + pickFile.launch(arrayOf("*/*")) + }, + ) + } + } + // The settings share what is left after the actions have taken what they + // need. A Row hands out intrinsic widths in order and clips whatever runs + // past the edge, so with these laid out first the arrival of Stop pushed + // Send off the screen entirely -- the app's central control, gone at the + // moment it is most in use. + Row( + verticalAlignment = Alignment.CenterVertically, + modifier = Modifier.weight(1f), + ) { + if (offeredModels.isNotEmpty()) { + PickerButton( + current = modelLabel(model), + // What the machine offers, plus the state a session is in when + // it has chosen none of them. The button has always been able + // to + // say "default"; until this the list could not, so leaving it + // was a one-way trip. + options = listOf(DEFAULT_MODEL) + offeredModels, + // Not set here. The button follows what the session reports it + // is set to, which arrives a moment later and is sometimes a + // different answer -- a name the CLI resolved, or no change at + // all on a provider whose model is fixed. Asked about first, + // unless there is nothing to lose by it -- see + // [ModelSwitchWarning]. + onPick = { chosen -> + if ( + modelLabel(chosen) == modelLabel(model) || + !worthWarningAbout(status, contextTokens, items) + ) { + act { setSessionModel(settings, summary.id, chosen) } + } else { + pendingModel = chosen + } + }, + ) + } PickerButton( - current = modelLabel(model), - // What the machine offers, plus the state a session is in when it - // has chosen none of them. The button has always been able to say - // "default"; until this the list could not, so leaving it was a - // one-way trip. - options = listOf(DEFAULT_MODEL) + offeredModels, - // Not set here. The button follows what the session reports it is - // set to, which arrives a moment later and is sometimes a different - // answer -- a name the CLI resolved, or no change at all on a - // provider whose model is fixed. Asked about first, unless there is - // nothing to lose by it -- see [ModelSwitchWarning]. + current = permissionMode, + options = PERMISSION_MODES, onPick = { chosen -> - if ( - modelLabel(chosen) == modelLabel(model) || - !worthWarningAbout(status, contextTokens, items) - ) { - act { setSessionModel(settings, summary.id, chosen) } - } else { - pendingModel = chosen - } + act { setSessionPermissionMode(settings, summary.id, chosen) } }, ) } - PickerButton( - current = permissionMode, - options = PERMISSION_MODES, - onPick = { chosen -> - act { setSessionPermissionMode(settings, summary.id, chosen) } - }, - ) - } - // The same filled shape as the button beside it, not an outlined one: these are - // two things you can do about the session, and weighting one as secondary said - // they were a primary action and its qualifier. What separates them is the - // colour and the mark, which is what they mean. - // - // Always here, rather than arriving with the turn as it used to. A control that - // comes and goes makes its own presence the signal, and a button always in the - // same place also cannot push Send off the end of the row by turning up. - val process = - when { - running -> ProcessAction.Pause - status == "exited" -> ProcessAction.Start - else -> ProcessAction.Stop - } - Button( - onClick = { - processInFlight = true - act(onDone = { processInFlight = false }) { - process.perform(settings, summary.id) + // The same filled shape as the button beside it, not an outlined one: these + // are two things you can do about the session, and weighting one as + // secondary said they were a primary action and its qualifier. What + // separates them is the colour and the mark, which is what they mean. + // + // Always here, rather than arriving with the turn as it used to. A control + // that comes and goes makes its own presence the signal, and a button + // always + // in the same place also cannot push Send off the end of the row by turning + // up. + val process = + when { + running -> ProcessAction.Pause + status == "exited" -> ProcessAction.Start + else -> ProcessAction.Stop } - }, - enabled = !processInFlight, - colors = actionButtonColors(process.colour()), - ) { - Glyph( - process.glyph, - colour = LocalContentColor.current, - modifier = Modifier.semantics { contentDescription = process.label }, - ) - } - Spacer(Modifier.width(8.dp)) - // The paper plane, with a clock on it while a turn is in flight: sending then - // queues the message for the next tool boundary rather than starting a turn of - // its own, and the two have to be told apart at a glance. The label says the - // same thing to a screen reader. - // - // Disabled while there is nothing to send, rather than pressable and silent: - // `send` has always returned early on an empty composer, so the button promised - // something it would not do. Disabled and not hidden, for the reason above. - Button( - onClick = { send() }, - enabled = input.text.isNotBlank() || pendingAttachments.isNotEmpty(), - colors = actionButtonColors(if (running) queueColor else sendColor), - ) { - Glyph( - if (running) QUEUE_GLYPH else SEND_GLYPH, - colour = LocalContentColor.current, - modifier = - Modifier.semantics { contentDescription = sendLabel(running) }, - ) + Button( + onClick = { + processInFlight = true + act(onDone = { processInFlight = false }) { + process.perform(settings, summary.id) + } + }, + enabled = !processInFlight, + colors = actionButtonColors(process.colour()), + ) { + Glyph( + process.glyph, + colour = LocalContentColor.current, + modifier = + Modifier.semantics { contentDescription = process.label }, + ) + } + Spacer(Modifier.width(8.dp)) + // The paper plane, with a clock on it while a turn is in flight: sending + // then queues the message for the next tool boundary rather than starting a + // turn of its own, and the two have to be told apart at a glance. The label + // says the same thing to a screen reader. + // + // Disabled while there is nothing to send, rather than pressable and + // silent: + // `send` has always returned early on an empty composer, so the button + // promised something it would not do. Disabled and not hidden, for the + // reason above. + Button( + onClick = { send() }, + enabled = input.text.isNotBlank() || pendingAttachments.isNotEmpty(), + colors = actionButtonColors(if (running) queueColor else sendColor), + ) { + Glyph( + if (running) QUEUE_GLYPH else SEND_GLYPH, + colour = LocalContentColor.current, + modifier = + Modifier.semantics { contentDescription = sendLabel(running) }, + ) + } } } } @@ -1852,7 +1933,7 @@ fun SessionScreen( // is the screen's business rather than any row's. See [SessionImageViewer]. fullImage?.let { ref -> SessionImageViewer(settings, summary.id, ref) { fullImage = null } } if (usageOpen) { - UsageDialog(feed = usageFeed, onDismiss = { usageOpen = false }) + usageFeed?.let { UsageDialog(feed = it, onDismiss = { usageOpen = false }) } } if (settingsOpen) { // Measured when the dialog opens rather than kept up to date: what the reader is being told @@ -1866,6 +1947,8 @@ fun SessionScreen( settings = settings, sessionId = summary.id, title = title, + effort = summary.effort.takeIf { summary.takesEffort }, + takesEffort = summary.takesEffort, cachedBytes = cachedBytes, // The purge finishes before the epoch moves, because the relaunched opening effect // reads the same directory and would otherwise draw what is about to be deleted. The @@ -2143,6 +2226,13 @@ private fun SessionStatusRow( /** Context the session is holding, or null where nothing has measured it. */ contextTokens: Long?, modifier: Modifier = Modifier, + /** + * Whether this row is for a subagent rather than a session, which changes only one word: + * "exited" reads as "finished" there too, the same as the subagent list's own card -- a + * subagent's process was always its parent's, so "exited" would read as a fault rather than the + * ordinary way one of these ends. + */ + subagent: Boolean = false, ) { DebugStats.count("status row recomposed") Row( @@ -2199,7 +2289,7 @@ private fun SessionStatusRow( Text( when (status) { "idle" -> "idle" - "exited" -> "exited" + "exited" -> if (subagent) "finished" else "exited" "awaitingInput" -> "your turn" "unknown" -> "can't tell" else -> status @@ -2270,7 +2360,7 @@ private const val ONE_TAP_MS = 250L * session is set to without spending a second line on saying it. */ @Composable -private fun PickerButton(current: String, options: List, onPick: (String) -> Unit) { +fun PickerButton(current: String, options: List, onPick: (String) -> Unit) { var open by remember { mutableStateOf(false) } // When an outside touch last closed the menu. // diff --git a/app/androidApp/src/main/kotlin/com/example/aiapp/SessionSettingsDialog.kt b/app/androidApp/src/main/kotlin/com/example/aiapp/SessionSettingsDialog.kt index ad44f41..1ecb917 100644 --- a/app/androidApp/src/main/kotlin/com/example/aiapp/SessionSettingsDialog.kt +++ b/app/androidApp/src/main/kotlin/com/example/aiapp/SessionSettingsDialog.kt @@ -6,8 +6,10 @@ import androidx.compose.foundation.layout.Spacer import androidx.compose.foundation.layout.fillMaxWidth import androidx.compose.foundation.layout.height import androidx.compose.foundation.layout.width +import androidx.compose.foundation.rememberScrollState import androidx.compose.foundation.text.KeyboardActions import androidx.compose.foundation.text.KeyboardOptions +import androidx.compose.foundation.verticalScroll import androidx.compose.material3.AlertDialog import androidx.compose.material3.CircularProgressIndicator import androidx.compose.material3.MaterialTheme @@ -26,6 +28,10 @@ import androidx.compose.ui.Alignment import androidx.compose.ui.Modifier import androidx.compose.ui.text.input.ImeAction import androidx.compose.ui.unit.dp +import java.time.Instant +import java.time.ZoneId +import java.time.format.DateTimeFormatter +import java.time.format.FormatStyle import kotlinx.coroutines.Dispatchers import kotlinx.coroutines.launch import kotlinx.coroutines.withContext @@ -54,6 +60,16 @@ fun SessionSettingsDialog( */ title: String, onRenamed: (String) -> Unit, + /** + * How hard the model thinks, as the session reports it, or null for the CLI's own default. + * + * Taken from the row this dialog was opened over rather than fetched, because unlike the + * notification switch there is nothing else that changes it: the level is this app's to set and + * the server does not resolve it into something else. + */ + effort: String?, + /** Whether a level does anything here; the row is left out entirely where it does not. */ + takesEffort: Boolean, /** * What this phone is holding of the conversation, or null while that is being measured -- see * the Reload row below, which is what would discard it. @@ -76,6 +92,8 @@ fun SessionSettingsDialog( ) { val scope = rememberCoroutineScope() var name by remember(sessionId) { mutableStateOf(title) } + var level by remember(sessionId) { mutableStateOf(effort) } + var effortError by remember { mutableStateOf(null) } var saving by remember { mutableStateOf(false) } var error by remember { mutableStateOf(null) } // Null until the server has been asked. The row this dialog was opened over is a snapshot of @@ -84,6 +102,16 @@ fun SessionSettingsDialog( // and a spinner sits beside it, which is what not knowing looks like. var notify by remember(sessionId) { mutableStateOf(null) } var notifyError by remember { mutableStateOf(null) } + // The same three-state shape the notification switch has, for the same reason: until the + // server has answered, the switch is disabled rather than showing a position nothing confirmed. + var autoResume by remember(sessionId) { mutableStateOf(null) } + var resumeMessage by remember(sessionId) { mutableStateOf(DEFAULT_RESUME_MESSAGE) } + // When the server next intends to ask whether the limit has lifted, or null when nothing is + // waiting. Read once with everything else: it moves on the server's schedule, not this + // screen's, and a figure that redrew itself here would be this app re-measuring what it was + // told. + var resumeAt by remember(sessionId) { mutableStateOf(null) } + var resumeError by remember { mutableStateOf(null) } // Where the session works. Null until the server has been asked, for the same reason the switch // above is. An empty answer is a session that was never given a directory, which is not the // same as one whose directory is unknown -- the field is only enabled once one of those is @@ -97,6 +125,9 @@ fun SessionSettingsDialog( try { val fresh = withContext(Dispatchers.IO) { fetchSession(settings, sessionId) } notify = fresh.notify + autoResume = fresh.autoResume + resumeMessage = fresh.autoResumeMessage + resumeAt = fresh.resumeAt cwd = fresh.cwd.orEmpty() typedCwd = fresh.cwd.orEmpty() } catch (e: ApiException) { @@ -104,6 +135,8 @@ fun SessionSettingsDialog( // instead of offering a position nothing confirmed. notifyError = e.message notify = null + resumeError = e.message + autoResume = null } } @@ -132,6 +165,26 @@ fun SessionSettingsDialog( } } + /** + * Chooses a thinking level, which ends the process the old level was launched with. + * + * Put back if the request is refused, for the reason the notification switch below gives: a + * control that stays where it was put after a refusal is stating something untrue. + */ + fun setEffort(chosen: String?) { + val was = level + level = chosen + effortError = null + scope.launch { + try { + withContext(Dispatchers.IO) { setSessionEffort(settings, sessionId, chosen) } + } catch (e: ApiException) { + level = was + effortError = e.message + } + } + } + // Moved optimistically so the switch answers the finger that moved it, and put back if the // request is refused -- a switch that waits for a round trip reads as broken on a slow tunnel, // and one that stays moved after a refusal lies. @@ -149,6 +202,39 @@ fun SessionSettingsDialog( } } + /** + * Turns auto-resume on or off, or changes what it would say. + * + * One request for both, because the server takes one: switching it on and typing the message + * are two halves of the same decision, and sending them separately would leave a moment where + * the session is armed with the old words. + * + * Put back if refused, like the notification switch. Turning it off also clears what was + * scheduled -- said here rather than only on the server, or the row would go on naming a time + * that no longer exists. + */ + fun setAutoResume(on: Boolean, message: String) { + val wasOn = autoResume + val wasMessage = resumeMessage + val wasAt = resumeAt + autoResume = on + resumeMessage = message + if (!on) resumeAt = null + resumeError = null + scope.launch { + try { + withContext(Dispatchers.IO) { + setSessionAutoResume(settings, sessionId, on, message) + } + } catch (e: ApiException) { + autoResume = wasOn + resumeMessage = wasMessage + resumeAt = wasAt + resumeError = e.message + } + } + } + // Nothing to do when the name has not changed, so the button says so rather than sending a // request whose success would look exactly like the failure of having typed nothing. val changed = name.trim().isNotEmpty() && name.trim() != title @@ -175,7 +261,10 @@ fun SessionSettingsDialog( onDismissRequest = onDismiss, title = { Text("Session settings") }, text = { - Column { + // Scrollable, because this dialog grew past a screenful: a Material dialog constrains + // its own height and clips what does not fit, so the last control on the list is one + // large system font away from being unreachable with nothing on screen to say so. + Column(Modifier.verticalScroll(rememberScrollState())) { OutlinedTextField( value = name, onValueChange = { name = it }, @@ -219,6 +308,70 @@ fun SessionSettingsDialog( ) } Spacer(Modifier.height(8.dp)) + Row( + verticalAlignment = Alignment.CenterVertically, + modifier = Modifier.fillMaxWidth(), + ) { + Text("Resume after a usage limit", modifier = Modifier.weight(1f)) + if (autoResume == null && resumeError == null) { + CircularProgressIndicator( + modifier = Modifier.width(16.dp).height(16.dp), + strokeWidth = 2.dp, + ) + Spacer(Modifier.width(8.dp)) + } + Switch( + checked = autoResume == true, + onCheckedChange = { setAutoResume(it, resumeMessage) }, + enabled = autoResume != null, + ) + } + // Disabled rather than hidden while the switch is off: a field that comes and goes + // makes its own presence the signal, and a visible one teaches what the switch will + // do. Committed on the keyboard's Done rather than on every keystroke, so typing a + // sentence is one request instead of one per letter. + OutlinedTextField( + value = resumeMessage, + onValueChange = { resumeMessage = it }, + label = { Text("Message to send") }, + // What an empty field means, in the field: the server's own word rather than a + // session poked with nothing to read. + placeholder = { Text(DEFAULT_RESUME_MESSAGE) }, + singleLine = true, + enabled = autoResume == true, + modifier = Modifier.fillMaxWidth(), + keyboardOptions = KeyboardOptions(imeAction = ImeAction.Done), + keyboardActions = + KeyboardActions(onDone = { setAutoResume(true, resumeMessage) }), + ) + // What it does and what it costs, in the order it happens. The last sentence is the + // one that matters: the time below is when the server will *ask*, not a promise + // about when the session speaks. + Text( + "When this session stops because the account is out of quota, the server " + + "checks the limit and sends this message once it has lifted. It checks " + + "again if the limit is still on.", + style = MaterialTheme.typography.bodySmall, + color = MaterialTheme.colorScheme.onSurfaceVariant, + ) + // Only where something is actually waiting. Absent is not a state worth a row: a + // session that has not hit a limit has nothing scheduled, which the reader can see + // from the switch. + resumeAt?.let { at -> + Text( + "Waiting now -- next check ${formatCheckTime(at)}.", + style = MaterialTheme.typography.bodySmall, + color = MaterialTheme.colorScheme.onSurfaceVariant, + ) + } + resumeError?.let { + Text( + it, + color = MaterialTheme.colorScheme.error, + style = MaterialTheme.typography.bodySmall, + ) + } + Spacer(Modifier.height(8.dp)) Row( verticalAlignment = Alignment.CenterVertically, modifier = Modifier.fillMaxWidth(), @@ -264,6 +417,44 @@ fun SessionSettingsDialog( style = MaterialTheme.typography.bodySmall, ) } + // Left out rather than disabled, the one place this dialog does that: a disabled + // control teaches what the thing can do, and a llama session cannot do this at all + // -- the row would be teaching something false about it. + if (takesEffort) { + Spacer(Modifier.height(8.dp)) + Row( + verticalAlignment = Alignment.CenterVertically, + modifier = Modifier.fillMaxWidth(), + ) { + Text("Thinking", modifier = Modifier.weight(1f)) + PickerButton( + current = level ?: DEFAULT_EFFORT, + // The level the CLI picks for itself is in the list as well as in the + // button, so leaving a level is not a one-way trip -- the same + // correction the model picker carries. + options = listOf(DEFAULT_EFFORT) + EFFORT_LEVELS, + onPick = { chosen -> + setEffort(chosen.takeIf { it != DEFAULT_EFFORT }) + }, + ) + } + // What it costs, said where it is about to be pressed, like Move above: the + // CLI reads the level when it launches and has no control request for + // changing one. + Text( + "Changing this stops the session's process. It starts again with the " + + "next message, or with Start.", + style = MaterialTheme.typography.bodySmall, + color = MaterialTheme.colorScheme.onSurfaceVariant, + ) + effortError?.let { + Text( + it, + color = MaterialTheme.colorScheme.error, + style = MaterialTheme.typography.bodySmall, + ) + } + } Spacer(Modifier.height(8.dp)) Row( verticalAlignment = Alignment.CenterVertically, @@ -351,3 +542,21 @@ fun SessionSettingsDialog( dismissButton = { TextButton(onClick = onDismiss) { Text("Close") } }, ) } + +/** + * When the server will next look, as a local time. + * + * A time rather than a countdown, for the reason the transcript's own limit row gives: this screen + * reads the figure once, and a span drawn from a value nothing refreshes goes stale while somebody + * is looking at it. + */ +private fun formatCheckTime(epochSeconds: Double): String = + try { + DateTimeFormatter.ofLocalizedTime(FormatStyle.SHORT) + .withZone(ZoneId.systemDefault()) + .format(Instant.ofEpochSecond(epochSeconds.toLong())) + } catch (_: Exception) { + // A time that cannot be read is not a time to show: the sentence above still says a check + // is coming, which is the part the reader can act on. + "soon" + } diff --git a/app/androidApp/src/main/kotlin/com/example/aiapp/SessionUsageBar.kt b/app/androidApp/src/main/kotlin/com/example/aiapp/SessionUsageBar.kt index e292ea3..fcd0814 100644 --- a/app/androidApp/src/main/kotlin/com/example/aiapp/SessionUsageBar.kt +++ b/app/androidApp/src/main/kotlin/com/example/aiapp/SessionUsageBar.kt @@ -70,13 +70,22 @@ class UsageFeed( /** Ask the backend again now. The dialog's refresh button; the poll does it on its own. */ val refresh: () -> Unit, ) { - /** What [setup]'s own limits came back as. See [usageFor] for why the states are these. */ - fun forSetup(setup: String): SessionUsage = - when (val state = snapshots) { + /** + * What meters [session], and what that meter came back as. See [usageFor] for the states. + * + * A session rather than a machine, because a machine is not what is metered: one machine runs + * the Claude CLI and an echo session side by side, and only the first of them spends anything. + */ + fun forSession(session: SessionSummary): SessionUsage { + // Settled without asking anybody: a session nothing meters has nothing to check, and + // "checking" is what the fetch's own states would say about it for as long as one is out. + val provider = session.usageProvider ?: return SessionUsage.NotMetered + return when (val state = snapshots) { is LoadState.Loading -> SessionUsage.Waiting is LoadState.Error -> SessionUsage.Unavailable(state.message) - is LoadState.Loaded -> usageFor(state.value, setup) + is LoadState.Loaded -> usageFor(state.value, session.setup, provider) } + } } /** @@ -158,9 +167,15 @@ fun SessionUsageBar(usage: SessionUsage, modifier: Modifier = Modifier) { } } - // Nothing at all for a machine that meters nothing: a row saying "unknown" there would report a - // problem about a setup somebody chose, on every screen, forever. - if (usage is SessionUsage.NotMetered) { + // Nothing at all for a session that meters nothing: a row saying "unknown" there would report + // a problem about a setup somebody chose, on every screen, forever. + // + // And nothing while the first fetch is out, which is a different silence. A request in flight + // is not a state to report -- and the session that meters nothing is exactly the one this + // cannot yet tell apart, so "5-hour usage: checking" appeared under an echo session for half a + // second and was then taken away. A row that has to be withdrawn is worse than one that + // arrives late. + if (usage is SessionUsage.NotMetered || usage is SessionUsage.Waiting) { return } @@ -171,9 +186,10 @@ fun SessionUsageBar(usage: SessionUsage, modifier: Modifier = Modifier) { // Words, not a colour and not an empty bar: every one of these is a different kind of // answer from "this much is used", and only words carry a difference in kind. when (val state = usage) { - SessionUsage.NotMetered -> Unit + // Both handled above, before the row exists at all. + SessionUsage.NotMetered, + SessionUsage.Waiting -> Unit is SessionUsage.Unavailable -> UsageNote("5-hour usage unknown -- ${state.why}") - SessionUsage.Waiting -> UsageNote("5-hour usage: checking") is SessionUsage.Known -> { val window = state.windows.firstOrNull { it.kind == "session" } if (window == null) { @@ -234,16 +250,22 @@ private fun fiveHourLabel(window: UsageWindow, now: OffsetDateTime): String { } /** - * One machine's snapshot, out of every machine's. + * One meter's snapshot, out of every machine's: [setup]'s row for [provider]. + * + * Both halves are needed to pick it. A machine can hold more than one meter -- the Claude CLI's + * account and, while a test has one set, an echo session's invented one -- and a snapshot is one + * service on one machine. * * Every way of having *failed* to get numbers is [SessionUsage.Unavailable] with the reason in it. * None of them may look like zero, and none may look like [SessionUsage.NotMetered], which is the * machine having no quota rather than the question going unanswered. */ -fun usageFor(snapshots: List, setup: String): SessionUsage { - // No snapshot at all means the backend never asked, which it only does for a machine with - // nothing metered on it. That is a different answer from having asked and failed. - val mine = snapshots.firstOrNull { it.setup == setup } ?: return SessionUsage.NotMetered +fun usageFor(snapshots: List, setup: String, provider: String): SessionUsage { + // No snapshot at all means the backend never asked, which it only does where there is nothing + // to ask about. That is a different answer from having asked and failed. + val mine = + snapshots.firstOrNull { it.setup == setup && it.provider == provider } + ?: return SessionUsage.NotMetered if (mine.state != "ok") { return SessionUsage.Unavailable(mine.detail ?: mine.state) } diff --git a/app/androidApp/src/main/kotlin/com/example/aiapp/SetupsScreen.kt b/app/androidApp/src/main/kotlin/com/example/aiapp/SetupsScreen.kt index 8bacde8..5b127cc 100644 --- a/app/androidApp/src/main/kotlin/com/example/aiapp/SetupsScreen.kt +++ b/app/androidApp/src/main/kotlin/com/example/aiapp/SetupsScreen.kt @@ -236,6 +236,7 @@ private fun AddSetupDialog( var address by remember { mutableStateOf("") } var identity by remember { mutableStateOf("") } var attachmentsDir by remember { mutableStateOf("") } + var modelsDir by remember { mutableStateOf("") } var tested by remember { mutableStateOf(null) } var testing by remember { mutableStateOf(false) } @@ -250,6 +251,7 @@ private fun AddSetupDialog( port = typedPort, identityFile = identity.trim().ifEmpty { null }, attachmentsDir = attachmentsDir.trim().ifEmpty { null }, + modelsDir = modelsDir.trim().ifEmpty { null }, ) } @@ -293,6 +295,14 @@ private fun AddSetupDialog( label = { Text("Folder for attached files (optional)") }, singleLine = true, ) + // Where that machine's GGUFs are, for a llama.cpp session on it. Blank means + // the same place this backend keeps its own downloads, read on that machine. + OutlinedTextField( + value = modelsDir, + onValueChange = { modelsDir = it }, + label = { Text("Folder for models (optional)") }, + singleLine = true, + ) tested?.let { Spacer(Modifier.height(8.dp)) Text(it, style = MaterialTheme.typography.bodySmall) diff --git a/app/androidApp/src/main/kotlin/com/example/aiapp/SpawnScreen.kt b/app/androidApp/src/main/kotlin/com/example/aiapp/SpawnScreen.kt index 88a2f36..4077cfa 100644 --- a/app/androidApp/src/main/kotlin/com/example/aiapp/SpawnScreen.kt +++ b/app/androidApp/src/main/kotlin/com/example/aiapp/SpawnScreen.kt @@ -61,18 +61,31 @@ fun SpawnScreen( // "auto" rather than "manual": on a phone every ask is a round trip to a question card, and // answering "allow Bash?" dozens of times per task is what this app exists to avoid. var permissionMode by remember { mutableStateOf("auto") } + // Null until the server has been asked, and null again if it answers "no level chosen" -- the + // two are told apart by [defaultsAsked], because a picker that shows a level before the answer + // arrives is one you can spawn at without having chosen it. + var effort by remember { mutableStateOf(null) } + var defaultsAsked by remember { mutableStateOf(false) } var busy by remember { mutableStateOf(false) } // Only the spawn's own failure. The fetch's lives in `options`: this one leaves a filled-in // form worth keeping, and that one leaves nothing to fill in. var spawnError by remember { mutableStateOf(null) } - // Downloaded models, for a llama provider to choose between. Kept separate from the setups: a - // Claude session needs none, so failing to list them must not stop the screen rendering. + // The models on the *chosen machine*, for a llama provider to choose between. Kept separate + // from the setups: a Claude session needs none, so failing to list them must not stop the + // screen rendering. Refetched when the machine changes, because a model is a file on one + // machine -- see [fetchSetupModels]. var models by remember { mutableStateOf>(emptyList()) } var modelKey by remember { mutableStateOf(null) } var contextSize by remember { mutableStateOf("") } var temperature by remember { mutableStateOf("") } LaunchedEffect(Unit) { + // Separate from the setups fetch below and deliberately not fatal: failing to learn the + // default must leave a screen you can still spawn from, so the picker stays on "default" + // and says so rather than the whole form refusing to draw. + runCatching { withContext(Dispatchers.IO) { fetchDefaultEffort(settings) } } + .onSuccess { effort = it } + defaultsAsked = true options = try { val fetched = withContext(Dispatchers.IO) { fetchSetups(settings) } @@ -83,9 +96,6 @@ fun SpawnScreen( } catch (e: ApiException) { LoadState.failed(e) } - models = - runCatching { withContext(Dispatchers.IO) { fetchModels(settings).local } } - .getOrDefault(emptyList()) } Column(Modifier.fillMaxSize().verticalScroll(rememberScrollState()).padding(16.dp)) { @@ -115,6 +125,17 @@ fun SpawnScreen( is LoadState.Loaded -> state.value } val setup = setups.firstOrNull { it.name == setupName } + // Whichever machine is chosen now, asked again when that changes. The old machine's list + // is dropped first rather than left on screen: a file name from another machine looks + // exactly like one from this one. + LaunchedEffect(setup?.id) { + models = emptyList() + modelKey = null + val id = setup?.id ?: return@LaunchedEffect + models = + runCatching { withContext(Dispatchers.IO) { fetchSetupModels(settings, id) } } + .getOrDefault(emptyList()) + } val current = setup?.providers?.firstOrNull { it.name == providerName } // Only the Claude CLI has models, a working directory and permission modes; keying the // extra fields on the kind rather than the provider name keeps a second Claude provider @@ -175,12 +196,13 @@ fun SpawnScreen( ) if (isLlama) { - // A llama session names one of the models this backend has downloaded, so the choice is - // that list rather than free text -- a name that is not on disk is a session that - // cannot start. + // A llama session names one of the models on the machine it will run on, so the + // choice is that list rather than free text -- a name that is not on that machine's + // disk is a session that cannot start. if (models.isEmpty()) { Text( - "No models downloaded yet. Get one from the Models screen first.", + "No models on ${setup?.name ?: "this machine"}. The Models screen downloads " + + "to the backend; another machine needs the file put there itself.", style = MaterialTheme.typography.bodyMedium, color = MaterialTheme.colorScheme.onSurfaceVariant, ) @@ -251,6 +273,19 @@ fun SpawnScreen( selected = permissionMode, onSelect = { permissionMode = it }, ) + Spacer(Modifier.height(16.dp)) + + // Says what it does to *later* spawns as well, because it does: the level chosen here + // is stored as the default, which is the whole way that default is set. A picker that + // quietly changed a global would be the same control with the fact left out. + ChipGroup( + label = "Thinking (kept as the default for new sessions)", + options = listOf(DEFAULT_EFFORT) + EFFORT_LEVELS, + // The CLI's own default is a level in the list, so this cannot be a one-way trip. + // Disabled-looking until the server has answered, for the reason above. + selected = if (defaultsAsked) effort ?: DEFAULT_EFFORT else null, + onSelect = { chosen -> effort = chosen.takeIf { it != DEFAULT_EFFORT } }, + ) } Spacer(Modifier.height(24.dp)) @@ -268,6 +303,13 @@ fun SpawnScreen( try { val spawned = withContext(Dispatchers.IO) { + // Stored before the spawn and not after it: choosing a level is + // an intent about new sessions in general, so a spawn that then + // fails must not also lose the choice. Non-fatal for the same + // reason the fetch above is -- the session is what was asked for. + if (isClaude) { + runCatching { setDefaultEffort(settings, effort) } + } spawnSession( settings, // The id, not the label: labels are editable and the server @@ -280,6 +322,7 @@ fun SpawnScreen( if (isLlama) modelKey else model.trim().takeIf { isClaude }, cwd = cwd.trim().takeIf { isClaude }, permissionMode = permissionMode.takeIf { isClaude }, + effort = effort.takeIf { isClaude }, // Sent only when set, so blank means "whatever llama.cpp does // by default" rather than a zero. params = diff --git a/app/androidApp/src/main/kotlin/com/example/aiapp/TranscriptAddress.kt b/app/androidApp/src/main/kotlin/com/example/aiapp/TranscriptAddress.kt new file mode 100644 index 0000000..ae2a97f --- /dev/null +++ b/app/androidApp/src/main/kotlin/com/example/aiapp/TranscriptAddress.kt @@ -0,0 +1,27 @@ +package com.example.aiapp + +/** + * Where one transcript lives: a session's own, or one of its subagents'. + * + * The single mechanism [fetchTranscript], [EventStream], [TranscriptSource] and + * [TranscriptCache.session] all take, rather than each growing its own branch between a session and + * a subagent -- see SUBAGENTS.md's "Phone" and "Wire shape". A caller that has only a session id + * builds one with the one-argument constructor; a subagent's screen supplies both ids. + */ +data class TranscriptAddress(val sessionId: String, val subagentId: String? = null) { + /** The URL segment naming this transcript, before `/transcript` or `/events`. */ + val urlPath: String + get() = + if (subagentId == null) "sessions/$sessionId" + else "sessions/$sessionId/subagents/$subagentId" + + /** + * Where this transcript's cache lives on the phone, relative to the cache root. + * + * A subagent's nests under its session's directory rather than sitting beside it, so deleting a + * session's cache directory takes its subagents' with it -- the same one-way door the server's + * own storage describes. + */ + val cachePath: String + get() = if (subagentId == null) sessionId else "$sessionId/subagents/$subagentId" +} diff --git a/app/androidApp/src/main/kotlin/com/example/aiapp/TranscriptCache.kt b/app/androidApp/src/main/kotlin/com/example/aiapp/TranscriptCache.kt index 606628e..d489f3a 100644 --- a/app/androidApp/src/main/kotlin/com/example/aiapp/TranscriptCache.kt +++ b/app/androidApp/src/main/kotlin/com/example/aiapp/TranscriptCache.kt @@ -35,8 +35,15 @@ class TranscriptCache( private val root: File, private val warn: (String) -> Unit = { Log.w("ai-app", it) }, ) { - /** The cache for one session, whether or not anything has been stored for it yet. */ - fun session(id: String): SessionCache = SessionCache(File(root, id), warn) + /** + * The cache for one transcript, whether or not anything has been stored for it yet. + * + * A subagent's [TranscriptAddress.cachePath] nests it under its session's directory, so + * deleting the session (below) takes its subagents' caches with it -- there is no separate + * purge for one. + */ + fun session(address: TranscriptAddress): SessionCache = + SessionCache(File(root, address.cachePath), warn) /** * Deletes every session directory not in [ids], called after a successful list fetch. The path diff --git a/app/androidApp/src/main/kotlin/com/example/aiapp/TranscriptItems.kt b/app/androidApp/src/main/kotlin/com/example/aiapp/TranscriptItems.kt index 5089ca2..96eefa1 100644 --- a/app/androidApp/src/main/kotlin/com/example/aiapp/TranscriptItems.kt +++ b/app/androidApp/src/main/kotlin/com/example/aiapp/TranscriptItems.kt @@ -175,6 +175,18 @@ sealed class TranscriptItem { val preTokens: Long?, val postTokens: Long?, ) : TranscriptItem() + + /** + * The account ran out of quota, so the turn stopped here. + * + * A divider rather than an error: nothing failed, and what a reader scrolling back needs from + * it is the same thing a clear or a compaction gives them -- why the conversation stops at this + * line. + * + * [resetsAt] is epoch seconds and null where the session was told nothing, which is a state the + * row has words for rather than a time it invents. + */ + data class LimitNote(override val seq: Long, val resetsAt: Double?) : TranscriptItem() } /** @@ -458,6 +470,7 @@ fun foldEvent(items: List, entry: SeqEvent): List items + TranscriptItem.LimitNote(entry.seq, event.resetsAt) is SessionEvent.Cleared -> items + TranscriptItem.ClearedNote(entry.seq) is SessionEvent.Compacted -> items + TranscriptItem.CompactedNote(entry.seq, event.preTokens, event.postTokens) diff --git a/app/androidApp/src/main/kotlin/com/example/aiapp/TranscriptSource.kt b/app/androidApp/src/main/kotlin/com/example/aiapp/TranscriptSource.kt index 8a7bdcd..10c7c32 100644 --- a/app/androidApp/src/main/kotlin/com/example/aiapp/TranscriptSource.kt +++ b/app/androidApp/src/main/kotlin/com/example/aiapp/TranscriptSource.kt @@ -18,7 +18,7 @@ import java.util.concurrent.atomic.AtomicReference */ class TranscriptSource( private val settings: ServerSettings, - private val sessionId: String, + private val address: TranscriptAddress, val cache: SessionCache, ) { private val stream = AtomicReference(null) @@ -65,7 +65,7 @@ class TranscriptSource( val tail = cache.tail() ?: return false // `before = seq + 1` is the newest event with seq <= the cursor, which is the event *at* // the cursor when the server still has one there. - val answer = fetchTranscript(settings, sessionId, before = tail.seq + 1, limit = 1) + val answer = fetchTranscript(settings, address, before = tail.seq + 1, limit = 1) val matches = answer.size == 1 && try { @@ -83,7 +83,7 @@ class TranscriptSource( */ suspend fun fetchOpening(): List { DebugStats.count("transcript page from server") - val page = fetchTranscript(settings, sessionId, limit = OPENING_WINDOW) + val page = fetchTranscript(settings, address, limit = OPENING_WINDOW) page.forEach { (line, entry) -> cache.append(line, entry.seq) } cache.flush() return page.map { it.second } @@ -108,7 +108,7 @@ class TranscriptSource( val page = fetchTranscript( settings, - sessionId, + address, before = before, limit = limit, coalesce = coalesce, @@ -131,7 +131,7 @@ class TranscriptSource( * well lose. */ fun follow(after: Long, onOpen: () -> Unit, onReset: () -> Unit, onEvent: (SeqEvent) -> Unit) { - val opened = EventStream(settings, sessionId) + val opened = EventStream(settings, address) stream.getAndSet(opened)?.close() try { opened.run(after, onOpen, onReset) { raw, entry -> diff --git a/app/androidApp/src/test/kotlin/com/example/aiapp/LimitRowTest.kt b/app/androidApp/src/test/kotlin/com/example/aiapp/LimitRowTest.kt new file mode 100644 index 0000000..e4c030f --- /dev/null +++ b/app/androidApp/src/test/kotlin/com/example/aiapp/LimitRowTest.kt @@ -0,0 +1,32 @@ +package com.example.aiapp + +import java.time.ZoneId +import kotlin.test.Test +import kotlin.test.assertEquals +import kotlin.test.assertTrue + +/** + * What the transcript says where a session ran out of quota. + * + * The pair worth a test is the one that reads the same when it goes wrong: a reset time that + * arrived and one that never did. The second must not turn into a plausible-looking time, because a + * reader has no way of telling an invented one from a reported one. + */ +class LimitRowTest { + private val utc = ZoneId.of("UTC") + + @Test + fun `a reported reset time is shown as a time`() { + // 2026-09-05T12:00:00Z. Asserted as a prefix and the clock reading rather than as the + // whole string: the platform's own short-time format is what this asks for, and it + // differs by JDK and locale down to which space character separates the meridiem. + val summary = limitSummary(1_788_609_600.0, utc) + assertTrue(summary.startsWith("Usage limit reached • resets "), summary) + assertTrue(summary.contains("12:00"), summary) + } + + @Test + fun `a limit with no reset time says only what is known`() { + assertEquals("Usage limit reached", limitSummary(null, utc)) + } +} diff --git a/app/androidApp/src/test/kotlin/com/example/aiapp/TranscriptCacheTest.kt b/app/androidApp/src/test/kotlin/com/example/aiapp/TranscriptCacheTest.kt index 04b4db9..e25341f 100644 --- a/app/androidApp/src/test/kotlin/com/example/aiapp/TranscriptCacheTest.kt +++ b/app/androidApp/src/test/kotlin/com/example/aiapp/TranscriptCacheTest.kt @@ -23,7 +23,7 @@ class TranscriptCacheTest { private fun cache() = TranscriptCache(File(temp, "v1/host_8443")) { said += it } - private fun session(id: String = "s") = cache().session(id) + private fun session(id: String = "s") = cache().session(TranscriptAddress(id)) private fun line(seq: Long, type: String = "toolStart") = """{"seq":$seq,"ts":1.5,"type":"$type","id":"x"}""" diff --git a/docs/DECISIONS.md b/docs/DECISIONS.md index ecdaf25..cef0361 100644 --- a/docs/DECISIONS.md +++ b/docs/DECISIONS.md @@ -60,6 +60,25 @@ marked **DEFERRED** are ones the agent chose not to decide alone. flagged here because it is the first half of something Iris explicitly asked to see before P1. +- **P0's iris half is also built and smoke-tested on the emulator, + 2026-09-05.** A new `bench` Cargo feature on `iris-android-app`, on top + of `transcript-screen`: the same checked-in fixture (`include_str!`, no + asset pipeline needed), the same 24-swipe scroll loop animated through + `List::scroll` and the same 400-event/20s streaming phase through + `fold_event`, "Run benchmark"/"Copy report" as named accessible + controls, and the same three added report fields (process CPU time, + peak RSS, battery current) via direct JNI calls + (`bench_jni.rs::PlatformHandle`) since `android_view` has no + `BatteryManager`/`ClipboardManager` wrapper of its own. One small public + API addition to get there: `AndroidAppState::platform_ready` (`IRIS.md`), + a default-no-op lifecycle hook handing an implementor a `JavaVM` + + `GlobalRef` it can call Java through from any thread. Packaged with a + new `release` build type on `iris-android-app`'s own Gradle project + (there was previously only `debug`), signed with the same key + `app/build-apk.sh` generates. Smoke run and the full report are in + RUST.md's P0 box; not attempted this pass: the real on-phone runs and + Iris's pass/fail call, which is the actual gate. + - **The intermittent touch-scroll dropout is root-caused and fixed: a missed `ACTION_DOWN` hit-test, not the previously-suspected coalesced first `ACTION_MOVE`.** Diagnosed by temporary logcat tracing of every diff --git a/docs/IRIS.md b/docs/IRIS.md index d51ed3f..cb1a2af 100644 --- a/docs/IRIS.md +++ b/docs/IRIS.md @@ -8,6 +8,29 @@ capability that moved. Small and trivial changes do not go here. An entry gives the date, what changed, why, and a short before/after where it helps judge the change without the session that made it. Newest first. +## 2026-09-05: `AndroidAppState::platform_ready` (RUST.md's P0 box, iris half) + +Added a second, optional lifecycle method to `iris::android::AndroidAppState` +(`iris/src/android/view.rs`), called once from `new_peer` right after `new`: + +```rust +fn platform_ready(&mut self, rsc: &mut AndroidRsc, vm: JavaVM, view: GlobalRef) {} +``` + +Default does nothing, so every existing implementor (`Client`, +`TranscriptClient`) is unaffected. It exists for a caller that needs to call +into Java itself beyond what a `RequestRedraw` handle already covers -- +P0's bench build (`iris-android-app`'s new `bench` feature, +`bench_client.rs`/`bench_jni.rs`) uses it to hold a `JavaVM` + `GlobalRef` +to the view so its "Copy report" control and once-a-second battery sampler +can call `BatteryManager`/`ClipboardManager` through the view's own +`Context`, from a background tokio task as well as the UI thread. `new` +itself was not extended with these two parameters: most implementors need +nothing here, and `new`'s job is building the widget tree, not holding a +platform handle. `vm`/`view` are independent handles from the ones +`new_peer` keeps for its own `RequestRedraw` (a fresh `get_java_vm`/ +`new_global_ref` each), so storing them has no effect on that mechanism. + ## 2026-09-05 (later still): `iris_core::device_limits()`, and iris no longer requests compute-shader limits New public function, `iris_core::device_limits() -> wgpu::Limits`. Why: diff --git a/docs/IRIS_TODO.md b/docs/IRIS_TODO.md index b9ea030..6a2c03f 100644 --- a/docs/IRIS_TODO.md +++ b/docs/IRIS_TODO.md @@ -105,6 +105,18 @@ order and what "done" looks like. Tick and date them in place. were exactly the same root cause measured two different ways. Frame 2 now reports 0 (see the numbers above); not a separate fix. +- [ ] **A read-only text display has no widget of its own — P0's bench + report area is a `TextEdit` standing in for one (2026-09-05).** The only + way to get selectable text on screen today is `.editable(...)` plus + `.attr::(())` (`Selectable` is only implemented for + `TextEdit`, `iris/src/attr.rs`), which also makes the field focusable — + tapping the bench report opens the soft keyboard over text nothing lets + you type into. Harmless for a bench-only debug screen (not fixed this + pass), but a real "selectable, not editable" text primitive would + remove the keyboard side effect and is worth having before another + screen wants the same thing (P1's own transcript rows already read + their content from a `TextEdit` for the same reason). + ## Build - [x] **Benchmarks**, not unit tests, run on demand (2026-09-05; a diff --git a/docs/PLAN.md b/docs/PLAN.md index 133ab15..c28c664 100644 --- a/docs/PLAN.md +++ b/docs/PLAN.md @@ -147,12 +147,37 @@ turn. Spawn: `claude -p --verbose --input-format stream-json --output-format stream-json --permission-mode ` in the chosen working directory, plus -`--model`. Wire-format notes are pinned against CLI 2.1.237 in -`session/claude.rs`'s module doc: permissions need the hidden -`--permission-prompt-tool stdio` flag, AskUserQuestion answers ride -`updatedInput.answers` keyed by question text, and `set_model`/`interrupt` +`--model` and, where one has been chosen, `--effort`. Wire-format notes are +pinned against CLI 2.1.237 in `session/claude.rs`'s module doc: permissions +need the hidden `--permission-prompt-tool stdio` flag, AskUserQuestion answers +ride `updatedInput.answers` keyed by question text, and `set_model`/`interrupt` are control requests. +**The thinking level is settled at launch** (added 2026-09-04, because it is +the largest saving available on a long session: output is about an eighth of +what a session costs and thinking is the bulk of output, against the ~1.5% that +is prose). The CLI's only two setting control requests are `set_model` and +`set_permission_mode` -- checked against the 2.1.258 binary -- so there is no +way to ask a running process to think differently. `set_session_effort` is +therefore shaped like `set_session_cwd` rather than like `set_session_model`: +it records the level and **stops the process**, and the next message or Start +launches one that has it. It lives in the session settings dialog beside the +working directory for that reason, not on the session bar beside the model and +the mode, which do take effect mid-turn. `None` is a level in its own right -- +the CLI's own default -- so the picker can return to it; a level this app named +as the default instead would be this app choosing one. + +**What a new session starts at is `Config::default_effort`**, applied in +`spawn_session` rather than filled in by the spawn screen, so it holds for an +import and a bare API call as well. It is set by the spawn screen's own +picker, whose label says so: one control, where new sessions are made, rather +than a settings page for a single value. It is not on a provider, because +providers are discovered and the next rediscovery would erase it, and not on +the phone, because a second device would then spawn at a level nobody there +chose. `GET`/`POST /defaults` carry it, as a struct rather than a bare value +so the permission mode -- still hardcoded to `auto` on the spawn screen -- can +move there without a second route. + **`--resume` only ever runs when nothing else has that session open.** That is the rule behind the import refusal, the single `ClaudeDriver::launch` entry point, and the `Exited` correction below; two CLIs on one session file @@ -169,10 +194,33 @@ deliberate and easy to undo by accident: when the process restarts. That leaves the Claude driver as the odd one out rather than this one — the CLI's memory is a cache in front of the same transcript. Resolve any inconsistency in this direction. -- **A llama session on an ssh host is refused.** The model is reached over - HTTP and forwarding that port is not built, so refusing beats silently - talking to the wrong machine. A transport is "run this" plus "reach this - port", and only the first half exists. +- **A llama session runs on whatever machine its setup names** (2026-09-04, + the last of phase 5). A transport is "run this" plus "reach this port", and + the second half is `Transport::reserve_port` — the port the server binds + *there* and the port that reaches it *here*, the same number locally — + carried by `Launch::reaching` onto the connection that already runs the + command. `llama-server` binds loopback on the far machine, so nothing is + served to its network. The far port is a guess from a range below the + ephemeral one, because no portable way to ask a machine for a free port + avoids racing the bind anyway; a collision is not silent, since the server + fails to bind and the readiness poll reports what its log said. +- **The model file lives on the machine that serves it** (2026-09-04). Each + setup names its own models directory (`SshConfig::models_dir`, default + `~/.local/share/ai-app/models` expanded *there*), and a spawn resolves the + key on that machine — one round trip answering "at /abs/path" or "missing", + so a model that is not there is refused at the spawn rather than becoming a + server that never becomes ready. The spawn screen offers + `GET /setups/{id}/models`, that machine's list, rather than `GET /models`, + which is this backend's downloads. Downloading *to* another machine is + deliberately not built: a multi-gigabyte transfer with no progress + anywhere, and the file gets there however anything else on that machine + did. +- **The readiness poll watches the process, not only the port.** A model that + will not load, a port already taken, a flag an older build does not know: + all exit within a second and none will ever answer `/health`, so waiting + out the 300s timeout turned the server's own account of the problem into + "gave up". The failure carries the tail of `llama-server.log`, which on a + remote session is the only copy anybody reading the phone can see. ### Models (2026-08-28) @@ -203,6 +251,13 @@ deliberate and easy to undo by accident: A driver says what to run; something above it turns that into a process. Otherwise transport knowledge sits inside a translator whose job is a wire format, and every future driver has to remember to do the same. +- **A forwarded launch gets a pty and every other one does not** (measured + 2026-09-04). Killing the ssh client ends a CLI because it closes the stdin + that CLI is reading; `llama-server` never reads its stdin, so the same kill + left it running on the far machine with the model loaded — one orphan per + stopped session. With `-tt` the far side takes SIGHUP when the connection + goes. Its log then arrives through a line discipline, which nothing parses. + `-T` stays everywhere else, where a pty would rewrite the JSONL. - **`command -v` follows ssh's non-login PATH**, which is narrower than an interactive shell's, so a binary somewhere unusual is invisible to discovery. Point `command` at an absolute path. @@ -499,6 +554,25 @@ rate-limited bucket). Poll at ≥180 s, only while a Claude session exists or the usage screen is open, and cache the last answer. It is undocumented, so `usage.rs` treats every field as optional and degrades rather than erroring. +**Per provider, not per machine (2026-09-04).** A machine is not what is +metered; the provider a session runs is. One machine offers echo, the Claude +CLI and a local model side by side, and only the second spends anything — so +pairing a session with a snapshot by machine alone drew the CLI's five-hour +window under every echo session on it, a quota that session cannot spend. A +session now names its meter (`usageProvider`, from +`DriverKind::usage_provider`, which `usage::providers_for` reads too, so the +two lists cannot disagree) and `GET /usage` is matched on machine *and* +provider. `None` is a session that meters nothing, and the phone draws +nothing at all for it — not a zero, and not "unknown". + +`DriverKind::Echo` names a meter of its own that exists only when a test has +asked for one: `/usage` in an echo session sets an invented answer +(`usage::Fixture`), and with none set there is no snapshot and no bar. That +is what makes those screens' states reachable — a number near the top, a +window between blocks with no reset time, a machine nobody logged into, one +that could not be reached — without spending real quota to arrange them, +which is why none of them had ever been looked at. + **Per machine, not per backend (2026-08-29).** The credential store that matters is the one on the machine the session runs on, because that is the account being billed — and in the layout this aims at, `ai-server` is on the @@ -522,6 +596,68 @@ always running. So absent means **not running**, and only a timestamp that arrives and cannot be parsed is unknown. `WindowEnd` in `ResetCountdown.kt` is the one rule both readers go through. +### Auto-resume (2026-09-05) + +**A session may pick itself back up when the account's usage limit lifts.** +Off unless somebody switched that session to it, because it spends quota the +moment quota exists and does so with nobody looking — that is not a thing a +default may decide. It sends one message, `continue` unless another was +typed, and then it is done; there is no retry loop around the conversation +itself. + +**Running out of quota is a state, not an error.** `Event::LimitReached` +carries the dialect's reset time where it gave one, and recognising it +belongs to the driver — the Claude CLI ends the turn with `is_error` and +`Claude AI usage limit reached|1788546972`, and nothing above the driver +matches on a string. The transcript draws it as a divider, like a clear or a +compaction: what a reader scrolling back wants from it is why the +conversation stops at that line. + +**The schedule is a plan to ask, never a plan to send.** Every reset time +available here is untrustworthy in the direction that matters: the dialect's +is written when the turn fails, and the endpoint's moves when the window +does. So the wait ends in a question to `usage.rs`, and only `ok` with no +window at 100% sends anything. A window still spent reschedules to *its own* +reset time — which is what makes a limit that lifts later than promised wait +longer, and one that lifts sooner resume sooner. A meter that cannot be +asked at all is a longer wait too, never a send: "we could not find out" +must not be able to produce the same action as "there is room". + +Bounded, because something has to be: a day after the limit was hit the wait +stops and says so in the session's own transcript. A machine that can never +be asked would otherwise be retried for ever with nothing on screen saying +so. + +The schedule is persisted on the session (`resume: Some(ScheduledResume)`), +not held in memory: a five-hour window routinely outlasts a backend restart, +and a wait forgotten across one is a session that silently never comes back. +`resume.rs` is the top layer — it holds the manager and the monitor and +neither holds it — which is what lets the decision be a pure function of a +snapshot and a clock. The pump reports limits downward on a broadcast, for +the reason `Shared` exists: the pump runs underneath the manager. + +**Exercised with echo, never with a real account.** `/limit [minutes]` in an +echo session reports the same event a real driver does, and `/usage` sets +what the meter answers — deliberately two commands, because the two +disagreeing is the state the whole design is about. The loop was driven end +to end that way on 2026-09-05: the wait moved from the dialect's two minutes +to the meter's seven when the meter changed its mind, and the message went +out on the first check after the meter came back under the limit. + +### Subagents (2026-09-05) + +**A subagent is a second transcript owned by a session, in the same event +model, with no process and no controls of its own.** Full design and wire +shape in `SUBAGENTS.md`, kept separate because the app half is being built +against it in parallel and it is the shared contract between the two. The +one-paragraph reason: a session's Task-tool helpers already speak the common +event model on the parent's own stdout (each line carrying +`parent_tool_use_id`), so giving each one its own small transcript — same +file format, same paging routes, same SSE stream, reused by addressing rather +than by copying — costs a routing step in the translator and a registry +(`session/subagent.rs`) rather than a second session type with a driver, a +process and a config entry it does not need. + ### HTTP surface **`routes.rs`'s module doc comment is the table.** REST for actions, one SSE @@ -800,8 +936,9 @@ Noticed and deliberately not fixed, so they are not re-found from scratch. Phases 1–3 (the skeleton pipe, the full Claude driver, the usage screen) done 2026-08-24. Phase 4 (llama.cpp: model browsing, downloads, and `llama-server` -through its OpenAI-compatible endpoint) and phase 5 (ssh) done 2026-08-28. -The file explorer and the transcript cache followed in September. What is +through its OpenAI-compatible endpoint) and phase 5 (ssh) done 2026-08-28, +except for the remote `llama-server` and its port forward, which landed +2026-09-04. The file explorer and the transcript cache followed in September. What is left is real-phone/WireGuard bring-up, which is operational rather than code. Each phase ended runnable and verified against the real thing. The backend diff --git a/docs/RUST.md b/docs/RUST.md index 1c44591..5f96ad8 100644 --- a/docs/RUST.md +++ b/docs/RUST.md @@ -3541,6 +3541,180 @@ device. this session was told not to touch `iris/`), and anything past the emulator — the actual on-phone runs and Iris's pass/fail call. + **iris half: done, 2026-09-05.** A `bench` Cargo feature on + `iris-android-app`, built on top of `transcript-screen` + (`bench = ["transcript-screen", "dep:libc", "dep:tokio"]`, + `iris/android-app/Cargo.toml`), gives `lib.rs`'s `ActiveClient` + priority a third `AndroidAppState` (`bench_client::BenchClient`) + over `TranscriptClient` when both features are listed together -- + matching the exact build command below, which lists both. + + **Fixture.** `include_str!("../../../app/bench-fixture/assets/ + transcript.jsonl")` (1,915,760 bytes) at compile time -- no asset + pipeline needed the way the Compose half's Gradle source set does. + `bench_client::parse_fixture` splits the same way `BenchFixture.kt` + does: the first 3,200 non-blank lines parsed as `serde_json::Value`s + and folded once through `client_core::transcript_fold::fold_page` + (the real fold a `/transcript` page goes through), the rest parsed + as `event_model::SeqEvent`s and held back as the streaming tail. + `build.rs` (transcript-screen's own) now exits early under `bench` + before requiring a live server's host/port/token/CA -- `BenchClient` + never calls `build_transport()`, so that requirement made no sense + for a build that talks to nothing. + + **"Run benchmark" (`.label("Run benchmark")`) and "Copy report" + (`.label("Copy report")`)** sit in a fixed bar above the transcript; + a selectable `TextEdit` (`.attr::(())`, the same + attribute the composer field uses) below it shows the report text. + Pressing "Run benchmark" resets `FrameReport`, then drives + `List::scroll` in ~60Hz steps (`ANIM_STEP_MS = 16`) to animate each + 900px/200ms swipe rather than jumping it -- iris's `List` has no + built-in tween the way `animateScrollBy(tween(...))` gives Compose, + so this is the one place the two backends' bench code has to differ + in shape rather than only in numbers -- through the same + `rsc.tasks.redraw_handle()` + manual `request_redraw()` per step + `transcript_client.rs` already established (a `Tasks::spawn`d + future's *automatic* redraw fires once, after the whole future + completes, which would show nothing moving until the run ends). + After the scroll loop, `List::jump_to_end()` pins to the newest + content (matching `stream-bench.sh`'s "Jump to latest" tap), then + 400 fixture events replay at 20/s through `fold_event` -- the same + fold path a live SSE frame takes in `transcript_client.rs`'s own + `apply_event` -- each one triggering `rebuild_transcript`'s full + `transcript_ui::build_tree` rebuild, same tradeoff as + `TranscriptClient`/`desktop-app`. A battery sampler runs + concurrently on its own `tokio::spawn`d task (not through + `ctx.update`, since a JNI battery read needs no widget-tree access), + attaching whichever thread it runs on via a stored `JavaVM` -- + `AndroidAppState::platform_ready` (new, `IRIS.md`) is what hands + `bench_client.rs` that `JavaVM` + a `GlobalRef` to the view, since + neither was reachable from `AndroidAppState::new` before this box. + + **Report fields.** `FrameStats`'s existing `Display` (frames, janky + %, p50/p90/p99, worst, and I5's own `cpu_p50`/`gpu_wait_p50` CPU/GPU + split) plus a `bench:`-shaped tail this box added: process CPU time + via `libc::getrusage(RUSAGE_SELF)` (user+system time; chosen over + parsing `/proc/self/stat` by hand to avoid assuming `USER_HZ`), peak + RSS from `/proc/self/status`'s `VmHWM` (same source `BenchRun.kt` + reads), and battery current sampled once a second via + `BatteryManager.getIntProperty(BATTERY_PROPERTY_CURRENT_NOW)` + through direct JNI calls (`bench_jni.rs`'s `PlatformHandle` -- + `android_view::context`'s own `Context`/`Resources` wrappers have no + `getSystemService`, so this calls it directly rather than growing + that crate's wrapper for two one-off calls). `0`/`Integer.MIN_VALUE` + read as "unavailable" rather than folded into the average, matching + `BatterySampler`'s own rule and UI_RULES.md's "never present an + inferred value as a measured one." The report is logged under the + existing `iris-android-app` logcat tag on a line starting `iris + bench report:` (grep-able the same way `transcript_client.rs`'s + "Frame report" control already is), shown in the on-screen + `TextEdit`, and copied to the system clipboard by "Copy report" + through `ClipboardManager.setPrimaryClip` (`bench_jni.rs`, same + `PlatformHandle`). + + **Build commands, all clean this pass:** + - `cargo fmt --all -- --check` (iris workspace) and + `cd iris/android-app && cargo fmt --all -- --check`: clean. + - `cargo clippy --workspace --all-targets` (iris workspace): clean + (only the pre-existing `wgpu`/`winit`/`naga` future-incompat + notice). + - `cargo test --workspace` (iris workspace): 39 + 8 + 10 = the same + pre-existing counts, all passing, unaffected by this box (it + touched no logic under test there beyond `AndroidAppState`'s new + default no-op method). + - `cargo ndk -t x86_64 -P 26 clippy --features "transcript-screen + force-gles bench" --lib -- -D warnings` (`iris/android-app`): + clean. + - `cargo ndk -t arm64-v8a -P 26 -o app/src/main/jniLibs/ build + --release --features "transcript-screen force-gles bench"`: + clean, `arm64-v8a/libmain.so` produced. The pre-existing "unused + dependency `tabs-ui`" Cargo advisory also appears on a plain + `--features transcript-screen` build with no `bench` (confirmed + by building that combination alone with fake env vars) -- not + something this box introduced, and not a clippy/rustc warning + (AGENTS.md's "keep the build clean" gate is `cargo clippy`, which + stays silent on it). + + **Packaging.** No `cargo xtask apk` exists for `iris/android-app` + yet (I2's own Gradle project is the only pipeline), so this reused + that split rather than inventing one: `cargo ndk --release` above + builds the cdylib straight into `app/src/main/jniLibs/`, then a new + `release` build type in `app/build.gradle` (there was previously + only `debug`) packages and signs it -- + `AI_APP_KEYSTORE=~/.config/ai-app/release.jks` + + `AI_APP_KEYSTORE_PASSWORD` (the same key `app/build-apk.sh` + generates for the Compose app) via `gradle :app:assembleRelease`, + with `applicationIdSuffix ".bench"` so it installs beside the plain + tabs demo rather than replacing it. `aapt2 dump badging` on the + result: `package: name='dev.iris.android.demo.bench'`, one native + library, `lib/arm64-v8a/libmain.so`. `apksigner verify + --print-certs` shows the same `CN=ai-app` certificate + `compose-bench-arm64.apk` is signed with. + + **Emulator smoke run, 2026-09-05.** This checkout's own AVD + (`ai-app-2`) was in use by the session recording I5's clean-scroll + comparison in this same file (its Compose app was in the + foreground, confirmed via `dumpsys window`/`dumpsys activity + processes` before touching anything) -- rather than contend for it + (AGENTS.md's "coordinate with peer agents"), a second, + differently-named AVD was created (`AVD_NAME=ai-app-2-bench emu + up`, `pixel_10`/`android-36`/`google_apis`/`x86_64`, cold boot, host + GPU, no `EMU_GPU=software`), with 12GB of the VM's memory still + available after both were up (this-machine-android's "two are + comfortable" guidance). Installed via `adb -s emulator-5556 install + -r`, launched, driven by `ui-trace record -s emulator-5556 --do + "tap 'Run benchmark'"` (the control resolved by its accessibility + label, per AGENTS.md's "no coordinate" rule), then read back over + `adb logcat`: + + iris bench report + frames=372 janky%=56.99 p50=19.5ms p90=219.5ms p99=284.5ms worst=369.3ms (measures redraw-start to after present() is called, not GPU/compositor completion) cpu_p50=0.4ms gpu_wait_p50=13.9ms (redraw-start-to-submit vs. submit-to-after-present) + scroll: 6 cycles (24 swipes), streamed 400/400 fixture events + process CPU time over this run: 24665ms + peak RSS: 224600kB + battery current: mean 900000µA over 21 samples (min 900000, max 900000) + + "Copy report" was pressed immediately after and logged `iris bench + report: copied to clipboard` (`ClipboardManager.setPrimaryClip` + succeeded). No crash (`adb logcat`'s `FATAL`/`AndroidRuntime` lines + checked -- only `ui-trace`'s own runtime, unrelated), process alive + throughout (`dumpsys activity processes`), 400/400 stream events + confirmed sent. + + Read this the same way the Compose half's own box already asks to + read its number: this is software-rasterised (well, GLES-over-virgl + under `force-gles`, per I5's "Where iris's frame time goes") + emulator output, "the harness runs end to end and produces every + field P0 asked for," not a phone number -- and the battery current + is again the emulator's fixed 900000µA mocked charger reporting a + constant, exactly what the Compose box's own run found, not a real + battery answering. `cpu_p50=0.4ms` (iris's own per-frame CPU work) + against a much larger `gpu_wait_p50`/`p50` again matches I5's "Where + iris's frame time goes" finding under real GPU rendering (`-gpu + host`, `force-gles`) -- the frame-time budget here is dominated by + the driver/compositor wait, not by iris's layout or primitive + building, though this run's `janky%`/`p90`/`p99` are considerably + worse than that earlier isolated pass, most likely the cost of this + AVD's very first cold boot plus running two emulators on this VM at + once (a fair comparison against Compose would need both apps run + back-to-back on the same freshly-booted device, not attempted this + pass since the second AVD was torn down immediately after per + AGENTS.md's "stop yours when you are done with it"). + + Copied to `~/host/bench/iris-bench-arm64.apk` (15,445,468 bytes) and + `~/host/bench/README.md`'s "iris" section filled in (install, open, + tap "Run benchmark", read the report from the on-screen text or + logcat, tap "Copy report", paste back). + + **Not done this pass**: the actual on-phone runs and Iris's + pass/fail call between the two reports (P0's own pass condition) -- + that needs Iris's phone, which this session has no access to. + `iris/src/android/view.rs` was touched (`AndroidAppState:: + platform_ready`, `new_peer`'s wiring) -- confirmed to not be one of + the three files the concurrent `device_limits()` work on this + branch was using (`iris/core/src/render/mod.rs`, + `iris/src/android/render.rs`, `iris/src/default/render.rs`). + - [ ] **P1 — session screen parity.** History paging backward (with the page-boundary healing `client-core` does not have yet, below), `TranscriptSource`-backed cache/server stitching, jump-to-latest, diff --git a/iris/android-app/Cargo.lock b/iris/android-app/Cargo.lock index 4f1a070..592dbfd 100644 --- a/iris/android-app/Cargo.lock +++ b/iris/android-app/Cargo.lock @@ -1768,9 +1768,11 @@ dependencies = [ "client-core", "event-model", "iris", + "libc", "log", "serde_json", "tabs-ui", + "tokio", "transcript-ui", ] diff --git a/iris/android-app/Cargo.toml b/iris/android-app/Cargo.toml index 776cbb4..d0551bd 100644 --- a/iris/android-app/Cargo.toml +++ b/iris/android-app/Cargo.toml @@ -32,6 +32,23 @@ transcript-ui = { path = "../transcript-ui", optional = true } client-core = { path = "../../client-core", optional = true } event-model = { path = "../../event-model", optional = true } serde_json = { version = "1", features = ["float_roundtrip"], optional = true } +# P0's bench build only (docs/RUST.md): `getrusage(RUSAGE_SELF)` for +# process CPU time, matching `libc::getrusage`'s mention in that box over +# parsing `/proc/self/stat` by hand and assuming `USER_HZ`. Already in the +# workspace's own dependency tree transitively (`iris/Cargo.lock`, pinned +# at 0.2.179) -- this makes it a direct dependency at the same version +# rather than a second, possibly-drifting resolution. +libc = { version = "0.2.179", optional = true } +# P0's bench build only: the scroll animation and the streaming phase are +# both a sequence of `sleep`s inside the async task `rsc.spawn_task` already +# runs on iris's own tokio runtime (`iris/src/task.rs`'s `Tasks::init`), and +# the battery sampler is a second, concurrent task on that same runtime +# (`tokio::spawn`) -- so this crate needs `tokio` directly rather than only +# through `iris`. `rt`+`time` only: no I/O, no macros, nothing this crate +# doesn't call. Version matches the one `iris`'s own dependency tree already +# resolves to (`iris/Cargo.lock`), so there is one copy of the runtime, not +# two. +tokio = { version = "1.53.1", features = ["rt", "time"], optional = true } [features] default = ["tabs-screen"] @@ -41,6 +58,14 @@ transcript-screen = ["dep:transcript-ui", "dep:client-core", "dep:event-model", # instead of SwiftShader's software Vulkan. See `iris/Cargo.toml`'s own doc # on the feature this forwards to. force-gles = ["iris/force-gles"] +# P0's iris half (docs/RUST.md, docs/AGENTS.md's "The rigs"): the same +# checked-in fixture, scroll loop and streaming phase the Compose `bench` +# build type drives, run here against `transcript-ui`'s real screen with no +# server. Depends on `transcript-screen` for `transcript-ui`/`client-core`/ +# `event-model` -- `lib.rs`'s `ActiveClient` selection gives this feature +# priority over `transcript-screen`'s own `TranscriptClient` when both are +# listed, which is how this crate's build command names both explicitly. +bench = ["transcript-screen", "dep:libc", "dep:tokio"] [profile.release] panic = "abort" diff --git a/iris/android-app/app/build.gradle b/iris/android-app/app/build.gradle index 4af67aa..b1f18e9 100644 --- a/iris/android-app/app/build.gradle +++ b/iris/android-app/app/build.gradle @@ -19,9 +19,39 @@ android { versionName = "1.0" } + // A release build must be signed, and the key is per machine rather than per repo -- same + // reasoning and the same key as `app/build-apk.sh` (the Compose app): it is what a phone + // recognises the app by, and a secret never lives in a checkout (the mount is shared with an + // untrusted VM). `build-apk.sh` generates this key once and points at it through the + // environment; without it a release build here is unsigned, which is fine for everything + // except installing. + def keystore = System.getenv("AI_APP_KEYSTORE") + signingConfigs { + if (keystore != null) { + release { + storeFile = file(keystore) + storePassword = System.getenv("AI_APP_KEYSTORE_PASSWORD") + keyAlias = "ai-app" + keyPassword = storePassword + } + } + } + buildTypes { debug { } + // P0's iris half (docs/RUST.md's P0 box): the build a phone actually runs. The `.so` + // itself is built separately with `cargo ndk --release --features "transcript-screen + // force-gles bench"` straight into src/main/jniLibs/ (this crate's own Cargo.toml) -- + // Gradle here only packages and signs whatever is already there, the same division as the + // debug/tabs-screen build this project started with. `applicationIdSuffix` keeps it + // installable beside a debug build of the tabs demo rather than replacing it. + release { + applicationIdSuffix ".bench" + if (keystore != null) { + signingConfig = signingConfigs.release + } + } } compileOptions { diff --git a/iris/android-app/build.rs b/iris/android-app/build.rs index 18704fe..f6cb3fb 100644 --- a/iris/android-app/build.rs +++ b/iris/android-app/build.rs @@ -22,6 +22,14 @@ fn main() { if std::env::var_os("CARGO_FEATURE_TRANSCRIPT_SCREEN").is_none() { return; } + // P0's bench build (docs/RUST.md) opens the checked-in fixture with no + // server at all -- `bench_client.rs` never references the `pinned` + // module this generates, so requiring a live server's host/port/token/ + // CA to build it (as plain `transcript-screen` does, below) would be a + // pointless requirement for a build that talks to nothing. + if std::env::var_os("CARGO_FEATURE_BENCH").is_some() { + return; + } println!("cargo:rerun-if-env-changed=AI_APP_TRANSCRIPT_HOST"); println!("cargo:rerun-if-env-changed=AI_APP_TRANSCRIPT_PORT"); println!("cargo:rerun-if-env-changed=AI_APP_TRANSCRIPT_TOKEN"); diff --git a/iris/android-app/src/bench_client.rs b/iris/android-app/src/bench_client.rs new file mode 100644 index 0000000..2b8c755 --- /dev/null +++ b/iris/android-app/src/bench_client.rs @@ -0,0 +1,418 @@ +//! P0's iris half (docs/RUST.md's P0 box, docs/AGENTS.md's "The rigs"): +//! the same fixture, scroll loop and streaming phase the Compose `bench` +//! build type's `BenchRun.kt`/`BenchFixture.kt` drive, run here against +//! `transcript-ui`'s real screen with no server -- a frame-time comparison +//! that measures the renderer rather than the data or the network. +//! +//! **Reuses `transcript_client.rs`'s shape** (folded items, a full +//! `transcript_ui::build_tree` rebuild per event) with the network half +//! replaced by the checked-in fixture, embedded with `include_str!` -- +//! `app/bench-fixture/assets/transcript.jsonl`, 1,915,760 bytes, generated +//! by `app/bench-fixture/generate.py` and never a real transcript (that +//! file's own README). The first 3,200 lines are the opening backlog, +//! folded once through `client_core::transcript_fold::fold_page` exactly +//! as a real `/transcript` page would be; the remaining ~400 are the +//! streaming tail, replayed one at a time through `fold_event` -- the same +//! fold path a live SSE reply arrives on -- by the "Run benchmark" +//! control below. + +use crate::bench_jni::PlatformHandle; +use android_view::jni::{JavaVM, objects::GlobalRef}; +use client_core::transcript_fold::{TranscriptItem, fold_event, fold_page, group_tool_runs}; +use event_model::SeqEvent; +use iris::android::{AndroidAppState, AndroidRsc, AndroidUiState, HasAndroidUiState}; +use iris::prelude::*; +use std::sync::Arc; +use std::sync::atomic::{AtomicBool, Ordering}; +use std::time::Duration; + +/// bench-fixture/README.md: the first `BACKLOG_COUNT` non-blank lines are +/// the opening window; the rest are the streaming tail. Kept in sync with +/// `BenchFixture.kt`'s identical constant by hand -- both read the same +/// checked-in file, so a mismatch would only mean the two apps' bench +/// builds open a different split of it, not a wrong-vs-right answer. +const BACKLOG_COUNT: usize = 3200; + +/// `BenchRun.kt`'s own constants -- kept identical so the two apps' bench +/// runs are the same gesture and the same load, which is the entire point +/// of a shared fixture and a shared scripted loop (P0's pass condition). +const CYCLES: usize = 6; +const SWIPE_PX: f32 = 900.0; +const SWIPE_MS: u64 = 200; +const SWIPE_PAUSE_MS: u64 = 500; +const STREAM_EVENTS_PER_SEC: u64 = 20; +const STREAM_SECONDS: u64 = 20; +/// One animation step's target cadence -- close enough to 60Hz that a +/// `List::scroll` swipe is many small moves rather than one jump, so +/// frames are actually rendered along the way (the point of animating it +/// at all rather than calling `scroll` once per swipe). +const ANIM_STEP_MS: u64 = 16; + +const FIXTURE_JSONL: &str = include_str!("../../../app/bench-fixture/assets/transcript.jsonl"); + +pub struct BenchClient { + ui_state: AndroidUiState, + content: WeakWidget, + report_display: WeakWidget, + screen: Option, + items: Vec, + /// The events not yet streamed -- consumed by `start_benchmark`'s own + /// clone, kept here only as the source a second run would need (the + /// button can be pressed more than once; `running` just stops overlap, + /// not repeat). + stream_tail: Vec, + platform: Option>, + last_report: Option, + running: bool, +} + +impl HasAndroidUiState for BenchClient { + fn android_state(&self) -> &AndroidUiState { + &self.ui_state + } + fn android_state_mut(&mut self) -> &mut AndroidUiState { + &mut self.ui_state + } +} + +/// Parses the fixture once: `serde_json::Value`s for the backlog +/// (`fold_page` takes a page of raw wire JSON, same as a real +/// `/transcript` response) and folded `SeqEvent`s for the tail (`fold_event` +/// takes one live wire event at a time, same as a real SSE frame). +fn parse_fixture() -> (Vec, Vec) { + let lines: Vec<&str> = FIXTURE_JSONL + .lines() + .filter(|line| !line.trim().is_empty()) + .collect(); + let mut backlog = Vec::with_capacity(BACKLOG_COUNT.min(lines.len())); + let mut stream_tail = Vec::new(); + for (i, line) in lines.iter().enumerate() { + let value: serde_json::Value = + serde_json::from_str(line).expect("bench fixture is generated JSON, always valid"); + if i < BACKLOG_COUNT { + backlog.push(value); + } else { + let event: SeqEvent = serde_json::from_value(value) + .expect("bench fixture event matches event-model's SeqEvent"); + stream_tail.push(event); + } + } + (backlog, stream_tail) +} + +fn placeholder(rsc: &mut Rsc, message: &str) -> StrongWidget { + wtext(message.to_string()) + .color(Color::WHITE) + .wrap(true) + .pad(16) + .add_strong(rsc) + .any() +} + +/// `getrusage(RUSAGE_SELF)`'s user+system time, in ms -- `None` only if +/// the syscall itself fails, which UI_RULES.md's "never present an +/// inferred value as a measured one" says to keep apart from a real (and +/// here, impossible) zero. +fn process_cpu_ms() -> Option { + // SAFETY: `rusage` is a plain-old-data struct `getrusage` fully + // initialises on success; on failure it is never read. + unsafe { + let mut usage: libc::rusage = std::mem::zeroed(); + if libc::getrusage(libc::RUSAGE_SELF, &mut usage) != 0 { + return None; + } + let user_ms = usage.ru_utime.tv_sec as u64 * 1000 + usage.ru_utime.tv_usec as u64 / 1000; + let sys_ms = usage.ru_stime.tv_sec as u64 * 1000 + usage.ru_stime.tv_usec as u64 / 1000; + Some(user_ms + sys_ms) + } +} + +/// `VmHWM` from `/proc/self/status` -- the process's peak RSS since it +/// started, in kB. Same source `BenchRun.kt`'s `peakRssLine` reads, so the +/// two reports' numbers mean the same thing. +fn peak_rss_kb() -> Option { + std::fs::read_to_string("/proc/self/status") + .ok()? + .lines() + .find_map(|line| line.strip_prefix("VmHWM:")) + .and_then(|rest| rest.trim().strip_suffix("kB")) + .and_then(|n| n.trim().parse().ok()) +} + +fn battery_line(samples: &[i32]) -> String { + if samples.is_empty() { + return " battery current: unavailable on this device".to_string(); + } + let mean = samples.iter().map(|&v| v as i64).sum::() / samples.len() as i64; + let min = samples.iter().min().unwrap(); + let max = samples.iter().max().unwrap(); + format!( + " battery current: mean {mean}\u{b5}A over {} samples (min {min}, max {max})", + samples.len() + ) +} + +impl AndroidAppState for BenchClient { + fn new(mut ui_state: AndroidUiState, rsc: &mut AndroidRsc) -> Self { + let content = WidgetPtr::new().add(rsc); + let loading = placeholder(rsc, "Loading fixture..."); + content(rsc).set(loading); + + let report_display = wtext("") + .editable(EditMode::MultiLine) + .text_align(Align::LEFT) + .wrap(true) + .size(14) + .color(Color::WHITE) + .attr::(()) + .label("Benchmark report") + .add(rsc); + + let controls = bench_controls(rsc); + let tree = ( + controls, + content.height(rest(2)), + report_display.height(rest(1)).pad(8), + ) + .span(Dir::DOWN) + .add_strong(rsc) + .any(); + ui_state.set_root(tree); + + let mut client = Self { + ui_state, + content, + report_display, + screen: None, + items: Vec::new(), + stream_tail: Vec::new(), + platform: None, + last_report: None, + running: false, + }; + + let (backlog, stream_tail) = parse_fixture(); + client.stream_tail = stream_tail; + match fold_page(&backlog) { + Ok(items) => { + client.items = items; + client.rebuild_transcript(rsc); + } + Err(message) => { + client.show_message(rsc, &format!("Couldn't fold the bench fixture: {message}")) + } + } + client + } + + fn platform_ready(&mut self, _rsc: &mut AndroidRsc, vm: JavaVM, view: GlobalRef) { + self.platform = Some(Arc::new(PlatformHandle::new(vm, view))); + } + + fn back_pressed(&mut self, _rsc: &mut AndroidRsc, _render: &mut UiRenderState) -> bool { + false + } +} + +type Rsc = AndroidRsc; + +fn bench_controls(rsc: &mut Rsc) -> WeakWidget { + let run_rect = rect(Color::rgb(40, 70, 40)) + .on( + CursorSense::click(), + |ctx: EventIdCtx<'_, Rsc, _, _>, rsc: &mut Rsc| { + ctx.state.start_benchmark(rsc); + }, + ) + .label("Run benchmark"); + let run = ( + run_rect, + wtext("Run benchmark").size(18).text_align(Align::CENTER), + ) + .stack() + .pad(8) + .add(rsc); + + let copy_rect = rect(Color::rgb(50, 50, 60)) + .on( + CursorSense::click(), + |ctx: EventIdCtx<'_, Rsc, _, _>, _rsc: &mut Rsc| { + ctx.state.copy_report(); + }, + ) + .label("Copy report"); + let copy = ( + copy_rect, + wtext("Copy report").size(18).text_align(Align::CENTER), + ) + .stack() + .pad(8) + .add(rsc); + + (run, copy).span(Dir::RIGHT).height(56).add(rsc) +} + +impl BenchClient { + fn show_message(&mut self, rsc: &mut Rsc, message: &str) { + let widget = placeholder(rsc, message); + (self.content)(rsc).set(widget); + self.screen = None; + } + + fn rebuild_transcript(&mut self, rsc: &mut Rsc) { + let rows = group_tool_runs(&self.items); + let (screen, tree) = transcript_ui::build_tree(rsc, rows); + (self.content)(rsc).set(tree); + self.screen = Some(screen); + } + + fn copy_report(&mut self) { + let Some(report) = &self.last_report else { + log::info!("iris bench report: nothing to copy -- run the benchmark first"); + return; + }; + let Some(platform) = &self.platform else { + log::info!("iris bench report: no platform handle, can't reach the clipboard"); + return; + }; + if platform.copy_to_clipboard("iris bench report", report) { + log::info!("iris bench report: copied to clipboard"); + } else { + log::info!("iris bench report: clipboard copy failed"); + } + } + + /// P0's scripted run: `BenchRun.kt`'s scroll loop, then its streaming + /// phase, then the report -- run in-process for the same reason that + /// file's own doc gives (no usable system tracing on a real phone, no + /// agent that can drive one). + fn start_benchmark(&mut self, rsc: &mut Rsc) { + if self.running { + log::info!("iris bench report: already running"); + return; + } + self.running = true; + self.android_state_mut().frame_report.reset(); + self.report_display.edit(rsc).set("Running benchmark..."); + + let redraw = rsc.tasks.redraw_handle(); + let platform = self.platform.clone(); + let stream_tail = self.stream_tail.clone(); + let cpu_start = process_cpu_ms(); + + rsc.spawn_task(async move |mut ctx| { + // The swipe loop: two drags toward newer content, two back -- + // a cycle returns to where it started, so the whole loop + // measures steady-state scrolling. `BenchRun.kt`'s own + // comment on this shape. + for _ in 0..CYCLES { + for delta in [SWIPE_PX, SWIPE_PX, -SWIPE_PX, -SWIPE_PX] { + animate_scroll(&mut ctx, &redraw, delta, SWIPE_MS).await; + tokio::time::sleep(Duration::from_millis(SWIPE_PAUSE_MS)).await; + } + } + + // Pinned to the newest end before streaming starts, matching + // `stream-bench.sh`'s "Jump to latest" tap. + ctx.update(|state: &mut BenchClient, rsc| { + if let Some(screen) = &state.screen { + (screen.list)(rsc).jump_to_end(); + } + }); + redraw.request_redraw(); + + // The battery sampler runs concurrently with the streaming + // phase, once a second, the same cadence `BatterySampler` uses + // on the Compose side -- via its own JNI-attached thread, not + // `ctx.update`, since a sample needs no widget-tree access. + let sampler_done = Arc::new(AtomicBool::new(false)); + let samples = Arc::new(std::sync::Mutex::new(Vec::::new())); + let sampler = platform.clone().map(|platform| { + let done = sampler_done.clone(); + let samples = samples.clone(); + tokio::spawn(async move { + while !done.load(Ordering::Relaxed) { + if let Some(value) = platform.battery_current_ua() { + samples.lock().unwrap().push(value); + } + tokio::time::sleep(Duration::from_secs(1)).await; + } + }) + }); + + let total = (STREAM_EVENTS_PER_SEC * STREAM_SECONDS) as usize; + let mut sent = 0usize; + for event in stream_tail.into_iter().take(total) { + ctx.update(move |state: &mut BenchClient, rsc| { + state.items = fold_event(&state.items, &event); + state.rebuild_transcript(rsc); + }); + redraw.request_redraw(); + sent += 1; + tokio::time::sleep(Duration::from_millis(1000 / STREAM_EVENTS_PER_SEC)).await; + } + // Lets the last few deltas land and draw before the report is + // read -- `BenchRun.kt`'s own closing delay. + tokio::time::sleep(Duration::from_millis(300)).await; + + sampler_done.store(true, Ordering::Relaxed); + if let Some(sampler) = sampler { + let _ = sampler.await; + } + let battery = battery_line(&samples.lock().unwrap()); + let cpu_line = match (cpu_start, process_cpu_ms()) { + (Some(start), Some(end)) => { + format!(" process CPU time over this run: {}ms", end.saturating_sub(start)) + } + _ => " process CPU time over this run: unavailable".to_string(), + }; + let rss_line = match peak_rss_kb() { + Some(kb) => format!(" peak RSS: {kb}kB"), + None => " peak RSS: unavailable (/proc/self/status unreadable)".to_string(), + }; + + ctx.update(move |state: &mut BenchClient, rsc| { + state.running = false; + let scroll_line = format!( + " scroll: {CYCLES} cycles ({} swipes), streamed {sent}/{total} fixture events", + CYCLES * 4 + ); + let frames_line = match state.android_state().frame_report.report() { + Some(stats) => format!("{stats}"), + None => "no frames recorded".to_string(), + }; + let report = format!( + "iris bench report\n{frames_line}\n{scroll_line}\n{cpu_line}\n{rss_line}\n{battery}" + ); + log::info!("iris bench report: {report}"); + state.report_display.edit(rsc).set(&report); + state.last_report = Some(report); + }); + redraw.request_redraw(); + }); + } +} + +/// Moves `List::scroll` by `total_px` over `duration_ms`, in ~60Hz steps, +/// so the swipe is many rendered frames rather than one jump -- the same +/// shape `animateScrollBy(SWIPE_PX, tween(SWIPE_MS))` gives on the Compose +/// side, in the one place the two backends have to differ (iris's `List` +/// has no built-in tween, so this drives it by hand). +async fn animate_scroll( + ctx: &mut iris::task::TaskCtx, + redraw: &Arc, + total_px: f32, + duration_ms: u64, +) { + let steps = (duration_ms / ANIM_STEP_MS).max(1); + let step_px = total_px / steps as f32; + for _ in 0..steps { + ctx.update(move |state: &mut BenchClient, rsc| { + if let Some(screen) = &state.screen { + (screen.list)(rsc).scroll(step_px); + } + }); + redraw.request_redraw(); + tokio::time::sleep(Duration::from_millis(ANIM_STEP_MS)).await; + } +} diff --git a/iris/android-app/src/bench_jni.rs b/iris/android-app/src/bench_jni.rs new file mode 100644 index 0000000..0a5ae4b --- /dev/null +++ b/iris/android-app/src/bench_jni.rs @@ -0,0 +1,134 @@ +//! JNI calls the `bench` feature needs that go through the shell's own +//! Java side rather than anything `iris`/`android-view` already wraps: +//! `BatteryManager.getIntProperty(BATTERY_PROPERTY_CURRENT_NOW)` for the +//! per-second battery sample, and `ClipboardManager.setPrimaryClip` for +//! the "Copy report" control (P0's iris half, docs/RUST.md). Neither is +//! part of `android_view::context`'s own `Context`/`Resources` wrappers +//! (that file's own `// TODO: more methods?`), so this calls them +//! directly rather than growing that crate's wrapper for two one-off +//! calls this crate alone needs. +//! +//! Holds its own `JavaVM` + `GlobalRef` to the view (handed in through +//! [`iris::android::AndroidAppState::platform_ready`]) so it can attach +//! whichever thread calls it -- the battery sampler runs on a background +//! tokio task, not the UI thread the rest of `IrisViewPeer`'s JNI calls +//! run on. `JavaVM::attach_current_thread` is safe to call from a thread +//! already attached (the `jni` crate detects it and does not double +//! attach), so no caller here needs to know or care which thread it is. + +use android_view::jni::{ + JNIEnv, JavaVM, + objects::{GlobalRef, JObject, JValue}, +}; + +/// `android.os.BatteryManager.BATTERY_PROPERTY_CURRENT_NOW` -- not exposed +/// as a constant anywhere reachable without the Android SDK jar, so named +/// here with its source rather than left as a bare `2`. +const BATTERY_PROPERTY_CURRENT_NOW: i32 = 2; + +pub struct PlatformHandle { + vm: JavaVM, + view: GlobalRef, +} + +impl PlatformHandle { + pub fn new(vm: JavaVM, view: GlobalRef) -> Self { + Self { vm, view } + } + + fn context<'e>(&self, env: &mut JNIEnv<'e>) -> Option> { + env.call_method( + self.view.as_obj(), + "getContext", + "()Landroid/content/Context;", + &[], + ) + .ok()? + .l() + .ok() + } + + fn system_service<'e>( + &self, + env: &mut JNIEnv<'e>, + context: &JObject<'e>, + name: &str, + ) -> Option> { + let jname = env.new_string(name).ok()?; + env.call_method( + context, + "getSystemService", + "(Ljava/lang/String;)Ljava/lang/Object;", + &[JValue::Object(jname.as_ref())], + ) + .ok()? + .l() + .ok() + } + + /// One sample of `BATTERY_PROPERTY_CURRENT_NOW`, in microamps. `None` + /// on any JNI failure, on a device with no `BatteryManager` service, + /// or when the platform itself answers "not supported" -- `0` or + /// `Integer.MIN_VALUE` are both documented SDK answers for that, and + /// both would read as a real (and wrong) measurement if folded into an + /// average rather than named apart. UI_RULES.md: never present an + /// inferred value as a measured one. + pub fn battery_current_ua(&self) -> Option { + let mut guard = self.vm.attach_current_thread().ok()?; + let env: &mut JNIEnv = &mut guard; + let context = self.context(env)?; + let battery_manager = self.system_service(env, &context, "batterymanager")?; + let value = env + .call_method( + &battery_manager, + "getIntProperty", + "(I)I", + &[JValue::Int(BATTERY_PROPERTY_CURRENT_NOW)], + ) + .ok()? + .i() + .ok()?; + if value == 0 || value == i32::MIN { + None + } else { + Some(value) + } + } + + /// Puts `text` on the system clipboard through `ClipboardManager` -- + /// `true` only if the whole JNI chain (service lookup, `ClipData`, + /// `setPrimaryClip`) succeeded. + pub fn copy_to_clipboard(&self, label: &str, text: &str) -> bool { + self.try_copy_to_clipboard(label, text).is_some() + } + + fn try_copy_to_clipboard(&self, label: &str, text: &str) -> Option<()> { + let mut guard = self.vm.attach_current_thread().ok()?; + let env: &mut JNIEnv = &mut guard; + let context = self.context(env)?; + let clipboard = self.system_service(env, &context, "clipboard")?; + let jlabel = env.new_string(label).ok()?; + let jtext = env.new_string(text).ok()?; + let clip = env + .call_static_method( + "android/content/ClipData", + "newPlainText", + "(Ljava/lang/CharSequence;Ljava/lang/CharSequence;)Landroid/content/ClipData;", + &[ + JValue::Object(jlabel.as_ref()), + JValue::Object(jtext.as_ref()), + ], + ) + .ok()? + .l() + .ok()?; + env.call_method( + &clipboard, + "setPrimaryClip", + "(Landroid/content/ClipData;)V", + &[JValue::Object(&clip)], + ) + .ok()?; + Some(()) + } +} diff --git a/iris/android-app/src/lib.rs b/iris/android-app/src/lib.rs index 6a8cc32..8825688 100644 --- a/iris/android-app/src/lib.rs +++ b/iris/android-app/src/lib.rs @@ -23,6 +23,18 @@ //! A build picks one screen or the other, never both, so `Client` and //! `TranscriptClient` are cfg-gated apart rather than switched at runtime -- //! there is no in-app navigation to switch *to* on either side yet. +//! +//! **`bench` feature (P0's iris half, docs/RUST.md):** a third +//! `AndroidAppState`, `bench_client::BenchClient`, on the same axis -- +//! `transcript_ui::build_tree` again, this time against the checked-in +//! fixture (`app/bench-fixture/assets/transcript.jsonl`) instead of a real +//! server, with a "Run benchmark" control that drives the same scroll loop +//! and streaming phase the Compose `bench` build type's `BenchRun.kt` +//! does. `bench` depends on `transcript-screen` (Cargo.toml) for +//! `transcript-ui`/`client-core`/`event-model`, so both features end up +//! enabled together -- `ActiveClient` below gives `bench` priority in that +//! case, the same way `transcript-screen` already takes priority over the +//! default `tabs-screen`. use android_view::{ Context, View, @@ -39,7 +51,11 @@ use iris::prelude::*; use log::LevelFilter; use std::ffi::c_void; -#[cfg(feature = "transcript-screen")] +#[cfg(feature = "bench")] +mod bench_client; +#[cfg(feature = "bench")] +mod bench_jni; +#[cfg(all(feature = "transcript-screen", not(feature = "bench")))] mod transcript_client; /// The app's `View` subclass, matching the Java side's package -- @@ -85,8 +101,10 @@ impl AndroidAppState for Client { #[cfg(not(feature = "transcript-screen"))] type ActiveClient = Client; -#[cfg(feature = "transcript-screen")] +#[cfg(all(feature = "transcript-screen", not(feature = "bench")))] type ActiveClient = transcript_client::TranscriptClient; +#[cfg(feature = "bench")] +type ActiveClient = bench_client::BenchClient; extern "system" fn new_view_peer<'local>( env: JNIEnv<'local>, diff --git a/iris/src/android/view.rs b/iris/src/android/view.rs index f735c21..ed3128b 100644 --- a/iris/src/android/view.rs +++ b/iris/src/android/view.rs @@ -4,7 +4,7 @@ use accesskit_android::Adapter as AccessAdapter; use android_view::{ AccessibilityNodeInfo, AccessibilityNodeProvider, Bundle, CallbackCtx, Context, InputConnection, KeyEvent, MotionEvent, Rect, View, ViewPeer, - jni::{JNIEnv, sys::jint}, + jni::{JNIEnv, JavaVM, objects::GlobalRef, sys::jint}, ndk::event::{Keycode, MotionAction}, }; // `marker::Sized` explicitly: `crate::prelude::*` below also brings in the @@ -107,6 +107,19 @@ pub trait AndroidAppState: HasAndroidUiState { fn back_pressed(&mut self, rsc: &mut AndroidRsc, render: &mut UiRenderState) -> bool { false } + /// Called once, right after `new`, with a fresh `JavaVM` handle and a + /// global reference to this app's own `View` -- for a caller that + /// needs to call into Java itself beyond what a [`RequestRedraw`] + /// handle already covers (P0's bench build calling + /// `BatteryManager`/`ClipboardManager` through the view's `Context`, + /// docs/RUST.md). Not folded into `new` itself: most implementors need + /// nothing here, and `new`'s job is building the widget tree, not + /// holding a platform handle -- the default does nothing. `vm`/`view` + /// are independent handles from the ones `new_peer` keeps for its own + /// `RequestRedraw` (a fresh `get_java_vm`/`new_global_ref` each), so + /// storing them has no effect on that mechanism. + #[allow(unused_variables)] + fn platform_ready(&mut self, rsc: &mut AndroidRsc, vm: JavaVM, view: GlobalRef) {} } /// The android-view analogue of `default::DefaultRsc` -- identical in @@ -561,7 +574,10 @@ pub fn new_peer<'local, State: AndroidAppState>( }; let shared = Rc::new(RefCell::new(Shared::default())); let ui_state = AndroidUiState::new(shared.clone()); - let state = State::new(ui_state, &mut rsc); + let mut state = State::new(ui_state, &mut rsc); + let platform_vm = env.get_java_vm().unwrap(); + let platform_view = env.new_global_ref(&view.0).unwrap(); + state.platform_ready(&mut rsc, platform_vm, platform_view); let peer = IrisViewPeer { rsc, render: UiRenderState::new(), diff --git a/server/Cargo.lock b/server/Cargo.lock index ae22ee9..66c4199 100644 --- a/server/Cargo.lock +++ b/server/Cargo.lock @@ -36,6 +36,7 @@ dependencies = [ "sha2", "tempfile", "thiserror", + "time", "tokio", "tokio-stream", "tower", diff --git a/server/Cargo.toml b/server/Cargo.toml index 5cdc805..43e590d 100644 --- a/server/Cargo.toml +++ b/server/Cargo.toml @@ -57,6 +57,12 @@ ureq = { version = "3", features = ["json"] } # both in the graph rustls refuses to auto-select one. rustls = "0.23" libc = "0.2.189" +# One ISO-8601 timestamp: the reset time on the invented rate-limit window +# an echo session's `/usage` puts up. Already in the tree behind the +# certificate machinery, so this is a direct name for what is compiled +# anyway rather than a new crate -- and the alternative was hand-rolling a +# civil-from-days conversion to print one line. +time = { version = "0.3", features = ["formatting"] } [dev-dependencies] tempfile = "3" diff --git a/server/src/auth.rs b/server/src/auth.rs index 7719bce..b8dc6ac 100644 --- a/server/src/auth.rs +++ b/server/src/auth.rs @@ -207,6 +207,20 @@ mod tests { /// like any other. #[tokio::test] async fn a_spooled_enrollment_is_adopted_on_first_use() { + // Under a subscriber, like every other exercise of this middleware. + // `tracing` caches a callsite's interest process-wide the first time it + // is reached, so the refusal at the end of this test -- reached with no + // subscriber on this thread -- could cache the rejection warning as + // never-enabled and make the tripwire above see an empty log. That + // failed about one full-suite run in ten, in the test that exists to + // notice a credential leak, which is the worst place for a flake. + let _guard = tracing::subscriber::set_default( + tracing_subscriber::fmt() + .with_max_level(tracing::Level::TRACE) + .with_writer(std::io::sink) + .finish(), + ); + let dir = tempfile::tempdir().expect("tempdir"); let manager = manager_with_token(dir.path(), "first"); let spooled = generate_token(); diff --git a/server/src/config.rs b/server/src/config.rs index 2cfe565..7e5fb8f 100644 --- a/server/src/config.rs +++ b/server/src/config.rs @@ -30,6 +30,18 @@ pub struct Config { pub tokens: Vec, pub setups: Vec, pub sessions: Vec, + /// What a new session's thinking level is when nothing chose one. + /// + /// Here rather than on a provider because providers are *discovered*: a + /// default written onto one would be erased by the next rediscovery, which + /// is the kind of setting that looks like it stuck until the day it did + /// not. Here rather than on the phone because a second device would then + /// spawn sessions the first one's owner did not expect. + /// + /// `None` is the CLI's own default, and stays reachable: this is a level + /// somebody chose, not a level this app picked for them. + #[serde(default, skip_serializing_if = "Option::is_none")] + pub default_effort: Option, } /// A machine, and the things it can run. @@ -78,6 +90,16 @@ pub struct ProviderConfig { pub models: Vec, } +impl ProviderConfig { + /// The executable to run for this provider: its override, or its kind's + /// default. + pub fn program(&self) -> &str { + self.command + .as_deref() + .unwrap_or(self.kind.default_program()) + } +} + /// How to reach a setup that isn't this machine, with the system `ssh` client /// -- so `~/.ssh/config`, agents and jump hosts all keep working, and there is /// one place to configure connections. A remote session is the identical @@ -94,6 +116,19 @@ pub struct SshConfig { /// Extra `-o` settings, each written as `Key=value`. #[serde(default, skip_serializing_if = "Vec::is_empty")] pub options: Vec, + /// Where this machine keeps the GGUF models it can serve, absent for + /// the same default this backend uses (`~/.local/share/ai-app/models` + /// -- `$XDG_DATA_HOME` is not read on the far side, since it is this + /// machine's environment that would answer). A `~` prefix is the + /// remote home. + /// + /// Here rather than on the provider because it is a fact about the + /// machine, and because a machine reached over ssh is where the model + /// has to be: a llama.cpp session serves the file from the machine + /// that runs `llama-server`, and this backend's own downloads are on + /// whichever machine that is only when they are the same one. + #[serde(default, skip_serializing_if = "Option::is_none")] + pub models_dir: Option, /// Where a file attached from the phone is put on this machine so the /// session can read it. Absent means the session's own working directory, /// or the login home for a session that has none. A `~` prefix is the @@ -145,6 +180,45 @@ impl DriverKind { } } + /// Which paid service meters a session of this kind, and `None` for one + /// that costs nothing. + /// + /// What decides which account -- if any -- a rate-limit bar is about is the + /// provider a session runs, not the machine it runs on: an echo session on + /// a machine that also has the Claude CLI was drawn with that CLI's + /// five-hour window, a quota it cannot spend. + /// + /// Echo names a meter of its own that exists only when a test has asked for + /// one (`/usage` in `session::echo`), which is how the bar's states are + /// reached without an account. With none set there is no snapshot, and the + /// phone draws nothing. + /// + /// The string is a [`crate::usage::UsageProvider::name`], and it is what + /// pairs a session with one of `GET /usage`'s snapshots -- so + /// `usage::providers_for` reads this rather than matching on kinds again. + pub fn usage_provider(self) -> Option<&'static str> { + match self { + Self::ClaudeCli => Some(crate::usage::CLAUDE), + Self::Echo => Some(crate::usage::ECHO), + Self::LlamaCpp => None, + } + } + + /// The executable a provider of this kind runs when it names none. + /// + /// Here rather than at each spawn site because it is not only the spawn + /// that runs it: `usage` runs the Claude CLI too, to have it refresh its + /// own OAuth token, and a default that disagreed with the driver's would + /// ask the wrong binary on a machine with two installs. + pub fn default_program(self) -> &'static str { + match self { + Self::ClaudeCli => "claude", + Self::LlamaCpp => "llama-server", + // Echo is translated in-process; nothing is spawned for it. + Self::Echo => "echo", + } + } + /// Whether the conversation exists outside this app, so that deleting the /// session here does not end it. /// @@ -162,6 +236,24 @@ impl DriverKind { Self::Echo | Self::LlamaCpp => false, } } + + /// Whether a thinking level means anything to this kind, so the phone can + /// offer the control only where it does something. + /// + /// Reported from here rather than decided on the phone, and asked of the + /// *kind* rather than branched on: the alternative is the session-type + /// `if` this app does not have anywhere else. `--effort` is the Claude + /// CLI's; a llama session's sampling is `params`, and echo does not think. + /// + /// It matters more than a control that would simply do nothing, because + /// choosing a level stops the process -- so on a session that cannot use + /// one it is a button whose only effect is the cost. + pub fn takes_effort(self) -> bool { + match self { + Self::ClaudeCli => true, + Self::Echo | Self::LlamaCpp => false, + } + } } #[derive(Debug, Clone, Serialize, Deserialize)] @@ -196,6 +288,17 @@ pub struct SessionConfig { /// the CLI stays the one authority on which modes exist. #[serde(skip_serializing_if = "Option::is_none")] pub permission_mode: Option, + /// How hard the model thinks, passed straight to `--effort`. A string for + /// the same reason `permission_mode` is: the CLI owns which levels exist. + /// + /// Unlike the model and the mode, there is no control request that changes + /// one -- checked against 2.1.258, whose only two are `set_model` and + /// `set_permission_mode` -- so this is settled at launch and `None` means + /// whatever the CLI's own default is. That is a state the phone has to be + /// able to *choose*, not just start in, which is why it is an option + /// rather than a level with a default written here. + #[serde(skip_serializing_if = "Option::is_none")] + pub effort: Option, /// Settings the driver interprets, chosen at spawn. /// /// Deliberately untyped: what a temperature or a context size means is the @@ -215,6 +318,30 @@ pub struct SessionConfig { /// turned off in one tap where one that never arrived is not diagnosable. #[serde(default = "notify_default")] pub notify: bool, + /// Whether a session stopped by the account's usage limit sends itself a + /// message once the limit lifts, instead of waiting for a person. + /// + /// Off unless somebody asked for it. It spends quota the moment it becomes + /// available and it does so while nobody is looking, which is exactly the + /// kind of thing that must not happen because a default said so. + #[serde(default, skip_serializing_if = "not_set")] + pub auto_resume: bool, + /// What that message says. `None` is [`DEFAULT_RESUME_MESSAGE`], and stays + /// reachable: it is this app's word, not one somebody chose, so clearing + /// the field goes back to it rather than sending an empty message. + #[serde(default, skip_serializing_if = "Option::is_none")] + pub auto_resume_message: Option, + /// The message this session owes itself once the limit lifts, and when to + /// try. Written when a limit is hit, moved when the wait turns out to be + /// wrong, and cleared when the message goes out or auto-resume is turned + /// off -- see [`ScheduledResume`]. + /// + /// Persisted rather than held in memory because the wait outlives the + /// process doing it: a five-hour window and a weekly one both routinely + /// outlast a backend restart, and a resume forgotten across one is a + /// session that silently never comes back. + #[serde(default, skip_serializing_if = "Option::is_none")] + pub resume: Option, /// Whether this session's process is stopped when the server exits, instead /// of being left running for the next start to adopt. /// @@ -231,6 +358,28 @@ pub struct SessionConfig { pub created: f64, } +/// A message owed to a session whose account ran out, and when to try sending +/// it. +/// +/// `since` is the whole reason this is a struct: the wait is rescheduled every +/// time the meter is asked and still says no, so `at` alone cannot say how long +/// this has been going on -- and something has to, or a machine that can never +/// be asked is retried until somebody notices. See `crate::resume`. +#[derive(Debug, Clone, Copy, Serialize, Deserialize)] +#[serde(rename_all = "camelCase")] +pub struct ScheduledResume { + /// Epoch seconds: when the limit is next worth checking. Never a promise + /// that the message goes out then -- the meter is asked first. + pub at: f64, + /// Epoch seconds the limit was hit. + pub since: f64, +} + +/// What an auto-resume says when nothing else was chosen. One word, because +/// the session already knows what it was doing and this is only the nudge that +/// lets it carry on. +pub const DEFAULT_RESUME_MESSAGE: &str = "continue"; + fn notify_default() -> bool { true } @@ -385,6 +534,7 @@ mod tests { port: Some(2222), identity_file: None, options: Vec::new(), + models_dir: None, attachments_dir: None, }), providers: vec![ProviderConfig { @@ -395,6 +545,7 @@ mod tests { }], }, ], + default_effort: Some("low".to_string()), sessions: vec![SessionConfig { id: "abc123".to_string(), setup: "vm".to_string(), @@ -403,8 +554,12 @@ mod tests { model: None, cwd: None, permission_mode: None, + effort: None, params: BTreeMap::new(), notify: true, + auto_resume: false, + auto_resume_message: None, + resume: None, throwaway: false, created: 1234.5, }], diff --git a/server/src/main.rs b/server/src/main.rs index fb6f479..3571b7b 100644 --- a/server/src/main.rs +++ b/server/src/main.rs @@ -16,6 +16,7 @@ mod config; mod files; mod media; mod models; +mod resume; mod routes; mod session; mod setups; @@ -264,7 +265,16 @@ async fn main() -> Result<()> { // No providers listed here any more: which machines can be asked, and about // what, comes from the setups at the moment the screen is opened -- so a // machine added from the phone reports its limits without a restart. - let monitor = Arc::new(usage::UsageMonitor::new()); + // The fixture is the manager's, because that is where the `/usage` command + // that sets it is typed; the monitor is what serves it. + let monitor = Arc::new(usage::UsageMonitor::new(manager.usage_fixture())); + + // The one thing in here that acts without a request behind it: a session + // switched to auto-resume waits out its account's usage limit and picks + // itself back up. Started whether or not any session has it on, because + // the setting is per session and changes from the phone -- see + // `resume::run`. + tokio::spawn(resume::run(Arc::clone(&manager), Arc::clone(&monitor))); // The bearer-token middleware wraps the entire router -- routes and fallback // alike -- here and only here, so a new route can't forget auth. diff --git a/server/src/models.rs b/server/src/models.rs index ccbbdf1..1b91dcb 100644 --- a/server/src/models.rs +++ b/server/src/models.rs @@ -29,6 +29,8 @@ use serde::Serialize; use wg_app_link::private; +use crate::session::transport::{Launch, Transport}; + /// Identifies this client to HuggingFace. They ask for one, and a request /// without it is more likely to be rate-limited. const USER_AGENT: &str = concat!("ai-server/", env!("CARGO_PKG_VERSION")); @@ -521,6 +523,83 @@ fn collect(root: &Path, dir: &Path, found: &mut Vec) { } } +/// Where a machine reached over ssh keeps its models, when its setup does +/// not say. +/// +/// The same place this backend puts its own downloads, written out rather +/// than derived: `$XDG_DATA_HOME` here describes *this* machine's +/// environment, and the far machine's is the far machine's business. A +/// setup whose models are elsewhere says so (`SshConfig::models_dir`). +const FAR_MODELS_DIR: &str = "~/.local/share/ai-app/models"; + +/// Which directory holds the models on the machine `transport` reaches. +/// +/// One answer, because two things ask: the list a spawn screen offers, +/// and the path a session hands `llama-server`. A machine that listed one +/// directory and served from another would offer models that then failed +/// to load, which reads as the model being broken. +pub fn dir_on(transport: &Transport, local: &Path) -> String { + match transport { + Transport::Here => local.to_string_lossy().into_owned(), + Transport::Ssh { ssh, .. } => ssh + .models_dir + .as_ref() + .map_or(FAR_MODELS_DIR.to_string(), |dir| { + dir.to_string_lossy().into_owned() + }), + } +} + +/// Every GGUF on the machine a setup names, which is the machine that +/// would have to serve it. +/// +/// The local half of this is [`ModelStore::list`], reading the same shape +/// off this machine's disk; a caller picks by transport, since a setup +/// with no ssh *is* this machine and asking a shell about it would be a +/// slower way to the same answer. What must not happen is offering this +/// backend's downloads for a session on another machine: the file has to +/// be where `llama-server` runs, and a list that says otherwise is a +/// claim about the wrong filesystem. +/// +/// `dir` is that machine's models directory, `~` included -- expanded on +/// the far side, which is the only place that knows what it is. A +/// directory that is not there is an empty list rather than a failure: a +/// machine that has never had a model put on it is an ordinary state, and +/// the same one as a machine whose directory exists and is empty. +pub async fn on_machine(transport: &Transport, dir: &str) -> Result> { + let script = "p=$1; case $p in \"~\") p=$HOME;; \"~/\"*) p=$HOME/${p#\"~/\"};; esac; \ + [ -d \"$p\" ] || exit 0; \ + find \"$p\" -type f -name '*.gguf' -printf '%s\\t%P\\0'"; + let launch = Launch::new( + "sh", + vec![ + "-c".to_string(), + script.to_string(), + "sh".to_string(), + dir.to_string(), + ], + None, + ); + let out = transport.capture(&launch).await?; + let mut found: Vec = out + .split('\0') + .filter(|record| !record.is_empty()) + // Two fields, and the name last, so a `\t` in a filename survives. + .filter_map(|record| record.split_once('\t')) + .filter_map(|(bytes, key)| { + let (repo, file) = key.rsplit_once('/')?; + Some(LocalModel { + key: key.to_string(), + repo: repo.to_string(), + file: file.to_string(), + bytes: bytes.trim().parse().unwrap_or(0), + }) + }) + .collect(); + found.sort_by(|a, b| a.key.cmp(&b.key)); + Ok(found) +} + /// A model repository on HuggingFace, as the browse screen shows it. #[derive(Debug, Clone, Serialize)] #[serde(rename_all = "camelCase")] diff --git a/server/src/resume.rs b/server/src/resume.rs new file mode 100644 index 0000000..2ef56d4 --- /dev/null +++ b/server/src/resume.rs @@ -0,0 +1,382 @@ +//! Auto-resume: picking a session back up when its account's usage limit +//! lifts. +//! +//! Off unless a session was switched to it, because this spends quota the +//! moment quota exists and does it while nobody is watching. What it does is +//! narrow on purpose: it sends one message -- "continue" unless something else +//! was typed -- to a session that stopped because the account ran out, and +//! then it is done. There is no retry loop around the conversation itself. +//! +//! **The schedule is a plan to ask, never a plan to send.** A reset time is +//! the one thing here that cannot be trusted: the dialect's is a hint written +//! when the turn failed, the endpoint's moves when the window moves, and both +//! are wrong across the case this exists for -- a limit that lifts later than +//! it said. So the wait ends in a *question* to [`crate::usage`], and only an +//! answer that says the limits no longer apply sends anything. Every other +//! answer, including one that cannot be got at all, becomes a new wait. +//! +//! This is the top layer: it holds the session manager and the usage monitor +//! and neither holds it. That is what lets the decision below be a pure +//! function of a snapshot and a clock, which is the whole of what is worth +//! testing here. + +use std::sync::Arc; +use std::time::Duration; + +use crate::session::{LimitHit, OwedResume, SessionManager, now}; +use crate::usage::{UsageMonitor, UsageSnapshot, UsageState}; + +/// How often to look at the schedule. Coarse deliberately: a wait measured in +/// hours does not deserve a fine-grained clock, and the meter behind it is +/// cached for three minutes anyway. +const TICK: Duration = Duration::from_secs(60); + +/// How close to a scheduled check is close enough to ask the meter. Anything +/// further out is left alone, so a session waiting five hours costs nothing +/// until the last few minutes of it. +const NEARLY: f64 = 300.0; + +/// How long to wait after an answer that decided nothing -- the machine could +/// not be asked, or it says the limit is still on with no reset time. +const BACKOFF: f64 = 300.0; + +/// The least time to wait before asking again, whatever a reset time says. A +/// window that claims to reset in the past would otherwise be asked about on +/// every tick. +const AT_LEAST: f64 = 60.0; + +/// How long after the limit was hit to stop waiting. +/// +/// Something has to bound it, or a machine that can never be asked -- an +/// unplugged laptop, a setup somebody edited away -- is retried for ever with +/// nothing on screen saying so. A day is past the longest window Claude +/// reports, so reaching this means the wait was never going to end on its own. +const GIVE_UP: f64 = 24.0 * 60.0 * 60.0; + +/// The percentage at which a window is spent. The API counts up to 100, so +/// this is an equality in all but name; written as a threshold because a +/// figure arriving slightly over is a full window, not a corrupt one. +const SPENT: f64 = 100.0; + +/// What to do about one owed resume, having asked the meter. +#[derive(Debug, Clone, Copy, PartialEq)] +pub enum Step { + /// The limits no longer apply: send the message. + Send, + /// Ask again at this epoch second. + WaitUntil(f64), + /// This has been waiting longer than anything real would take. + GiveUp, +} + +/// Runs the schedule until the server stops. +/// +/// Two things wake it: the tick, and a session reporting that it has just run +/// out. The second is not an optimisation -- a limit hit is what *creates* a +/// schedule, and a tick that happened a moment before it would otherwise leave +/// the session unrecorded until the next one. +pub async fn run(manager: Arc, monitor: Arc) { + let mut limits = manager.subscribe_limits(); + loop { + tokio::select! { + _ = tokio::time::sleep(TICK) => {} + hit = limits.recv() => match hit { + Ok(LimitHit { session_id, resets_at }) => note(&manager, &session_id, resets_at), + // Lagged: some reports were dropped, and a session that hit a + // limit while this was busy has no schedule. Nothing is lost + // for good -- the sweep below reads the config, and the + // session will report again the next time it is poked -- but + // it is worth saying, because until then that session waits + // for a person. + Err(tokio::sync::broadcast::error::RecvError::Lagged(missed)) => { + tracing::warn!("auto-resume missed {missed} limit reports"); + } + Err(tokio::sync::broadcast::error::RecvError::Closed) => return, + }, + } + sweep(&manager, &monitor).await; + } +} + +/// Records a limit against the session that hit it, if it is one that resumes. +pub(crate) fn note(manager: &SessionManager, session_id: &str, resets_at: Option) { + match manager.note_limit(session_id, resets_at) { + Ok(true) => tracing::info!("session {session_id} hit its usage limit; auto-resume is on"), + Ok(false) => {} + Err(err) => tracing::error!("couldn't schedule a resume for {session_id}: {err:#}"), + } +} + +/// One pass over everything owed a message. +async fn sweep(manager: &SessionManager, monitor: &Arc) { + let at = now(); + for owed in manager.owed_resumes() { + if owed.scheduled.at - at > NEARLY { + continue; + } + // Asked per session rather than once for the whole sweep: the answer + // is cached per machine and per meter, so several sessions on one + // account share one fetch, and a machine nobody is waiting on is not + // dialled at all. + let snapshot = snapshot_for(Arc::clone(monitor), manager, &owed).await; + match decide(snapshot.as_ref(), &owed, now()) { + Step::Send => match manager.resume_now(&owed.session_id) { + Ok(message) => tracing::info!( + "the limit on {} has lifted; sent \"{message}\" to {}", + owed.setup, + owed.session_id + ), + Err(err) => { + tracing::error!("couldn't resume {}: {err:#}", owed.session_id) + } + }, + Step::WaitUntil(next) => { + if let Err(err) = manager.reschedule_resume(&owed.session_id, next) { + tracing::error!( + "couldn't move {}'s resume to {next}: {err:#}", + owed.session_id + ); + } + } + Step::GiveUp => { + // About the machine rather than in the state's own words: the + // detail is in the log, and what lands in the transcript has + // to read on a phone. + let why = match snapshot.as_ref().map(|snapshot| &snapshot.state) { + Some(UsageState::Ok) => "the limit has not lifted in a day".to_string(), + _ => format!("{} could not be asked for a day", owed.setup), + }; + if let Err(err) = manager.abandon_resume(&owed.session_id, &why) { + tracing::error!("couldn't clear {}'s resume: {err:#}", owed.session_id); + } + } + } + } +} + +/// The numbers for the machine and the meter this session is billed against, +/// and `None` when nothing reports on it. +/// +/// Blocking work, so it goes to a blocking thread: the fetch behind it reads a +/// credential file over ssh and then makes an HTTP call. +async fn snapshot_for( + monitor: Arc, + manager: &SessionManager, + owed: &OwedResume, +) -> Option { + let setups: Vec<_> = manager + .setups() + .into_iter() + .filter(|setup| setup.id == owed.setup) + .collect(); + if setups.is_empty() { + return None; + } + let provider = owed.provider; + tokio::task::spawn_blocking(move || { + monitor + .snapshots(&setups) + .into_iter() + .find(|snapshot| snapshot.provider == provider) + }) + .await + .unwrap_or_default() +} + +/// What one owed resume should do, given what the meter said and the time. +/// +/// A pure function of the two, which is what makes the rule inspectable: every +/// answer that is not "the limits no longer apply" is a longer wait, and the +/// only thing that ends the waiting other than success is the clock. +/// +/// The reset time comes from the *snapshot* rather than from the schedule, so +/// a window that turns out to reset later than the dialect said pushes the +/// check back, and one that resets sooner pulls it forward. That is the case +/// the whole design is about: the first answer was a guess, this one is a +/// measurement. +pub fn decide(snapshot: Option<&UsageSnapshot>, owed: &OwedResume, at: f64) -> Step { + let step = match snapshot { + // The meter answered with numbers, which is the only answer that can + // send anything. + Some(snapshot) if snapshot.state == UsageState::Ok => { + let spent: Vec<&crate::usage::UsageWindow> = snapshot + .windows + .iter() + .filter(|window| window.percent >= SPENT) + .collect(); + if spent.is_empty() { + Step::Send + } else { + // The earliest of the spent windows: it is the first moment + // the situation can have changed, and if the others are still + // full this comes straight back here. + match spent + .iter() + .filter_map(|window| epoch_of(window.resets_at.as_deref())) + .min_by(f64::total_cmp) + { + Some(resets) => Step::WaitUntil(resets), + // Spent with no reset time anybody could read. Not a + // reason to send: what is known is that the limit is on. + None => Step::WaitUntil(at + BACKOFF), + } + } + } + // Logged out, unreachable, or the endpoint refused us -- and nothing + // at all, which is a session whose machine or provider has gone. None + // of them says the limit has lifted, and sending on any of them is + // exactly the "inferred value presented as a measured one" this is + // built to avoid. + _ => Step::WaitUntil(at + BACKOFF), + }; + match step { + // Waiting past the point where a real window would have reset means + // whatever is wrong is not going to fix itself. + Step::WaitUntil(_) if at - owed.scheduled.since > GIVE_UP => Step::GiveUp, + Step::WaitUntil(next) => Step::WaitUntil(next.max(at + AT_LEAST)), + other => other, + } +} + +/// An RFC-3339 timestamp as epoch seconds, and `None` for one that is absent +/// or unreadable -- the same two answers the phone's countdown makes, kept +/// apart from each other nowhere here because both mean "this cannot decide +/// when to ask". +fn epoch_of(resets_at: Option<&str>) -> Option { + let text = resets_at?; + time::OffsetDateTime::parse(text, &time::format_description::well_known::Rfc3339) + .ok() + .map(|at| at.unix_timestamp() as f64) +} + +#[cfg(test)] +mod tests { + use super::*; + use crate::config::ScheduledResume; + use crate::usage::UsageWindow; + + fn owed(since: f64) -> OwedResume { + OwedResume { + session_id: "s1".to_string(), + setup: "local".to_string(), + provider: crate::usage::CLAUDE, + scheduled: ScheduledResume { at: since, since }, + } + } + + fn snapshot(state: UsageState, windows: Vec) -> UsageSnapshot { + UsageSnapshot { + provider: crate::usage::CLAUDE.to_string(), + setup: "local".to_string(), + setup_name: "this machine".to_string(), + state, + windows, + fetched_at: 0.0, + } + } + + fn window(percent: f64, resets_at: Option<&str>) -> UsageWindow { + UsageWindow { + kind: "session".to_string(), + label: "5-hour window".to_string(), + percent, + resets_at: resets_at.map(str::to_string), + active: true, + } + } + + #[test] + fn a_meter_with_room_in_it_is_the_only_thing_that_sends() { + let clear = snapshot(UsageState::Ok, vec![window(41.0, None)]); + assert_eq!(decide(Some(&clear), &owed(0.0), 100.0), Step::Send); + } + + #[test] + fn a_window_still_spent_moves_the_check_to_its_own_reset_time() { + // The case the feature exists for: the wait was scheduled for one + // time, the limit is still on, and the endpoint now names another. + let at = 1_788_546_972.0; + let later = "2026-09-05T12:00:00+00:00"; + let spent = snapshot(UsageState::Ok, vec![window(100.0, Some(later))]); + assert_eq!( + decide(Some(&spent), &owed(at - 60.0), at), + Step::WaitUntil(epoch_of(Some(later)).expect("parses")) + ); + } + + #[test] + fn a_reset_time_already_past_still_waits_a_little() { + let at = 1_788_546_972.0; + let spent = snapshot( + UsageState::Ok, + vec![window(100.0, Some("2020-01-01T00:00:00+00:00"))], + ); + assert_eq!( + decide(Some(&spent), &owed(at - 60.0), at), + Step::WaitUntil(at + AT_LEAST) + ); + } + + #[test] + fn the_earliest_spent_window_is_the_one_worth_waiting_on() { + let at = 1_788_546_972.0; + let soon = "2026-09-05T12:00:00+00:00"; + let far = "2026-09-09T12:00:00+00:00"; + let mut weekly = window(100.0, Some(far)); + weekly.kind = "weekly_all".to_string(); + let spent = snapshot(UsageState::Ok, vec![window(100.0, Some(soon)), weekly]); + assert_eq!( + decide(Some(&spent), &owed(at - 60.0), at), + Step::WaitUntil(epoch_of(Some(soon)).expect("parses")) + ); + } + + #[test] + fn a_meter_that_could_not_be_asked_never_sends() { + let at = 1_788_546_972.0; + for state in [ + UsageState::NotLoggedIn, + UsageState::Unreachable { + detail: "no route".to_string(), + }, + UsageState::Failed { + detail: "429".to_string(), + }, + ] { + let broken = snapshot(state.clone(), Vec::new()); + assert_eq!( + decide(Some(&broken), &owed(at - 60.0), at), + Step::WaitUntil(at + BACKOFF), + "{state:?}" + ); + } + // And no snapshot at all -- a machine or provider edited away under a + // session that was waiting on it. + assert_eq!( + decide(None, &owed(at - 60.0), at), + Step::WaitUntil(at + BACKOFF) + ); + } + + #[test] + fn waiting_longer_than_any_real_window_gives_up_rather_than_retrying_for_ever() { + let at = 1_788_546_972.0; + let broken = snapshot( + UsageState::Unreachable { + detail: "no route".to_string(), + }, + Vec::new(), + ); + assert_eq!( + decide(Some(&broken), &owed(at - GIVE_UP - 1.0), at), + Step::GiveUp + ); + // A meter that answers is still allowed to send on the same tick: the + // ceiling bounds waiting, not resuming. + let clear = snapshot(UsageState::Ok, vec![window(3.0, None)]); + assert_eq!( + decide(Some(&clear), &owed(at - GIVE_UP - 1.0), at), + Step::Send + ); + } +} diff --git a/server/src/routes.rs b/server/src/routes.rs index 6be359c..aac707e 100644 --- a/server/src/routes.rs +++ b/server/src/routes.rs @@ -7,6 +7,7 @@ //! POST /setups add {name, ssh?} -- providers are discovered //! POST /setups/probe dry run {ssh?}: what would be found there //! GET /setups/{id} one machine, for refetching after a change +//! GET /setups/{id}/models GGUFs on that machine, for a llama session //! GET /setups/{id}/dir?path=P entries of directory P, and P resolved //! GET /setups/{id}/file?path=P content of file P, or why not //! PUT /setups/{id}/file {path, content, ifSha256} -> new size/mtime/sha256 @@ -28,6 +29,12 @@ //! GET /sessions/{id}/transcript a page of history: ?before=N (newest when absent), //! ?limit=N, ?coalesce=true to count rows not deltas, //! ?after=N to floor it at what the caller already holds +//! GET /sessions/{id}/subagents [{id, title, status, created, lastActivity}], oldest +//! first -- see SUBAGENTS.md +//! GET /sessions/{id}/subagents/{sub}/transcript exactly the transcript route above, +//! against that subagent's own transcript +//! GET /sessions/{id}/subagents/{sub}/events?after=N exactly the events route above, +//! against that subagent's own stream //! POST /sessions/{id}/message {text, attachmentIds?} //! (starts the process first if it has exited) //! POST /sessions/{id}/unqueue {messageId} -- take back one not read yet @@ -41,6 +48,8 @@ //! which starts again in the new one //! POST /sessions/{id}/model {model} //! POST /sessions/{id}/permission-mode {permissionMode} +//! POST /sessions/{id}/effort {effort} -- null for the CLI's default; +//! settled at launch, so this stops the process //! POST /sessions/{id}/command {text} -- /compact, /clear, /rename x, or the dialect's own //! (starts the process first if it has exited) //! POST /sessions/{id}/compact @@ -49,8 +58,12 @@ //! DELETE /sessions/{id} kill process, delete transcript + files //! (?deleteForeign=true removes the machine's own copy too) //! POST /sessions/{id}/notify {notify} -- announce this one or not +//! POST /sessions/{id}/auto-resume {autoResume, message?} -- carry on by itself +//! once the account's usage limit lifts //! GET /notifications SSE: every session's attention-wanting //! moments, live only (see `notifications`) +//! GET /defaults {effort} -- what a new session starts at +//! POST /defaults {effort} -- null for the CLI's own default //! GET /usage cached usage windows per provider //! GET /models downloaded GGUFs, and what is being fetched //! GET /models/search?q=Q HuggingFace repositories matching Q @@ -92,6 +105,7 @@ use tokio_stream::wrappers::{BroadcastStream, ReceiverStream}; use crate::session::driver::{SessionCommand, Unqueued}; use crate::session::pending::Operation; +use crate::session::subagent::{Subagent, SubagentInfo}; use crate::session::transcript::{CATCH_UP_LIMIT, CatchUp, SeqEvent, catch_up}; use crate::session::{LiveSession, SessionInfo, SessionManager, SpawnSpec}; @@ -109,6 +123,8 @@ pub fn router(manager: Arc) -> Router { "/setups/{id}", get(read_setup).put(update_setup).delete(delete_setup), ) + // The models on the machine a setup names, for a llama session there. + .route("/setups/{id}/models", get(setup_models)) // The filesystem of the machine a setup names. Under the setup // rather than under a session because a filesystem is a property of // a machine; a session only says where to start looking. @@ -121,6 +137,15 @@ pub fn router(manager: Arc) -> Router { .route("/sessions/{id}", get(read_session).delete(delete_session)) .route("/sessions/{id}/events", get(events)) .route("/sessions/{id}/transcript", get(transcript)) + .route("/sessions/{id}/subagents", get(list_subagents)) + .route( + "/sessions/{id}/subagents/{sub}/transcript", + get(subagent_transcript), + ) + .route( + "/sessions/{id}/subagents/{sub}/events", + get(subagent_events), + ) .route("/sessions/{id}/message", post(message)) .route("/sessions/{id}/unqueue", post(unqueue)) .route("/sessions/{id}/answer", post(answer)) @@ -131,7 +156,10 @@ pub fn router(manager: Arc) -> Router { .route("/sessions/{id}/cwd", post(set_cwd)) .route("/sessions/{id}/model", post(set_model)) .route("/sessions/{id}/permission-mode", post(set_permission_mode)) + .route("/sessions/{id}/effort", post(set_effort)) + .route("/defaults", get(defaults).post(set_defaults)) .route("/sessions/{id}/notify", post(set_notify)) + .route("/sessions/{id}/auto-resume", post(set_auto_resume)) .route("/notifications", get(notifications)) .route("/sessions/{id}/compact", post(compact)) .route("/sessions/{id}/command", post(command)) @@ -197,6 +225,18 @@ fn lookup(manager: &SessionManager, id: &str) -> Result, ApiErr .ok_or_else(|| ApiError::NotFound(format!("no session {id}"))) } +/// A session's subagent by id -- the second half of the lookup every +/// `/sessions/{id}/subagents/{sub}/...` route needs. `Arc` because reopening +/// one from disk (a subagent this process has not touched yet) inserts it +/// into the registry, and a route holding a borrow across that would be +/// holding the registry's lock the whole request. +fn lookup_subagent(session: &LiveSession, sub: &str) -> Result, ApiError> { + session + .subagents() + .get(sub) + .ok_or_else(|| ApiError::NotFound(format!("no subagent {sub}"))) +} + async fn list_sessions(State(manager): State>) -> axum::Json> { axum::Json(manager.sessions()) } @@ -291,6 +331,9 @@ struct SshRequest { /// Where attached files land on that machine; see `SshConfig`. #[serde(default)] attachments_dir: Option, + /// Where that machine keeps its GGUF models; see `SshConfig`. + #[serde(default)] + models_dir: Option, } impl SshRequest { @@ -320,6 +363,14 @@ impl SshRequest { .map(str::trim) .filter(|dir| !dir.is_empty()) .map(std::path::PathBuf::from), + // The same rule, and for the same reason: this directory is + // on the other machine, so a `~` in it is that machine's home. + models_dir: self + .models_dir + .as_deref() + .map(str::trim) + .filter(|dir| !dir.is_empty()) + .map(std::path::PathBuf::from), }) } } @@ -495,6 +546,27 @@ struct PathQuery { path: String, } +/// The models **that machine** has, which is the list a llama.cpp session +/// on it can choose from. +/// +/// Not `GET /models`, which is this backend's own downloads: those are on +/// the machine a session runs on only when they are the same machine. A +/// spawn screen offering this backend's list for a remote setup would be +/// naming files that are not there, and the session would fail at the +/// point of loading rather than at the point of choosing. +async fn setup_models( + State(manager): State>, + UrlPath(id): UrlPath, +) -> Result>, ApiError> { + let setup = setup_by_id(&manager, &id)?; + let transport = crate::session::transport::Transport::for_setup(&setup); + let dir = crate::models::dir_on(&transport, manager.models_dir()); + crate::models::on_machine(&transport, &dir) + .await + .map(axum::Json) + .map_err(from_machine) +} + /// What is in a directory, and what that directory resolved to. async fn list_dir( State(manager): State>, @@ -609,6 +681,8 @@ struct SpawnRequest { cwd: Option, #[serde(default)] permission_mode: Option, + #[serde(default)] + effort: Option, /// Whatever the chosen driver understands -- llama.cpp's context size and /// sampling. Opaque here on purpose: see `SessionConfig::params`. #[serde(default)] @@ -796,6 +870,7 @@ async fn start_import( model: body.model.clone(), cwd: None, permission_mode: body.permission_mode.clone(), + effort: body.effort.clone(), params: std::collections::BTreeMap::new(), import: Some(session.clone()), }; @@ -828,6 +903,8 @@ struct ImportRequest { model: Option, #[serde(default)] permission_mode: Option, + #[serde(default)] + effort: Option, } /// Runs `work` on the server, marked as in flight for as long as it takes. @@ -978,6 +1055,7 @@ async fn spawn(manager: &Arc, body: SpawnRequest) -> Result, +} + +async fn defaults(State(manager): State>) -> axum::Json { + axum::Json(Defaults { + effort: manager.default_effort(), + }) +} + +/// Sets what a new session's thinking level is. Applied when a session is +/// spawned, so nothing already running changes underneath anybody. +async fn set_defaults( + State(manager): State>, + axum::Json(body): axum::Json, +) -> Result { + manager + .set_default_effort(body.effort.as_deref()) + .map_err(bad_request)?; + Ok(StatusCode::NO_CONTENT) +} + +#[derive(Deserialize)] +#[serde(rename_all = "camelCase")] +#[serde(deny_unknown_fields)] +struct EffortRequest { + /// Absent or null is the CLI's own default, which is a choice somebody can + /// make rather than only a state to start in. + #[serde(default)] + effort: Option, +} + +/// Records how hard this session thinks, and stops the process so the next one +/// is launched with it -- `--effort` has no control request behind it. See +/// [`SessionManager::set_session_effort`]. +async fn set_effort( + State(manager): State>, + UrlPath(id): UrlPath, + axum::Json(body): axum::Json, +) -> Result { + manager + .set_session_effort(&id, body.effort.as_deref()) + .map_err(bad_request)?; + Ok(StatusCode::NO_CONTENT) +} + async fn set_permission_mode( State(manager): State>, UrlPath(id): UrlPath, @@ -1339,6 +1471,33 @@ async fn set_notify( Ok(StatusCode::NO_CONTENT) } +#[derive(Deserialize)] +#[serde(rename_all = "camelCase", deny_unknown_fields)] +struct AutoResumeRequest { + auto_resume: bool, + /// What to send when the limit lifts. Absent -- and empty, which is what a + /// cleared field sends -- means this app's own default word, which is a + /// choice a caller has to be able to make rather than only start in. + #[serde(default)] + message: Option, +} + +/// Turns auto-resume on or off, and sets what it would say. +/// +/// One request for both, because they are one decision: switching it on +/// without saying what to send is the ordinary case, and changing the words +/// while it is off is how somebody sets it up before it is needed. +async fn set_auto_resume( + State(manager): State>, + UrlPath(id): UrlPath, + axum::Json(body): axum::Json, +) -> Result { + manager + .set_session_auto_resume(&id, body.auto_resume, body.message.as_deref()) + .map_err(bad_request)?; + Ok(StatusCode::NO_CONTENT) +} + #[derive(Deserialize)] #[serde(deny_unknown_fields)] struct CommandRequest { @@ -1616,8 +1775,36 @@ async fn transcript( Query(query): Query, ) -> Result>, ApiError> { let session = lookup(&manager, &id)?; + transcript_page(session.transcript_path(), &id, query) +} + +/// Exactly [`transcript`]'s route and answer, against one subagent's own +/// transcript instead of its session's -- see `SUBAGENTS.md`'s wire shape. +async fn subagent_transcript( + State(manager): State>, + UrlPath((id, sub)): UrlPath<(String, String)>, + Query(query): Query, +) -> Result>, ApiError> { + let session = lookup(&manager, &id)?; + let subagent = lookup_subagent(&session, &sub)?; + transcript_page( + &subagent.transcript_path(), + &format!("{id}/subagents/{sub}"), + query, + ) +} + +/// A page of history at `path`, newest first to open with -- the one +/// implementation [`transcript`] and [`subagent_transcript`] share, since a +/// subagent's transcript is read exactly the way a session's is. `label` is +/// only for the debug line below. +fn transcript_page( + path: &Path, + label: &str, + query: TranscriptQuery, +) -> Result>, ApiError> { let events = crate::session::transcript::read_window( - session.transcript_path(), + path, query.before, query.after, query.limit, @@ -1629,7 +1816,7 @@ async fn transcript( // for events and draws rows, and the ratio between them is a property of // the conversation. `RUST_LOG=ai_server=debug`. tracing::debug!( - session = %id, + session = %label, before = ?query.before, after = ?query.after, limit = query.limit, @@ -1647,23 +1834,74 @@ async fn events( headers: HeaderMap, ) -> Result>>, ApiError> { let session = lookup(&manager, &id)?; - let cursor = headers - .get("last-event-id") - .and_then(|value| value.to_str().ok()) - .and_then(|value| value.parse().ok()) - .unwrap_or(query.after); - + let cursor = cursor_of(&headers, query.after); // Subscribe before reading the file so nothing can land in the gap // between replay and live; overlap is deduplicated by seq. let live = session.subscribe(); - let (tx, stream) = mpsc::channel(64); - tokio::spawn(stream_session( + Ok(sse_stream( session.transcript_path().to_path_buf(), cursor, live, - tx, - )); - Ok(Sse::new(ReceiverStream::new(stream).map(Ok)).keep_alive(KeepAlive::default())) + )) +} + +/// Exactly [`events`]'s route and answer, against one subagent's own stream +/// instead of its session's -- see `SUBAGENTS.md`'s wire shape. +async fn subagent_events( + State(manager): State>, + UrlPath((id, sub)): UrlPath<(String, String)>, + Query(query): Query, + headers: HeaderMap, +) -> Result>>, ApiError> { + let session = lookup(&manager, &id)?; + let subagent = lookup_subagent(&session, &sub)?; + let cursor = cursor_of(&headers, query.after); + let live = subagent.subscribe(); + Ok(sse_stream(subagent.transcript_path(), cursor, live)) +} + +/// The cursor an SSE reconnect resumes from: the native `Last-Event-ID` +/// takes precedence over the query parameter, same cursor either way. +fn cursor_of(headers: &HeaderMap, query_after: u64) -> u64 { + headers + .get("last-event-id") + .and_then(|value| value.to_str().ok()) + .and_then(|value| value.parse().ok()) + .unwrap_or(query_after) +} + +/// Spawns the backlog-then-live task and wraps it as the response, the one +/// piece [`events`] and [`subagent_events`] share. +fn sse_stream( + transcript: PathBuf, + cursor: u64, + live: broadcast::Receiver, +) -> Sse>> { + let (tx, stream) = mpsc::channel(64); + tokio::spawn(stream_session(transcript, cursor, live, tx)); + Sse::new(ReceiverStream::new(stream).map(Ok)).keep_alive(KeepAlive::default()) +} + +/// `GET /sessions/{id}/subagents`: every subagent this session has started, +/// oldest first, with a status read from its own transcript -- see +/// `SUBAGENTS.md`'s wire shape. A subagent whose last status is `Running` is +/// reported `Unknown` instead when the session itself is not running: its +/// process was the session's, and a session with none has nothing left to +/// ask. +async fn list_subagents( + State(manager): State>, + UrlPath(id): UrlPath, +) -> Result>, ApiError> { + let session = lookup(&manager, &id)?; + // Anything but `Exited` or `Unknown` has a process behind it, which is + // what decides whether a subagent still reading `Running` from its own + // transcript can be believed -- see `SUBAGENTS.md`'s wire shape. + let running = !matches!( + session.status(), + crate::session::driver::SessionStatus::Exited + | crate::session::driver::SessionStatus::Unknown + ); + Ok(axum::Json(session.subagents().list(running))) } /// Every session's attention-wanting moments, on one stream. diff --git a/server/src/session/claude.rs b/server/src/session/claude.rs index 5adfc93..ac88f6b 100644 --- a/server/src/session/claude.rs +++ b/server/src/session/claude.rs @@ -59,6 +59,7 @@ use tokio::sync::mpsc; use super::driver::{AttachmentRef, Driver, Event, EventSink, SessionStatus, Unqueued}; use super::process; +use super::subagent::Subagents; use super::transport::{Launch, Streams, Transport}; use crate::config::{ProviderConfig, SessionConfig}; use translate::{AnswerOutcome, Setting, Translator, starts_a_model_call}; @@ -213,8 +214,12 @@ impl ClaudeDriver { transport: &Transport, session_dir: &Path, sink: EventSink, + subagents: Arc, ) -> Result { - let state = Arc::new(Mutex::new(Translator::new(session_dir.to_path_buf()))); + let state = Arc::new(Mutex::new(Translator::new( + session_dir.to_path_buf(), + subagents, + ))); let queue = Arc::new(Mutex::new(Queue::default())); let reading = Arc::new(AtomicBool::new(true)); @@ -351,6 +356,12 @@ impl ClaudeDriver { if let Some(mode) = &meta.permission_mode { push("--permission-mode", mode); } + // Launch-only: see `SessionConfig::effort`. Omitted entirely when + // unset, so the CLI's own default is what an unchosen session gets + // rather than a level this app decided to call the default. + if let Some(effort) = &meta.effort { + push("--effort", effort); + } // Named at birth, so this session is the same session in the CLI's own // picker and in what other agents see. // @@ -386,7 +397,7 @@ impl ClaudeDriver { let stdout = create_log(&session_dir.join(STDOUT_LOG))?; let stderr = create_log(&session_dir.join(STDERR_LOG))?; - let program = provider.command.as_deref().unwrap_or("claude"); + let program = provider.program(); let launch = Launch::new(program, args, meta.cwd.as_deref()); let child = transport.spawn( &launch, @@ -1163,7 +1174,10 @@ mod tests { /// this" and "the transcript records that". fn events_from_lines(lines: &[&str]) -> Vec { let dir = tempfile::tempdir().expect("temp dir"); - let state = Arc::new(Mutex::new(Translator::new(dir.path().to_path_buf()))); + let state = Arc::new(Mutex::new(Translator::new( + dir.path().to_path_buf(), + Arc::new(Subagents::new(dir.path().to_path_buf())), + ))); let queue = Arc::new(Mutex::new(Queue::default())); let (sink, mut out) = mpsc::unbounded_channel::(); for line in lines { @@ -1187,7 +1201,10 @@ mod tests { interject: impl FnOnce(&Arc>), ) -> Vec { let dir = tempfile::tempdir().expect("temp dir"); - let state = Arc::new(Mutex::new(Translator::new(dir.path().to_path_buf()))); + let state = Arc::new(Mutex::new(Translator::new( + dir.path().to_path_buf(), + Arc::new(Subagents::new(dir.path().to_path_buf())), + ))); let queue = Arc::new(Mutex::new(Queue::default())); let (sink, mut out) = mpsc::unbounded_channel::(); let mut interject = Some(interject); @@ -1481,7 +1498,10 @@ mod tests { // doing. let dir = tempfile::tempdir().expect("tempdir"); let (sink, mut received) = mpsc::unbounded_channel(); - let state = Arc::new(Mutex::new(Translator::new(dir.path().to_path_buf()))); + let state = Arc::new(Mutex::new(Translator::new( + dir.path().to_path_buf(), + Arc::new(Subagents::new(dir.path().to_path_buf())), + ))); let queue = Arc::new(Mutex::new(Queue::default())); let text = r#"{"type":"stream_event","event":{"type":"content_block_delta","index":0,"delta":{"type":"text_delta","text":"working"}},"parent_tool_use_id":null}"#; @@ -1526,7 +1546,10 @@ mod tests { // session back to work. let dir = tempfile::tempdir().expect("tempdir"); let (sink, mut received) = mpsc::unbounded_channel(); - let state = Arc::new(Mutex::new(Translator::new(dir.path().to_path_buf()))); + let state = Arc::new(Mutex::new(Translator::new( + dir.path().to_path_buf(), + Arc::new(Subagents::new(dir.path().to_path_buf())), + ))); let queue = Arc::new(Mutex::new(Queue::default())); queue.lock().unwrap().close(&sink, "the session ended"); diff --git a/server/src/session/claude/translate.rs b/server/src/session/claude/translate.rs index 3baae30..72964e4 100644 --- a/server/src/session/claude/translate.rs +++ b/server/src/session/claude/translate.rs @@ -11,10 +11,12 @@ use std::collections::HashMap; use std::path::{Path, PathBuf}; +use std::sync::{Arc, Mutex}; use serde_json::{Value, json}; use super::super::driver::{Event, QuestionOption, SessionStatus, context_tokens}; +use super::super::subagent::Subagents; /// Whether this line is the CLI opening a fresh model call. /// @@ -92,10 +94,20 @@ pub(super) struct Translator { /// reports none rather than repeating the previous turn's. context: Option, session_dir: PathBuf, + /// This session's subagents, shared with every child translator below -- + /// see `SUBAGENTS.md`. One registry per session, so a subagent started + /// through this translator or any of its children lands in the same + /// place a route reads it back from. + subagents: Arc, + /// One translator per subagent id, holding *its* streaming and + /// tool-tracking state -- separate from the parent's because tool ids + /// are unique but a `stream_event`'s content-block index is not, and + /// parallel subagents interleave their deltas on one stdout. + children: HashMap>>, } impl Translator { - pub(super) fn new(session_dir: PathBuf) -> Self { + pub(super) fn new(session_dir: PathBuf, subagents: Arc) -> Self { Self { session_id: None, pending: HashMap::new(), @@ -103,6 +115,8 @@ impl Translator { interrupting: false, context: None, session_dir, + subagents, + children: HashMap::new(), } } @@ -122,13 +136,77 @@ impl Translator { pub(super) fn translate(&mut self, message: &Value) -> Vec { // Events from subagents (Task tool internals) carry a // parent_tool_use_id; the transcript shows the Task tool's own - // start/end instead of every nested step. - if message - .get("parent_tool_use_id") - .is_some_and(|id| !id.is_null()) - { - return Vec::new(); + // start/end instead of every nested step. Routed into that + // subagent's own transcript rather than dropped -- see + // `SUBAGENTS.md`. + if let Some(parent_id) = message.get("parent_tool_use_id").and_then(Value::as_str) { + return self.translate_child(parent_id, message); } + self.dispatch(message) + } + + /// A line belonging to a subagent rather than to this translator's own + /// session. Always returns nothing to the *caller*: everything it + /// produces goes into the subagent's own transcript instead. + fn translate_child(&mut self, id: &str, message: &Value) -> Vec { + match self.subagents.get(id) { + Some(subagent) if !subagent.is_open() => { + // Not stale: the Task tool runs in the background by + // default, so a finished subagent can still be sent another + // message later (SendMessage) and start working again. A + // line arriving after `finish` means exactly that, not a + // conversation that is over -- see `SUBAGENTS.md`. + self.subagents.reopen(id); + } + Some(_) => {} + None => { + // Nobody has heard of this id yet: the Task call itself + // either has not been seen or never will be. Started here + // with the best title available -- the tool name of this + // first line -- since SUBAGENTS.md's real title only + // arrives with the Task call. + self.subagents.start(id, &fallback_title(message), None); + } + } + let child = self + .children + .entry(id.to_string()) + .or_insert_with(|| { + Arc::new(Mutex::new(Translator::new( + self.session_dir.clone(), + Arc::clone(&self.subagents), + ))) + }) + .clone(); + let events = child.lock().unwrap().dispatch(message); + for event in events { + // The subagent's own vocabulary is Running/Exited/Unknown, never + // Idle -- a background Task is either working or it has ended, + // never merely "between turns" the way a session is. Dropped + // here rather than never produced, so a `result` line's own + // `Idle` (dispatch's ordinary end-of-turn event, for a subagent + // dialect that ever sends one) is caught the same way a + // `message_delta` would be. + if !matches!( + event, + Event::Status { + state: SessionStatus::Idle + } + ) { + self.subagents.record(id, event); + } + } + // What actually ends a subagent's turn: not the parent's + // `tool_result`, which for a background Task arrives at launch + // ("Async agent launched...") long before the work is done -- see + // `SUBAGENTS.md`. + if ends_a_turn(message) { + self.subagents.finish(id); + } + Vec::new() + } + + fn dispatch(&mut self, message: &Value) -> Vec { match message.get("type").and_then(Value::as_str) { Some("system") => self.translate_system(message), // The CLI's own announcement that `/clear` took effect, sent just @@ -224,12 +302,12 @@ impl Translator { .and_then(Value::as_bool) .unwrap_or(false) { - events.push(Event::Error { - message: message - .get("result") - .and_then(Value::as_str) - .unwrap_or("the turn ended with an error") - .to_string(), + let said = message.get("result").and_then(Value::as_str); + events.push(match said.and_then(usage_limit) { + Some(resets_at) => Event::LimitReached { resets_at }, + None => Event::Error { + message: said.unwrap_or("the turn ended with an error").to_string(), + }, }); } let context = self.context.take(); @@ -379,22 +457,47 @@ impl Translator { content .iter() .filter(|block| block.get("type").and_then(Value::as_str) == Some("tool_use")) - .map(|block| Event::ToolStart { - id: block + .map(|block| { + let id = block .get("id") .and_then(Value::as_str) .unwrap_or_default() - .to_string(), - tool: block + .to_string(); + let tool = block .get("name") .and_then(Value::as_str) .unwrap_or_default() - .to_string(), - input: block.get("input").cloned().unwrap_or(Value::Null), + .to_string(); + let input = block.get("input").cloned().unwrap_or(Value::Null); + // A subagent this call is about to start -- see + // `SUBAGENTS.md`'s lifecycle #1. The parent's own transcript + // still shows only the Task call itself, below. + if tool == "Task" || tool == "Agent" { + self.start_subagent_from_task(&id, &input); + } + Event::ToolStart { id, tool, input } }) .collect() } + /// Starts the subagent a Task call names, with the title and prompt + /// SUBAGENTS.md describes: the call's `description`, then + /// `()` when one is given, falling back to the tool's own + /// name when there is no description to build one from. + fn start_subagent_from_task(&self, id: &str, input: &Value) { + let description = text_field(input, "description"); + let subagent_type = text_field(input, "subagent_type"); + let prompt = input.get("prompt").and_then(Value::as_str); + let title = match (description, subagent_type) { + (Some(description), Some(subagent_type)) => { + format!("{description} ({subagent_type})") + } + (Some(description), None) => description, + (None, _) => "Task".to_string(), + }; + self.subagents.start(id, &title, prompt); + } + fn translate_control_request(&mut self, message: &Value) -> Vec { let request = &message["request"]; if request.get("subtype").and_then(Value::as_str) != Some("can_use_tool") { @@ -598,11 +701,90 @@ impl Translator { id: about.clone(), output: texts.join("\n"), }); + // Deliberately does *not* finish a subagent `about` might name: + // the Task tool runs in the background by default, so this + // `tool_result` -- "Async agent launched..." -- arrives at + // launch, long before the subagent's own work is done. What + // ends it is its own turn ending, handled in `translate_child`. } events } } +/// The title to start a subagent under when its own first line arrives +/// before (or without) its Task call ever being seen: the tool name of that +/// first line, which is the only thing known about it yet. `"subagent"` for +/// a line this cannot even find a tool name in, such as one that opens with +/// something other than a tool call. +fn fallback_title(message: &Value) -> String { + message["message"]["content"] + .as_array() + .into_iter() + .flatten() + .find(|block| block.get("type").and_then(Value::as_str) == Some("tool_use")) + .and_then(|block| block.get("name")) + .and_then(Value::as_str) + .unwrap_or("subagent") + .to_string() +} + +/// Whether this line is a subagent's *own* turn ending -- the only thing +/// that does, per `SUBAGENTS.md`: not the parent's `tool_result`, which for +/// a background Task arrives at launch rather than at completion. +/// +/// Checked on the raw line rather than on what `dispatch` returns, so this +/// never has to touch the shared `translate_stream_event`/`dispatch` code a +/// top-level session's own turn-ending also goes through -- a subagent's +/// idea of "ended" must not change when a real session's does. +/// +/// `message_delta` is the raw API's own signal, carrying the stop reason: +/// `end_turn` is genuinely done, `tool_use` means the model is about to call +/// one and there is more coming. A `result` line is the CLI's own shape for +/// a top-level turn; a subagent has not been observed to send one, but +/// SUBAGENTS.md counts it too in case a future CLI version does. +fn ends_a_turn(message: &Value) -> bool { + match message.get("type").and_then(Value::as_str) { + Some("stream_event") => { + let event = &message["event"]; + event.get("type").and_then(Value::as_str) == Some("message_delta") + && event["delta"].get("stop_reason").and_then(Value::as_str) == Some("end_turn") + } + Some("result") => true, + _ => false, + } +} + +/// Whether a failed turn failed because the account is out of quota, and when +/// the CLI said the limit lifts. +/// +/// The wording is the CLI's: a turn stopped by the limit ends with `is_error` +/// and a result of `Claude AI usage limit reached|1788546972`, the reset being +/// epoch seconds after a pipe. Matched on the sentence rather than on a code +/// because the CLI sends none, so this is deliberately loose about everything +/// but the four words. +/// +/// The two `None`s mean different things and both are real. The outer one is +/// "some other failure". The inner one is "the limit is reached and the CLI did +/// not say until when" -- which is not a reason to invent a time: `crate::resume` +/// asks the usage endpoint before sending anything, and that answer is the one +/// that decides. +/// +/// Milliseconds are accepted as well as seconds and told apart by magnitude, +/// since a wrong guess would schedule a resume tens of thousands of years out +/// and look exactly like auto-resume being broken. +fn usage_limit(result: &str) -> Option> { + if !result.to_ascii_lowercase().contains("usage limit reached") { + return None; + } + let stamp = result + .rsplit('|') + .next() + .and_then(|tail| tail.trim().parse::().ok()) + .filter(|stamp| *stamp > 0.0) + .map(|stamp| if stamp > 1e11 { stamp / 1000.0 } else { stamp }); + Some(stamp) +} + /// A string field that is there and not empty, or `None`. The CLI omits these /// rather than sending them empty, but a caller that sends `""` means the same /// thing and should not produce a description that draws as a blank line. @@ -665,10 +847,18 @@ mod tests { .collect() } + /// A fresh, empty subagent registry over the same temp dir a test's + /// translator writes into -- every test here is about the parent's own + /// events, so what a registry does with a subagent is `subagent.rs`'s + /// tests to make, not these. + fn test_subagents(dir: &tempfile::TempDir) -> Arc { + Arc::new(Subagents::new(dir.path().to_path_buf())) + } + #[test] fn captures_the_resume_token_and_the_settings_from_init() { let dir = tempfile::tempdir().expect("tempdir"); - let mut translator = Translator::new(dir.path().to_path_buf()); + let mut translator = Translator::new(dir.path().to_path_buf(), test_subagents(&dir)); let events = translate_lines( &mut translator, &[ @@ -690,7 +880,7 @@ mod tests { #[test] fn a_setting_is_reported_when_the_cli_accepts_it_and_not_before() { let dir = tempfile::tempdir().expect("tempdir"); - let mut translator = Translator::new(dir.path().to_path_buf()); + let mut translator = Translator::new(dir.path().to_path_buf(), test_subagents(&dir)); // What `set_model` does: remember, send, and say nothing yet. translator.expect_setting("req-a".to_string(), Setting::Model("sonnet".to_string())); @@ -767,7 +957,7 @@ mod tests { // The line it sends just after answering `set_permission_mode`, which is // also how a mode changed from the terminal arrives. let dir = tempfile::tempdir().expect("tempdir"); - let mut translator = Translator::new(dir.path().to_path_buf()); + let mut translator = Translator::new(dir.path().to_path_buf(), test_subagents(&dir)); let events = translate_lines( &mut translator, &[ @@ -787,7 +977,7 @@ mod tests { fn streams_text_deltas_and_skips_the_consolidated_copy() { // Real lines (trimmed) from the 2.1.237 probe. let dir = tempfile::tempdir().expect("tempdir"); - let mut translator = Translator::new(dir.path().to_path_buf()); + let mut translator = Translator::new(dir.path().to_path_buf(), test_subagents(&dir)); let events = translate_lines( &mut translator, &[ @@ -807,7 +997,7 @@ mod tests { #[test] fn tool_use_and_result_become_tool_events() { let dir = tempfile::tempdir().expect("tempdir"); - let mut translator = Translator::new(dir.path().to_path_buf()); + let mut translator = Translator::new(dir.path().to_path_buf(), test_subagents(&dir)); let events = translate_lines( &mut translator, &[ @@ -834,7 +1024,7 @@ mod tests { #[test] fn subagent_events_are_not_duplicated_into_the_transcript() { let dir = tempfile::tempdir().expect("tempdir"); - let mut translator = Translator::new(dir.path().to_path_buf()); + let mut translator = Translator::new(dir.path().to_path_buf(), test_subagents(&dir)); let events = translate_lines( &mut translator, &[ @@ -844,10 +1034,243 @@ mod tests { assert!(events.is_empty()); } + /// A child line does not just vanish from the parent -- it lands in its + /// own subagent's transcript, with that transcript's own sequence + /// numbers, starting at 1 like any other. + #[test] + fn a_child_line_lands_in_its_own_subagents_transcript() { + let dir = tempfile::tempdir().expect("tempdir"); + let subagents = test_subagents(&dir); + let mut translator = Translator::new(dir.path().to_path_buf(), Arc::clone(&subagents)); + translate_lines( + &mut translator, + &[ + r#"{"type":"assistant","message":{"content":[{"type":"tool_use","id":"toolu_c1","name":"Bash","input":{"command":"echo hi"}}]},"parent_tool_use_id":"toolu_parent"}"#, + ], + ); + let subagent = subagents.get("toolu_parent").expect("subagent started"); + let lines = crate::session::transcript::read_after(&subagent.transcript_path(), 0) + .expect("read subagent transcript"); + assert_eq!(lines[0].seq, 1); + assert_eq!( + lines[0].event, + Event::Status { + state: SessionStatus::Running + } + ); + assert!( + lines.iter().any( + |entry| matches!(&entry.event, Event::ToolStart { tool, .. } if tool == "Bash") + ) + ); + } + + /// The title and prompt shown for a subagent come from the Task call + /// that started it, not from anything guessed at its first line. + #[test] + fn the_subagent_takes_its_title_and_prompt_from_the_task_call() { + let dir = tempfile::tempdir().expect("tempdir"); + let subagents = test_subagents(&dir); + let mut translator = Translator::new(dir.path().to_path_buf(), Arc::clone(&subagents)); + translate_lines( + &mut translator, + &[ + r#"{"type":"assistant","message":{"content":[{"type":"tool_use","id":"toolu_task","name":"Task","input":{"description":"Investigate the bug","prompt":"Find why X fails","subagent_type":"general-purpose"}}]},"parent_tool_use_id":null}"#, + ], + ); + let rows = subagents.list(true); + assert_eq!(rows.len(), 1); + assert_eq!(rows[0].title, "Investigate the bug (general-purpose)"); + let subagent = subagents.get(&rows[0].id).expect("subagent"); + let lines = crate::session::transcript::read_after(&subagent.transcript_path(), 0) + .expect("read subagent transcript"); + assert!(lines.iter().any( + |entry| matches!(&entry.event, Event::UserMessage { text, .. } if text == "Find why X fails") + )); + } + + /// The parent's `tool_result` for the Task id is what ends the + /// subagent -- SUBAGENTS.md's lifecycle #3 -- and nothing else does. + #[test] + fn the_parents_tool_result_does_not_finish_the_subagent() { + // The Task tool runs in the background by default: this + // `tool_result` is "Async agent launched...", arriving the moment + // the subagent *starts*, while it goes on working for however long + // its own turn takes. Finishing it here was the bug -- a running + // background agent read as "finished" with its transcript truncated + // at launch. + let dir = tempfile::tempdir().expect("tempdir"); + let subagents = test_subagents(&dir); + let mut translator = Translator::new(dir.path().to_path_buf(), Arc::clone(&subagents)); + translate_lines( + &mut translator, + &[ + r#"{"type":"assistant","message":{"content":[{"type":"tool_use","id":"toolu_task2","name":"Task","input":{"description":"helper"}}]},"parent_tool_use_id":null}"#, + ], + ); + let subagent = subagents.get("toolu_task2").expect("subagent started"); + assert!(subagent.is_open()); + translate_lines( + &mut translator, + &[ + r#"{"type":"user","message":{"role":"user","content":[{"type":"tool_result","tool_use_id":"toolu_task2","content":"Async agent launched","is_error":false}]},"parent_tool_use_id":null}"#, + ], + ); + assert!(subagent.is_open()); + } + + /// What actually ends a subagent: the raw API's own `message_delta` + /// saying its turn stopped with `end_turn`. Never written into the + /// subagent's own transcript as `Idle` -- its vocabulary has no such + /// state. + #[test] + fn the_subagents_own_end_turn_finishes_it() { + let dir = tempfile::tempdir().expect("tempdir"); + let subagents = test_subagents(&dir); + let mut translator = Translator::new(dir.path().to_path_buf(), Arc::clone(&subagents)); + translate_lines( + &mut translator, + &[ + r#"{"type":"assistant","message":{"content":[{"type":"tool_use","id":"toolu_task3","name":"Task","input":{"description":"helper"}}]},"parent_tool_use_id":null}"#, + r#"{"type":"stream_event","event":{"type":"message_delta","delta":{"stop_reason":"end_turn"}},"parent_tool_use_id":"toolu_task3"}"#, + ], + ); + let subagent = subagents.get("toolu_task3").expect("subagent started"); + assert!(!subagent.is_open()); + let lines = crate::session::transcript::read_after(&subagent.transcript_path(), 0) + .expect("read subagent transcript"); + assert!( + !lines + .iter() + .any(|entry| matches!(&entry.event, Event::Status { state } if *state == SessionStatus::Idle)), + "a subagent's transcript must never carry Idle: {lines:?}" + ); + assert_eq!( + lines.last().unwrap().event, + Event::Status { + state: SessionStatus::Exited + } + ); + } + + /// `stop_reason: "tool_use"` is the model about to call a tool, with + /// more of the turn still coming -- not an end. + #[test] + fn a_stop_reason_of_tool_use_does_not_finish_the_subagent() { + let dir = tempfile::tempdir().expect("tempdir"); + let subagents = test_subagents(&dir); + let mut translator = Translator::new(dir.path().to_path_buf(), Arc::clone(&subagents)); + translate_lines( + &mut translator, + &[ + r#"{"type":"assistant","message":{"content":[{"type":"tool_use","id":"toolu_task4","name":"Task","input":{"description":"helper"}}]},"parent_tool_use_id":null}"#, + r#"{"type":"stream_event","event":{"type":"message_delta","delta":{"stop_reason":"tool_use"}},"parent_tool_use_id":"toolu_task4"}"#, + ], + ); + assert!( + subagents + .get("toolu_task4") + .expect("subagent started") + .is_open() + ); + } + + /// A background Task can be sent another message long after its first + /// turn ended -- a further child line for it reopens rather than being + /// dropped, and the same transcript and child translator carry on. + #[test] + fn a_line_after_finish_reopens_the_subagent_rather_than_being_dropped() { + let dir = tempfile::tempdir().expect("tempdir"); + let subagents = test_subagents(&dir); + let mut translator = Translator::new(dir.path().to_path_buf(), Arc::clone(&subagents)); + translate_lines( + &mut translator, + &[ + r#"{"type":"assistant","message":{"content":[{"type":"tool_use","id":"toolu_task5","name":"Task","input":{"description":"helper"}}]},"parent_tool_use_id":null}"#, + r#"{"type":"stream_event","event":{"type":"message_delta","delta":{"stop_reason":"end_turn"}},"parent_tool_use_id":"toolu_task5"}"#, + ], + ); + let subagent = subagents.get("toolu_task5").expect("subagent started"); + assert!(!subagent.is_open()); + + translate_lines( + &mut translator, + &[ + r#"{"type":"assistant","message":{"content":[{"type":"tool_use","id":"toolu_more","name":"Bash","input":{}}]},"parent_tool_use_id":"toolu_task5"}"#, + ], + ); + assert!(subagent.is_open()); + let lines = crate::session::transcript::read_after(&subagent.transcript_path(), 0) + .expect("read subagent transcript"); + // Running, [prompt], Exited, Running (reopened), then the new line's + // own ToolStart -- the same transcript throughout, not a new one. + assert!( + lines.iter().any( + |entry| matches!(&entry.event, Event::ToolStart { tool, .. } if tool == "Bash") + ) + ); + assert_eq!( + lines + .iter() + .filter(|entry| matches!( + &entry.event, + Event::Status { + state: SessionStatus::Running + } + )) + .count(), + 2, + "expected one Running at creation and one at the reopen: {lines:?}" + ); + } + + /// Two subagents running at once keep two separate transcripts: tool ids + /// are unique but a `stream_event`'s content-block index is not, so + /// sharing translation state between them would cross their streams. + #[test] + fn two_parallel_subagents_keep_separate_transcripts() { + let dir = tempfile::tempdir().expect("tempdir"); + let subagents = test_subagents(&dir); + let mut translator = Translator::new(dir.path().to_path_buf(), Arc::clone(&subagents)); + translate_lines( + &mut translator, + &[ + r#"{"type":"assistant","message":{"content":[{"type":"tool_use","id":"toolu_a","name":"Bash","input":{}}]},"parent_tool_use_id":"toolu_task_a"}"#, + r#"{"type":"assistant","message":{"content":[{"type":"tool_use","id":"toolu_b","name":"Read","input":{}}]},"parent_tool_use_id":"toolu_task_b"}"#, + ], + ); + let a = subagents.get("toolu_task_a").expect("subagent a"); + let b = subagents.get("toolu_task_b").expect("subagent b"); + let a_events = crate::session::transcript::read_after(&a.transcript_path(), 0) + .expect("read a's transcript"); + let b_events = crate::session::transcript::read_after(&b.transcript_path(), 0) + .expect("read b's transcript"); + assert!( + a_events.iter().any( + |entry| matches!(&entry.event, Event::ToolStart { tool, .. } if tool == "Bash") + ) + ); + assert!( + b_events.iter().any( + |entry| matches!(&entry.event, Event::ToolStart { tool, .. } if tool == "Read") + ) + ); + assert!( + !a_events.iter().any( + |entry| matches!(&entry.event, Event::ToolStart { tool, .. } if tool == "Read") + ) + ); + assert!( + !b_events.iter().any( + |entry| matches!(&entry.event, Event::ToolStart { tool, .. } if tool == "Bash") + ) + ); + } + #[test] fn a_permission_request_becomes_an_allow_deny_question() { let dir = tempfile::tempdir().expect("tempdir"); - let mut translator = Translator::new(dir.path().to_path_buf()); + let mut translator = Translator::new(dir.path().to_path_buf(), test_subagents(&dir)); let events = translate_lines( &mut translator, &[ @@ -896,7 +1319,7 @@ mod tests { #[test] fn denying_a_permission_sends_deny() { let dir = tempfile::tempdir().expect("tempdir"); - let mut translator = Translator::new(dir.path().to_path_buf()); + let mut translator = Translator::new(dir.path().to_path_buf(), test_subagents(&dir)); translate_lines( &mut translator, &[ @@ -914,7 +1337,7 @@ mod tests { // The real 2.1.237 shape, verified live: answers go back inside // updatedInput, keyed by the question text. let dir = tempfile::tempdir().expect("tempdir"); - let mut translator = Translator::new(dir.path().to_path_buf()); + let mut translator = Translator::new(dir.path().to_path_buf(), test_subagents(&dir)); let events = translate_lines( &mut translator, &[ @@ -966,7 +1389,7 @@ mod tests { // in the event: a phone that had to read this dialect's tool input to // find them would be the only place that knew how. let dir = tempfile::tempdir().expect("tempdir"); - let mut translator = Translator::new(dir.path().to_path_buf()); + let mut translator = Translator::new(dir.path().to_path_buf(), test_subagents(&dir)); let events = translate_lines( &mut translator, &[ @@ -1014,7 +1437,7 @@ mod tests { #[test] fn images_in_tool_results_are_saved_and_referenced() { let dir = tempfile::tempdir().expect("tempdir"); - let mut translator = Translator::new(dir.path().to_path_buf()); + let mut translator = Translator::new(dir.path().to_path_buf(), test_subagents(&dir)); // A 1x1 PNG, the smallest real payload worth round-tripping. let png = "iVBORw0KGgoAAAANSUhEUgAAAAEAAAABCAYAAAAfFcSJAAAADUlEQVR42mNk+M9QDwADhgGAWjR9awAAAABJRU5ErkJggg=="; let line = format!( @@ -1043,7 +1466,7 @@ mod tests { #[test] fn a_turn_result_reports_usage_and_returns_to_idle() { let dir = tempfile::tempdir().expect("tempdir"); - let mut translator = Translator::new(dir.path().to_path_buf()); + let mut translator = Translator::new(dir.path().to_path_buf(), test_subagents(&dir)); let events = translate_lines( &mut translator, &[ @@ -1078,7 +1501,7 @@ mod tests { #[test] fn a_turn_started_by_another_agent_records_who_and_what() { let dir = tempfile::tempdir().expect("tempdir"); - let mut translator = Translator::new(dir.path().to_path_buf()); + let mut translator = Translator::new(dir.path().to_path_buf(), test_subagents(&dir)); let events = translate_lines( &mut translator, &[ @@ -1112,7 +1535,7 @@ mod tests { #[test] fn an_ordinary_turn_carries_no_peer_note() { let dir = tempfile::tempdir().expect("tempdir"); - let mut translator = Translator::new(dir.path().to_path_buf()); + let mut translator = Translator::new(dir.path().to_path_buf(), test_subagents(&dir)); let events = translate_lines( &mut translator, &[ @@ -1137,7 +1560,7 @@ mod tests { #[test] fn the_context_is_what_the_last_message_held_not_the_turn_added_up() { let dir = tempfile::tempdir().expect("tempdir"); - let mut translator = Translator::new(dir.path().to_path_buf()); + let mut translator = Translator::new(dir.path().to_path_buf(), test_subagents(&dir)); let events = translate_lines( &mut translator, &[ @@ -1177,7 +1600,7 @@ mod tests { // Note the snake_case keys -- the CLI's transcript file writes the same // records in camelCase. let dir = tempfile::tempdir().expect("tempdir"); - let mut translator = Translator::new(dir.path().to_path_buf()); + let mut translator = Translator::new(dir.path().to_path_buf(), test_subagents(&dir)); let events = translate_lines( &mut translator, &[ @@ -1207,7 +1630,7 @@ mod tests { #[test] fn a_failed_compaction_says_why_and_leaves_the_turn_running() { let dir = tempfile::tempdir().expect("tempdir"); - let mut translator = Translator::new(dir.path().to_path_buf()); + let mut translator = Translator::new(dir.path().to_path_buf(), test_subagents(&dir)); let events = translate_lines( &mut translator, &[ @@ -1234,7 +1657,7 @@ mod tests { #[test] fn a_boundary_without_counts_says_so_rather_than_inventing_them() { let dir = tempfile::tempdir().expect("tempdir"); - let mut translator = Translator::new(dir.path().to_path_buf()); + let mut translator = Translator::new(dir.path().to_path_buf(), test_subagents(&dir)); let events = translate_lines( &mut translator, &[ @@ -1254,7 +1677,7 @@ mod tests { #[test] fn an_error_result_surfaces_the_message() { let dir = tempfile::tempdir().expect("tempdir"); - let mut translator = Translator::new(dir.path().to_path_buf()); + let mut translator = Translator::new(dir.path().to_path_buf(), test_subagents(&dir)); let events = translate_lines( &mut translator, &[ @@ -1275,6 +1698,55 @@ mod tests { ); } + /// Running out of quota is a state, not a failure of the work. + /// + /// The naive reading -- an error result like any other -- is what shipped + /// before this: the transcript said "Claude AI usage limit reached|…" in + /// red, which is neither readable nor actionable, and nothing above the + /// driver could tell it apart from a broken tool call. + #[test] + fn a_turn_stopped_by_the_usage_limit_says_so_and_carries_the_reset() { + let dir = tempfile::tempdir().expect("tempdir"); + let mut translator = Translator::new(dir.path().to_path_buf(), test_subagents(&dir)); + let events = translate_lines( + &mut translator, + &[ + r#"{"type":"result","subtype":"error_during_execution","is_error":true,"result":"Claude AI usage limit reached|1788546972","usage":{}}"#, + ], + ); + assert_eq!( + events[0], + Event::LimitReached { + resets_at: Some(1_788_546_972.0) + } + ); + } + + #[test] + fn a_limit_the_cli_gave_no_reset_for_is_reported_without_one() { + let dir = tempfile::tempdir().expect("tempdir"); + let mut translator = Translator::new(dir.path().to_path_buf(), test_subagents(&dir)); + let events = translate_lines( + &mut translator, + &[ + r#"{"type":"result","subtype":"error_during_execution","is_error":true,"result":"Claude AI usage limit reached","usage":{}}"#, + ], + ); + // Not a time this side invented: the meter is asked before anything is + // sent, and a made-up reset would only decide when to ask. + assert_eq!(events[0], Event::LimitReached { resets_at: None }); + } + + #[test] + fn a_reset_in_milliseconds_is_not_read_as_the_year_58000() { + assert_eq!( + usage_limit("Claude AI usage limit reached|1788546972000"), + Some(Some(1_788_546_972.0)) + ); + // And anything that is not the limit stays an ordinary failure. + assert_eq!(usage_limit("something broke"), None); + } + /// Pressing Stop is not a failure, and the CLI cannot tell you which it was. /// /// An interrupted turn arrives as exactly the same shape a broken one does, @@ -1286,7 +1758,7 @@ mod tests { #[test] fn a_turn_stopped_on_purpose_is_not_an_error() { let dir = tempfile::tempdir().expect("tempdir"); - let mut translator = Translator::new(dir.path().to_path_buf()); + let mut translator = Translator::new(dir.path().to_path_buf(), test_subagents(&dir)); let stopped_result = r#"{"type":"result","subtype":"error_during_execution","is_error":true,"result":"Interrupted by user","usage":{}}"#; translator.expect_interrupt(); @@ -1318,7 +1790,7 @@ mod tests { #[test] fn replayed_and_synthetic_user_text_is_skipped() { let dir = tempfile::tempdir().expect("tempdir"); - let mut translator = Translator::new(dir.path().to_path_buf()); + let mut translator = Translator::new(dir.path().to_path_buf(), test_subagents(&dir)); let events = translate_lines( &mut translator, &[ diff --git a/server/src/session/echo.rs b/server/src/session/echo.rs index 8822b1f..c899524 100644 --- a/server/src/session/echo.rs +++ b/server/src/session/echo.rs @@ -26,6 +26,20 @@ //! - `/error [text]` -- a failure, which is otherwise awkward to cause. //! - `/peer [text]`, `/peer-turn` -- a message from another agent, in the //! in-place and the live shapes. +//! - `/usage [what]` -- an invented rate-limit answer, or `/usage off` to take +//! it away. An echo session meters nothing, so it draws no usage bar until +//! this is set; what it exists for is the states that bar can be in, which +//! otherwise cost real quota to reach. `/usage 42`, `/usage 95 20`, +//! `/usage 42 never`, `/usage notloggedin`, `/usage unreachable`, +//! `/usage failed`. The vocabulary is `usage::Fixture`'s, where the states +//! live. +//! - `/limit [minutes]` -- a turn that stops because the account is out of +//! quota, saying the limit lifts in `minutes` (default 5, and `never` for a +//! limit with no stated reset). What it exists for is auto-resume, which is +//! otherwise reachable only by actually exhausting somebody's account: pair +//! it with `/usage 100 5` for a meter that agrees, and then `/usage 20` for +//! the moment the limit lifts. The wait itself is decided by the meter, so +//! those two commands are the whole rig. //! - `/compact` -- a compaction, start to finish. //! - `/stream N` -- one long answer in N small pieces, 50ms apart: the shape a //! real model's reply arrives in, and the one where the row a reader is @@ -34,6 +48,10 @@ //! and height the app draws, in one session, which is what a scrolling //! problem needs in order to be reproduced twice the same way. //! - `/table [columns]` -- a markdown table with cells too long for one line. +//! - `/subagent [n]` -- n subagents at once (default 1), each named +//! "helper k", its prompt recorded as its own first user message: a +//! streamed reply, one Bash call, then it finishes about three seconds +//! later, the same lifecycle a real Task call has -- see `SUBAGENTS.md`. //! //! `/slow` earns its place: a queued message, a Stop button and a spinner are //! states that only exist mid-turn, and the obvious way to get one -- ask a @@ -48,6 +66,7 @@ use std::time::Duration; use super::driver::{ AttachmentRef, Driver, Event, EventSink, QuestionOption, SessionStatus, Unqueued, }; +use super::subagent::Subagents; /// Delay between streamed deltas -- long enough that streaming is visibly /// streaming, short enough that tests waiting on a full turn stay fast. @@ -85,11 +104,19 @@ pub struct EchoDriver { /// because `/ask` puts up to four on one tool call, and the turn resumes /// when the last is answered rather than the first. pending_questions: Mutex>, + /// The invented rate-limit answer `/usage` sets, shared with the usage + /// monitor that serves it. An echo session meters nothing, so this is unset + /// until a test asks for something -- see [`crate::usage::Fixture`]. + usage: crate::usage::Fixture, /// A pretend context, so the status row has something that behaves the way /// a real one does: it grows with each turn, drops to what the compaction /// says it recovered, and a clear leaves it unmeasured. What is real is /// which way the numbers move. context: Arc, + /// This session's subagents -- see `SUBAGENTS.md`. `/subagent` is the + /// test rig for the same registry the claude driver routes real Task + /// calls into. + subagents: Arc, } impl EchoDriver { @@ -311,6 +338,110 @@ impl EchoDriver { return; } + // Answered here rather than in the turn below, because it is not a + // turn: nothing is generated, and what is being exercised is the + // *other* screens -- the bar under the header, the button beside it and + // the dialog it opens, which read the usage route, not this transcript. + if let Some(rest) = text.strip_prefix("/usage") { + if announce { + self.emit(Event::MessageTaken { + id: None, + text: text.clone(), + attachments, + }); + } + let said = self.usage.command(rest); + self.emit(Event::AssistantText { + delta: format!("{said}\n"), + }); + self.emit(Event::Status { + state: SessionStatus::Idle, + }); + return; + } + + // A turn that ends the way a real one does when the account runs out: + // the same event a real driver reports, so what acts on it -- the + // transcript row and `crate::resume` -- is exercised rather than + // imitated. The meter it should agree with is `/usage`'s fixture, + // deliberately separate: the two disagreeing is a state worth being + // able to produce, since it is what a stale reset time looks like. + if let Some(rest) = text.strip_prefix("/limit") { + if announce { + self.emit(Event::MessageTaken { + id: None, + text: text.clone(), + attachments, + }); + } + let rest = rest.trim(); + let resets_at = match rest { + "never" | "none" => None, + "" => Some(super::now() + 5.0 * 60.0), + minutes => Some(super::now() + minutes.parse::().unwrap_or(5.0) * 60.0), + }; + self.emit(Event::Status { + state: SessionStatus::Running, + }); + self.emit(Event::AssistantText { + delta: "Working on it".to_string(), + }); + self.emit(Event::LimitReached { resets_at }); + self.emit(Event::Status { + state: SessionStatus::Idle, + }); + return; + } + + // `n` subagents at once, each with its own transcript in the + // registry a real Task call routes into -- see `SUBAGENTS.md`. The + // parent's own Task calls end when their subagent does, three + // seconds later, which is long enough to see the running state on + // the phone before it finishes. + if let Some(rest) = text.strip_prefix("/subagent") { + let n = rest.trim().parse::().unwrap_or(1).clamp(1, 8); + if announce { + self.emit(Event::MessageTaken { + id: None, + text: text.clone(), + attachments, + }); + } + self.emit(Event::Status { + state: SessionStatus::Running, + }); + let sink = self.sink.clone(); + let subagents = Arc::clone(&self.subagents); + tokio::spawn(async move { + let mut helpers = Vec::new(); + for k in 1..=n { + let id = format!("echo-subagent-{k}-{}", super::random_hex()); + let title = format!("helper {k}"); + let prompt = format!( + "You are helper {k} of {n}. Say a few words, run a command, then stop." + ); + let _ = sink.send(Event::ToolStart { + id: id.clone(), + tool: "Task".to_string(), + input: serde_json::json!({ + "description": title, + "prompt": prompt, + "subagent_type": "general-purpose", + }), + }); + subagents.start(&id, &title, Some(&prompt)); + helpers.push((id, sink.clone(), Arc::clone(&subagents))); + } + for (id, sink, subagents) in helpers { + tokio::spawn(run_helper(id, sink, subagents)); + } + let _ = sink.send(Event::Status { + state: SessionStatus::Idle, + }); + }); + return; + } + // The same word the real CLI takes, so a phone drives both the same way. // `Driver::compact` is what the manager's route calls; this is the typed // path onto it. @@ -606,7 +737,12 @@ impl EchoDriver { }); } - pub fn new(sink: EventSink, session_dir: PathBuf) -> Self { + pub fn new( + sink: EventSink, + session_dir: PathBuf, + usage: crate::usage::Fixture, + subagents: Arc, + ) -> Self { let driver = Self { sink, pending_questions: Mutex::new(Vec::new()), @@ -614,6 +750,8 @@ impl EchoDriver { busy: Arc::new(AtomicBool::new(false)), queued: Arc::new(Mutex::new(Vec::new())), session_dir, + usage, + subagents, }; driver.emit(Event::Status { state: SessionStatus::Idle, @@ -735,6 +873,51 @@ async fn write_beat(sink: &EventSink, session_dir: &Path, beat: usize) { tokio::time::sleep(Duration::from_millis(120)).await; } +/// One `/subagent` helper: a few streamed words, one Bash call, then +/// `Status::Exited` about three seconds after it started -- long enough that +/// its `Running` state can be seen on the phone before it finishes. The +/// parent's own Task call for it ends at the same moment, the same way a +/// real Task's `tool_result` ends it. +async fn run_helper(id: String, sink: EventSink, subagents: Arc) { + let start = tokio::time::Instant::now(); + for word in "Working on it now.".split_inclusive(' ') { + subagents.record( + &id, + Event::AssistantText { + delta: word.to_string(), + }, + ); + tokio::time::sleep(DELTA_DELAY).await; + } + let tool_id = format!("{id}-bash"); + subagents.record( + &id, + Event::ToolStart { + id: tool_id.clone(), + tool: "Bash".to_string(), + input: serde_json::json!({ "command": "echo helper done" }), + }, + ); + tokio::time::sleep(DELTA_DELAY).await; + subagents.record( + &id, + Event::ToolEnd { + id: tool_id, + output: "helper done".to_string(), + }, + ); + let target = Duration::from_secs(3); + let elapsed = start.elapsed(); + if elapsed < target { + tokio::time::sleep(target - elapsed).await; + } + subagents.finish(&id); + let _ = sink.send(Event::ToolEnd { + id, + output: "subagent finished".to_string(), + }); +} + /// A message written during a turn and waiting for it to end: the id of the /// `MessageQueued` that announced it, what it said, and what was attached. All /// three, because all three are what the `MessageTaken` at the other end owes. diff --git a/server/src/session/llama.rs b/server/src/session/llama.rs index 0b8a18e..e70e041 100644 --- a/server/src/session/llama.rs +++ b/server/src/session/llama.rs @@ -5,9 +5,19 @@ //! //! **It is spawned but not spoken to over stdio.** The process is started //! through the same [`Transport`] as any other and then reached over HTTP on a -//! loopback port. A remote llama-server would need its port forwarded as well -//! as its command wrapped, which is not built, so a session on an ssh host is -//! refused rather than silently talking to the wrong machine. +//! loopback port. That is the second half of what a transport is -- "run this" +//! plus "reach this port" -- and it is what lets a session run on another +//! machine: [`Transport::reserve_port`] hands back a port the server binds +//! *there* and one that reaches it *here*, and the ssh connection carrying the +//! command carries the tunnel between them. The far `llama-server` binds +//! loopback only, so a model is never served to that machine's network. +//! +//! **The model file is the far machine's, not this one's.** A remote setup +//! names its own models directory (`SshConfig::models_dir`, defaulting to where +//! this backend keeps its downloads), and the file is looked for *there* -- so +//! a session naming a model that machine does not have says so, instead of +//! starting a server that will never load one. Downloading to another machine +//! is not built; the model gets there however anything else does. //! //! **The server is stateless between requests**, so the whole conversation goes //! with every one. It is rebuilt from the session's transcript rather than kept @@ -67,6 +77,7 @@ impl LlamaDriver { /// in a different currency: two servers holding the same model is twice the /// memory, and the second would bind a different port while the phone kept /// talking to the first. + #[allow(clippy::too_many_arguments)] pub fn launch( meta: &SessionConfig, provider: &ProviderConfig, @@ -75,17 +86,15 @@ impl LlamaDriver { transcript: &Path, session_dir: &Path, sink: EventSink, + // llama.cpp has no notion of a Task call, so this is accepted only + // to keep one shape across every driver's launch -- see + // `SUBAGENTS.md`'s "Server layout". + _subagents: Arc, ) -> Result { - if !matches!(transport, Transport::Here) { - bail!( - "llama.cpp sessions can only run on this machine for now: the model is served \ - over HTTP, and forwarding that port to another host isn't built yet." - ); - } let model = meta.model.as_deref().context( "a llama.cpp session needs a model -- one of the downloaded ones, by its key", )?; - let path = model_path(models_dir, model)?; + let path = model_on(transport, models_dir, model)?; // Already loaded and still running: keep talking to it. The health poll // below confirms it is really answering, so adopting a pid whose server @@ -111,14 +120,21 @@ impl LlamaDriver { )); } - let port = free_port().context("finding a port for llama-server")?; + // Where it listens on its own machine, and where that is reached + // from here -- the same number when that machine is this one. + let forward = transport + .reserve_port() + .context("finding a port for llama-server")?; let mut args: Vec = vec![ "-m".into(), - path.to_string_lossy().into_owned(), + path.clone(), + // Loopback there, whichever machine there is: what reaches it + // from outside that machine is the ssh tunnel and nothing + // else. "--host".into(), "127.0.0.1".into(), "--port".into(), - port.to_string(), + forward.there.to_string(), ]; // Settings that belong to the server because they decide how the model // is loaded; the sampling ones ride on each request instead, so changing @@ -134,8 +150,8 @@ impl LlamaDriver { } } - let program = provider.command.as_deref().unwrap_or("llama-server"); - let launch = Launch::new(program, args, meta.cwd.as_deref()); + let program = provider.program(); + let launch = Launch::new(program, args, meta.cwd.as_deref()).reaching(forward); // Its output goes to files, not pipes. Not only so the process can // outlive this server: nothing ever read those pipes, so a chatty // llama-server filled the 64 KB buffer and blocked mid-load with no sign @@ -152,8 +168,12 @@ impl LlamaDriver { .id() .context("llama-server exited before it could be recorded")?; tracing::info!( - "session {} running {program} for {model} on 127.0.0.1:{port} as pid {pid}", - meta.id + "session {} running {program} for {model} {} on 127.0.0.1:{} there, \ + reached at 127.0.0.1:{} here, as pid {pid}", + meta.id, + transport.describe(), + forward.there, + forward.here, ); // Reaped so it does not become a zombie while this server is still its // parent; the health poll and the record are what say whether the @@ -163,12 +183,18 @@ impl LlamaDriver { let _ = child.wait().await; }); - let record = process::Record::of(pid, process::Detail::Http { port }) + // The *near* port, because that is the one anything reaching this + // server has to dial -- including a later run of this backend, + // which adopts the record without knowing which machine the server + // is on. For a remote session the recorded pid is the ssh + // client's, which is the process this machine owns and which holds + // the tunnel open for exactly as long as the far server lives. + let record = process::Record::of(pid, process::Detail::Http { port: forward.here }) .context("llama-server was gone before its start time could be read")?; process::write(session_dir, &record); Ok(Self::attached( - format!("http://127.0.0.1:{port}"), + format!("http://127.0.0.1:{}", forward.here), meta, model, transcript, @@ -202,7 +228,7 @@ impl LlamaDriver { let endpoint = endpoint.clone(); let model = model.to_string(); let session_dir = session_dir.to_path_buf(); - std::thread::spawn(move || match wait_until_ready(&endpoint) { + std::thread::spawn(move || match wait_until_ready(&endpoint, &session_dir) { Ok(()) => { tracing::info!("{model} loaded and answering at {endpoint}"); let _ = sink.send(Event::Status { @@ -499,17 +525,70 @@ fn model_path(models_dir: &Path, key: &str) -> Result { Ok(path) } -/// An unused loopback port, by asking the OS for one and letting it go. Racy in -/// principle, but nothing on this machine is hunting for ports, and the -/// alternative -- parsing the port back out of the server's log -- couples us to -/// its output format for no real gain. -fn free_port() -> Result { - let listener = std::net::TcpListener::bind("127.0.0.1:0")?; - Ok(listener.local_addr()?.port()) +/// The model file's path **on the machine that will serve it**, confirmed to be +/// there. +/// +/// One function rather than a local check and hope for the other case: the same +/// question has to be asked of two filesystems. The remote answer is measured +/// for the reason the local one is -- a missing file otherwise becomes a +/// `llama-server` that starts, fails to load, and reports as a session that +/// never became ready, which reads as the machine being slow. +/// +/// One blocking round trip on a remote spawn, which is what the spawn is +/// already paying to start ssh. The alternative is a path built here from a `~` +/// this machine cannot expand. +fn model_on(transport: &Transport, models_dir: &Path, key: &str) -> Result { + let Transport::Ssh { name, .. } = transport else { + return Ok(model_path(models_dir, key)?.to_string_lossy().into_owned()); + }; + // The same directory the spawn screen listed for this machine, and one + // function for the same reason: a list from one place and a load from + // another is a model that appears and then fails. + let dir = crate::models::dir_on(transport, models_dir); + // Checked here rather than in the script: `..` in a key would walk out of + // the models directory on a machine this server can start processes on, + // and the phone is where the key comes from. + for part in key.split('/') { + if part.is_empty() || part == "." || part == ".." { + bail!("\"{key}\" is not a model key this can resolve"); + } + } + let path = format!("{}/{key}", dir.trim_end_matches('/')); + // `$HOME` on the far side, which is the only machine that knows what it is, + // and the resolved path printed back so the launch hands `llama-server` + // something absolute. "Not there" is answered rather than failed, because a + // machine that could not be asked at all has to say so in its own words -- + // it would otherwise arrive as this same sentence about a missing model. + let script = "p=$1; case $p in \"~\") p=$HOME;; \"~/\"*) p=$HOME/${p#\"~/\"};; esac; \ + [ -f \"$p\" ] && printf 'at\\t%s\\n' \"$p\" || printf 'missing\\n'" + .to_string(); + let launch = Launch::new( + "sh", + vec!["-c".to_string(), script, "sh".to_string(), path.clone()], + None, + ); + let answer = transport + .capture_blocking(&launch) + .with_context(|| format!("couldn't ask {name} where its models are"))?; + match answer.trim().split_once('\t') { + Some(("at", resolved)) => Ok(resolved.to_string()), + _ => bail!( + "{name} has no model at {path}. A llama.cpp session serves the file from the \ + machine it runs on, so the model has to be on {name} -- what this backend has \ + downloaded is somewhere else." + ), + } } /// Polls until the server says it is ready, or gives up. -fn wait_until_ready(endpoint: &str) -> Result<()> { +/// +/// Watches the process as well as the port, because the two failures need +/// different words and one of them is common: a model that will not load, +/// a port already taken on the far machine, a `llama-server` too old for +/// a flag. All of those exit within a second and none of them will ever +/// answer `/health`, so waiting out the timeout turns a server that said +/// exactly what was wrong into "gave up after 300s". +fn wait_until_ready(endpoint: &str, session_dir: &Path) -> Result<()> { let deadline = std::time::Instant::now() + READY_TIMEOUT; let url = format!("{endpoint}/health"); loop { @@ -518,13 +597,48 @@ fn wait_until_ready(endpoint: &str) -> Result<()> { { return Ok(()); } + // `None` is the session having been stopped or deleted while this + // waited, which is nobody's fault and still not worth waiting on. + match process::recorded(session_dir) { + Some((_, process::Liveness::Alive | process::Liveness::Unknown)) => {} + Some((_, process::Liveness::Dead)) | None => { + bail!("it exited before it answered.{}", log_tail(session_dir)); + } + } if std::time::Instant::now() > deadline { - bail!("gave up after {}s", READY_TIMEOUT.as_secs()); + bail!( + "gave up after {}s.{}", + READY_TIMEOUT.as_secs(), + log_tail(session_dir) + ); } std::thread::sleep(std::time::Duration::from_millis(250)); } } +/// The end of `llama-server`'s own log, for a failure message. +/// +/// Its account of what went wrong is the useful half -- "failed to load +/// model", "bind: Address already in use" -- and on a remote session it +/// is the only half, since nobody reading the phone can open a file on +/// that machine. Bounded, because this ends up in an event a phone draws. +fn log_tail(session_dir: &Path) -> String { + let Ok(text) = std::fs::read_to_string(session_dir.join(SERVER_LOG)) else { + return String::new(); + }; + let tail: Vec<&str> = text.lines().rev().take(LOG_TAIL_LINES).collect(); + if tail.is_empty() { + return String::new(); + } + format!( + " It last said: {}", + tail.into_iter().rev().collect::>().join(" / ") + ) +} + +/// How much of that log to carry into a message somebody reads on a phone. +const LOG_TAIL_LINES: usize = 6; + /// One streamed completion: posts the conversation, emits each delta as it /// arrives. Emits rather than returns, because the transcript those events land /// in is what the next turn reads back. diff --git a/server/src/session/mod.rs b/server/src/session/mod.rs index adcedd2..8d1adc5 100644 --- a/server/src/session/mod.rs +++ b/server/src/session/mod.rs @@ -15,6 +15,7 @@ pub mod import; pub mod llama; pub mod pending; pub mod process; +pub mod subagent; pub mod transcript; pub mod transport; @@ -28,7 +29,8 @@ use serde::Serialize; use tokio::sync::{broadcast, mpsc}; use crate::config::{ - Config, DriverKind, ProviderConfig, SessionConfig, SetupConfig, SshConfig, TokenEntry, + Config, DEFAULT_RESUME_MESSAGE, DriverKind, ProviderConfig, ScheduledResume, SessionConfig, + SetupConfig, SshConfig, TokenEntry, }; use claude::ClaudeDriver; use driver::{ @@ -36,6 +38,7 @@ use driver::{ }; use echo::EchoDriver; use llama::LlamaDriver; +use subagent::Subagents; use transcript::{SeqEvent, Transcript}; use transport::Transport; @@ -49,6 +52,44 @@ const EVENT_BUFFER: usize = 256; /// -- the newest "your turn" is the one still true. const NOTIFICATION_BUFFER: usize = 64; +/// Fan-out buffer for limit reports. One per session per rate-limit window, +/// so a handful a day across everything -- but sized like the notifications +/// above rather than at 1, because the only subscriber is a task that may be +/// mid-tick when several arrive. +const LIMIT_BUFFER: usize = 64; + +/// How long after a limit with no stated reset to ask the meter about it. +/// Short, because the meter is the authority and this is only how soon it is +/// worth the first question. +const FIRST_CHECK: f64 = 60.0; + +/// A session that stopped because its account is out of quota, as the pump +/// saw it. +/// +/// Broadcast downward rather than acted on here, for the reason `Shared` +/// gives: the pump runs underneath the manager and reaching back up would +/// invert that. `crate::resume` is the one subscriber, and what it does with +/// this is decided by the session's own `auto_resume`. +/// The two channels a pump reports on, which carry what this layer records +/// but does not act on: what a phone should be told, and what +/// `crate::resume` should schedule. +/// +/// One struct because they travel together through every launch and every +/// pump, and a second one arriving should not be a third parameter on both. +#[derive(Clone)] +pub struct Announcements { + notifications: broadcast::Sender, + limits: broadcast::Sender, +} + +#[derive(Debug, Clone)] +pub struct LimitHit { + pub session_id: String, + /// Epoch seconds the dialect said the limit lifts, and `None` where it + /// said nothing. Only ever a hint -- see [`Event::LimitReached`]. + pub resets_at: Option, +} + pub fn now() -> f64 { SystemTime::now() .duration_since(UNIX_EPOCH) @@ -63,6 +104,8 @@ pub struct SpawnSpec { pub model: Option, pub cwd: Option, pub permission_mode: Option, + /// See `SessionConfig::effort`. + pub effort: Option, /// Driver-interpreted settings; see `SessionConfig::params`. pub params: std::collections::BTreeMap, } @@ -94,6 +137,54 @@ pub enum NotificationKind { Finished, } +/// What a session's auto-resume setting looks like from outside: on or off, +/// what it would say, and when it next intends to check. +/// +/// One struct rather than three parameters on [`LiveSession::info`], and read +/// from the config rather than from the launch snapshot beside it, for the +/// reason `cwd` is: all three change under a running session. +#[derive(Debug, Clone)] +pub struct AutoResumeView { + pub on: bool, + pub message: String, + pub at: Option, +} + +impl AutoResumeView { + fn of(meta: &SessionConfig) -> Self { + Self { + on: meta.auto_resume, + message: resume_message(meta), + at: meta.resume.map(|scheduled| scheduled.at), + } + } +} + +/// A session with a message owed to it once its account has quota again -- +/// see [`SessionManager::owed_resumes`]. +/// +/// Carries no message: what to send is read under the lock at the moment it is +/// sent (see [`SessionManager::resume_now`]), because a wait lasts hours and +/// the words can be edited from the phone inside one. +#[derive(Debug, Clone)] +pub struct OwedResume { + pub session_id: String, + /// The machine whose account ran out, which is the one to ask. + pub setup: String, + /// Which meter reports on it -- a `crate::usage::UsageProvider::name`, the + /// same pairing `SessionInfo::usage_provider` uses. + pub provider: &'static str, + pub scheduled: ScheduledResume, +} + +/// What a session's auto-resume says, with the default filled in. One place, +/// so the phone is shown the words that would actually be sent. +fn resume_message(meta: &SessionConfig) -> String { + meta.auto_resume_message + .clone() + .unwrap_or_else(|| DEFAULT_RESUME_MESSAGE.to_string()) +} + /// One row of `GET /sessions`. #[derive(Debug, Clone, Serialize)] #[serde(rename_all = "camelCase")] @@ -118,6 +209,16 @@ pub struct SessionInfo { /// were confirming. #[serde(skip_serializing_if = "Option::is_none")] pub permission_mode: Option, + /// How hard it thinks; see `SessionConfig::effort`. Reported for the same + /// reason the mode is, and absent where nothing has been chosen -- which + /// the phone draws as the CLI's default rather than as a level. + #[serde(skip_serializing_if = "Option::is_none")] + pub effort: Option, + /// Whether a level means anything here; see `DriverKind::takes_effort`. + /// Reported beside the level because absent-and-irrelevant and + /// absent-and-unchosen are different answers, and only one of them is a + /// control worth drawing. + pub takes_effort: bool, /// Whether this session continues one the machine already had. /// Reported because it changes what deleting *means*: an imported /// session's real transcript belongs to the CLI and survives, so @@ -135,13 +236,42 @@ pub struct SessionInfo { /// that happens to be big" are different answers. #[serde(skip_serializing_if = "Option::is_none")] pub max_image_edge: Option, + /// Which of `GET /usage`'s snapshots reports on this session, and + /// absent where nothing meters it -- see + /// [`DriverKind::usage_provider`]. + /// + /// Reported for the same reason `keeps_own_transcript` is: it is a + /// fact about the provider's *kind*, and the phone has only its name. + /// Pairing by machine alone was the bug it exists to fix -- one + /// machine runs echo and the Claude CLI, so every echo session drew + /// the CLI's five-hour window as if it were its own. + #[serde(skip_serializing_if = "Option::is_none")] + pub usage_provider: Option<&'static str>, /// Whether this session announces itself -- reported for the same /// reason `permission_mode` is: a switch that guesses its own position /// is how you turn something off while believing you are reading it. pub notify: bool, + /// Whether this session sends itself a message when its account's usage + /// limit lifts, and what that message says. Reported for the same reason + /// `notify` is. + pub auto_resume: bool, + /// The words that would be sent, with the default already filled in -- + /// the phone shows what would actually happen rather than an empty field + /// meaning "something". + pub auto_resume_message: String, + /// Epoch seconds this session next intends to check whether the limit has + /// lifted, and absent when nothing is waiting. A measurement rather than + /// a promise: what decides is the meter, asked at that moment. + #[serde(skip_serializing_if = "Option::is_none")] + pub resume_at: Option, pub status: SessionStatus, pub last_activity: f64, pub created: f64, + /// How many subagents this session has started, from a directory + /// listing rather than reading each one's status -- see + /// `GET /sessions/{id}/subagents` for that. 0 when it has none, not + /// absent: every session can say this without asking anything. + pub subagents: usize, } /// What is running a session at this moment, and `None` when nothing is. @@ -170,6 +300,10 @@ pub struct LiveSession { events: broadcast::Sender, transcript_path: PathBuf, shared: Arc, + /// This session's subagents -- see `SUBAGENTS.md`. Built once at launch + /// and handed to whichever driver replaces it across a stop/start, so a + /// subagent started before a Stop is still there to read after a Start. + subagents: Arc, } /// Commands waiting for the session to be between turns. @@ -379,6 +513,19 @@ impl LiveSession { &self.transcript_path } + pub fn subagents(&self) -> &Arc { + &self.subagents + } + + /// What this session is doing right now, as the pump last recorded it -- + /// the same word `SessionInfo::status` reports. Read here rather than + /// only through `SessionManager::sessions` for + /// `GET /sessions/{id}/subagents`, which needs exactly this and nothing + /// else `SessionInfo` carries. + pub fn status(&self) -> SessionStatus { + *self.shared.status.lock().unwrap() + } + /// The session's directory (attachments in, produced files out live in /// `attachments/` and `files/` under it). pub fn dir(&self) -> &Path { @@ -424,8 +571,10 @@ impl LiveSession { &self, setup_name: &str, cwd: Option<&Path>, + effort: Option<&str>, imported: bool, kind: Option, + resume: AutoResumeView, ) -> SessionInfo { SessionInfo { id: self.meta.id.clone(), @@ -435,15 +584,25 @@ impl LiveSession { title: self.shared.title.lock().unwrap().clone(), model: self.shared.model.lock().unwrap().clone(), permission_mode: self.shared.permission_mode.lock().unwrap().clone(), + // From the config rather than from `shared`, like the cwd beside + // it: neither can change under a running process, so there is no + // live value for one to disagree with. + effort: effort.map(str::to_string), + takes_effort: kind.is_some_and(DriverKind::takes_effort), context_tokens: *self.shared.context_tokens.lock().unwrap(), notify: *self.shared.notify.lock().unwrap(), + auto_resume: resume.on, + auto_resume_message: resume.message, + resume_at: resume.at, max_image_edge: kind.and_then(DriverKind::max_image_edge), + usage_provider: kind.and_then(DriverKind::usage_provider), imported, keeps_own_transcript: kind.is_some_and(DriverKind::keeps_own_transcript), cwd: cwd.map(Path::to_path_buf), status: *self.shared.status.lock().unwrap(), last_activity: *self.shared.last_activity.lock().unwrap(), created: self.meta.created, + subagents: subagent::count(self.dir()), } } } @@ -461,8 +620,9 @@ pub struct SessionManager { /// Downloaded GGUF models, shared by every session that names one, /// which is why they sit beside the session directories. models_dir: PathBuf, - /// Where every session's pump sends what a phone should be told about. - notifications: broadcast::Sender, + /// Where every session's pump reports what this layer does not act on -- + /// see [`Announcements`]. + announce: Announcements, /// Imports and deletes running against a machine's Claude Code /// sessions: like the notifications, state the phone reads but does not /// own. @@ -470,6 +630,11 @@ pub struct SessionManager { /// What to mark sessions spawned here as -- see /// [`SessionManager::marking_new_sessions_throwaway`]. spawn_throwaway: bool, + /// The invented rate-limit answer an echo session's `/usage` sets, + /// shared with the usage monitor that serves it. Held here because + /// every echo driver this manager builds is handed a clone -- see + /// [`SessionManager::reporting_usage_fixture`]. + usage_fixture: crate::usage::Fixture, inner: RwLock, } @@ -485,6 +650,16 @@ impl SessionManager { wg_app_link::private::create_dir(&data_dir)?; let (notifications, _) = broadcast::channel(NOTIFICATION_BUFFER); + let (limits, _) = broadcast::channel(LIMIT_BUFFER); + let announce = Announcements { + notifications, + limits, + }; + // Made here rather than passed in, and handed *out* to the usage + // monitor by whoever wires the two together: every echo driver + // this manager builds gets a clone, including the ones built + // below, so it has to exist before the first session does. + let usage_fixture = crate::usage::Fixture::new(); let mut live = HashMap::new(); for meta in &config.sessions { // One unlaunchable session -- a corrupt transcript, an @@ -495,9 +670,12 @@ impl SessionManager { meta.clone(), &setup, &provider, - &data_dir, - &models_dir, - notifications.clone(), + Env { + data_dir: &data_dir, + models_dir: &models_dir, + usage: &usage_fixture, + }, + announce.clone(), // Nothing is started here; see `Launching`. Launching::Restart, ) @@ -514,14 +692,44 @@ impl SessionManager { config_path, data_dir, models_dir, - notifications, + announce, pending: Arc::new(pending::Registry::default()), spawn_throwaway: false, + usage_fixture, inner: RwLock::new(Inner { config, live }), }; Ok(manager) } + /// Where this backend's own model downloads live. The machine a + /// session runs on may keep its elsewhere -- see `models::dir_on`. + pub fn models_dir(&self) -> &Path { + &self.models_dir + } + + /// What this manager lends a session it launches. Borrowed from the + /// manager rather than cloned, so there is one answer to where things + /// are kept. + fn env(&self) -> Env<'_> { + Env { + data_dir: &self.data_dir, + models_dir: &self.models_dir, + usage: &self.usage_fixture, + } + } + + /// The invented rate-limit answer this manager's echo sessions set + /// with `/usage`, for the usage monitor to serve. + /// + /// Handed out rather than taken in because the drivers built inside + /// the constructor need it, and because the direction is the one the + /// layering allows: `usage` sits below the session layer, so a + /// session can hold one of its types while it holds nothing of a + /// session's. + pub fn usage_fixture(&self) -> crate::usage::Fixture { + self.usage_fixture.clone() + } + /// Marks every session spawned from here on as one whose process is /// stopped when this server exits. Set from `--throwaway-sessions`, /// which a debug build defaults to on. It decides only what a *new* @@ -847,8 +1055,10 @@ impl SessionManager { Some(session) => session.info( label_of(&inner.config, &meta.setup), meta.cwd.as_deref(), + meta.effort.as_deref(), import::read_cursor(&self.data_dir.join(&meta.id)).is_some(), kind_of(&inner.config, &meta.setup, &meta.provider), + AutoResumeView::of(meta), ), None => SessionInfo { id: meta.id.clone(), @@ -858,10 +1068,18 @@ impl SessionManager { title: meta.title.clone(), model: meta.model.clone(), permission_mode: meta.permission_mode.clone(), + effort: meta.effort.clone(), + takes_effort: kind_of(&inner.config, &meta.setup, &meta.provider) + .is_some_and(DriverKind::takes_effort), context_tokens: None, max_image_edge: kind_of(&inner.config, &meta.setup, &meta.provider) .and_then(DriverKind::max_image_edge), + usage_provider: kind_of(&inner.config, &meta.setup, &meta.provider) + .and_then(DriverKind::usage_provider), notify: meta.notify, + auto_resume: meta.auto_resume, + auto_resume_message: resume_message(meta), + resume_at: meta.resume.map(|scheduled| scheduled.at), imported: import::read_cursor(&self.data_dir.join(&meta.id)).is_some(), keeps_own_transcript: keeps_own_transcript( &inner.config, @@ -872,6 +1090,7 @@ impl SessionManager { status: status_of_unlaunched(&self.data_dir.join(&meta.id)), last_activity: meta.created, created: meta.created, + subagents: subagent::count(&self.data_dir.join(&meta.id)), }, }) .collect() @@ -880,8 +1099,15 @@ impl SessionManager { /// Every session's attention-wanting moments, on one stream. One /// connection for the whole backend rather than one per session: the /// phone subscribes while showing no session at all. + /// Every session running out of quota, on one stream -- the other half of + /// [`SessionManager::owed_resumes`]. Subscribed to by `crate::resume`, so + /// a limit hit is acted on when it happens rather than at the next tick. + pub fn subscribe_limits(&self) -> broadcast::Receiver { + self.announce.limits.subscribe() + } + pub fn subscribe_notifications(&self) -> broadcast::Receiver { - self.notifications.subscribe() + self.announce.notifications.subscribe() } /// Imports and deletes running against importable sessions -- see @@ -959,10 +1185,29 @@ impl SessionManager { model: spec.model, cwd: spec.cwd, permission_mode: spec.permission_mode, + // Applied here rather than on the spawn screen, so it holds + // however a session was made -- the phone, an import, or a bare + // API call -- instead of only where somebody remembered to fill it + // in. And only where the driver reads one: a llama session storing + // a level it never passes to anything is a config file that + // answers a question about itself wrongly. + effort: spec.effort.or_else(|| { + provider + .kind + .takes_effort() + .then(|| inner.config.default_effort.clone()) + .flatten() + }), params: spec.params, // On by default. Not offered at spawn: a session's first turn // is exactly the one somebody is waiting for. notify: true, + // Off, and not offered at spawn either -- for the opposite + // reason: this one spends quota with nobody watching, so it is + // asked for on a session somebody already has, never inherited. + auto_resume: false, + auto_resume_message: None, + resume: None, // Recorded on the session rather than remembered here, so // whichever server is running when the time comes knows what to // do with it -- see `SessionConfig::throwaway`. @@ -974,9 +1219,8 @@ impl SessionManager { meta.clone(), &setup, &provider, - &self.data_dir, - &self.models_dir, - self.notifications.clone(), + self.env(), + self.announce.clone(), Launching::Asked(seed), )?; let mut candidate = inner.config.clone(); @@ -994,8 +1238,10 @@ impl SessionManager { let info = session.info( &setup.name, session.meta.cwd.as_deref(), + session.meta.effort.as_deref(), import::read_cursor(&self.data_dir.join(&id)).is_some(), Some(provider.kind), + AutoResumeView::of(&session.meta), ); inner.live.insert(id, session); Ok(info) @@ -1058,6 +1304,176 @@ impl SessionManager { Ok(()) } + /// Turns auto-resume on or off for one session, and sets what it will + /// say. + /// + /// Turning it off cancels anything already scheduled, which is the path + /// out of the state the previous call put the session in: a message left + /// owed by a switch somebody has since turned off would arrive hours + /// later with nothing on screen to explain it. + /// + /// An empty message is not a message -- it is what a cleared field sends + /// -- so it means [`DEFAULT_RESUME_MESSAGE`] rather than a session poked + /// with nothing to read. + pub fn set_session_auto_resume( + &self, + id: &str, + auto_resume: bool, + message: Option<&str>, + ) -> Result<()> { + let mut inner = self.inner.write().unwrap(); + if !inner.config.sessions.iter().any(|meta| meta.id == id) { + bail!("no session {id}"); + } + let mut candidate = inner.config.clone(); + for meta in candidate.sessions.iter_mut().filter(|meta| meta.id == id) { + meta.auto_resume = auto_resume; + meta.auto_resume_message = message + .map(str::trim) + .filter(|message| !message.is_empty()) + .map(str::to_string); + if !auto_resume { + meta.resume = None; + } + } + candidate.save(&self.config_path)?; + inner.config = candidate; + Ok(()) + } + + /// Records that a session ran out of quota, and when to look again. + /// + /// Does nothing for a session that does not auto-resume, and nothing for + /// one already waiting: a turn that fails twice against the same window + /// is the same wait, and taking the second report would push the check + /// back every time the session was poked. + /// + /// `resets_at` is the dialect's hint and is used only to decide when to + /// *ask*; [`crate::resume`] asks the meter before anything is sent. A + /// session told nothing is checked shortly, since the meter is the + /// authority either way. + pub fn note_limit(&self, id: &str, resets_at: Option) -> Result { + let mut inner = self.inner.write().unwrap(); + let meta = inner + .config + .sessions + .iter() + .find(|meta| meta.id == id) + .with_context(|| format!("no session {id}"))?; + if !meta.auto_resume || meta.resume.is_some() { + return Ok(false); + } + let at = now(); + let scheduled = ScheduledResume { + at: resets_at.unwrap_or(at + FIRST_CHECK), + since: at, + }; + let mut candidate = inner.config.clone(); + for meta in candidate.sessions.iter_mut().filter(|meta| meta.id == id) { + meta.resume = Some(scheduled); + } + candidate.save(&self.config_path)?; + inner.config = candidate; + Ok(true) + } + + /// Every session with a message owed to it, oldest schedule first. + /// + /// Carries what deciding needs rather than a session id to look things up + /// by, so the scheduler holds no lock while it makes a network call: the + /// machine and the meter to ask, and the words to send. + pub fn owed_resumes(&self) -> Vec { + let inner = self.inner.read().unwrap(); + let mut owed: Vec = inner + .config + .sessions + .iter() + .filter_map(|meta| { + let scheduled = meta.resume?; + Some(OwedResume { + session_id: meta.id.clone(), + setup: meta.setup.clone(), + provider: kind_of(&inner.config, &meta.setup, &meta.provider)? + .usage_provider()?, + scheduled, + }) + }) + .collect(); + owed.sort_by(|a, b| a.scheduled.at.total_cmp(&b.scheduled.at)); + owed + } + + /// Moves a scheduled check later (or earlier), leaving everything else + /// about it alone -- including when the limit was hit, which is what + /// bounds the retrying. + pub fn reschedule_resume(&self, id: &str, at: f64) -> Result<()> { + let mut inner = self.inner.write().unwrap(); + let mut candidate = inner.config.clone(); + for meta in candidate.sessions.iter_mut().filter(|meta| meta.id == id) { + if let Some(scheduled) = meta.resume.as_mut() { + scheduled.at = at; + } + } + candidate.save(&self.config_path)?; + inner.config = candidate; + Ok(()) + } + + /// Sends the message this session is owed and clears the schedule. + /// + /// Cleared first, and saved before the message goes out: a send that + /// fails leaves nothing owed, where a schedule left standing by a failed + /// send is one that fires again on the next tick and every tick after. + /// The session is started if it has none, exactly as any other message + /// does. + pub fn resume_now(&self, id: &str) -> Result { + let message = { + let mut inner = self.inner.write().unwrap(); + let meta = inner + .config + .sessions + .iter() + .find(|meta| meta.id == id) + .with_context(|| format!("no session {id}"))?; + let message = resume_message(meta); + let mut candidate = inner.config.clone(); + for meta in candidate.sessions.iter_mut().filter(|meta| meta.id == id) { + meta.resume = None; + } + candidate.save(&self.config_path)?; + inner.config = candidate; + message + }; + self.send_message(id, message.clone(), Vec::new())?; + Ok(message) + } + + /// Gives up on a scheduled resume, and says so in the transcript. + /// + /// In the transcript because that is where somebody looking at this + /// session will be: a wait that quietly stopped waiting is + /// indistinguishable from one still going, and the session is sitting + /// there having said nothing since the limit was hit. + pub fn abandon_resume(&self, id: &str, why: &str) -> Result<()> { + { + let mut inner = self.inner.write().unwrap(); + let mut candidate = inner.config.clone(); + for meta in candidate.sessions.iter_mut().filter(|meta| meta.id == id) { + meta.resume = None; + } + candidate.save(&self.config_path)?; + inner.config = candidate; + } + if let Some(session) = self.session(id) { + let _ = session.sink.send(Event::Error { + message: format!( + "auto-resume gave up on this session: {why}. Send it something to carry on." + ), + }); + } + Ok(()) + } + /// Renames a session: persisted, shown, and passed on to whatever is /// running it. /// @@ -1191,6 +1607,68 @@ impl SessionManager { Ok(()) } + /// Sets how hard this session's model thinks. + /// + /// Shaped like [`SessionManager::set_session_cwd`] rather than like + /// [`SessionManager::set_session_model`], because `--effort` is a launch + /// flag with no control request behind it: the running process cannot be + /// asked, so the choice is recorded and the process ended, and the next + /// thing said to the session starts one that has it. Announcing it as a + /// settled change instead would put a level on the phone that the process + /// still running underneath was not using. + /// + /// `None` clears it, which is a level in its own right -- the CLI's own + /// default -- and the reason this takes an option rather than a string. + /// What a new session's thinking level is when nothing chose one, and the + /// setting of it. See `Config::default_effort`; `None` is the CLI's own. + /// + /// Only the default: a session already spawned keeps the level it was + /// given, because changing what running conversations do from a screen + /// about *new* ones is not something anybody asked for by setting a + /// default. + pub fn default_effort(&self) -> Option { + self.inner.read().unwrap().config.default_effort.clone() + } + + pub fn set_default_effort(&self, effort: Option<&str>) -> Result<()> { + let effort = effort.map(str::trim).filter(|level| !level.is_empty()); + let mut inner = self.inner.write().unwrap(); + let mut candidate = inner.config.clone(); + candidate.default_effort = effort.map(str::to_string); + candidate.save(&self.config_path)?; + inner.config = candidate; + Ok(()) + } + + pub fn set_session_effort(&self, id: &str, effort: Option<&str>) -> Result<()> { + let effort = effort.map(str::trim).filter(|level| !level.is_empty()); + { + let mut inner = self.inner.write().unwrap(); + if !inner.config.sessions.iter().any(|meta| meta.id == id) { + bail!("no session {id}"); + } + let mut candidate = inner.config.clone(); + for meta in candidate.sessions.iter_mut().filter(|meta| meta.id == id) { + meta.effort = effort.map(str::to_string); + } + candidate.save(&self.config_path)?; + inner.config = candidate; + } + // Saved before the process is touched, for the reason `set_session_cwd` + // gives: a process that will not stop must not leave the session + // recorded as something nothing agrees with. + let dir = self.data_dir.join(id); + if let Some(record) = process::live(&dir) { + tracing::info!( + "session {id} effort now {} -- stopping pid {}", + effort.unwrap_or("default"), + record.pid + ); + process::stop(&record, process::STOP_GRACE); + } + Ok(()) + } + /// Ends this session's process, leaving the session -- its transcript, /// its place in the list, everything a phone is watching -- exactly /// where it is. [`SessionManager::start_session`] is the way back. @@ -1372,10 +1850,11 @@ impl SessionManager { &meta, &setup, &provider, - &self.models_dir, + self.env(), session.dir(), session.transcript_path(), &session.sink, + session.subagents(), )?); } // Nothing is live for this one -- a session whose launch failed @@ -1385,9 +1864,8 @@ impl SessionManager { meta, &setup, &provider, - &self.data_dir, - &self.models_dir, - self.notifications.clone(), + self.env(), + self.announce.clone(), Launching::Asked(None), )?; inner.live.insert(id.to_string(), session); @@ -1760,6 +2238,19 @@ enum Launching { Restart, } +/// What the server around a session lends it: where sessions and models +/// are kept, and the usage fixture an echo session's `/usage` sets. +/// +/// One parameter rather than three because they travel together through +/// every launch path and none of them is a fact about the session -- +/// they are this server's belongings, handed down. +#[derive(Clone, Copy)] +struct Env<'a> { + data_dir: &'a Path, + models_dir: &'a Path, + usage: &'a crate::usage::Fixture, +} + /// Creates the session directory, opens its transcript (continuing the /// sequence numbering if one exists), settles what the session is doing, /// and spawns the event pump -- with a driver behind it where there is a @@ -1768,12 +2259,11 @@ fn launch( meta: SessionConfig, setup: &SetupConfig, provider: &ProviderConfig, - data_dir: &Path, - models_dir: &Path, - notifications: broadcast::Sender, + env: Env<'_>, + announce: Announcements, why: Launching, ) -> Result> { - let dir = data_dir.join(&meta.id); + let dir = env.data_dir.join(&meta.id); wg_app_link::private::create_dir(&dir)?; let transcript_path = dir.join("transcript.jsonl"); let mut transcript = Transcript::open(&transcript_path)?; @@ -1826,6 +2316,11 @@ fn launch( let (sink, source) = mpsc::unbounded_channel(); let (events, _) = broadcast::channel(EVENT_BUFFER); + // Built once per session, here, rather than per driver: a subagent + // started before a Stop has to still be there to read after a Start, + // and only `launch` runs once across that boundary -- `start_if_exited` + // replaces the driver alone. + let subagents = Arc::new(subagent::Subagents::new(dir.clone())); let shared = Arc::new(Shared { // What it was last known to be doing, not an assumption. A driver // that has something to say corrects this within its first poll. @@ -1892,10 +2387,11 @@ fn launch( &meta, setup, provider, - models_dir, + env, &dir, &transcript_path, &sink, + &subagents, ) }) .transpose()?, @@ -1914,7 +2410,8 @@ fn launch( Arc::clone(&shared), events.clone(), Arc::clone(&commands), - notifications, + announce, + Arc::clone(&subagents), )); Ok(Arc::new(LiveSession { @@ -1925,6 +2422,7 @@ fn launch( events, transcript_path, shared, + subagents, })) } @@ -1935,25 +2433,35 @@ fn launch( /// what [`SessionManager::start_session`] builds. That path replaces the /// driver and nothing else, so it has to construct one the same way rather /// than becoming a second answer to "what runs this". +#[allow(clippy::too_many_arguments)] fn make_driver( meta: &SessionConfig, setup: &SetupConfig, provider: &ProviderConfig, - models_dir: &Path, + env: Env<'_>, dir: &Path, transcript_path: &Path, sink: &EventSink, + subagents: &Arc, ) -> Result> { Ok(match provider.kind { - DriverKind::Echo => Arc::new(EchoDriver::new(sink.clone(), dir.to_path_buf())), + DriverKind::Echo => Arc::new(EchoDriver::new( + sink.clone(), + dir.to_path_buf(), + env.usage.clone(), + Arc::clone(subagents), + )), + // llama.cpp has no notion of a Task call, so it takes the registry + // and never touches it -- see `SUBAGENTS.md`'s "Server layout". DriverKind::LlamaCpp => Arc::new(LlamaDriver::launch( meta, provider, &Transport::for_setup(setup), - models_dir, + env.models_dir, transcript_path, dir, sink.clone(), + Arc::clone(subagents), )?), DriverKind::ClaudeCli => Arc::new(ClaudeDriver::launch( meta, @@ -1961,6 +2469,7 @@ fn make_driver( &Transport::for_setup(setup), dir, sink.clone(), + Arc::clone(subagents), )?), }) } @@ -2025,6 +2534,7 @@ fn notification_for( } } +#[allow(clippy::too_many_arguments)] async fn pump( id: String, mut transcript: Transcript, @@ -2032,7 +2542,8 @@ async fn pump( shared: Arc, events: broadcast::Sender, commands: Arc, - notifications: broadcast::Sender, + announce: Announcements, + subagents: Arc, ) { // Messages the session has been given and not started reading, which is // what makes a turn ending not the same thing as the work ending. @@ -2111,7 +2622,7 @@ async fn pump( { // No subscribers is the ordinary case -- nobody has // the app open -- and it is not an error. - let _ = notifications.send(Notification { + let _ = announce.notifications.send(Notification { session_id: id.clone(), title: shared.title.lock().unwrap().clone(), kind, @@ -2119,6 +2630,16 @@ async fn pump( }); } } + if let Event::LimitReached { resets_at } = &entry.event { + // Sent whether or not this session auto-resumes: whether + // to act is the manager's decision, and it is the one + // holding the setting. No subscribers is the ordinary + // case -- nothing waits on this in the tests. + let _ = announce.limits.send(LimitHit { + session_id: id.clone(), + resets_at: *resets_at, + }); + } *shared.last_activity.lock().unwrap() = ts; *shared.written.lock().unwrap() += 1; // The turn's own first line, kept for whatever arrives at the @@ -2144,7 +2665,12 @@ async fn pump( } => commands.take_one(), Event::Status { state: SessionStatus::Exited, - } => commands.abandon("this session's process has exited"), + } => { + commands.abandon("this session's process has exited"); + // The process behind every open subagent was this + // session's own -- see `SUBAGENTS.md`'s lifecycle #4. + subagents.finish_all(); + } // The two ends of a message's wait. A `UserMessage` with // no id never waited -- it was sent between turns, and // counting it would take the total below zero. @@ -2195,6 +2721,7 @@ mod tests { model: None, cwd: None, permission_mode: None, + effort: None, } } @@ -2283,6 +2810,8 @@ mod tests { driver: Arc::new(Mutex::new(Some(Arc::new(EchoDriver::new( sink.clone(), dir.path().to_path_buf(), + crate::usage::Fixture::new(), + Arc::new(subagent::Subagents::new(dir.path().to_path_buf())), ))))), sink, waiting: Mutex::new(VecDeque::new()), @@ -2470,6 +2999,99 @@ mod tests { ); } + /// The whole of the server's half of auto-resume, driven by echo: a + /// limit is reported, the session that asked for it is scheduled, and the + /// one that did not is left alone. + /// + /// Echo rather than the Claude CLI on purpose -- reaching this state for + /// real means exhausting an account, and the event both drivers report is + /// the same one. + #[tokio::test] + async fn a_limit_schedules_a_resume_only_where_one_was_asked_for() { + let dir = tempfile::tempdir().expect("tempdir"); + let config_path = dir.path().join("config.ron"); + let data_dir = dir.path().join("sessions"); + seed_echo_only(&config_path); + let manager = SessionManager::new(config_path, data_dir.clone(), data_dir.join("models")) + .expect("manager"); + let quiet = manager.spawn_session(echo_spec()).expect("spawn"); + let resuming = manager.spawn_session(echo_spec()).expect("spawn"); + manager + .set_session_auto_resume(&resuming.id, true, Some("carry on")) + .expect("on"); + + let mut limits = manager.subscribe_limits(); + for id in [&quiet.id, &resuming.id] { + manager + .session(id) + .expect("live") + .send_message("/limit 10".to_string(), Vec::new()); + } + // Both report; only one is owed anything. Drained rather than slept + // through, so the assertions below cannot run before the events they + // are about. + for _ in 0..2 { + let hit = tokio::time::timeout(Duration::from_secs(5), limits.recv()) + .await + .expect("a limit within five seconds") + .expect("channel open"); + crate::resume::note(&manager, &hit.session_id, hit.resets_at); + } + + let owed = manager.owed_resumes(); + assert_eq!( + owed.iter().map(|owed| &owed.session_id).collect::>(), + vec![&resuming.id], + "a session nobody switched on was scheduled anyway" + ); + // The dialect's hint decides when to *ask*, so it is what was written + // down -- ten minutes out, not the minute a session told nothing gets. + assert!( + owed[0].scheduled.at - now() > FIRST_CHECK, + "the reset time the session reported was ignored" + ); + + // Turning it off is the way out of the state turning it on created. + manager + .set_session_auto_resume(&resuming.id, false, None) + .expect("off"); + assert!( + manager.owed_resumes().is_empty(), + "a message stayed owed after auto-resume was switched off" + ); + } + + /// What the phone reads back, which is what its switch and its text field + /// are drawn from. + #[tokio::test] + async fn a_session_reports_its_auto_resume_setting_and_its_default_words() { + let dir = tempfile::tempdir().expect("tempdir"); + let config_path = dir.path().join("config.ron"); + let data_dir = dir.path().join("sessions"); + seed_echo_only(&config_path); + let manager = SessionManager::new(config_path, data_dir.clone(), data_dir.join("models")) + .expect("manager"); + let info = manager.spawn_session(echo_spec()).expect("spawn"); + assert!(!info.auto_resume); + // The default is reported rather than left empty: the field shows + // what would actually be sent. + assert_eq!(info.auto_resume_message, DEFAULT_RESUME_MESSAGE); + assert_eq!(info.resume_at, None); + + // An empty message is what a cleared field sends, and means the + // default rather than a session poked with nothing to read. + manager + .set_session_auto_resume(&info.id, true, Some(" ")) + .expect("on"); + let fresh = manager + .sessions() + .into_iter() + .find(|session| session.id == info.id) + .expect("listed"); + assert!(fresh.auto_resume); + assert_eq!(fresh.auto_resume_message, DEFAULT_RESUME_MESSAGE); + } + /// The switch reaches the running pump, not just the config file. The /// failure is silent in the direction that matters: a /// `set_session_notify(false)` writing only the config looks correct on @@ -2496,7 +3118,23 @@ mod tests { assert_eq!(first.session_id, info.id); // The title travels with it, because the phone may have no screen // open to look one up on. - assert_eq!(first.title, session.info("m", None, false, None).title); + assert_eq!( + first.title, + session + .info( + "m", + None, + None, + false, + None, + AutoResumeView { + on: false, + message: DEFAULT_RESUME_MESSAGE.to_string(), + at: None, + }, + ) + .title + ); manager.set_session_notify(&info.id, false).expect("off"); // Subscribed before the message, or the turn can finish in the gap @@ -3349,6 +3987,128 @@ mod tests { std::fs::write(path, rewritten).expect("write transcript"); } + /// A new session takes the stored default, and an explicit choice still + /// wins over it. + /// + /// Applied where the session is made rather than on the spawn screen, so + /// it holds for an import and a bare API call too -- a default that only + /// worked from one screen would be a default somebody had already set and + /// would reasonably believe was in force. + #[tokio::test] + async fn a_new_session_starts_at_the_stored_default_thinking_level() { + let dir = tempfile::tempdir().expect("tempdir"); + let config_path = dir.path().join("config.ron"); + let data_dir = dir.path().join("sessions"); + // Seeded with both kinds, because half of what this asks is that a + // driver which does not read a level is not given one. + let cli = seed_stand_in_cli(&config_path, dir.path()); + let manager = SessionManager::new( + config_path.clone(), + data_dir.clone(), + data_dir.join("models"), + ) + .expect("manager"); + assert_eq!( + manager.default_effort(), + None, + "nothing is set to begin with" + ); + + manager + .set_default_effort(Some("low")) + .expect("store the default"); + let took = manager.spawn_session(stand_in_spec(&cli)).expect("spawn"); + assert_eq!( + took.effort.as_deref(), + Some("low"), + "a new session takes it" + ); + + let chosen = manager + .spawn_session(SpawnSpec { + effort: Some("max".to_string()), + ..stand_in_spec(&cli) + }) + .expect("spawn"); + assert_eq!( + chosen.effort.as_deref(), + Some("max"), + "an explicit choice is not overwritten by the default" + ); + + // The case this change had no reason to touch: echo does not read a + // level, so storing one on it would be a config file describing a + // session in terms of something that never reaches it. + let echo = manager.spawn_session(echo_spec()).expect("spawn echo"); + assert_eq!( + echo.effort, None, + "a driver that does not take a level is not given the default" + ); + + // Clearing it is reachable, so the CLI's own default can be restored. + manager.set_default_effort(None).expect("clear the default"); + let cleared = manager.spawn_session(stand_in_spec(&cli)).expect("spawn"); + assert_eq!(cleared.effort, None, "and then new sessions choose nothing"); + + for id in [took.id, chosen.id, echo.id, cleared.id] { + manager.delete_session(&id).expect("delete"); + } + } + + /// A thinking level is stored and the process **ended**, because `--effort` + /// is read when the CLI launches and has no control request behind it. A + /// session left running would go on thinking at the old level underneath a + /// phone showing the new one -- the failure this app has already had once + /// with the model, and the one a stop makes impossible rather than + /// unlikely. + /// + /// Clearing it back to the CLI's own default is exercised too: that is a + /// level somebody can choose, not only one to start in, so a picker that + /// could not return to it would make leaving a level a one-way trip. + #[tokio::test] + async fn choosing_a_thinking_level_stores_it_and_ends_the_process() { + let dir = tempfile::tempdir().expect("tempdir"); + let config_path = dir.path().join("config.ron"); + let data_dir = dir.path().join("sessions"); + // The stand-in CLI rather than the echo driver: what is under test is + // that a *process* is ended, and an echo session has none to end. + let provider = seed_stand_in_cli(&config_path, dir.path()); + let manager = SessionManager::new( + config_path.clone(), + data_dir.clone(), + data_dir.join("models"), + ) + .expect("manager"); + let info = manager + .spawn_session(stand_in_spec(&provider)) + .expect("spawn"); + assert_eq!(info.effort, None, "nothing is chosen at spawn"); + let record = process::live(&data_dir.join(&info.id)).expect("the session has a process"); + + manager + .set_session_effort(&info.id, Some("low")) + .expect("store the level"); + assert_eq!( + manager.sessions()[0].effort.as_deref(), + Some("low"), + "stored, so the next start is launched with it" + ); + process::wait_gone(&[record], process::STOP_GRACE); + + // Blank is the same answer as unchosen; normalized here so a caller + // clearing the field cannot store a level the CLI would reject. + manager + .set_session_effort(&info.id, Some(" ")) + .expect("clear the level"); + assert_eq!( + manager.sessions()[0].effort, + None, + "the CLI's own default has to be reachable again" + ); + + manager.delete_session(&info.id).expect("delete"); + } + /// A setting changed on a session with nothing running is recorded as the /// session's own, rather than refused because there is no driver. The /// config already took it, so the refusal was about the driver while @@ -3699,4 +4459,69 @@ mod tests { let seen = collect_turn(&mut rx).await; assert!(seen.first().expect("events").seq > last_seq); } + + /// `/subagent 2` is the test rig for `SUBAGENTS.md`'s whole feature: + /// each helper gets its own transcript with its prompt as its first + /// user message, `SessionInfo::subagents` counts them from the + /// directory, and each finishes on its own a few seconds later. + #[tokio::test] + async fn subagent_helpers_get_their_own_transcripts_and_finish() { + let dir = tempfile::tempdir().expect("tempdir"); + let config_path = dir.path().join("config.ron"); + let data_dir = dir.path().join("sessions"); + seed_echo_only(&config_path); + let manager = SessionManager::new(config_path, data_dir.clone(), data_dir.join("models")) + .expect("manager"); + let info = manager.spawn_session(echo_spec()).expect("spawn"); + let session = manager.session(&info.id).expect("live"); + + session.send_message("/subagent 2".to_string(), Vec::new()); + // Both helpers exist as soon as their Task calls go out, well before + // either finishes. + let deadline = tokio::time::Instant::now() + Duration::from_secs(2); + loop { + if session.subagents().list(true).len() == 2 { + break; + } + assert!( + tokio::time::Instant::now() < deadline, + "both helpers should have started by now" + ); + tokio::time::sleep(Duration::from_millis(20)).await; + } + let rows = session.subagents().list(true); + let mut titles: Vec<&str> = rows.iter().map(|row| row.title.as_str()).collect(); + titles.sort_unstable(); + assert_eq!(titles, ["helper 1", "helper 2"]); + assert!(rows.iter().all(|row| row.status == SessionStatus::Running)); + assert_eq!(manager.sessions()[0].subagents, 2); + + // Each subagent's own transcript opens with its prompt. + let first = session.subagents().get(&rows[0].id).expect("subagent"); + let events = + transcript::read_after(&first.transcript_path(), 0).expect("read subagent transcript"); + assert!( + events + .iter() + .any(|entry| matches!(&entry.event, Event::UserMessage { text, .. } if text.contains("helper"))) + ); + + // Each finishes on its own about three seconds after it started. + let deadline = tokio::time::Instant::now() + Duration::from_secs(5); + loop { + if session + .subagents() + .list(true) + .iter() + .all(|row| row.status == SessionStatus::Exited) + { + break; + } + assert!( + tokio::time::Instant::now() < deadline, + "both helpers should have finished by now" + ); + tokio::time::sleep(Duration::from_millis(50)).await; + } + } } diff --git a/server/src/session/subagent.rs b/server/src/session/subagent.rs new file mode 100644 index 0000000..a6a23e5 --- /dev/null +++ b/server/src/session/subagent.rs @@ -0,0 +1,507 @@ +//! A session's subagents -- see `SUBAGENTS.md`. +//! +//! **A subagent is a second transcript owned by a session, in the same event +//! model, with no process and no controls.** It shares the transcript file +//! format, the paging routes, and the SSE stream with a session by +//! addressing, not by copying: `Transcript`, `read_window` and `catch_up` +//! work on a subagent's file unchanged. +//! +//! Storage is `/subagents//{meta.json,transcript.jsonl}`, +//! where `` is the Task tool_use id that started it -- unique, stable +//! across a backend restart, and already the key the parent side uses. Only +//! ids matching [`is_subagent_id`] are ever turned into a path. + +use std::collections::HashMap; +use std::fs; +use std::path::{Path, PathBuf}; +use std::sync::{Arc, Mutex}; + +use anyhow::{Context, Result}; +use serde::{Deserialize, Serialize}; +use tokio::sync::broadcast; + +use super::driver::{Event, SessionStatus}; +use super::transcript::{SeqEvent, Transcript}; + +/// Fan-out buffer for one subagent's SSE subscribers. Smaller than a +/// session's: a subagent's whole conversation is usually a handful of tool +/// calls, not an hours-long session. +const EVENT_BUFFER: usize = 64; + +/// Whether `id` is safe to become a path segment under a session's +/// `subagents/` directory. Mirrors `import::is_session_id`'s reasoning: the +/// id arrives as a value inside JSON the CLI sent, and it becomes a +/// directory name, so a `/` or `..` in it must never be trusted. +fn is_subagent_id(id: &str) -> bool { + !id.is_empty() + && id.len() <= 200 + && id + .bytes() + .all(|b| b.is_ascii_alphanumeric() || b == b'_' || b == b'-') +} + +/// What a subagent's directory holds beside its transcript. Small and +/// separate from `Subagent` itself because this is exactly what survives a +/// backend restart on disk, and nothing else does. +#[derive(Debug, Clone, Serialize, Deserialize)] +struct Meta { + title: String, + /// Epoch seconds. Absent from `SubagentInfo`'s sort key deliberately: + /// `list` sorts by this rather than by directory order, which a + /// filesystem does not promise. + created: f64, +} + +/// One row of `GET /sessions/{id}/subagents`. +#[derive(Debug, Clone, Serialize)] +#[serde(rename_all = "camelCase")] +pub struct SubagentInfo { + pub id: String, + pub title: String, + pub status: SessionStatus, + pub created: f64, + pub last_activity: f64, +} + +/// One subagent: its own transcript and broadcast, same shape as a +/// session's but with no driver behind it. +pub struct Subagent { + dir: PathBuf, + transcript: Mutex, + events: broadcast::Sender, + /// Mirrors the transcript's last `Status` event, kept live rather than + /// read back from `Transcript::last_status` -- that answers "as of + /// opening" (see its own doc comment) and never moves for an append made + /// through *this* object, which is every append a live subagent ever + /// makes. Without this, `finish` immediately after `start` in the same + /// process read the file's stale opening status and reported itself + /// still open. + status: Mutex, +} + +impl Subagent { + pub fn transcript_path(&self) -> PathBuf { + self.dir.join("transcript.jsonl") + } + + pub fn subscribe(&self) -> broadcast::Receiver { + self.events.subscribe() + } + + /// Whether this subagent's last recorded status is not `Exited` -- + /// what decides whether a further child line reopens it (see + /// `Subagents::reopen`) rather than continuing straight through. See + /// `SUBAGENTS.md`'s lifecycle. + pub fn is_open(&self) -> bool { + *self.status.lock().unwrap() != SessionStatus::Exited + } + + fn append(&self, event: Event) { + let mut transcript = self.transcript.lock().unwrap(); + match transcript.append(event, super::now()) { + Ok(entry) => { + if let Event::Status { state } = &entry.event { + *self.status.lock().unwrap() = *state; + } + // No subscribers is fine; the transcript already has it, + // same as a session's pump. + let _ = self.events.send(entry); + } + Err(err) => tracing::error!("subagent transcript append failed: {err:#}"), + } + } +} + +/// Every subagent one session has started, keyed by the Task tool_use id +/// that names it. +/// +/// Lives beside a session's driver rather than inside it: a claude driver +/// holds an `Arc` to this and routes child lines into it; echo uses it for +/// its `/subagent` rig; llama ignores it, since it has no notion of a Task +/// call. One instance per live session, built at launch and handed to +/// whichever driver replaces it across a stop/start. +pub struct Subagents { + /// The session's own directory; subagents live under `/subagents`. + dir: PathBuf, + live: Mutex>>, +} + +impl Subagents { + pub fn new(session_dir: PathBuf) -> Self { + Self { + dir: session_dir, + live: Mutex::new(HashMap::new()), + } + } + + fn subagents_dir(&self) -> PathBuf { + self.dir.join("subagents") + } + + /// Opens the subagent named `id`, creating it if this is the first + /// anyone has heard of it -- on disk as well as in memory, so a + /// subagent from before a backend restart is reopened rather than + /// recreated. `title`/`prompt` are used only at creation: reopening an + /// existing one keeps its original title and never repeats the prompt + /// into its transcript a second time. + fn open_or_create(&self, id: &str, title: &str, prompt: Option<&str>) -> Result> { + let dir = self.subagents_dir().join(id); + let meta_path = dir.join("meta.json"); + let existed = meta_path.is_file(); + let meta = if existed { + let text = fs::read_to_string(&meta_path) + .with_context(|| format!("read {}", meta_path.display()))?; + serde_json::from_str::(&text).context("parse subagent meta")? + } else { + wg_app_link::private::create_dir(&dir)?; + let meta = Meta { + title: title.to_string(), + created: super::now(), + }; + wg_app_link::private::write_file( + &meta_path, + serde_json::to_string(&meta) + .context("serialize subagent meta")? + .as_bytes(), + )?; + meta + }; + let mut transcript = Transcript::open(&dir.join("transcript.jsonl"))?; + if !existed { + // First lines, in order: the subagent is running the moment it + // exists, and its prompt -- when known -- is genuinely its first + // user turn. Written once, here, so a reopen never repeats them. + transcript.append( + Event::Status { + state: SessionStatus::Running, + }, + meta.created, + )?; + if let Some(prompt) = prompt { + transcript.append( + Event::UserMessage { + id: None, + text: prompt.to_string(), + attachments: Vec::new(), + }, + meta.created, + )?; + } + } + // A freshly created subagent is running by construction (its only + // lines so far are `Status::Running` and maybe its prompt); a + // reopened one takes whatever the file last said, since this + // `Transcript` has not been appended to yet in this process. + let status = if existed { + transcript.last_status().unwrap_or(SessionStatus::Running) + } else { + SessionStatus::Running + }; + let (events, _) = broadcast::channel(EVENT_BUFFER); + Ok(Arc::new(Subagent { + dir, + transcript: Mutex::new(transcript), + events, + status: Mutex::new(status), + })) + } + + /// Starts a subagent unless one is already known by this id -- see + /// `SUBAGENTS.md`'s lifecycle: created at the Task call or at the first + /// child line, whichever comes first, and never twice. A bad id is + /// refused rather than turned into a path. + pub fn start(&self, id: &str, title: &str, prompt: Option<&str>) { + if !is_subagent_id(id) { + tracing::debug!("refusing to start a subagent with a bad id {id:?}"); + return; + } + let mut live = self.live.lock().unwrap(); + if live.contains_key(id) { + return; + } + match self.open_or_create(id, title, prompt) { + Ok(subagent) => { + live.insert(id.to_string(), subagent); + } + Err(err) => tracing::error!("couldn't start subagent {id}: {err:#}"), + } + } + + /// The subagent named `id`, reopening it from disk on first use in this + /// process if one is there. `None` for an id nothing has ever started -- + /// deliberately not created here, since a route or a routing decision is + /// not the Task call that is supposed to be the only way one begins. + pub fn get(&self, id: &str) -> Option> { + if !is_subagent_id(id) { + return None; + } + if let Some(existing) = self.live.lock().unwrap().get(id).cloned() { + return Some(existing); + } + if !self.subagents_dir().join(id).join("meta.json").is_file() { + return None; + } + // Title and prompt are ignored: the directory already exists, so + // `open_or_create` reads its own meta rather than using either. + match self.open_or_create(id, "", None) { + Ok(subagent) => { + self.live + .lock() + .unwrap() + .insert(id.to_string(), Arc::clone(&subagent)); + Some(subagent) + } + Err(err) => { + tracing::error!("couldn't reopen subagent {id}: {err:#}"); + None + } + } + } + + /// Appends one event to a subagent's own transcript. A no-op, with a + /// debug log, for an id nothing was started under -- a child line for a + /// subagent this registry never opened is dropped rather than guessed + /// at. + pub fn record(&self, id: &str, event: Event) { + match self.live.lock().unwrap().get(id).cloned() { + Some(subagent) => subagent.append(event), + None => tracing::debug!("dropping an event for unknown subagent {id}"), + } + } + + /// The subagent's own turn ended: its `Status::Exited`. Called from + /// `translate_child` on the subagent's own `end_turn`, never on the + /// parent's `tool_result` -- a background Task's `tool_result` arrives + /// at launch, not at completion, so it says nothing about whether this + /// is over. A no-op for an id that is not a subagent's or is already + /// closed. + pub fn finish(&self, id: &str) { + if let Some(subagent) = self.live.lock().unwrap().get(id).cloned() + && subagent.is_open() + { + subagent.append(Event::Status { + state: SessionStatus::Exited, + }); + } + } + + /// A line arrived for a subagent that had already finished: it is + /// working again, not stale -- a background Task can be sent another + /// message long after its first turn ended. Appends `Status::Running` + /// so the list stops reporting it as finished; a no-op if it was not + /// actually closed, so a caller need not check first. + pub fn reopen(&self, id: &str) { + if let Some(subagent) = self.live.lock().unwrap().get(id).cloned() + && !subagent.is_open() + { + subagent.append(Event::Status { + state: SessionStatus::Running, + }); + } + } + + /// The parent session's process is gone, so nothing still open here has + /// a process behind it either -- see `SUBAGENTS.md`'s lifecycle #4. + pub fn finish_all(&self) { + let subagents: Vec> = self.live.lock().unwrap().values().cloned().collect(); + for subagent in subagents { + if subagent.is_open() { + subagent.append(Event::Status { + state: SessionStatus::Exited, + }); + } + } + } + + /// Every subagent under this session's directory, oldest first -- + /// `GET /sessions/{id}/subagents`. Read straight from disk rather than + /// from `live`, so a subagent from before this process started (or one + /// this run has not yet touched) still shows up; one file read per + /// subagent, which is fine at the handful a session usually has. + /// + /// `session_running` is what turns a subagent whose last status is + /// `Running` into `Unknown`: its process was the session's, and the + /// session has none. + pub fn list(&self, session_running: bool) -> Vec { + let mut rows: Vec = match fs::read_dir(self.subagents_dir()) { + Ok(entries) => entries + .filter_map(Result::ok) + .filter_map(|entry| info_of(&entry.path(), session_running)) + .collect(), + // No directory is no subagents, not a fault worth reporting. + Err(_) => Vec::new(), + }; + rows.sort_by(|a, b| { + a.created + .partial_cmp(&b.created) + .unwrap_or(std::cmp::Ordering::Equal) + }); + rows + } +} + +fn info_of(subagent_dir: &Path, session_running: bool) -> Option { + let id = subagent_dir.file_name()?.to_str()?.to_string(); + let meta_path = subagent_dir.join("meta.json"); + let text = fs::read_to_string(&meta_path).ok()?; + let meta: Meta = serde_json::from_str(&text).ok()?; + let transcript = Transcript::open(&subagent_dir.join("transcript.jsonl")).ok()?; + // The subagent's first line is always `Status::Running`, written before + // this directory is discoverable at all, so `None` here is not a state a + // reader can actually observe -- but it is not this function's place to + // invent one, so a status this build does not expect to see falls back + // to the word the lifecycle promises it started in. + let last_status = transcript.last_status().unwrap_or(SessionStatus::Running); + let status = if last_status == SessionStatus::Running && !session_running { + SessionStatus::Unknown + } else { + last_status + }; + Some(SubagentInfo { + id, + title: meta.title, + status, + created: meta.created, + last_activity: transcript.last_activity().unwrap_or(meta.created), + }) +} + +/// How many subagents a session has, for `SessionInfo::subagents`: a +/// directory listing, so the session list stays cheap and only the +/// dedicated route pays for reading a status out of each one. +pub fn count(session_dir: &Path) -> usize { + fs::read_dir(session_dir.join("subagents")) + .map(|entries| entries.filter_map(Result::ok).count()) + .unwrap_or(0) +} + +#[cfg(test)] +mod tests { + use super::*; + + #[test] + fn a_bad_id_is_refused_rather_than_turned_into_a_path() { + let dir = tempfile::tempdir().expect("tempdir"); + let subagents = Subagents::new(dir.path().to_path_buf()); + subagents.start("../../etc", "escape", None); + assert!(subagents.get("../../etc").is_none()); + assert!(!dir.path().join("subagents").exists()); + } + + #[test] + fn starting_twice_keeps_the_first_title_and_does_not_repeat_the_prompt() { + let dir = tempfile::tempdir().expect("tempdir"); + let subagents = Subagents::new(dir.path().to_path_buf()); + subagents.start("toolu_1", "first title", Some("do the thing")); + subagents.start("toolu_1", "second title", Some("do the thing")); + + let rows = subagents.list(true); + assert_eq!(rows.len(), 1); + assert_eq!(rows[0].title, "first title"); + + let events = crate::session::transcript::read_after( + &subagents.get("toolu_1").unwrap().transcript_path(), + 0, + ) + .expect("read"); + assert_eq!( + events + .iter() + .filter(|e| matches!(e.event, Event::UserMessage { .. })) + .count(), + 1 + ); + } + + #[test] + fn a_reopened_subagent_continues_its_own_transcript() { + let dir = tempfile::tempdir().expect("tempdir"); + { + let subagents = Subagents::new(dir.path().to_path_buf()); + subagents.start("toolu_2", "helper", Some("go")); + subagents.record( + "toolu_2", + Event::AssistantText { + delta: "working".to_string(), + }, + ); + } + // A fresh registry, the way a backend restart builds one. + let subagents = Subagents::new(dir.path().to_path_buf()); + let subagent = subagents.get("toolu_2").expect("reopened"); + assert!(subagent.is_open()); + subagents.record( + "toolu_2", + Event::AssistantText { + delta: " more".to_string(), + }, + ); + let events = + crate::session::transcript::read_after(&subagent.transcript_path(), 0).expect("read"); + // Status, UserMessage, two AssistantText deltas, seq continuing. + assert_eq!(events.len(), 4); + assert_eq!(events.last().unwrap().seq, 4); + } + + #[test] + fn finishing_appends_exited_and_further_lines_are_droppable() { + let dir = tempfile::tempdir().expect("tempdir"); + let subagents = Subagents::new(dir.path().to_path_buf()); + subagents.start("toolu_3", "helper", None); + subagents.finish("toolu_3"); + let subagent = subagents.get("toolu_3").unwrap(); + assert!(!subagent.is_open()); + // On disk too, not only in the live cache `is_open` reads. + assert_eq!( + Transcript::open(&subagent.transcript_path()) + .expect("reopen") + .last_status(), + Some(SessionStatus::Exited) + ); + + // Finishing an id that was never a subagent is a no-op, not a panic. + subagents.finish("never-started"); + } + + #[test] + fn finish_all_closes_only_what_is_still_open() { + let dir = tempfile::tempdir().expect("tempdir"); + let subagents = Subagents::new(dir.path().to_path_buf()); + subagents.start("toolu_4", "one", None); + subagents.start("toolu_5", "two", None); + subagents.finish("toolu_4"); + subagents.finish_all(); + + let rows = subagents.list(false); + assert_eq!(rows.len(), 2); + for row in rows { + assert_eq!(row.status, SessionStatus::Exited); + } + } + + #[test] + fn a_subagent_still_running_when_the_session_is_not_reports_unknown() { + let dir = tempfile::tempdir().expect("tempdir"); + let subagents = Subagents::new(dir.path().to_path_buf()); + subagents.start("toolu_6", "helper", None); + + assert_eq!(subagents.list(true)[0].status, SessionStatus::Running); + assert_eq!(subagents.list(false)[0].status, SessionStatus::Unknown); + } + + #[test] + fn list_is_oldest_first_and_the_count_matches_the_directory() { + let dir = tempfile::tempdir().expect("tempdir"); + let subagents = Subagents::new(dir.path().to_path_buf()); + assert_eq!(count(dir.path()), 0); + subagents.start("toolu_a", "a", None); + std::thread::sleep(std::time::Duration::from_millis(2)); + subagents.start("toolu_b", "b", None); + let rows = subagents.list(true); + assert_eq!( + rows.iter().map(|r| r.id.as_str()).collect::>(), + ["toolu_a", "toolu_b"] + ); + assert_eq!(count(dir.path()), 2); + } +} diff --git a/server/src/session/transport.rs b/server/src/session/transport.rs index 5271faf..5e20ff4 100644 --- a/server/src/session/transport.rs +++ b/server/src/session/transport.rs @@ -11,9 +11,13 @@ //! `crate::ssh`'s: this module decides *which* transport, that one knows what a //! correct ssh invocation is. //! -//! Known second operation, not built because nothing needs it yet: a managed -//! `llama-server` is spawned as a process but then spoken to over HTTP, so a -//! remote one needs a forwarded port (`ssh -L`) as well. +//! A transport is therefore two operations rather than one: **run this** and +//! **reach this port**. The second is what a managed `llama-server` needs -- it +//! is spawned as a process and then spoken to over HTTP -- and it is a no-op +//! locally, where the port a program binds is already one this machine can +//! dial. Over ssh it is an `-L` tunnel on the same connection that runs the +//! command, so the model server binds loopback on the far machine and is never +//! exposed to its network. See [`Transport::reserve_port`]. use std::path::{Path, PathBuf}; use std::process::Stdio; @@ -22,14 +26,26 @@ use anyhow::{Context, Result}; use tokio::process::Child; use crate::config::SshConfig; +pub use crate::ssh::Forward; /// What a driver needs run in order to exist as a process. Deliberately just -/// the three things every transport can carry; anything a particular machine -/// needs is the transport's own configuration, not something a driver states. +/// what every transport can carry -- the command, where it runs, and a port the +/// caller needs to reach; anything a particular machine needs is the +/// transport's own configuration, not something a driver states. pub struct Launch { pub program: String, pub args: Vec, pub cwd: Option, + /// A port this program will listen on, and the port that reaches it + /// from here -- see [`Transport::reserve_port`], which is the only + /// thing that should produce one. + /// + /// On the launch rather than in [`Transport::spawn`]'s signature + /// because it is part of what is being run: a caller that needs to + /// reach the process it is starting says so once, where it says + /// everything else about it, and every transport reads it the same + /// way. + pub forward: Option, } impl Launch { @@ -38,8 +54,16 @@ impl Launch { program: program.into(), args, cwd: cwd.map(Path::to_path_buf), + forward: None, } } + + /// Says that this program serves `forward.there`, and that the caller + /// will reach it at `forward.here`. + pub fn reaching(mut self, forward: Forward) -> Self { + self.forward = Some(forward); + self + } } /// How a launched process's standard streams are connected. @@ -102,6 +126,7 @@ impl Transport { &launch.program, &launch.args, launch.cwd.as_deref(), + launch.forward, )); match streams { Streams::Piped => { @@ -153,12 +178,15 @@ impl Transport { Self::Here => None, Self::Ssh { ssh, .. } => Some(ssh), }; - let output = - crate::ssh::command(host, &launch.program, &launch.args, launch.cwd.as_deref()) - .output() - .with_context(|| { - format!("couldn't run \"{}\" {}", launch.program, self.describe()) - })?; + let output = crate::ssh::command( + host, + &launch.program, + &launch.args, + launch.cwd.as_deref(), + launch.forward, + ) + .output() + .with_context(|| format!("couldn't run \"{}\" {}", launch.program, self.describe()))?; if !output.status.success() { let stderr = String::from_utf8_lossy(&output.stderr).trim().to_string(); anyhow::bail!(if stderr.is_empty() { @@ -217,6 +245,34 @@ impl Transport { }) } + /// Picks a port for a launched program to serve on, and the port that + /// reaches it from here. + /// + /// The "reach this port" half of what a transport is. Locally there is + /// one port and the OS chooses it, by binding and letting go -- racy + /// in principle, and nothing on this machine is hunting for ports. + /// + /// Over ssh the near end is chosen the same way and the far end is a + /// guess, because there is no portable way to ask a machine for a free + /// port that does not race with binding it anyway. It is taken from + /// [`FAR_PORTS`], below the range Linux hands out to outgoing + /// connections, so a collision means something else deliberately + /// listening there. That is not silent: the program fails to bind and + /// exits, and `session::llama` reports what its log said rather than + /// waiting out its readiness timeout. + pub fn reserve_port(&self) -> Result { + let listener = std::net::TcpListener::bind("127.0.0.1:0") + .context("asking this machine for a free port")?; + let here = listener.local_addr()?.port(); + Ok(match self { + Self::Here => Forward { there: here, here }, + Self::Ssh { .. } => Forward { + there: rand::random_range(FAR_PORTS), + here, + }, + }) + } + /// How to say where this runs, for a log line a person reads. pub fn describe(&self) -> String { match self { @@ -226,6 +282,11 @@ impl Transport { } } +/// Where a port on another machine is guessed from: high enough to be out +/// of the way of services, and below the 32768-60999 Linux hands out to +/// outgoing connections, which is where a guess would most often collide. +const FAR_PORTS: std::ops::Range = 20000..30000; + /// What a command is given on its standard input. /// /// Three cases rather than an `Option` because they are three genuinely diff --git a/server/src/setups.rs b/server/src/setups.rs index 219850b..1b2a853 100644 --- a/server/src/setups.rs +++ b/server/src/setups.rs @@ -25,7 +25,10 @@ use crate::session::transport::{Launch, Transport}; /// session stores. const PROBES: &[(&str, &str, DriverKind)] = &[ ("claude-cli", "claude", DriverKind::ClaudeCli), - ("local-llama", "llama-server", DriverKind::LlamaCpp), + // Named for the program rather than for where it runs: it runs + // wherever the setup is, and "local" was true only while a llama + // session could not be spawned on another machine. + ("llama-cpp", "llama-server", DriverKind::LlamaCpp), ]; /// Models offered for a discovered Claude CLI. A shortcut list for the spawn diff --git a/server/src/ssh.rs b/server/src/ssh.rs index 179ee80..ea4a17d 100644 --- a/server/src/ssh.rs +++ b/server/src/ssh.rs @@ -14,6 +14,22 @@ use std::process::Command; use crate::config::SshConfig; +/// A port on the machine a command runs on, and the port that reaches it from +/// the backend. +/// +/// The second half of what a transport is (PLAN.md's SSH section): "run this" +/// plus "reach this port". Locally the two numbers are one and nothing is +/// forwarded; over ssh the connection carries an `-L` tunnel, so a model server +/// binds loopback on the far machine and is never exposed to its network. +#[derive(Debug, Clone, Copy, PartialEq, Eq)] +pub struct Forward { + /// What the launched program should listen on, on its own machine. + pub there: u16, + /// What this machine connects to. The same number as `there` when the + /// program runs here. + pub here: u16, +} + /// Options forced onto every connection. `BatchMode` makes a missing key fail /// immediately with a readable message instead of hanging on a password prompt /// nothing can answer; the keepalives turn a silently dropped link into a @@ -39,6 +55,7 @@ pub fn command( program: &str, args: &[String], cwd: Option<&Path>, + forward: Option, ) -> Command { let Some(ssh) = remote else { let mut command = Command::new(program); @@ -55,9 +72,31 @@ pub fn command( }; let mut command = Command::new("ssh"); - // -T: no pty. This carries JSONL, and a pty would rewrite it (echo, CRLF - // translation, ^C handling) into something the parser can't read. - command.arg("-T"); + if let Some(forward) = forward { + // A forwarded process is not spoken to over stdio, and that changes how + // it is shut down. Everything else here is a CLI reading its stdin, so + // killing the ssh client ends it; a `llama-server` never reads its own, + // so the same kill left it running on the far machine with the model + // loaded -- measured 2026-09-04, an orphan per stopped session. A pty + // is what makes sshd hang the far side up. `-tt` because this client + // has no terminal to inherit one from. The cost is a log that arrives + // through a line discipline, which nothing parses. + command.arg("-tt"); + // Loopback at both ends: the far side binds 127.0.0.1, so what it + // serves is reachable only through this connection. + command.args([ + "-L", + &format!("127.0.0.1:{}:127.0.0.1:{}", forward.here, forward.there), + ]); + // Without this a forward that cannot be set up is a warning on stderr + // and a session that runs anyway, answering nothing -- which would + // arrive as "the model never became ready". + command.args(["-o", "ExitOnForwardFailure=yes"]); + } else { + // -T: no pty. This carries JSONL, and a pty would rewrite it (echo, + // CRLF translation, ^C handling) into something the parser can't read. + command.arg("-T"); + } for option in SSH_OPTIONS { command.args(["-o", option]); } @@ -181,6 +220,7 @@ mod tests { port: None, identity_file: None, options: vec![], + models_dir: None, attachments_dir: None, } } @@ -192,6 +232,7 @@ mod tests { "claude", &args(["-p", "--verbose"]), Some(Path::new("/tmp/x")), + None, ); assert_eq!(argv(&command), ["claude", "-p", "--verbose"]); assert_eq!(command.get_current_dir(), Some(Path::new("/tmp/x"))); @@ -204,6 +245,7 @@ mod tests { port: Some(2222), identity_file: Some("/home/me/.ssh/id_ai".into()), options: vec!["StrictHostKeyChecking=accept-new".to_string()], + models_dir: None, attachments_dir: None, }; let rendered = argv(&command( @@ -211,6 +253,7 @@ mod tests { "claude", &args(["-p", "--model", "haiku"]), Some(Path::new("/home/bob/work")), + None, )); assert_eq!(rendered[0], "ssh"); @@ -231,12 +274,59 @@ mod tests { #[test] fn a_remote_command_without_a_cwd_just_execs() { let ssh = bare_host(); - let rendered = argv(&command(Some(&ssh), "claude", &args(["-p"]), None)); + let rendered = argv(&command(Some(&ssh), "claude", &args(["-p"]), None, None)); assert_eq!(rendered.last().unwrap(), "exec 'claude' '-p'"); // No -i means no IdentitiesOnly: ~/.ssh/config decides instead. assert!(!rendered.contains(&"IdentitiesOnly=yes".to_string())); } + /// The second half of a transport: the connection that runs the command also + /// carries the port that reaches it. + /// + /// Both ends are pinned to loopback, which is what keeps a model server off + /// the far machine's network -- asserted rather than trusted, because + /// dropping the addresses is a one-word edit that still works on a machine + /// nobody else can reach. + #[test] + fn a_forwarded_port_rides_the_same_connection_as_the_command() { + let ssh = bare_host(); + let rendered = argv(&command( + Some(&ssh), + "llama-server", + &args(["--port", "24242"]), + None, + Some(Forward { + there: 24242, + here: 41000, + }), + )); + let forward = rendered + .iter() + .position(|arg| arg == "-L") + .expect("a forward"); + assert_eq!(rendered[forward + 1], "127.0.0.1:41000:127.0.0.1:24242"); + assert!(rendered.contains(&"ExitOnForwardFailure=yes".to_string())); + // The half that is easy to lose: without a pty the far process outlives + // the connection, because nothing closes a stdin it never reads. + assert!(rendered.contains(&"-tt".to_string())); + assert!(!rendered.contains(&"-T".to_string())); + // Options come before the host, or ssh reads them as part of the + // remote command. + assert!(forward < rendered.len() - 2); + assert_eq!( + rendered.last().unwrap(), + "exec 'llama-server' '--port' '24242'" + ); + + // Nothing forwarded is nothing added: every other session is one of + // these, and an -L on it would bind a port for no reason. + let plain = argv(&command(Some(&ssh), "claude", &args(["-p"]), None, None)); + assert!(!plain.contains(&"-L".to_string())); + // And a session that *is* spoken to over stdio keeps its raw pipe. + assert!(plain.contains(&"-T".to_string())); + assert!(!plain.contains(&"-tt".to_string())); + } + /// The one character quoting must not swallow. A working directory typed as /// `~/repos/ai-app` was arriving as the literal directory `~`, and the /// remote shell reported it missing -- which reads as the path being wrong @@ -277,7 +367,13 @@ mod tests { assert_eq!(expand_home(Path::new("/tmp/~/x")), Path::new("/tmp/~/x")); assert_eq!(expand_home(Path::new("~user/x")), Path::new("~user/x")); - let local = command(None, "claude", &args(["-p"]), Some(Path::new("~/work"))); + let local = command( + None, + "claude", + &args(["-p"]), + Some(Path::new("~/work")), + None, + ); assert_eq!(local.get_current_dir(), Some(home.join("work").as_path())); } @@ -301,7 +397,7 @@ mod tests { // tries to close the quote and start a new command. let ssh = bare_host(); let evil = Path::new("/tmp/'; touch /tmp/pwned; '"); - let rendered = argv(&command(Some(&ssh), "claude", &[], Some(evil))); + let rendered = argv(&command(Some(&ssh), "claude", &[], Some(evil), None)); let script = rendered.last().unwrap(); assert_eq!( script, diff --git a/server/src/usage.rs b/server/src/usage.rs index 4bcc94a..1f14a9e 100644 --- a/server/src/usage.rs +++ b/server/src/usage.rs @@ -30,13 +30,13 @@ //! already has over that machine. use std::collections::HashMap; -use std::sync::Mutex; +use std::sync::{Arc, Mutex}; use std::time::{Duration, Instant}; use serde::Serialize; use serde_json::Value; -use crate::config::{DriverKind, SetupConfig}; +use crate::config::SetupConfig; use crate::session::transport::{Launch, Transport}; const USAGE_URL: &str = "https://api.anthropic.com/api/oauth/usage"; @@ -107,10 +107,30 @@ pub struct UsageSnapshot { pub fetched_at: f64, } +/// The name of each meter, said in one place because two lists have to +/// agree on it: [`UsageSnapshot::provider`], which is what `GET /usage` +/// labels a row with, and [`crate::config::DriverKind::usage_provider`], +/// which is how a session says which of those rows is about it. +pub const CLAUDE: &str = "claude"; +/// The invented one, for testing the screens that draw these -- see +/// [`Fixture`]. +pub const ECHO: &str = "echo"; + pub trait UsageProvider: Send + Sync { fn name(&self) -> &'static str; /// Blocking -- call off the async workers. fn fetch(&self) -> UsageSnapshot; + /// How long an answer from this one may be reused. + /// + /// A property of the provider rather than of the cache, because what + /// sets it is what asking costs: [`ClaudeUsage`] makes a network call + /// against an endpoint that rate-limits impatient callers, and the + /// fixture below reads a mutex. Caching the fixture for three minutes + /// would mean a test setting a number and watching the old one for + /// most of that, which reads exactly like the command not working. + fn poll_interval(&self) -> Duration { + MIN_POLL_INTERVAL + } } /// Reads the numbers behind Claude Code's `/usage` from one machine, using the @@ -121,6 +141,10 @@ pub struct ClaudeUsage { pub setup_name: String, /// How to reach that machine. `Here` for the backend's own. pub transport: Transport, + /// The CLI to run there, for the one thing this asks of it: refreshing its + /// own expired token. The provider's, so a machine with the CLI somewhere + /// odd is asked at the same path its sessions run. + pub program: String, } /// Where Claude Code keeps its credentials, as a shell word rather than a path: @@ -170,7 +194,7 @@ impl ClaudeUsage { impl UsageProvider for ClaudeUsage { fn name(&self) -> &'static str { - "claude" + CLAUDE } fn fetch(&self) -> UsageSnapshot { @@ -178,36 +202,111 @@ impl UsageProvider for ClaudeUsage { Ok(token) => token, Err(state) => return self.snapshot(state, Vec::new()), }; - let text = match ureq::get(USAGE_URL) + let body = match self.call(&token) { + Ok(body) => body, + Err(Refused::Other(detail)) => { + return self.snapshot(UsageState::Failed { detail }, Vec::new()); + } + Err(Refused::Unauthorized) => match self.after_cli_refresh(&token) { + Ok(body) => body, + Err(state) => return self.snapshot(state, Vec::new()), + }, + }; + self.snapshot(UsageState::Ok, parse_windows(&body)) + } +} + +/// Why one call to the usage endpoint did not produce numbers. +/// +/// 401 is apart from the rest because it is the only one with a way out: the +/// endpoint answered, and it means the access token has expired rather than +/// that anything is broken. +enum Refused { + Unauthorized, + Other(String), +} + +impl ClaudeUsage { + /// One call to the endpoint with one token. + fn call(&self, token: &str) -> Result { + let text = ureq::get(USAGE_URL) .header("Authorization", &format!("Bearer {token}")) .header("anthropic-beta", "oauth-2025-04-20") .header("User-Agent", USER_AGENT) .call() .and_then(|mut response| response.body_mut().read_to_string()) - { - Ok(text) => text, - Err(err) => { - // The error string can embed the URL but never the token. - return self.snapshot( - UsageState::Failed { - detail: format!("usage endpoint unreachable: {err}"), - }, - Vec::new(), - ); - } - }; - let body: Value = match serde_json::from_str(&text) { - Ok(body) => body, - Err(err) => { - return self.snapshot( - UsageState::Failed { - detail: format!("usage endpoint sent non-JSON: {err}"), - }, - Vec::new(), - ); - } - }; - self.snapshot(UsageState::Ok, parse_windows(&body)) + // The error string can embed the URL but never the token. + .map_err(|err| match err { + ureq::Error::StatusCode(401) => Refused::Unauthorized, + other => Refused::Other(why(&other)), + })?; + serde_json::from_str(&text) + .map_err(|err| Refused::Other(format!("usage endpoint sent non-JSON: {err}"))) + } + + /// Have the machine's own CLI refresh its token, then ask once more. + /// + /// **The CLI does the refresh, never this.** Anthropic's OAuth rotates the + /// refresh token, so whoever refreshes second presents a dead one and the + /// machine is logged out until somebody runs `/login` on it -- and the + /// machine we would be refreshing on is usually one with a live session of + /// its own. Running the CLI keeps it the only writer of + /// `.credentials.json`. + /// + /// `doctor` rather than the `auth status` it reads like, measured against + /// this CLI (2.1.258) on 2026-09-05 with a deliberately invalid token: + /// `auth status` reports `loggedIn: true` off the file alone and never + /// touches the network, so it would have refreshed nothing while looking + /// like it had. `doctor` resolves the account, which is what makes it + /// refresh, and it spends no quota. The same probe showed what a *failed* + /// refresh does -- the CLI blanks both tokens -- so this must stay on the + /// 401 path, where the access token is already dead, and never be used to + /// refresh speculatively. + /// + /// Only a token that actually changed is retried, so a CLI that refreshed + /// nothing costs one call rather than two, and this cannot become a loop. + fn after_cli_refresh(&self, stale: &str) -> Result { + let launch = Launch::new(&self.program, vec!["doctor".to_string()], None); + if let Err(err) = self.transport.capture_blocking(&launch) { + return Err(UsageState::Failed { + detail: format!( + "the Claude login on {} has expired, and `{} doctor` couldn't be run there to refresh it: {err:#}", + self.setup_name, self.program + ), + }); + } + let fresh = self.access_token()?; + if fresh == stale { + return Err(self.still_expired()); + } + self.call(&fresh).map_err(|err| match err { + Refused::Unauthorized => self.still_expired(), + Refused::Other(detail) => UsageState::Failed { detail }, + }) + } + + /// A login the CLI could not renew: the one state here somebody has to act + /// on, so it says where and what to run. + fn still_expired(&self) -> UsageState { + UsageState::Failed { + detail: format!( + "the Claude login on {} has expired and could not be refreshed; run `{} /login` there", + self.setup_name, self.program + ), + } + } +} + +/// What a failed call to the usage endpoint should say. +/// +/// A status is not a network fault and must not be reported as one: the +/// endpoint answered. 401 never reaches here -- it has its own way out in +/// [`ClaudeUsage::after_cli_refresh`] -- so what is left is a refusal nobody +/// on this side can fix. +fn why(err: &ureq::Error) -> String { + match err { + ureq::Error::StatusCode(code) => format!("usage endpoint refused the request: HTTP {code}"), + other => format!("usage endpoint unreachable: {other}"), } } @@ -278,24 +377,260 @@ fn parse_windows(body: &Value) -> Vec { .collect() } +/// An invented answer, so the screens that draw these can be exercised +/// without an account. +/// +/// Every state the usage bar and the usage dialog can be in is otherwise +/// reachable only by spending somebody's quota or by breaking a machine: +/// a number near the top, a machine nobody has logged into, one that +/// cannot be reached, a window between blocks with no reset time. Those +/// are exactly the states worth looking at, and the ones nobody looks at +/// because arranging them costs real turns. An echo session sets this +/// with `/usage` (see `session::echo`), which is the same bargain the +/// rest of that driver makes: the fixture is invented, what is real is +/// the path it travels. +/// +/// Shared by the session layer, which writes it, and [`UsageMonitor`], +/// which reads it. Empty until something sets it, and an empty fixture +/// produces no snapshot at all -- an echo session meters nothing, and +/// nothing is what the phone should draw. +#[derive(Clone, Default)] +pub struct Fixture { + said: Arc>>, +} + +/// What a meter answered: which of the four states it is in, and whatever +/// windows go with it. Empty for every state but [`UsageState::Ok`]. +type Reported = (UsageState, Vec); + +/// How long the invented five-hour window has left, when nothing says. +const FIXTURE_MINUTES: i64 = 125; + +impl Fixture { + pub fn new() -> Self { + Self::default() + } + + fn is_set(&self) -> bool { + self.said.lock().unwrap().is_some() + } + + fn read(&self) -> Option { + self.said.lock().unwrap().clone() + } + + /// Acts on the words typed after `/usage`, and says what it did. + /// + /// The vocabulary lives here rather than in the echo driver because + /// these are this module's states: a driver spelling them out would + /// be a second place that has to learn about a fifth one. + pub fn command(&self, words: &str) -> String { + let mut words = words.split_whitespace(); + let Some(first) = words.next() else { + return match self.read() { + Some((state, windows)) => format!("usage fixture: {}", describe(&state, &windows)), + None => "usage fixture: unset, so this session meters nothing. \ + `/usage 42` puts up a five-hour window at 42%." + .to_string(), + }; + }; + let rest: Vec<&str> = words.collect(); + let detail = || { + if rest.is_empty() { + "set by /usage".to_string() + } else { + rest.join(" ") + } + }; + let (state, windows) = match first { + "off" | "none" | "clear" => { + *self.said.lock().unwrap() = None; + return "usage fixture cleared: this session meters nothing again".to_string(); + } + "notloggedin" | "logged-out" => (UsageState::NotLoggedIn, Vec::new()), + "unreachable" => (UsageState::Unreachable { detail: detail() }, Vec::new()), + "failed" => (UsageState::Failed { detail: detail() }, Vec::new()), + percent => match percent.parse::() { + Ok(percent) => ( + UsageState::Ok, + fixture_windows(percent.clamp(0.0, 100.0), rest.first().copied()), + ), + Err(_) => { + return format!( + "\"{percent}\" is not one of this fixture's answers. Say a percentage \ + (`/usage 42`, optionally with `90` minutes left, `never` for a window \ + between blocks, or `unreadable` for a reset time that cannot be read), \ + or one of `notloggedin`, `unreachable`, `failed`, `off`." + ); + } + }, + }; + let said = describe(&state, &windows); + *self.said.lock().unwrap() = Some((state, windows)); + format!("usage fixture set: {said}") + } +} + +/// The three windows Claude reports today, invented around one number. +/// +/// Three rather than one because the bar under a session header reads the +/// five-hour window and the dialog behind the button draws all of them, +/// and a fixture with one window leaves half the screen untested. The +/// weekly ones are derived from the same figure so that the worst of them +/// -- which is what colours the button -- is still the one asked for. +fn fixture_windows(percent: f64, reset: Option<&str>) -> Vec { + let resets_at = match reset { + // The state a real response is in between blocks: there is no + // window running, so there is nothing to reset. It is not a + // missing value, and the phone words it differently. + Some("never") | Some("none") => None, + // A timestamp that arrives and cannot be read, which is the one + // case that really is "we could not find out". + Some("unreadable") | Some("bad") => Some("whenever it feels like it".to_string()), + other => Some(reset_in( + other + .and_then(|word| word.parse().ok()) + .unwrap_or(FIXTURE_MINUTES), + )), + }; + vec![ + UsageWindow { + kind: "session".to_string(), + label: "5-hour window".to_string(), + percent, + resets_at: resets_at.clone(), + active: true, + }, + UsageWindow { + kind: "weekly_all".to_string(), + label: "Weekly (all models)".to_string(), + percent: percent / 2.0, + resets_at: resets_at.as_ref().map(|_| reset_in(FIXTURE_MINUTES * 40)), + active: false, + }, + UsageWindow { + kind: "weekly_scoped".to_string(), + label: "Weekly (Echo)".to_string(), + percent: percent / 4.0, + resets_at: resets_at.as_ref().map(|_| reset_in(FIXTURE_MINUTES * 40)), + active: false, + }, + ] +} + +/// `minutes` from now, in the format the real endpoint sends. +fn reset_in(minutes: i64) -> String { + let at = time::OffsetDateTime::now_utc() + time::Duration::minutes(minutes); + at.format(&time::format_description::well_known::Rfc3339) + // Formatting a timestamp cannot fail for any input this builds; + // saying so beats a fixture that silently has no reset time. + .unwrap_or_else(|_| "unformattable".to_string()) +} + +/// One line naming what a fixture is currently claiming, for the reply +/// the echo session writes back. +fn describe(state: &UsageState, windows: &[UsageWindow]) -> String { + match state { + UsageState::Ok => match windows.first() { + Some(window) => format!( + "{}% of the five-hour window, {}", + window.percent, + match &window.resets_at { + Some(at) => format!("resetting at {at}"), + None => "with no reset time (the between-blocks state)".to_string(), + } + ), + None => "no windows at all".to_string(), + }, + UsageState::NotLoggedIn => "nobody is logged in on this machine".to_string(), + UsageState::Unreachable { detail } => format!("machine unreachable ({detail})"), + UsageState::Failed { detail } => format!("the meter failed ({detail})"), + } +} + +/// The fixture, as a provider, so it travels the same route and the same +/// cache as a real meter rather than being spliced in at the screen. +struct EchoUsage { + setup: String, + setup_name: String, + fixture: Fixture, +} + +impl UsageProvider for EchoUsage { + fn name(&self) -> &'static str { + ECHO + } + + fn fetch(&self) -> UsageSnapshot { + let (state, windows) = self + .fixture + .read() + // Only ever built for a fixture that is set; a race with + // `/usage off` between the two reads lands here, and "the + // machine could not be asked" is the honest word for it. + .unwrap_or(( + UsageState::Unreachable { + detail: "the usage fixture was cleared".to_string(), + }, + Vec::new(), + )); + UsageSnapshot { + provider: self.name().to_string(), + setup: self.setup.clone(), + setup_name: self.setup_name.clone(), + state, + windows, + fetched_at: crate::session::now(), + } + } + + /// Read from memory, and set by somebody who is about to look at the + /// screen it changes. + fn poll_interval(&self) -> Duration { + Duration::ZERO + } +} + /// Which paid services a machine can be asked about. /// /// Derived from what the setup says it can run, so a machine with no Claude /// provider is not asked about Claude limits -- it has none, and a row saying -/// so would be a fact about nothing. A second service later adds a branch here -/// and an impl beside [`ClaudeUsage`], not a screen. -fn providers_for(setup: &SetupConfig) -> Vec> { +/// so would be a fact about nothing. +/// +/// Which meter a provider has is [`DriverKind::usage_provider`]'s answer rather +/// than a second match on kinds here, because the phone pairs a session with +/// one of these rows by that same name: two lists that disagreed would leave a +/// session looking for a snapshot nothing produces. A second service later is a +/// name there and an impl beside [`ClaudeUsage`], not a screen. +fn providers_for(setup: &SetupConfig, fixture: &Fixture) -> Vec> { let mut found: Vec> = Vec::new(); - if setup - .providers - .iter() - .any(|provider| provider.kind == DriverKind::ClaudeCli) - { - found.push(Box::new(ClaudeUsage { - setup: setup.id.clone(), - setup_name: setup.name.clone(), - transport: Transport::for_setup(setup), - })); + for provider in &setup.providers { + let Some(name) = provider.kind.usage_provider() else { + continue; + }; + // A machine offering two Claude providers has one account, not + // two: the meter belongs to the machine and the service, which is + // exactly what the cache is keyed by. + if found.iter().any(|already| already.name() == name) { + continue; + } + match name { + CLAUDE => found.push(Box::new(ClaudeUsage { + setup: setup.id.clone(), + setup_name: setup.name.clone(), + transport: Transport::for_setup(setup), + program: provider.program().to_string(), + })), + // Nothing at all until a test has asked for something: an + // echo session costs nothing, so the honest answer is no row + // rather than a row saying zero. + ECHO if fixture.is_set() => found.push(Box::new(EchoUsage { + setup: setup.id.clone(), + setup_name: setup.name.clone(), + fixture: fixture.clone(), + })), + _ => {} + } } found } @@ -313,11 +648,18 @@ type Cached = HashMap<(String, &'static str), (Instant, UsageSnapshot)>; /// machine per service per [`MIN_POLL_INTERVAL`], however often the phone asks. pub struct UsageMonitor { cache: Mutex, + /// The invented meter an echo session can put up; empty unless one + /// has. Shared with the session layer, which is where the command + /// that sets it is typed -- see [`Fixture`]. + fixture: Fixture, } impl UsageMonitor { - pub fn new() -> Self { - Self::default() + pub fn new(fixture: Fixture) -> Self { + Self { + cache: Mutex::new(Cached::new()), + fixture, + } } /// One snapshot per machine that offers a paid service, in the order the @@ -329,10 +671,10 @@ impl UsageMonitor { pub fn snapshots(&self, setups: &[SetupConfig]) -> Vec { let mut fresh = Vec::new(); for setup in setups { - for provider in providers_for(setup) { + for provider in providers_for(setup, &self.fixture) { let key = (setup.id.clone(), provider.name()); if let Some((fetched, snapshot)) = self.cache.lock().unwrap().get(&key) - && fetched.elapsed() < MIN_POLL_INTERVAL + && fetched.elapsed() < provider.poll_interval() { // Cached numbers, but the machine's *name* is read fresh: a // rename should show immediately rather than waiting out a @@ -367,6 +709,44 @@ impl UsageMonitor { #[cfg(test)] mod tests { use super::*; + use crate::config::DriverKind; + + #[test] + fn a_refusal_the_endpoint_answered_is_not_reported_as_an_unreachable_one() { + assert!(why(&ureq::Error::StatusCode(500)).contains("HTTP 500")); + assert!(why(&ureq::Error::HostNotFound).contains("unreachable")); + } + + #[test] + fn an_expired_login_says_where_to_log_in_rather_than_naming_the_network() { + let provider = ClaudeUsage { + setup: "far".to_string(), + setup_name: "somewhere else".to_string(), + transport: Transport::for_setup(&unreachable_setup()), + program: "/opt/claude".to_string(), + }; + // The machine cannot be reached, so the refresh attempt fails there + // rather than at the endpoint -- and the message still has to name the + // machine and the command, since that is all anybody gets to act on. + let UsageState::Failed { detail } = provider + .after_cli_refresh("stale") + .expect_err("an unreachable machine cannot refresh anything") + else { + panic!("an expired login is a fault to report, not a logged-out machine"); + }; + assert!(detail.contains("somewhere else"), "{detail}"); + assert!(detail.contains("/opt/claude doctor"), "{detail}"); + assert!( + !detail.contains("stale"), + "the token must never be quoted back" + ); + + let UsageState::Failed { detail } = provider.still_expired() else { + panic!("still expired is a fault"); + }; + assert!(detail.contains("/opt/claude /login"), "{detail}"); + assert!(!detail.contains("unreachable"), "{detail}"); + } #[test] fn parses_the_limits_array_defensively() { @@ -404,6 +784,7 @@ mod tests { port: None, identity_file: None, options: vec!["ConnectTimeout=1".to_string()], + models_dir: None, attachments_dir: None, }), providers: vec![crate::config::ProviderConfig { @@ -421,6 +802,7 @@ mod tests { setup: "far".to_string(), setup_name: "somewhere else".to_string(), transport: Transport::for_setup(&unreachable_setup()), + program: "claude".to_string(), }; let snapshot = provider.fetch(); // The distinction the old single `error` string could not make: this @@ -475,9 +857,62 @@ mod tests { models: vec![], }]; // A machine with no Claude on it has no Claude limits, and a row - // reporting on it would be a fact about nothing. - assert!(providers_for(&echo_only).is_empty()); - assert_eq!(providers_for(&unreachable_setup()).len(), 1); + // reporting on it would be a fact about nothing. Echo included: + // an echo session spends nothing, so until a fixture says + // otherwise there is no meter to report. + let unset = Fixture::new(); + assert!(providers_for(&echo_only, &unset).is_empty()); + assert_eq!(providers_for(&unreachable_setup(), &unset).len(), 1); + + // And with one set, that machine has exactly the invented meter + // -- under the name the session's `usageProvider` will name. + let fixture = Fixture::new(); + fixture.command("42"); + let found = providers_for(&echo_only, &fixture); + assert_eq!(found.len(), 1); + assert_eq!(found[0].name(), ECHO); + assert_eq!(DriverKind::Echo.usage_provider(), Some(ECHO)); + assert_eq!(DriverKind::ClaudeCli.usage_provider(), Some(CLAUDE)); + // A local model costs nothing to run, so it meters nothing. + assert_eq!(DriverKind::LlamaCpp.usage_provider(), None); + } + + /// The states the fixture exists to make reachable, and the one thing + /// it must not do: invent a reset time for a window that has none. + #[test] + fn the_fixture_says_each_state_the_screens_have_to_draw() { + let fixture = Fixture::new(); + assert!(fixture.read().is_none(), "unset until somebody sets it"); + + fixture.command("42 90"); + let (state, windows) = fixture.read().expect("set"); + assert_eq!(state, UsageState::Ok); + assert_eq!(windows[0].kind, "session"); + assert_eq!(windows[0].percent, 42.0); + assert!(windows[0].resets_at.is_some()); + + // Between blocks: no reset time, which the phone words as the + // window not running rather than as a time it could not read. + fixture.command("42 never"); + assert_eq!(fixture.read().expect("set").1[0].resets_at, None); + + fixture.command("unreachable no route to host"); + assert!(matches!( + fixture.read().expect("set").0, + UsageState::Unreachable { detail } if detail == "no route to host", + )); + + fixture.command("off"); + assert!(fixture.read().is_none()); + + // A word it does not know changes nothing and says what it takes. + fixture.command("42"); + let refused = fixture.command("sideways"); + assert!( + refused.contains("not one of this fixture's answers"), + "{refused}" + ); + assert_eq!(fixture.read().expect("still set").1[0].percent, 42.0); } #[test]