From 127b25e60ad595c9bd4d02446a9b301bbf814ffd Mon Sep 17 00:00:00 2001 From: iris <2+iris@noreply.localhost> Date: Fri, 4 Sep 2026 17:45:32 -0400 Subject: [PATCH 01/12] Meter a session by its provider, and let llama.cpp run over ssh The rate-limit bar answered a question about an account, and picked the answer by machine. One machine runs echo, the Claude CLI and a local model side by side, so every echo session on it drew the CLI's five-hour window: a quota that session cannot spend and could never run down. A session now names its meter (`usageProvider`, from `DriverKind::usage_provider`, which `usage::providers_for` reads too so the two lists cannot disagree), and the phone matches on machine *and* provider. Nothing meters echo or llama, and nothing at all is drawn -- including while the first fetch is out, since "checking" under a session that turns out to meter nothing is a row the screen then withdraws. Echo gets a meter it can be *told* about instead: `/usage 42`, `/usage 95 20`, `/usage 42 never`, `/usage notloggedin`, `/usage unreachable`, `/usage failed`, `/usage off`. Those states cost real quota to arrange, which is why none of them had been looked at. And llama.cpp runs wherever a setup says, which was the last of phase 5. `Transport::reserve_port` is the second half of what a transport is -- "run this" plus "reach this port" -- returning the port the server binds there and the port that reaches it here, and `Launch::reaching` puts the `-L` tunnel on the connection that already carries the command. Three things that came out of building it: - A forwarded launch gets a pty and every other one keeps `-T`. Killing the ssh client ends a CLI by closing the stdin it reads; llama-server never reads its stdin, so the same kill left it running on the far machine with the model loaded -- one orphan per stopped session. - The model is looked for on the machine that will serve it, at that machine's own models directory, so `GET /setups/{id}/models` is what the spawn screen offers rather than the backend's own downloads. - The readiness poll watches the process, not only the port: a model that will not load exits in a second and would otherwise have been reported as "gave up after 300s". The failure carries the log's tail. Exercised end to end against this VM over ssh to itself: spawn, load, answer, outlive a backend restart, be adopted, answer again, and stop -- with both the ssh client and the far llama-server gone afterwards. The local path, the Claude bar and the spawn screen checked on the emulator. Co-Authored-By: Claude Opus 5 --- AGENTS.md | 88 ++++- PLAN.md | 64 ++- .../src/main/kotlin/com/example/aiapp/Api.kt | 39 ++ .../kotlin/com/example/aiapp/SessionScreen.kt | 2 +- .../com/example/aiapp/SessionUsageBar.kt | 49 ++- .../kotlin/com/example/aiapp/SetupsScreen.kt | 10 + .../kotlin/com/example/aiapp/SpawnScreen.kt | 32 +- server/Cargo.lock | 1 + server/Cargo.toml | 6 + server/src/config.rs | 46 +++ server/src/main.rs | 4 +- server/src/models.rs | 79 ++++ server/src/routes.rs | 34 ++ server/src/session/echo.rs | 38 +- server/src/session/llama.rs | 177 +++++++-- server/src/session/mod.rs | 110 ++++-- server/src/session/transport.rs | 86 ++++- server/src/setups.rs | 5 +- server/src/ssh.rs | 121 +++++- server/src/usage.rs | 364 ++++++++++++++++-- 20 files changed, 1212 insertions(+), 143 deletions(-) diff --git a/AGENTS.md b/AGENTS.md index 1f49443..d46e571 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -203,15 +203,44 @@ images both ways), and the usage screen. host, `session::transport` turns that into an `ssh host …` invocation, and the driver never learns which it got. -**Phase 4 (llama.cpp)** works end to end, phone included (2026-08-28). -Models are browsed and downloaded from HuggingFace (`models.rs`, resumable -and verified), and `session::llama` runs one through `llama-server` over -its OpenAI-compatible streaming endpoint. Two things are deliberate and -easy to undo by accident: the conversation is rebuilt from the +**Phase 4 (llama.cpp)** works end to end, phone included (2026-08-28), +**on any machine a setup names** (2026-09-04). Models are browsed and +downloaded from HuggingFace (`models.rs`, resumable and verified), and +`session::llama` runs one through `llama-server` over its +OpenAI-compatible streaming endpoint. The conversation is rebuilt from the **transcript** rather than kept in the driver, because driver memory is -invisible to a second device; and a llama session is refused on an ssh -host, because the model is reached over HTTP and forwarding that port is -not built. +invisible to a second device -- deliberate, and easy to undo by accident. + +A remote llama session is the same command through the same transport +plus the second half of what a transport is: `Transport::reserve_port` +hands back a port the server binds *there* and a port that reaches it +*here*, and the ssh connection carrying the command carries the `-L` +tunnel between them (`llama-server` binds loopback on the far machine, so +nothing is served to its network). Three things that came out of building +it, each of which is easy to get wrong again: + +- **A forwarded launch gets a pty (`-tt`); every other one keeps `-T`.** + Killing the ssh client ends a CLI because it closes the stdin that CLI + is reading. `llama-server` never reads its stdin, so the same kill left + it running on the far machine holding the model in memory -- measured + 2026-09-04, one orphan per stopped session. A pty is what makes sshd + hang the far side up. Its log then arrives through a line discipline, + which nothing parses. +- **The model is looked for on the machine that will serve it**, at that + machine's own models directory (`SshConfig::models_dir`, defaulting to + `~/.local/share/ai-app/models` expanded *there*). What this backend has + downloaded is on that machine only when they are the same machine, so + `GET /setups/{id}/models` is what the spawn screen offers rather than + `GET /models`, and a model that is not there is refused at the spawn + with a sentence saying so. Downloading *to* another machine is not + built; the file gets there however anything else does. +- **A readiness poll watches the process, not only the port.** A model + that will not load, a port already taken, a flag an older build does not + know: all of them exit within a second and none will ever answer + `/health`, so waiting out the 300s timeout turned the server's own + account of the problem into "gave up". The failure now carries the last + few lines of `llama-server.log`, which on a remote session is the only + copy anybody reading the phone can see. Setups — machines, each carrying what it can run — are added, renamed, re-probed and removed from the app; providers are **discovered by asking @@ -226,12 +255,25 @@ symlink into `~/.local/bin` before a setup finds it. The escape hatch for anything odder is editing `config.ron` on the backend, deliberately the one authority the phone does not have. -**Testing llama.cpp here:** the prebuilt CPU build lives outside the repo -at `~/.local/opt/llama.cpp` (the 15 MB `ubuntu-x64` release asset). It -needs its own directory on `LD_LIBRARY_PATH`, so start the server as -`LD_LIBRARY_PATH=~/.local/opt/llama.cpp ai-server …` and point a provider's -`command` at `~/.local/opt/llama.cpp/llama-server`. A 0.6B Q8_0 answers at -usable speed on this VM's 8 cores. **Do not test with a 2-bit quant**: the +**llama.cpp is set up in this VM** (2026-09-04) and needs nothing typed: +the prebuilt CPU build is at `~/.local/opt/llama.cpp` (the 15 MB +`ubuntu-x64` release asset), symlinked as `/usr/local/bin/llama-server` so +that **discovery finds it over ssh too** -- `~/.local/bin` is not on the +PATH a non-interactive ssh session gets, which is why the symlink is +there and not only in `~/.local/bin`. It resolves its own libraries +through `$ORIGIN`, so no `LD_LIBRARY_PATH` is needed. One model is +downloaded, `unsloth/Qwen3-0.6B-GGUF/Qwen3-0.6B-Q8_0.gguf` (639 MB, under +`~/.local/share/ai-app/models`), which answers at usable speed on this +VM's 8 cores. + +**And the ssh path is exercisable here**, because this VM can ssh to +itself: the key is `~/.config/ai-app/ssh-self` (its public half is in +`~/.ssh/authorized_keys`, labelled removable), and a setup naming +`bob@127.0.0.1` with that `identityFile` plus +`options: ["StrictHostKeyChecking=no", "UserKnownHostsFile=/tmp/ai-app-known-hosts"]` +discovers `claude-cli` and `llama-cpp` on it. That is the whole rig for +"does a remote llama session work", since the far machine is this one and +the model file is the same file. **Do not test with a 2-bit quant**: the IQ2_XXS of that model produces fluent nonsense, which reads exactly like a broken driver — `llama-cli` produces the same from the file directly, which is how to tell the two apart in a hurry. @@ -292,6 +334,24 @@ first if a remote spawn ever mangles an argument. Dev Updater lists every variant under `build/outputs/apk`, so pick `release` there; a phone still holding the debug build has to uninstall it first, since the two are signed differently. +- **A rate-limit bar belongs to a session's provider, not to its + machine.** One machine offers echo, the Claude CLI and a local model at + once and only the CLI spends anything, so a session says which meter + reports on it (`usageProvider`, from `DriverKind::usage_provider`, which + `usage::providers_for` reads too so the two lists cannot disagree) and + the phone matches a snapshot on machine *and* provider. Nothing meters + a llama or echo session, and the phone draws **nothing at all** for one + -- not a zero, and not "unknown". The bar also draws nothing while the + first fetch is out: "checking" under a session that turns out to meter + nothing is a row the screen then has to withdraw. +- **`/usage` in an echo session puts up an invented meter**, which is how + those screens' states are reached without spending quota: + `/usage 42`, `/usage 95 20` (minutes left), `/usage 42 never` (the + between-blocks window with no reset time), `/usage 42 unreadable`, + `/usage notloggedin`, `/usage unreachable`, `/usage failed`, + `/usage off`. The vocabulary is `usage::Fixture`'s, since those are its + states; with none set an echo session meters nothing, which is the + ordinary case. - **A row something is happening to is dimmed, drained of colour, inert, and says which operation in a word** -- `BusyItem`, used by both the session list and the import list so the appearance is learned once. The diff --git a/PLAN.md b/PLAN.md index f83f603..7cdcd55 100644 --- a/PLAN.md +++ b/PLAN.md @@ -684,7 +684,22 @@ host) and **hosts**. The manager runs at most one llama-server per prompt-replayed by pi against the new endpoint). - Remote llama-server output is only reachable from the backend host, and binds localhost on the remote side with an SSH local port forward - (`ssh -L`) held by the manager — no LAN-exposed inference ports. + (`ssh -L`) held by the manager — no LAN-exposed inference ports. Built + 2026-09-04, held by the session's own ssh client rather than by a + manager: there is one server per session (not per `(host, model)`), so + the process that runs it is the process that owns the tunnel, and the + two die together. +- **The model file lives on the machine that serves it** (2026-09-04). + Each setup names its own models directory (`SshConfig::models_dir`, + default `~/.local/share/ai-app/models` expanded on that machine), and a + spawn resolves the key there — one round trip that answers "at + /abs/path" or "missing", so a model that is not there is refused at the + spawn instead of becoming a server that never becomes ready. The spawn + screen offers `GET /setups/{id}/models`, which is that machine's list, + rather than `GET /models`, which is the backend's downloads. Downloading + *to* another machine is deliberately not built: it would be a + multi-gigabyte transfer with no progress anywhere, and the file gets + there however anything else on that machine got there. ### SSH @@ -712,7 +727,23 @@ host) and **hosts**. The manager runs at most one llama-server per as a process but then spoken to over HTTP, so a remote one needs a forwarded port (`ssh -L`) as well as a spawned process. A transport is therefore "run this" plus "reach this port", and the second operation is - a no-op locally. + a no-op locally. **Built 2026-09-04**: `Transport::reserve_port` returns + a `Forward { there, here }` — the port the program binds on its own + machine and the port that reaches it from the backend, the same number + when that machine is this one — and `Launch::reaching` carries it, so + the connection that runs the command also carries the tunnel. The far + end is a guess from a range below the ephemeral one, because no + portable way to ask a machine for a free port avoids racing with the + bind anyway; a collision is not silent, since the program fails to bind + and the readiness poll reports what its log said. +- **A forwarded launch gets a pty and every other one does not** (measured + 2026-09-04). Killing the ssh client ends a CLI because it closes the + stdin that CLI is reading; `llama-server` never reads its stdin, so the + same kill left it running on the far machine with the model loaded — + one orphan per stopped session. With `-tt` the far side takes SIGHUP + when the connection goes. Its log then arrives through a line + discipline, which nothing parses. `-T` stays everywhere else, where a + pty would rewrite the JSONL. - Images need no file transfer, contrary to what this section said before: `attachment_block` base64s an uploaded image into the stream-json message itself, and produced images come back the same way @@ -746,6 +777,26 @@ optional and degrades rather than erroring. Structure it as one `UsageProvider` per paid service so a second service later is a new impl, not a parallel screen (rule 9). +**Per provider, not per machine (decided 2026-09-04).** A machine is not +what is metered; the provider a session runs is. One machine offers echo, +the Claude CLI and a local model side by side, and only the second of them +spends anything — so pairing a session with a snapshot by machine alone +drew the CLI's five-hour window under every echo session on it, reporting +a quota that session cannot spend and could never run down. A session now +names its meter (`usageProvider`, from `DriverKind::usage_provider`, which +`usage::providers_for` also reads so the two lists cannot disagree), and +`GET /usage` is matched on machine *and* provider. `None` is a session +that meters nothing, and the phone draws nothing at all for it — not a +zero, and not "unknown". + +`DriverKind::Echo` names a meter of its own, and it exists only when a +test has asked for one: `/usage` in an echo session sets an invented +answer (`usage::Fixture`), and with none set there is no snapshot and no +bar. That is what makes the states of those screens reachable — a number +near the top, a window between blocks with no reset time, a machine +nobody logged into, one that could not be reached — without spending real +quota to arrange them, which is why none of them had ever been looked at. + **Per machine, not per backend (decided 2026-08-29).** The credential store that matters is the one on the machine the session runs on, because that is the account being billed. Reading this machine's was right only while the @@ -789,7 +840,8 @@ POST /sessions/:id/compact (llama sessions) POST /sessions/:id/attachments multipart upload → id (referenced by /message) GET /sessions/:id/files/:ref images the session produced or was sent DELETE /sessions/:id kill process, release llama-server, delete transcript+files -GET /usage cached usage windows +GET /usage cached usage windows, per machine and provider +GET /setups/:id/models GGUFs on that machine, for a llama session there GET /setups/:id/dir?path=P entries of directory P, and P resolved GET /setups/:id/file?path=P content of file P, or why not PUT /setups/:id/file {path, content, ifSha256}; 409 if it moved on @@ -1140,8 +1192,10 @@ window just fills. shell-quoted). Attachment shipping turned out to be unnecessary for images — they ride the stdio JSONL as base64 in both directions, so nothing needs `scp` — and was built on 2026-09-03 for files, which - are attached by path (see "Transport" above). Still outstanding: - remote llama-server with its port forward, which comes with phase 4. + are attached by path (see "Transport" above). Remote llama-server with + its port forward landed 2026-09-04 — see "Transport" above for the + forward and the pty, and "llama-server management" for where the model + file has to be. Two things learned doing it: a remote session inherits ssh's non-login PATH, which is narrower than an interactive shell's (point `command` at an absolute path if a CLI isn't found), and the remote command is run diff --git a/app/androidApp/src/main/kotlin/com/example/aiapp/Api.kt b/app/androidApp/src/main/kotlin/com/example/aiapp/Api.kt index c47e2c3..7528178 100644 --- a/app/androidApp/src/main/kotlin/com/example/aiapp/Api.kt +++ b/app/androidApp/src/main/kotlin/com/example/aiapp/Api.kt @@ -193,6 +193,17 @@ data class SessionSummary( * server because that is where a provider's kind is known -- see `uploadPickedImage`. */ val maxImageEdge: Int?, + /** + * Which of `GET /usage`'s snapshots is about this session, and null where nothing meters it. + * + * The rate-limit bar answers a question about an *account*, and what decides which account -- + * if any -- is the provider this session runs, not the machine it runs on. Pairing by machine + * alone drew the Claude CLI's five-hour window under every echo session on a machine that also + * has the CLI: a quota that session cannot spend and could never run down. Decided by the + * server for the same reason [maxImageEdge] is -- it is a fact about the provider's kind, and + * this app has only its name. + */ + val usageProvider: String?, val status: String, val lastActivity: Double, ) @@ -213,6 +224,7 @@ private fun parseSession(session: JSONObject) = contextTokens = if (session.has("contextTokens")) session.getLong("contextTokens") else null, maxImageEdge = session.optInt("maxImageEdge", 0).takeIf { it > 0 }, + usageProvider = session.optString("usageProvider").ifEmpty { null }, status = session.getString("status"), lastActivity = session.getDouble("lastActivity"), ) @@ -402,6 +414,12 @@ data class SshDetails( * Where files attached from here land on that machine; null for the session's own directory. */ val attachmentsDir: String? = null, + /** + * Where that machine keeps its GGUF models; null for the same place the backend keeps its own + * (`~/.local/share/ai-app/models`, read on that machine). A llama.cpp session serves the file + * from the machine it runs on, so this is where its models are looked for and listed. + */ + val modelsDir: String? = null, ) private fun SshDetails.toJson() = @@ -409,6 +427,7 @@ private fun SshDetails.toJson() = if (port != null) put("port", port) if (!identityFile.isNullOrBlank()) put("identityFile", identityFile) if (!attachmentsDir.isNullOrBlank()) put("attachmentsDir", attachmentsDir) + if (!modelsDir.isNullOrBlank()) put("modelsDir", modelsDir) } /** What a machine turns out to have, without saving anything. */ @@ -1107,6 +1126,26 @@ private fun parseDownload(o: JSONObject) = error = if (o.has("error")) o.getString("error") else null, ) +/** + * The models on one machine, which is the list a llama.cpp session there can choose from. + * + * Not [fetchModels], which is what the *backend* has downloaded. A session serves its model from + * the machine it runs on, so for a machine reached over ssh those are two different lists -- and + * offering the backend's would name files that are not there, turning a choice that cannot work + * into a session that fails when it tries to load one. + */ +fun fetchSetupModels(settings: ServerSettings, setupId: String): List = + requestFromServer(settings, "/setups/${setupId.urlEncoded()}/models") { connection -> + JSONArray(connection.inputStream.bufferedReader().readText()).mapObjects { m -> + LocalModel( + key = m.getString("key"), + repo = m.getString("repo"), + file = m.getString("file"), + bytes = m.getLong("bytes"), + ) + } + } + fun fetchModels(settings: ServerSettings): Models = requestFromServer(settings, "/models") { connection -> val body = JSONObject(connection.inputStream.bufferedReader().readText()) diff --git a/app/androidApp/src/main/kotlin/com/example/aiapp/SessionScreen.kt b/app/androidApp/src/main/kotlin/com/example/aiapp/SessionScreen.kt index b108c10..75ada87 100644 --- a/app/androidApp/src/main/kotlin/com/example/aiapp/SessionScreen.kt +++ b/app/androidApp/src/main/kotlin/com/example/aiapp/SessionScreen.kt @@ -1193,7 +1193,7 @@ fun SessionScreen( // One poll for the machines' limits, read by everything on this screen that reports them: // the bar under the header, the colour of the button that opens the dialog, and the dialog. val usageFeed = rememberUsageFeed(settings) - val usage = usageFeed.forSetup(summary.setup) + val usage = usageFeed.forSession(summary) RecordFrames() var usageOpen by remember { mutableStateOf(false) } var settingsOpen by remember { mutableStateOf(false) } diff --git a/app/androidApp/src/main/kotlin/com/example/aiapp/SessionUsageBar.kt b/app/androidApp/src/main/kotlin/com/example/aiapp/SessionUsageBar.kt index c2b3695..5b61d62 100644 --- a/app/androidApp/src/main/kotlin/com/example/aiapp/SessionUsageBar.kt +++ b/app/androidApp/src/main/kotlin/com/example/aiapp/SessionUsageBar.kt @@ -74,13 +74,22 @@ class UsageFeed( /** Ask the backend again now. The dialog's refresh button; the poll does it on its own. */ val refresh: () -> Unit, ) { - /** What [setup]'s own limits came back as. See [usageFor] for why the states are these. */ - fun forSetup(setup: String): SessionUsage = - when (val state = snapshots) { + /** + * What meters [session], and what that meter came back as. See [usageFor] for the states. + * + * A session rather than a machine, because a machine is not what is metered: one machine runs + * the Claude CLI and an echo session side by side, and only the first of them spends anything. + */ + fun forSession(session: SessionSummary): SessionUsage { + // Settled without asking anybody: a session nothing meters has nothing to check, and + // "checking" is what the fetch's own states would say about it for as long as one is out. + val provider = session.usageProvider ?: return SessionUsage.NotMetered + return when (val state = snapshots) { is LoadState.Loading -> SessionUsage.Waiting is LoadState.Error -> SessionUsage.Unavailable(state.message) - is LoadState.Loaded -> usageFor(state.value, setup) + is LoadState.Loaded -> usageFor(state.value, session.setup, provider) } + } } /** @@ -165,9 +174,16 @@ fun SessionUsageBar(usage: SessionUsage, modifier: Modifier = Modifier) { } } - // Nothing at all for a machine that meters nothing: a row saying "unknown" there would + // Nothing at all for a session that meters nothing: a row saying "unknown" there would // report a problem about a setup somebody chose, on every screen, forever. - if (usage is SessionUsage.NotMetered) { + // + // And nothing while the first fetch is out, which is not the same kind of silence. A + // request in flight is not a state to report -- and the session that meters nothing is + // exactly the one this cannot yet tell apart, so "5-hour usage: checking" appeared under + // an echo session for half a second and was then taken away. A row that has to be + // withdrawn is worse than one that arrives late, and this is the only state here whose + // wrongness is a matter of timing rather than of fact. + if (usage is SessionUsage.NotMetered || usage is SessionUsage.Waiting) { return } @@ -178,9 +194,10 @@ fun SessionUsageBar(usage: SessionUsage, modifier: Modifier = Modifier) { // Words, not a colour and not an empty bar: every one of these is a different kind of // answer from "this much is used", and only words carry a difference in kind. when (val state = usage) { - SessionUsage.NotMetered -> Unit + // Both handled above, before the row exists at all. + SessionUsage.NotMetered, + SessionUsage.Waiting -> Unit is SessionUsage.Unavailable -> UsageNote("5-hour usage unknown -- ${state.why}") - SessionUsage.Waiting -> UsageNote("5-hour usage: checking") is SessionUsage.Known -> { val window = state.windows.firstOrNull { it.kind == "session" } if (window == null) { @@ -242,17 +259,23 @@ private fun fiveHourLabel(window: UsageWindow, now: OffsetDateTime): String { } /** - * One machine's snapshot, out of every machine's. + * One meter's snapshot, out of every machine's: [setup]'s row for [provider]. + * + * Both halves are needed to pick it. A machine can hold more than one meter -- the Claude CLI's + * account and, while a test has one set, an echo session's invented one -- and a snapshot is one + * service on one machine. * * Every way of having *failed* to get numbers is [SessionUsage.Unavailable] with the reason in it: * a machine nobody logged into, one that could not be reached, a snapshot that came back empty. * None of them may look like zero, and none may look like [SessionUsage.NotMetered], which is the * machine having no quota rather than the question going unanswered. */ -fun usageFor(snapshots: List, setup: String): SessionUsage { - // No snapshot at all means the backend never asked, which it only does for a machine with - // nothing metered on it. That is a different answer from having asked and failed. - val mine = snapshots.firstOrNull { it.setup == setup } ?: return SessionUsage.NotMetered +fun usageFor(snapshots: List, setup: String, provider: String): SessionUsage { + // No snapshot at all means the backend never asked, which it only does where there is nothing + // to ask about. That is a different answer from having asked and failed. + val mine = + snapshots.firstOrNull { it.setup == setup && it.provider == provider } + ?: return SessionUsage.NotMetered if (mine.state != "ok") { return SessionUsage.Unavailable(mine.detail ?: mine.state) } diff --git a/app/androidApp/src/main/kotlin/com/example/aiapp/SetupsScreen.kt b/app/androidApp/src/main/kotlin/com/example/aiapp/SetupsScreen.kt index 23c2dc0..5fb4561 100644 --- a/app/androidApp/src/main/kotlin/com/example/aiapp/SetupsScreen.kt +++ b/app/androidApp/src/main/kotlin/com/example/aiapp/SetupsScreen.kt @@ -237,6 +237,7 @@ private fun AddSetupDialog( var address by remember { mutableStateOf("") } var identity by remember { mutableStateOf("") } var attachmentsDir by remember { mutableStateOf("") } + var modelsDir by remember { mutableStateOf("") } var tested by remember { mutableStateOf(null) } var testing by remember { mutableStateOf(false) } @@ -251,6 +252,7 @@ private fun AddSetupDialog( port = typedPort, identityFile = identity.trim().ifEmpty { null }, attachmentsDir = attachmentsDir.trim().ifEmpty { null }, + modelsDir = modelsDir.trim().ifEmpty { null }, ) } @@ -296,6 +298,14 @@ private fun AddSetupDialog( label = { Text("Folder for attached files (optional)") }, singleLine = true, ) + // Where that machine's GGUFs are, for a llama.cpp session on it. Blank means + // the same place this backend keeps its own downloads, read on that machine. + OutlinedTextField( + value = modelsDir, + onValueChange = { modelsDir = it }, + label = { Text("Folder for models (optional)") }, + singleLine = true, + ) tested?.let { Spacer(Modifier.height(8.dp)) Text(it, style = MaterialTheme.typography.bodySmall) diff --git a/app/androidApp/src/main/kotlin/com/example/aiapp/SpawnScreen.kt b/app/androidApp/src/main/kotlin/com/example/aiapp/SpawnScreen.kt index 699c81f..172c861 100644 --- a/app/androidApp/src/main/kotlin/com/example/aiapp/SpawnScreen.kt +++ b/app/androidApp/src/main/kotlin/com/example/aiapp/SpawnScreen.kt @@ -70,9 +70,10 @@ fun SpawnScreen( // one leaves a filled-in form worth keeping, and that one leaves // nothing to fill in. var spawnError by remember { mutableStateOf(null) } - // Downloaded models, for a llama provider to choose between. Fetched - // beside the setups but kept separate: a Claude session needs none, so - // failing to list them must not stop the screen rendering. + // The models on the *chosen machine*, for a llama provider to choose between. Kept separate + // from the setups: a Claude session needs none, so failing to list them must not stop the + // screen rendering. Refetched when the machine changes, because a model is a file on one + // machine -- see [fetchSetupModels]. var models by remember { mutableStateOf>(emptyList()) } var modelKey by remember { mutableStateOf(null) } var contextSize by remember { mutableStateOf("") } @@ -89,9 +90,6 @@ fun SpawnScreen( } catch (e: ApiException) { LoadState.failed(e) } - models = - runCatching { withContext(Dispatchers.IO) { fetchModels(settings).local } } - .getOrDefault(emptyList()) } Column(Modifier.fillMaxSize().verticalScroll(rememberScrollState()).padding(16.dp)) { @@ -122,6 +120,17 @@ fun SpawnScreen( is LoadState.Loaded -> state.value } val setup = setups.firstOrNull { it.name == setupName } + // Whichever machine is chosen now, asked again when that changes. The old machine's list + // is dropped first rather than left on screen: a file name from another machine looks + // exactly like one from this one. + LaunchedEffect(setup?.id) { + models = emptyList() + modelKey = null + val id = setup?.id ?: return@LaunchedEffect + models = + runCatching { withContext(Dispatchers.IO) { fetchSetupModels(settings, id) } } + .getOrDefault(emptyList()) + } val current = setup?.providers?.firstOrNull { it.name == providerName } // Only the Claude CLI has models, a working directory and // permission modes; keying the extra fields on the kind rather @@ -183,13 +192,14 @@ fun SpawnScreen( ) if (isLlama) { - // A llama session names one of the models this backend has - // downloaded, so the choice is that list rather than free - // text -- there is nothing sensible to type here, and a name - // that is not on disk is a session that cannot start. + // A llama session names one of the models on the machine it will run on, so the + // choice is that list rather than free text -- there is nothing sensible to type + // here, and a name that is not on that machine's disk is a session that cannot + // start. if (models.isEmpty()) { Text( - "No models downloaded yet. Get one from the Models screen first.", + "No models on ${setup?.name ?: "this machine"}. The Models screen downloads " + + "to the backend; another machine needs the file put there itself.", style = MaterialTheme.typography.bodyMedium, color = MaterialTheme.colorScheme.onSurfaceVariant, ) diff --git a/server/Cargo.lock b/server/Cargo.lock index ad173ae..a6d5600 100644 --- a/server/Cargo.lock +++ b/server/Cargo.lock @@ -35,6 +35,7 @@ dependencies = [ "sha2", "tempfile", "thiserror", + "time", "tokio", "tokio-stream", "tower", diff --git a/server/Cargo.toml b/server/Cargo.toml index 6c941ef..b993455 100644 --- a/server/Cargo.toml +++ b/server/Cargo.toml @@ -47,6 +47,12 @@ ureq = { version = "3", features = ["json"] } # both in the graph rustls refuses to auto-select one. rustls = "0.23" libc = "0.2.189" +# One ISO-8601 timestamp: the reset time on the invented rate-limit window +# an echo session's `/usage` puts up. Already in the tree behind the +# certificate machinery, so this is a direct name for what is compiled +# anyway rather than a new crate -- and the alternative was hand-rolling a +# civil-from-days conversion to print one line. +time = { version = "0.3", features = ["formatting"] } [dev-dependencies] tempfile = "3" diff --git a/server/src/config.rs b/server/src/config.rs index 9760d8f..81b26a7 100644 --- a/server/src/config.rs +++ b/server/src/config.rs @@ -106,6 +106,19 @@ pub struct SshConfig { /// Extra `-o` settings, each written as `Key=value`. #[serde(default, skip_serializing_if = "Vec::is_empty")] pub options: Vec, + /// Where this machine keeps the GGUF models it can serve, absent for + /// the same default this backend uses (`~/.local/share/ai-app/models` + /// -- `$XDG_DATA_HOME` is not read on the far side, since it is this + /// machine's environment that would answer). A `~` prefix is the + /// remote home. + /// + /// Here rather than on the provider because it is a fact about the + /// machine, and because a machine reached over ssh is where the model + /// has to be: a llama.cpp session serves the file from the machine + /// that runs `llama-server`, and this backend's own downloads are on + /// whichever machine that is only when they are the same one. + #[serde(default, skip_serializing_if = "Option::is_none")] + pub models_dir: Option, /// Where a file attached from the phone is put on this machine so the /// session can read it. Absent means the session's own working /// directory, or the login home for a session that has none. A `~` @@ -167,6 +180,38 @@ impl DriverKind { } } + /// Which paid service meters a session of this kind, and `None` for + /// one that costs nothing. + /// + /// The rate-limit bars answer a question about an *account*, and what + /// decides which account -- if any -- is the provider a session runs, + /// not the machine it runs on. Those were the same thing only for as + /// long as a machine ran one kind of session: an echo session on a + /// laptop that also has the Claude CLI was drawn with that CLI's + /// five-hour window under its header, reporting a quota it cannot + /// spend and could not run down. A llama.cpp session is the same + /// story with the model on the far side. + /// + /// [`DriverKind::Echo`] names a meter of its own, which exists only + /// when a test has asked for one (`/usage` in `session::echo`). That + /// is what makes the bar's states -- a number, a machine nobody + /// logged into, one that could not be reached -- reachable without an + /// account and without spending a turn on somebody else's. With no + /// fixture set there is no snapshot for it, which the phone draws as + /// nothing at all. + /// + /// The string is a [`crate::usage::UsageProvider::name`], and it is + /// what pairs a session with one of the snapshots `GET /usage` + /// returns; the two lists have to agree, so `usage::providers_for` + /// reads this rather than matching on kinds a second time. + pub fn usage_provider(self) -> Option<&'static str> { + match self { + Self::ClaudeCli => Some(crate::usage::CLAUDE), + Self::Echo => Some(crate::usage::ECHO), + Self::LlamaCpp => None, + } + } + /// Whether the conversation exists outside this app, so that deleting /// the session here does not end it. /// @@ -442,6 +487,7 @@ mod tests { port: Some(2222), identity_file: None, options: Vec::new(), + models_dir: None, attachments_dir: None, }), providers: vec![ProviderConfig { diff --git a/server/src/main.rs b/server/src/main.rs index f85d4f5..5c0a586 100644 --- a/server/src/main.rs +++ b/server/src/main.rs @@ -284,7 +284,9 @@ async fn main() -> Result<()> { // -- so a machine added from the phone reports its limits without a // restart, and the backend's own account stops standing in for every // machine's. - let monitor = Arc::new(usage::UsageMonitor::new()); + // The fixture is the manager's, because that is where the `/usage` + // command that sets it is typed; the monitor is what serves it. + let monitor = Arc::new(usage::UsageMonitor::new(manager.usage_fixture())); // The bearer-token middleware wraps the entire router -- routes and // fallback alike -- here and only here, so a new route can't forget diff --git a/server/src/models.rs b/server/src/models.rs index 6f3ae7a..5434bd2 100644 --- a/server/src/models.rs +++ b/server/src/models.rs @@ -35,6 +35,8 @@ use serde::Serialize; use wg_app_link::private; +use crate::session::transport::{Launch, Transport}; + /// Identifies this client to HuggingFace. They ask for one, and a request /// without it is more likely to be rate-limited. const USER_AGENT: &str = concat!("ai-server/", env!("CARGO_PKG_VERSION")); @@ -551,6 +553,83 @@ fn collect(root: &Path, dir: &Path, found: &mut Vec) { } } +/// Where a machine reached over ssh keeps its models, when its setup does +/// not say. +/// +/// The same place this backend puts its own downloads, written out rather +/// than derived: `$XDG_DATA_HOME` here describes *this* machine's +/// environment, and the far machine's is the far machine's business. A +/// setup whose models are elsewhere says so (`SshConfig::models_dir`). +const FAR_MODELS_DIR: &str = "~/.local/share/ai-app/models"; + +/// Which directory holds the models on the machine `transport` reaches. +/// +/// One answer, because two things ask: the list a spawn screen offers, +/// and the path a session hands `llama-server`. A machine that listed one +/// directory and served from another would offer models that then failed +/// to load, which reads as the model being broken. +pub fn dir_on(transport: &Transport, local: &Path) -> String { + match transport { + Transport::Here => local.to_string_lossy().into_owned(), + Transport::Ssh { ssh, .. } => ssh + .models_dir + .as_ref() + .map_or(FAR_MODELS_DIR.to_string(), |dir| { + dir.to_string_lossy().into_owned() + }), + } +} + +/// Every GGUF on the machine a setup names, which is the machine that +/// would have to serve it. +/// +/// The local half of this is [`ModelStore::list`], reading the same shape +/// off this machine's disk; a caller picks by transport, since a setup +/// with no ssh *is* this machine and asking a shell about it would be a +/// slower way to the same answer. What must not happen is offering this +/// backend's downloads for a session on another machine: the file has to +/// be where `llama-server` runs, and a list that says otherwise is a +/// claim about the wrong filesystem. +/// +/// `dir` is that machine's models directory, `~` included -- expanded on +/// the far side, which is the only place that knows what it is. A +/// directory that is not there is an empty list rather than a failure: a +/// machine that has never had a model put on it is an ordinary state, and +/// the same one as a machine whose directory exists and is empty. +pub async fn on_machine(transport: &Transport, dir: &str) -> Result> { + let script = "p=$1; case $p in \"~\") p=$HOME;; \"~/\"*) p=$HOME/${p#\"~/\"};; esac; \ + [ -d \"$p\" ] || exit 0; \ + find \"$p\" -type f -name '*.gguf' -printf '%s\\t%P\\0'"; + let launch = Launch::new( + "sh", + vec![ + "-c".to_string(), + script.to_string(), + "sh".to_string(), + dir.to_string(), + ], + None, + ); + let out = transport.capture(&launch).await?; + let mut found: Vec = out + .split('\0') + .filter(|record| !record.is_empty()) + // Two fields, and the name last, so a `\t` in a filename survives. + .filter_map(|record| record.split_once('\t')) + .filter_map(|(bytes, key)| { + let (repo, file) = key.rsplit_once('/')?; + Some(LocalModel { + key: key.to_string(), + repo: repo.to_string(), + file: file.to_string(), + bytes: bytes.trim().parse().unwrap_or(0), + }) + }) + .collect(); + found.sort_by(|a, b| a.key.cmp(&b.key)); + Ok(found) +} + /// A model repository on HuggingFace, as the browse screen shows it. #[derive(Debug, Clone, Serialize)] #[serde(rename_all = "camelCase")] diff --git a/server/src/routes.rs b/server/src/routes.rs index 6cbdf15..a9986d9 100644 --- a/server/src/routes.rs +++ b/server/src/routes.rs @@ -7,6 +7,7 @@ //! POST /setups add {name, ssh?} -- providers are discovered //! POST /setups/probe dry run {ssh?}: what would be found there //! GET /setups/{id} one machine, for refetching after a change +//! GET /setups/{id}/models GGUFs on that machine, for a llama session //! GET /setups/{id}/dir?path=P entries of directory P, and P resolved //! GET /setups/{id}/file?path=P content of file P, or why not //! PUT /setups/{id}/file {path, content, ifSha256} -> new size/mtime/sha256 @@ -102,6 +103,7 @@ pub fn router(manager: Arc) -> Router { // `crate::files`. Under the setup rather than under a session // because a filesystem is a property of a machine; a session only // says where to start looking. + .route("/setups/{id}/models", get(setup_models)) .route("/setups/{id}/dir", get(list_dir).post(create_dir)) .route( "/setups/{id}/file", @@ -285,6 +287,9 @@ struct SshRequest { /// Where attached files land on that machine; see `SshConfig`. #[serde(default)] attachments_dir: Option, + /// Where that machine keeps its GGUF models; see `SshConfig`. + #[serde(default)] + models_dir: Option, } impl SshRequest { @@ -315,6 +320,14 @@ impl SshRequest { .map(str::trim) .filter(|dir| !dir.is_empty()) .map(std::path::PathBuf::from), + // The same rule, and for the same reason: this directory is + // on the other machine, so a `~` in it is that machine's home. + models_dir: self + .models_dir + .as_deref() + .map(str::trim) + .filter(|dir| !dir.is_empty()) + .map(std::path::PathBuf::from), }) } } @@ -499,6 +512,27 @@ struct PathQuery { path: String, } +/// The models **that machine** has, which is the list a llama.cpp session +/// on it can choose from. +/// +/// Not `GET /models`, which is this backend's own downloads: those are on +/// the machine a session runs on only when they are the same machine. A +/// spawn screen offering this backend's list for a remote setup would be +/// naming files that are not there, and the session would fail at the +/// point of loading rather than at the point of choosing. +async fn setup_models( + State(manager): State>, + UrlPath(id): UrlPath, +) -> Result>, ApiError> { + let setup = setup_by_id(&manager, &id)?; + let transport = crate::session::transport::Transport::for_setup(&setup); + let dir = crate::models::dir_on(&transport, manager.models_dir()); + crate::models::on_machine(&transport, &dir) + .await + .map(axum::Json) + .map_err(from_machine) +} + /// What is in a directory, and what that directory resolved to. async fn list_dir( State(manager): State>, diff --git a/server/src/session/echo.rs b/server/src/session/echo.rs index e731b99..8e8189b 100644 --- a/server/src/session/echo.rs +++ b/server/src/session/echo.rs @@ -28,6 +28,13 @@ //! - `/error [text]` -- a failure, which is otherwise awkward to cause. //! - `/peer [text]` -- a message from another agent, which otherwise takes //! two live sessions and one of them deciding to write. +//! - `/usage [what]` -- puts up an invented rate-limit answer, or takes +//! it away again (`/usage off`). An echo session meters nothing, so it +//! draws no usage bar at all until this is set; what it exists for is +//! the states the bar can be in, which otherwise cost real quota to +//! reach. `/usage 42`, `/usage 95 20`, `/usage 42 never`, +//! `/usage notloggedin`, `/usage unreachable`, `/usage failed`. The +//! vocabulary is `usage::Fixture`'s, which is where the states live. //! - `/compact` -- a compaction, start to finish. Typed rather than //! pressed, because the real dialects take it as a typed command too and //! the phone no longer has a button for it. @@ -106,6 +113,11 @@ pub struct EchoDriver { /// way AskUserQuestion does, and the turn resumes when the last of /// them is answered rather than the first. pending_questions: Mutex>, + /// The invented rate-limit answer `/usage` sets, shared with the + /// usage monitor that serves it. An echo session meters nothing, so + /// this is unset until a test asks for something -- see + /// [`crate::usage::Fixture`]. + usage: crate::usage::Fixture, /// A pretend context, so the status row has something that behaves the /// way a real one does: it grows with each turn, drops to what the /// compaction says it recovered, and a clear leaves it unmeasured. The @@ -345,6 +357,29 @@ impl EchoDriver { return; } + // Answered here rather than in the turn below, because it is not + // a turn: nothing is generated, and what is being exercised is + // the *other* screens -- the bar under the header, the button + // beside it and the dialog it opens, all of which read the usage + // route rather than this transcript. + if let Some(rest) = text.strip_prefix("/usage") { + if announce { + self.emit(Event::MessageTaken { + id: None, + text: text.clone(), + attachments, + }); + } + let said = self.usage.command(rest); + self.emit(Event::AssistantText { + delta: format!("{said}\n"), + }); + self.emit(Event::Status { + state: SessionStatus::Idle, + }); + return; + } + // The same word the real CLI takes, so a phone drives both the same // way. `Driver::compact` is what the manager's own route calls; // this is the typed path onto it. @@ -651,7 +686,7 @@ impl EchoDriver { }); } - pub fn new(sink: EventSink, session_dir: PathBuf) -> Self { + pub fn new(sink: EventSink, session_dir: PathBuf, usage: crate::usage::Fixture) -> Self { let driver = Self { sink, pending_questions: Mutex::new(Vec::new()), @@ -659,6 +694,7 @@ impl EchoDriver { busy: Arc::new(AtomicBool::new(false)), queued: Arc::new(Mutex::new(Vec::new())), session_dir, + usage, }; driver.emit(Event::Status { state: SessionStatus::Idle, diff --git a/server/src/session/llama.rs b/server/src/session/llama.rs index dca5eb3..0e86154 100644 --- a/server/src/session/llama.rs +++ b/server/src/session/llama.rs @@ -7,10 +7,24 @@ //! //! **It is spawned but not spoken to over stdio.** The process is started //! through the same [`Transport`] as any other, and then reached over -//! HTTP on a loopback port. That is the case the transport's doc comment -//! flags: a remote llama-server would need its port forwarded as well as -//! its command wrapped, which is not built, so a session on an ssh host -//! is refused rather than silently talking to the wrong machine. +//! HTTP on a loopback port. That is the second half of what a transport +//! is -- "run this" plus "reach this port" -- and it is what lets a +//! session run on another machine: [`Transport::reserve_port`] hands back +//! a port the server binds *there* and a port that reaches it *here*, and +//! the ssh connection carrying the command carries the tunnel between +//! them. The far `llama-server` binds loopback only, so a model is never +//! served to that machine's network. +//! +//! **The model file is the far machine's, not this one's.** A session +//! serves a GGUF from the machine that runs `llama-server`, so a remote +//! setup names its own models directory (`SshConfig::models_dir`, +//! defaulting to the same place this backend keeps its own downloads). +//! What this backend has downloaded is on that machine only when they are +//! the same machine -- so the file is looked for *there*, and a session +//! that names a model the machine does not have says so instead of +//! starting a server that will never load one. Downloading to another +//! machine is not built; the model gets there however anything else +//! gets there. //! //! **The server is stateless between requests**, so the whole //! conversation goes with every one. It is rebuilt from the session's @@ -85,16 +99,10 @@ impl LlamaDriver { session_dir: &Path, sink: EventSink, ) -> Result { - if !matches!(transport, Transport::Here) { - bail!( - "llama.cpp sessions can only run on this machine for now: the model is served \ - over HTTP, and forwarding that port to another host isn't built yet." - ); - } let model = meta.model.as_deref().context( "a llama.cpp session needs a model -- one of the downloaded ones, by its key", )?; - let path = model_path(models_dir, model)?; + let path = model_on(transport, models_dir, model)?; // Already loaded and still running: keep talking to it. The // health poll below is what confirms it is really answering, so @@ -120,14 +128,21 @@ impl LlamaDriver { )); } - let port = free_port().context("finding a port for llama-server")?; + // Where it listens on its own machine, and where that is reached + // from here -- the same number when that machine is this one. + let forward = transport + .reserve_port() + .context("finding a port for llama-server")?; let mut args: Vec = vec![ "-m".into(), - path.to_string_lossy().into_owned(), + path.clone(), + // Loopback there, whichever machine there is: what reaches it + // from outside that machine is the ssh tunnel and nothing + // else. "--host".into(), "127.0.0.1".into(), "--port".into(), - port.to_string(), + forward.there.to_string(), ]; // Settings that belong to the server because they decide how the // model is loaded; the sampling ones ride on each request instead, @@ -144,7 +159,7 @@ impl LlamaDriver { } let program = provider.command.as_deref().unwrap_or("llama-server"); - let launch = Launch::new(program, args, meta.cwd.as_deref()); + let launch = Launch::new(program, args, meta.cwd.as_deref()).reaching(forward); // Its output goes to files, not pipes. Not only so the process can // outlive this server: nothing ever read those pipes, so a chatty // llama-server filled the 64 KB buffer and blocked mid-load with @@ -161,8 +176,12 @@ impl LlamaDriver { .id() .context("llama-server exited before it could be recorded")?; tracing::info!( - "session {} running {program} for {model} on 127.0.0.1:{port} as pid {pid}", - meta.id + "session {} running {program} for {model} {} on 127.0.0.1:{} there, \ + reached at 127.0.0.1:{} here, as pid {pid}", + meta.id, + transport.describe(), + forward.there, + forward.here, ); // Reaped so it does not become a zombie while this server is still // its parent; the health poll and the record are what actually say @@ -173,12 +192,18 @@ impl LlamaDriver { let _ = child.wait().await; }); - let record = process::Record::of(pid, process::Detail::Http { port }) + // The *near* port, because that is the one anything reaching this + // server has to dial -- including a later run of this backend, + // which adopts the record without knowing which machine the server + // is on. For a remote session the recorded pid is the ssh + // client's, which is the process this machine owns and which holds + // the tunnel open for exactly as long as the far server lives. + let record = process::Record::of(pid, process::Detail::Http { port: forward.here }) .context("llama-server was gone before its start time could be read")?; process::write(session_dir, &record); Ok(Self::attached( - format!("http://127.0.0.1:{port}"), + format!("http://127.0.0.1:{}", forward.here), meta, model, transcript, @@ -212,7 +237,7 @@ impl LlamaDriver { let endpoint = endpoint.clone(); let model = model.to_string(); let session_dir = session_dir.to_path_buf(); - std::thread::spawn(move || match wait_until_ready(&endpoint) { + std::thread::spawn(move || match wait_until_ready(&endpoint, &session_dir) { Ok(()) => { tracing::info!("{model} loaded and answering at {endpoint}"); let _ = sink.send(Event::Status { @@ -518,19 +543,76 @@ fn model_path(models_dir: &Path, key: &str) -> Result { Ok(path) } -/// An unused loopback port, by asking the OS for one and letting it go. +/// The model file's path **on the machine that will serve it**, confirmed +/// to be there. /// -/// Racy in principle: something else could take it between here and -/// llama-server binding. In practice nothing on this machine is hunting -/// for ports, and the alternative -- parsing the port back out of the -/// server's log -- couples us to its output format for no real gain. -fn free_port() -> Result { - let listener = std::net::TcpListener::bind("127.0.0.1:0")?; - Ok(listener.local_addr()?.port()) +/// Local and remote answer the same question and it has to be asked of +/// two different filesystems, which is why this is one function rather +/// than a check beside the local path and hope for the other case. The +/// remote answer is measured for the same reason the local one is: a +/// missing file otherwise becomes a `llama-server` that starts, fails to +/// load, and reports as a session that never became ready -- which reads +/// as the machine being slow. +/// +/// One blocking round trip on a remote spawn, which is the same cost the +/// spawn is already paying to start ssh. The alternative is a path built +/// here from a `~` this machine cannot expand. +fn model_on(transport: &Transport, models_dir: &Path, key: &str) -> Result { + let Transport::Ssh { name, .. } = transport else { + return Ok(model_path(models_dir, key)?.to_string_lossy().into_owned()); + }; + // The same directory the spawn screen listed for this machine, and + // for the same reason it is one function: a list from one place and a + // load from another is a model that appears and then fails. + let dir = crate::models::dir_on(transport, models_dir); + // Checked here rather than in the script: `..` in a key would walk + // out of the models directory on a machine this server can start + // processes on, and the phone is where the key comes from. + for part in key.split('/') { + if part.is_empty() || part == "." || part == ".." { + bail!("\"{key}\" is not a model key this can resolve"); + } + } + let path = format!("{}/{key}", dir.trim_end_matches('/')); + // `$HOME` on the far side, which is the only machine that knows what + // it is -- and the resolved path is printed back so the launch below + // hands `llama-server` something absolute. + // + // "the file is not there" is answered rather than failed, because the + // two are different things to a reader and only one of them is a + // fault: a machine that could not be asked at all has to say so in + // its own words, and it would otherwise arrive as this same sentence + // about a missing model. + let script = "p=$1; case $p in \"~\") p=$HOME;; \"~/\"*) p=$HOME/${p#\"~/\"};; esac; \ + [ -f \"$p\" ] && printf 'at\\t%s\\n' \"$p\" || printf 'missing\\n'" + .to_string(); + let launch = Launch::new( + "sh", + vec!["-c".to_string(), script, "sh".to_string(), path.clone()], + None, + ); + let answer = transport + .capture_blocking(&launch) + .with_context(|| format!("couldn't ask {name} where its models are"))?; + match answer.trim().split_once('\t') { + Some(("at", resolved)) => Ok(resolved.to_string()), + _ => bail!( + "{name} has no model at {path}. A llama.cpp session serves the file from the \ + machine it runs on, so the model has to be on {name} -- what this backend has \ + downloaded is somewhere else." + ), + } } /// Polls until the server says it is ready, or gives up. -fn wait_until_ready(endpoint: &str) -> Result<()> { +/// +/// Watches the process as well as the port, because the two failures need +/// different words and one of them is common: a model that will not load, +/// a port already taken on the far machine, a `llama-server` too old for +/// a flag. All of those exit within a second and none of them will ever +/// answer `/health`, so waiting out the timeout turns a server that said +/// exactly what was wrong into "gave up after 300s". +fn wait_until_ready(endpoint: &str, session_dir: &Path) -> Result<()> { let deadline = std::time::Instant::now() + READY_TIMEOUT; let url = format!("{endpoint}/health"); loop { @@ -539,13 +621,48 @@ fn wait_until_ready(endpoint: &str) -> Result<()> { { return Ok(()); } + // `None` is the session having been stopped or deleted while this + // waited, which is nobody's fault and still not worth waiting on. + match process::recorded(session_dir) { + Some((_, process::Liveness::Alive | process::Liveness::Unknown)) => {} + Some((_, process::Liveness::Dead)) | None => { + bail!("it exited before it answered.{}", log_tail(session_dir)); + } + } if std::time::Instant::now() > deadline { - bail!("gave up after {}s", READY_TIMEOUT.as_secs()); + bail!( + "gave up after {}s.{}", + READY_TIMEOUT.as_secs(), + log_tail(session_dir) + ); } std::thread::sleep(std::time::Duration::from_millis(250)); } } +/// The end of `llama-server`'s own log, for a failure message. +/// +/// Its account of what went wrong is the useful half -- "failed to load +/// model", "bind: Address already in use" -- and on a remote session it +/// is the only half, since nobody reading the phone can open a file on +/// that machine. Bounded, because this ends up in an event a phone draws. +fn log_tail(session_dir: &Path) -> String { + let Ok(text) = std::fs::read_to_string(session_dir.join(SERVER_LOG)) else { + return String::new(); + }; + let tail: Vec<&str> = text.lines().rev().take(LOG_TAIL_LINES).collect(); + if tail.is_empty() { + return String::new(); + } + format!( + " It last said: {}", + tail.into_iter().rev().collect::>().join(" / ") + ) +} + +/// How much of that log to carry into a message somebody reads on a phone. +const LOG_TAIL_LINES: usize = 6; + /// One streamed completion: posts the conversation, emits each delta as it /// arrives. Emits rather than returns: the transcript those events land /// in is what the next turn reads back, so there is nothing to hand up. diff --git a/server/src/session/mod.rs b/server/src/session/mod.rs index f21241e..10d21ed 100644 --- a/server/src/session/mod.rs +++ b/server/src/session/mod.rs @@ -163,6 +163,17 @@ pub struct SessionInfo { /// answers and only one of them stays true. #[serde(skip_serializing_if = "Option::is_none")] pub max_image_edge: Option, + /// Which of `GET /usage`'s snapshots reports on this session, and + /// absent where nothing meters it -- see + /// [`DriverKind::usage_provider`]. + /// + /// Reported for the same reason `keeps_own_transcript` is: it is a + /// fact about the provider's *kind*, and the phone has only its name. + /// Pairing by machine alone was the bug it exists to fix -- one + /// machine runs echo and the Claude CLI, so every echo session drew + /// the CLI's five-hour window as if it were its own. + #[serde(skip_serializing_if = "Option::is_none")] + pub usage_provider: Option<&'static str>, /// Whether this session announces itself -- reported for the same /// reason `permission_mode` is: a switch that guesses its own position /// is how you turn something off while believing you are reading it. @@ -549,6 +560,7 @@ impl LiveSession { context_tokens: *self.shared.context_tokens.lock().unwrap(), notify: *self.shared.notify.lock().unwrap(), max_image_edge: kind.and_then(DriverKind::max_image_edge), + usage_provider: kind.and_then(DriverKind::usage_provider), imported, keeps_own_transcript: kind.is_some_and(DriverKind::keeps_own_transcript), cwd: cwd.map(Path::to_path_buf), @@ -585,6 +597,11 @@ pub struct SessionManager { /// [`SessionManager::marking_new_sessions_throwaway`] and /// [`SessionConfig::throwaway`]. spawn_throwaway: bool, + /// The invented rate-limit answer an echo session's `/usage` sets, + /// shared with the usage monitor that serves it. Held here because + /// every echo driver this manager builds is handed a clone -- see + /// [`SessionManager::reporting_usage_fixture`]. + usage_fixture: crate::usage::Fixture, inner: RwLock, } @@ -603,6 +620,11 @@ impl SessionManager { wg_app_link::private::create_dir(&data_dir)?; let (notifications, _) = broadcast::channel(NOTIFICATION_BUFFER); + // Made here rather than passed in, and handed *out* to the usage + // monitor by whoever wires the two together: every echo driver + // this manager builds gets a clone, including the ones built + // below, so it has to exist before the first session does. + let usage_fixture = crate::usage::Fixture::new(); let mut live = HashMap::new(); for meta in &config.sessions { // One unlaunchable session -- a corrupt transcript, an @@ -614,8 +636,11 @@ impl SessionManager { meta.clone(), &setup, &provider, - &data_dir, - &models_dir, + Env { + data_dir: &data_dir, + models_dir: &models_dir, + usage: &usage_fixture, + }, notifications.clone(), // Nothing is started here. See `Launching`: a restart // picks up the processes that are still running and @@ -638,11 +663,41 @@ impl SessionManager { notifications, pending: Arc::new(pending::Registry::default()), spawn_throwaway: false, + usage_fixture, inner: RwLock::new(Inner { config, live }), }; Ok(manager) } + /// Where this backend's own model downloads live. The machine a + /// session runs on may keep its elsewhere -- see `models::dir_on`. + pub fn models_dir(&self) -> &Path { + &self.models_dir + } + + /// What this manager lends a session it launches. Borrowed from the + /// manager rather than cloned, so there is one answer to where things + /// are kept. + fn env(&self) -> Env<'_> { + Env { + data_dir: &self.data_dir, + models_dir: &self.models_dir, + usage: &self.usage_fixture, + } + } + + /// The invented rate-limit answer this manager's echo sessions set + /// with `/usage`, for the usage monitor to serve. + /// + /// Handed out rather than taken in because the drivers built inside + /// the constructor need it, and because the direction is the one the + /// layering allows: `usage` sits below the session layer, so a + /// session can hold one of its types while it holds nothing of a + /// session's. + pub fn usage_fixture(&self) -> crate::usage::Fixture { + self.usage_fixture.clone() + } + /// Marks every session spawned from here on as one whose process is /// stopped when this server exits -- see [`SessionConfig::throwaway`] /// and [`SessionManager::stop_throwaway_sessions`]. @@ -1035,6 +1090,8 @@ impl SessionManager { context_tokens: None, max_image_edge: kind_of(&inner.config, &meta.setup, &meta.provider) .and_then(DriverKind::max_image_edge), + usage_provider: kind_of(&inner.config, &meta.setup, &meta.provider) + .and_then(DriverKind::usage_provider), notify: meta.notify, imported: import::read_cursor(&self.data_dir.join(&meta.id)).is_some(), keeps_own_transcript: keeps_own_transcript( @@ -1162,8 +1219,7 @@ impl SessionManager { meta.clone(), &setup, &provider, - &self.data_dir, - &self.models_dir, + self.env(), self.notifications.clone(), Launching::Asked(seed), )?; @@ -1605,7 +1661,7 @@ impl SessionManager { &meta, &setup, &provider, - &self.models_dir, + self.env(), session.dir(), session.transcript_path(), &session.sink, @@ -1619,8 +1675,7 @@ impl SessionManager { meta, &setup, &provider, - &self.data_dir, - &self.models_dir, + self.env(), self.notifications.clone(), Launching::Asked(None), )?; @@ -2020,6 +2075,19 @@ enum Launching { Restart, } +/// What the server around a session lends it: where sessions and models +/// are kept, and the usage fixture an echo session's `/usage` sets. +/// +/// One parameter rather than three because they travel together through +/// every launch path and none of them is a fact about the session -- +/// they are this server's belongings, handed down. +#[derive(Clone, Copy)] +struct Env<'a> { + data_dir: &'a Path, + models_dir: &'a Path, + usage: &'a crate::usage::Fixture, +} + /// Creates the session directory, opens its transcript (continuing the /// sequence numbering if one exists), settles what the session is doing, /// and spawns the event pump -- with a driver behind it where there is a @@ -2028,12 +2096,11 @@ fn launch( meta: SessionConfig, setup: &SetupConfig, provider: &ProviderConfig, - data_dir: &Path, - models_dir: &Path, + env: Env<'_>, notifications: broadcast::Sender, why: Launching, ) -> Result> { - let dir = data_dir.join(&meta.id); + let dir = env.data_dir.join(&meta.id); wg_app_link::private::create_dir(&dir)?; let transcript_path = dir.join("transcript.jsonl"); let mut transcript = Transcript::open(&transcript_path)?; @@ -2163,17 +2230,7 @@ fn launch( let driver = Arc::new(Mutex::new( driving - .then(|| { - make_driver( - &meta, - setup, - provider, - models_dir, - &dir, - &transcript_path, - &sink, - ) - }) + .then(|| make_driver(&meta, setup, provider, env, &dir, &transcript_path, &sink)) .transpose()?, )); @@ -2216,18 +2273,22 @@ fn make_driver( meta: &SessionConfig, setup: &SetupConfig, provider: &ProviderConfig, - models_dir: &Path, + env: Env<'_>, dir: &Path, transcript_path: &Path, sink: &EventSink, ) -> Result> { Ok(match provider.kind { - DriverKind::Echo => Arc::new(EchoDriver::new(sink.clone(), dir.to_path_buf())), + DriverKind::Echo => Arc::new(EchoDriver::new( + sink.clone(), + dir.to_path_buf(), + env.usage.clone(), + )), DriverKind::LlamaCpp => Arc::new(LlamaDriver::launch( meta, provider, &Transport::for_setup(setup), - models_dir, + env.models_dir, transcript_path, dir, sink.clone(), @@ -2584,6 +2645,7 @@ mod tests { driver: Arc::new(Mutex::new(Some(Arc::new(EchoDriver::new( sink.clone(), dir.path().to_path_buf(), + crate::usage::Fixture::new(), ))))), sink, waiting: Mutex::new(VecDeque::new()), diff --git a/server/src/session/transport.rs b/server/src/session/transport.rs index cf2262b..aad4f4e 100644 --- a/server/src/session/transport.rs +++ b/server/src/session/transport.rs @@ -13,11 +13,14 @@ //! this module decides *which* transport, that one knows what a correct //! ssh invocation is. //! -//! Known second operation, not built because nothing needs it yet: a -//! managed `llama-server` is spawned as a process but then spoken to over -//! HTTP, so a remote one needs a forwarded port (`ssh -L`) as well. A -//! transport is eventually "run this" plus "reach this port", where the -//! second is a no-op locally. See PLAN.md's SSH section. +//! A transport is therefore two operations rather than one: **run this**, +//! and **reach this port**. The second is what a managed `llama-server` +//! needs -- it is spawned as a process and then spoken to over HTTP -- and +//! it is a no-op locally, where the port a program binds is already a port +//! this machine can dial. Over ssh it is an `-L` tunnel carried by the +//! same connection that runs the command, so the model server binds +//! loopback on the far machine and is never exposed to its network. See +//! [`Transport::reserve_port`] and PLAN.md's SSH section. use std::path::{Path, PathBuf}; use std::process::Stdio; @@ -26,16 +29,28 @@ use anyhow::{Context, Result}; use tokio::process::Child; use crate::config::SshConfig; +pub use crate::ssh::Forward; /// What a driver needs run in order to exist as a process. /// -/// Deliberately just the three things every transport can carry. Anything -/// a particular machine needs -- a port, a key, extra ssh options -- is +/// Deliberately just what every transport can carry: the command, where +/// it runs, and a port the caller needs to reach. Anything a particular +/// machine needs -- a key, extra ssh options, which address to dial -- is /// the transport's own configuration, not something a driver states. pub struct Launch { pub program: String, pub args: Vec, pub cwd: Option, + /// A port this program will listen on, and the port that reaches it + /// from here -- see [`Transport::reserve_port`], which is the only + /// thing that should produce one. + /// + /// On the launch rather than in [`Transport::spawn`]'s signature + /// because it is part of what is being run: a caller that needs to + /// reach the process it is starting says so once, where it says + /// everything else about it, and every transport reads it the same + /// way. + pub forward: Option, } impl Launch { @@ -44,8 +59,16 @@ impl Launch { program: program.into(), args, cwd: cwd.map(Path::to_path_buf), + forward: None, } } + + /// Says that this program serves `forward.there`, and that the caller + /// will reach it at `forward.here`. + pub fn reaching(mut self, forward: Forward) -> Self { + self.forward = Some(forward); + self + } } /// How a launched process's standard streams are connected. @@ -113,6 +136,7 @@ impl Transport { &launch.program, &launch.args, launch.cwd.as_deref(), + launch.forward, )); match streams { Streams::Piped => { @@ -169,12 +193,15 @@ impl Transport { Self::Here => None, Self::Ssh { ssh, .. } => Some(ssh), }; - let output = - crate::ssh::command(host, &launch.program, &launch.args, launch.cwd.as_deref()) - .output() - .with_context(|| { - format!("couldn't run \"{}\" {}", launch.program, self.describe()) - })?; + let output = crate::ssh::command( + host, + &launch.program, + &launch.args, + launch.cwd.as_deref(), + launch.forward, + ) + .output() + .with_context(|| format!("couldn't run \"{}\" {}", launch.program, self.describe()))?; if !output.status.success() { let stderr = String::from_utf8_lossy(&output.stderr).trim().to_string(); anyhow::bail!(if stderr.is_empty() { @@ -237,6 +264,34 @@ impl Transport { }) } + /// Picks a port for a launched program to serve on, and the port that + /// reaches it from here. + /// + /// The "reach this port" half of what a transport is. Locally there is + /// one port and the OS chooses it, by binding and letting go -- racy + /// in principle, and nothing on this machine is hunting for ports. + /// + /// Over ssh the near end is chosen the same way and the far end is a + /// guess, because there is no portable way to ask a machine for a free + /// port that does not race with binding it anyway. It is taken from + /// [`FAR_PORTS`], below the range Linux hands out to outgoing + /// connections, so a collision means something else deliberately + /// listening there. That is not silent: the program fails to bind and + /// exits, and `session::llama` reports what its log said rather than + /// waiting out its readiness timeout. + pub fn reserve_port(&self) -> Result { + let listener = std::net::TcpListener::bind("127.0.0.1:0") + .context("asking this machine for a free port")?; + let here = listener.local_addr()?.port(); + Ok(match self { + Self::Here => Forward { there: here, here }, + Self::Ssh { .. } => Forward { + there: rand::random_range(FAR_PORTS), + here, + }, + }) + } + /// How to say where this runs, for a log line a person reads. pub fn describe(&self) -> String { match self { @@ -246,6 +301,11 @@ impl Transport { } } +/// Where a port on another machine is guessed from: high enough to be out +/// of the way of services, and below the 32768-60999 Linux hands out to +/// outgoing connections, which is where a guess would most often collide. +const FAR_PORTS: std::ops::Range = 20000..30000; + /// What a command is given on its standard input. /// /// Three cases rather than an `Option` because they are three diff --git a/server/src/setups.rs b/server/src/setups.rs index 0937bdb..24e3437 100644 --- a/server/src/setups.rs +++ b/server/src/setups.rs @@ -30,7 +30,10 @@ use crate::session::transport::{Launch, Transport}; /// what the phone shows and what a session stores. const PROBES: &[(&str, &str, DriverKind)] = &[ ("claude-cli", "claude", DriverKind::ClaudeCli), - ("local-llama", "llama-server", DriverKind::LlamaCpp), + // Named for the program rather than for where it runs: it runs + // wherever the setup is, and "local" was true only while a llama + // session could not be spawned on another machine. + ("llama-cpp", "llama-server", DriverKind::LlamaCpp), ]; /// Models offered for a discovered Claude CLI. A shortcut list for the diff --git a/server/src/ssh.rs b/server/src/ssh.rs index e47a511..f811227 100644 --- a/server/src/ssh.rs +++ b/server/src/ssh.rs @@ -15,6 +15,23 @@ use std::process::Command; use crate::config::SshConfig; +/// A port on the machine a command runs on, and the port that reaches it +/// from the backend. +/// +/// The second half of what a transport is (PLAN.md's SSH section): "run +/// this" plus "reach this port". Locally the two numbers are the same one +/// and nothing is forwarded; over ssh the connection carries an `-L` +/// tunnel, so a model server binds loopback on the far machine and is +/// never exposed to its network. +#[derive(Debug, Clone, Copy, PartialEq, Eq)] +pub struct Forward { + /// What the launched program should listen on, on its own machine. + pub there: u16, + /// What this machine connects to. The same number as `there` when the + /// program runs here. + pub here: u16, +} + /// Options forced onto every connection. `BatchMode` makes a missing key /// fail immediately with a readable message instead of hanging on a /// password prompt that nothing can answer; the keepalives turn a silently @@ -45,6 +62,7 @@ pub fn command( program: &str, args: &[String], cwd: Option<&Path>, + forward: Option, ) -> Command { let Some(ssh) = remote else { let mut command = Command::new(program); @@ -64,9 +82,42 @@ pub fn command( }; let mut command = Command::new("ssh"); - // -T: no pty. This carries JSONL, and a pty would rewrite it (echo, - // CRLF translation, ^C handling) into something the parser can't read. - command.arg("-T"); + if let Some(forward) = forward { + // A forwarded process is not spoken to over stdio, and that + // changes how it has to be shut down. Everything else here is a + // CLI reading its stdin, so killing the ssh client closes that + // stdin and the far process ends; a `llama-server` never reads + // its own, so the same kill left it running on the far machine + // holding the model in memory -- measured 2026-09-04, an orphan + // per stopped session. A pty is what makes sshd hang the far side + // up: when the connection goes, the master closes and the session + // takes SIGHUP. `-tt` because this client has no terminal of its + // own to inherit one from. + // + // The cost is that its log arrives through a line discipline + // (CRLF, and whatever the program does when it thinks it is on a + // terminal). Nothing parses that log, so it is a fair trade for a + // process that reliably goes away. + command.arg("-tt"); + // Loopback on both ends: the far side binds 127.0.0.1, so the + // port it serves is reachable only through this connection and + // never from that machine's network -- and the near end is bound + // to this host alone for the same reason. + command.args([ + "-L", + &format!("127.0.0.1:{}:127.0.0.1:{}", forward.here, forward.there), + ]); + // Without this a forward that cannot be set up is a warning on + // stderr and a session that runs anyway, answering nothing: the + // failure would arrive as "the model never became ready", which + // is the wrong thing to go looking at. + command.args(["-o", "ExitOnForwardFailure=yes"]); + } else { + // -T: no pty. This carries JSONL, and a pty would rewrite it + // (echo, CRLF translation, ^C handling) into something the parser + // can't read. + command.arg("-T"); + } for option in SSH_OPTIONS { command.args(["-o", option]); } @@ -200,6 +251,7 @@ mod tests { port: None, identity_file: None, options: vec![], + models_dir: None, attachments_dir: None, } } @@ -211,6 +263,7 @@ mod tests { "claude", &args(["-p", "--verbose"]), Some(Path::new("/tmp/x")), + None, ); assert_eq!(argv(&command), ["claude", "-p", "--verbose"]); assert_eq!(command.get_current_dir(), Some(Path::new("/tmp/x"))); @@ -223,6 +276,7 @@ mod tests { port: Some(2222), identity_file: Some("/home/me/.ssh/id_ai".into()), options: vec!["StrictHostKeyChecking=accept-new".to_string()], + models_dir: None, attachments_dir: None, }; let rendered = argv(&command( @@ -230,6 +284,7 @@ mod tests { "claude", &args(["-p", "--model", "haiku"]), Some(Path::new("/home/bob/work")), + None, )); assert_eq!(rendered[0], "ssh"); @@ -250,12 +305,60 @@ mod tests { #[test] fn a_remote_command_without_a_cwd_just_execs() { let ssh = bare_host(); - let rendered = argv(&command(Some(&ssh), "claude", &args(["-p"]), None)); + let rendered = argv(&command(Some(&ssh), "claude", &args(["-p"]), None, None)); assert_eq!(rendered.last().unwrap(), "exec 'claude' '-p'"); // No -i means no IdentitiesOnly: ~/.ssh/config decides instead. assert!(!rendered.contains(&"IdentitiesOnly=yes".to_string())); } + /// The second half of a transport: the connection that runs the + /// command also carries the port that reaches it. + /// + /// Both ends are pinned to loopback, which is the property that keeps + /// a model server off the far machine's network -- asserted here + /// rather than trusted, because dropping the addresses is a one-word + /// edit that still works on a machine nobody else can reach. + #[test] + fn a_forwarded_port_rides_the_same_connection_as_the_command() { + let ssh = bare_host(); + let rendered = argv(&command( + Some(&ssh), + "llama-server", + &args(["--port", "24242"]), + None, + Some(Forward { + there: 24242, + here: 41000, + }), + )); + let forward = rendered + .iter() + .position(|arg| arg == "-L") + .expect("a forward"); + assert_eq!(rendered[forward + 1], "127.0.0.1:41000:127.0.0.1:24242"); + assert!(rendered.contains(&"ExitOnForwardFailure=yes".to_string())); + // The half that is easy to lose: without a pty the far process + // outlives the connection, because nothing closes a stdin it + // never reads. + assert!(rendered.contains(&"-tt".to_string())); + assert!(!rendered.contains(&"-T".to_string())); + // Options come before the host, or ssh reads them as part of the + // remote command. + assert!(forward < rendered.len() - 2); + assert_eq!( + rendered.last().unwrap(), + "exec 'llama-server' '--port' '24242'" + ); + + // Nothing forwarded is nothing added: every other session is one + // of these, and an -L on it would bind a port for no reason. + let plain = argv(&command(Some(&ssh), "claude", &args(["-p"]), None, None)); + assert!(!plain.contains(&"-L".to_string())); + // And a session that *is* spoken to over stdio keeps its raw pipe. + assert!(plain.contains(&"-T".to_string())); + assert!(!plain.contains(&"-tt".to_string())); + } + /// The one character quoting must not swallow. /// /// A working directory typed as `~/repos/ai-app` was arriving as the @@ -300,7 +403,13 @@ mod tests { assert_eq!(expand_home(Path::new("/tmp/~/x")), Path::new("/tmp/~/x")); assert_eq!(expand_home(Path::new("~user/x")), Path::new("~user/x")); - let local = command(None, "claude", &args(["-p"]), Some(Path::new("~/work"))); + let local = command( + None, + "claude", + &args(["-p"]), + Some(Path::new("~/work")), + None, + ); assert_eq!(local.get_current_dir(), Some(home.join("work").as_path())); } @@ -324,7 +433,7 @@ mod tests { // that tries to close the quote and start a new command. let ssh = bare_host(); let evil = Path::new("/tmp/'; touch /tmp/pwned; '"); - let rendered = argv(&command(Some(&ssh), "claude", &[], Some(evil))); + let rendered = argv(&command(Some(&ssh), "claude", &[], Some(evil), None)); let script = rendered.last().unwrap(); assert_eq!( script, diff --git a/server/src/usage.rs b/server/src/usage.rs index 929f855..a63acc7 100644 --- a/server/src/usage.rs +++ b/server/src/usage.rs @@ -35,13 +35,13 @@ //! on it). use std::collections::HashMap; -use std::sync::Mutex; +use std::sync::{Arc, Mutex}; use std::time::{Duration, Instant}; use serde::Serialize; use serde_json::Value; -use crate::config::{DriverKind, SetupConfig}; +use crate::config::SetupConfig; use crate::session::transport::{Launch, Transport}; const USAGE_URL: &str = "https://api.anthropic.com/api/oauth/usage"; @@ -116,10 +116,30 @@ pub struct UsageSnapshot { pub fetched_at: f64, } +/// The name of each meter, said in one place because two lists have to +/// agree on it: [`UsageSnapshot::provider`], which is what `GET /usage` +/// labels a row with, and [`crate::config::DriverKind::usage_provider`], +/// which is how a session says which of those rows is about it. +pub const CLAUDE: &str = "claude"; +/// The invented one, for testing the screens that draw these -- see +/// [`Fixture`]. +pub const ECHO: &str = "echo"; + pub trait UsageProvider: Send + Sync { fn name(&self) -> &'static str; /// Blocking -- call off the async workers. fn fetch(&self) -> UsageSnapshot; + /// How long an answer from this one may be reused. + /// + /// A property of the provider rather than of the cache, because what + /// sets it is what asking costs: [`ClaudeUsage`] makes a network call + /// against an endpoint that rate-limits impatient callers, and the + /// fixture below reads a mutex. Caching the fixture for three minutes + /// would mean a test setting a number and watching the old one for + /// most of that, which reads exactly like the command not working. + fn poll_interval(&self) -> Duration { + MIN_POLL_INTERVAL + } } /// Reads the numbers behind Claude Code's `/usage` from one machine, using @@ -182,7 +202,7 @@ impl ClaudeUsage { impl UsageProvider for ClaudeUsage { fn name(&self) -> &'static str { - "claude" + CLAUDE } fn fetch(&self) -> UsageSnapshot { @@ -293,24 +313,260 @@ fn parse_windows(body: &Value) -> Vec { .collect() } +/// An invented answer, so the screens that draw these can be exercised +/// without an account. +/// +/// Every state the usage bar and the usage dialog can be in is otherwise +/// reachable only by spending somebody's quota or by breaking a machine: +/// a number near the top, a machine nobody has logged into, one that +/// cannot be reached, a window between blocks with no reset time. Those +/// are exactly the states worth looking at, and the ones nobody looks at +/// because arranging them costs real turns. An echo session sets this +/// with `/usage` (see `session::echo`), which is the same bargain the +/// rest of that driver makes: the fixture is invented, what is real is +/// the path it travels. +/// +/// Shared by the session layer, which writes it, and [`UsageMonitor`], +/// which reads it. Empty until something sets it, and an empty fixture +/// produces no snapshot at all -- an echo session meters nothing, and +/// nothing is what the phone should draw. +#[derive(Clone, Default)] +pub struct Fixture { + said: Arc>>, +} + +/// What a meter answered: which of the four states it is in, and whatever +/// windows go with it. Empty for every state but [`UsageState::Ok`]. +type Reported = (UsageState, Vec); + +/// How long the invented five-hour window has left, when nothing says. +const FIXTURE_MINUTES: i64 = 125; + +impl Fixture { + pub fn new() -> Self { + Self::default() + } + + fn is_set(&self) -> bool { + self.said.lock().unwrap().is_some() + } + + fn read(&self) -> Option { + self.said.lock().unwrap().clone() + } + + /// Acts on the words typed after `/usage`, and says what it did. + /// + /// The vocabulary lives here rather than in the echo driver because + /// these are this module's states: a driver spelling them out would + /// be a second place that has to learn about a fifth one. + pub fn command(&self, words: &str) -> String { + let mut words = words.split_whitespace(); + let Some(first) = words.next() else { + return match self.read() { + Some((state, windows)) => format!("usage fixture: {}", describe(&state, &windows)), + None => "usage fixture: unset, so this session meters nothing. \ + `/usage 42` puts up a five-hour window at 42%." + .to_string(), + }; + }; + let rest: Vec<&str> = words.collect(); + let detail = || { + if rest.is_empty() { + "set by /usage".to_string() + } else { + rest.join(" ") + } + }; + let (state, windows) = match first { + "off" | "none" | "clear" => { + *self.said.lock().unwrap() = None; + return "usage fixture cleared: this session meters nothing again".to_string(); + } + "notloggedin" | "logged-out" => (UsageState::NotLoggedIn, Vec::new()), + "unreachable" => (UsageState::Unreachable { detail: detail() }, Vec::new()), + "failed" => (UsageState::Failed { detail: detail() }, Vec::new()), + percent => match percent.parse::() { + Ok(percent) => ( + UsageState::Ok, + fixture_windows(percent.clamp(0.0, 100.0), rest.first().copied()), + ), + Err(_) => { + return format!( + "\"{percent}\" is not one of this fixture's answers. Say a percentage \ + (`/usage 42`, optionally with `90` minutes left, `never` for a window \ + between blocks, or `unreadable` for a reset time that cannot be read), \ + or one of `notloggedin`, `unreachable`, `failed`, `off`." + ); + } + }, + }; + let said = describe(&state, &windows); + *self.said.lock().unwrap() = Some((state, windows)); + format!("usage fixture set: {said}") + } +} + +/// The three windows Claude reports today, invented around one number. +/// +/// Three rather than one because the bar under a session header reads the +/// five-hour window and the dialog behind the button draws all of them, +/// and a fixture with one window leaves half the screen untested. The +/// weekly ones are derived from the same figure so that the worst of them +/// -- which is what colours the button -- is still the one asked for. +fn fixture_windows(percent: f64, reset: Option<&str>) -> Vec { + let resets_at = match reset { + // The state a real response is in between blocks: there is no + // window running, so there is nothing to reset. It is not a + // missing value, and the phone words it differently. + Some("never") | Some("none") => None, + // A timestamp that arrives and cannot be read, which is the one + // case that really is "we could not find out". + Some("unreadable") | Some("bad") => Some("whenever it feels like it".to_string()), + other => Some(reset_in( + other + .and_then(|word| word.parse().ok()) + .unwrap_or(FIXTURE_MINUTES), + )), + }; + vec![ + UsageWindow { + kind: "session".to_string(), + label: "5-hour window".to_string(), + percent, + resets_at: resets_at.clone(), + active: true, + }, + UsageWindow { + kind: "weekly_all".to_string(), + label: "Weekly (all models)".to_string(), + percent: percent / 2.0, + resets_at: resets_at.as_ref().map(|_| reset_in(FIXTURE_MINUTES * 40)), + active: false, + }, + UsageWindow { + kind: "weekly_scoped".to_string(), + label: "Weekly (Echo)".to_string(), + percent: percent / 4.0, + resets_at: resets_at.as_ref().map(|_| reset_in(FIXTURE_MINUTES * 40)), + active: false, + }, + ] +} + +/// `minutes` from now, in the format the real endpoint sends. +fn reset_in(minutes: i64) -> String { + let at = time::OffsetDateTime::now_utc() + time::Duration::minutes(minutes); + at.format(&time::format_description::well_known::Rfc3339) + // Formatting a timestamp cannot fail for any input this builds; + // saying so beats a fixture that silently has no reset time. + .unwrap_or_else(|_| "unformattable".to_string()) +} + +/// One line naming what a fixture is currently claiming, for the reply +/// the echo session writes back. +fn describe(state: &UsageState, windows: &[UsageWindow]) -> String { + match state { + UsageState::Ok => match windows.first() { + Some(window) => format!( + "{}% of the five-hour window, {}", + window.percent, + match &window.resets_at { + Some(at) => format!("resetting at {at}"), + None => "with no reset time (the between-blocks state)".to_string(), + } + ), + None => "no windows at all".to_string(), + }, + UsageState::NotLoggedIn => "nobody is logged in on this machine".to_string(), + UsageState::Unreachable { detail } => format!("machine unreachable ({detail})"), + UsageState::Failed { detail } => format!("the meter failed ({detail})"), + } +} + +/// The fixture, as a provider, so it travels the same route and the same +/// cache as a real meter rather than being spliced in at the screen. +struct EchoUsage { + setup: String, + setup_name: String, + fixture: Fixture, +} + +impl UsageProvider for EchoUsage { + fn name(&self) -> &'static str { + ECHO + } + + fn fetch(&self) -> UsageSnapshot { + let (state, windows) = self + .fixture + .read() + // Only ever built for a fixture that is set; a race with + // `/usage off` between the two reads lands here, and "the + // machine could not be asked" is the honest word for it. + .unwrap_or(( + UsageState::Unreachable { + detail: "the usage fixture was cleared".to_string(), + }, + Vec::new(), + )); + UsageSnapshot { + provider: self.name().to_string(), + setup: self.setup.clone(), + setup_name: self.setup_name.clone(), + state, + windows, + fetched_at: crate::session::now(), + } + } + + /// Read from memory, and set by somebody who is about to look at the + /// screen it changes. + fn poll_interval(&self) -> Duration { + Duration::ZERO + } +} + /// Which paid services a machine can be asked about. /// /// Derived from what the setup says it can run, so a machine with no /// Claude provider is not asked about Claude limits -- it has none, and a -/// row saying so would be a fact about nothing. A second service later -/// adds a branch here and an impl beside [`ClaudeUsage`], not a screen. -fn providers_for(setup: &SetupConfig) -> Vec> { +/// row saying so would be a fact about nothing. +/// +/// Which meter a provider has is [`DriverKind::usage_provider`]'s answer +/// rather than a second match on kinds here, because the phone pairs a +/// session with one of these rows by that same name: two lists that +/// disagree would leave a session looking for a snapshot nothing +/// produces, and nothing on screen could say why. A second service later +/// is a name there and an impl beside [`ClaudeUsage`], not a screen. +fn providers_for(setup: &SetupConfig, fixture: &Fixture) -> Vec> { let mut found: Vec> = Vec::new(); - if setup - .providers - .iter() - .any(|provider| provider.kind == DriverKind::ClaudeCli) - { - found.push(Box::new(ClaudeUsage { - setup: setup.id.clone(), - setup_name: setup.name.clone(), - transport: Transport::for_setup(setup), - })); + for provider in &setup.providers { + let Some(name) = provider.kind.usage_provider() else { + continue; + }; + // A machine offering two Claude providers has one account, not + // two: the meter belongs to the machine and the service, which is + // exactly what the cache is keyed by. + if found.iter().any(|already| already.name() == name) { + continue; + } + match name { + CLAUDE => found.push(Box::new(ClaudeUsage { + setup: setup.id.clone(), + setup_name: setup.name.clone(), + transport: Transport::for_setup(setup), + })), + // Nothing at all until a test has asked for something: an + // echo session costs nothing, so the honest answer is no row + // rather than a row saying zero. + ECHO if fixture.is_set() => found.push(Box::new(EchoUsage { + setup: setup.id.clone(), + setup_name: setup.name.clone(), + fixture: fixture.clone(), + })), + _ => {} + } } found } @@ -330,11 +586,18 @@ type Cached = HashMap<(String, &'static str), (Instant, UsageSnapshot)>; #[derive(Default)] pub struct UsageMonitor { cache: Mutex, + /// The invented meter an echo session can put up; empty unless one + /// has. Shared with the session layer, which is where the command + /// that sets it is typed -- see [`Fixture`]. + fixture: Fixture, } impl UsageMonitor { - pub fn new() -> Self { - Self::default() + pub fn new(fixture: Fixture) -> Self { + Self { + cache: Mutex::new(Cached::new()), + fixture, + } } /// One snapshot per machine that offers a paid service, in the order @@ -346,10 +609,10 @@ impl UsageMonitor { pub fn snapshots(&self, setups: &[SetupConfig]) -> Vec { let mut fresh = Vec::new(); for setup in setups { - for provider in providers_for(setup) { + for provider in providers_for(setup, &self.fixture) { let key = (setup.id.clone(), provider.name()); if let Some((fetched, snapshot)) = self.cache.lock().unwrap().get(&key) - && fetched.elapsed() < MIN_POLL_INTERVAL + && fetched.elapsed() < provider.poll_interval() { // Cached numbers, but the machine's *name* is read // fresh: a rename should show immediately rather than @@ -386,6 +649,7 @@ impl UsageMonitor { #[cfg(test)] mod tests { use super::*; + use crate::config::DriverKind; #[test] fn parses_the_limits_array_defensively() { @@ -423,6 +687,7 @@ mod tests { port: None, identity_file: None, options: vec!["ConnectTimeout=1".to_string()], + models_dir: None, attachments_dir: None, }), providers: vec![crate::config::ProviderConfig { @@ -494,9 +759,62 @@ mod tests { models: vec![], }]; // A machine with no Claude on it has no Claude limits, and a row - // reporting on it would be a fact about nothing. - assert!(providers_for(&echo_only).is_empty()); - assert_eq!(providers_for(&unreachable_setup()).len(), 1); + // reporting on it would be a fact about nothing. Echo included: + // an echo session spends nothing, so until a fixture says + // otherwise there is no meter to report. + let unset = Fixture::new(); + assert!(providers_for(&echo_only, &unset).is_empty()); + assert_eq!(providers_for(&unreachable_setup(), &unset).len(), 1); + + // And with one set, that machine has exactly the invented meter + // -- under the name the session's `usageProvider` will name. + let fixture = Fixture::new(); + fixture.command("42"); + let found = providers_for(&echo_only, &fixture); + assert_eq!(found.len(), 1); + assert_eq!(found[0].name(), ECHO); + assert_eq!(DriverKind::Echo.usage_provider(), Some(ECHO)); + assert_eq!(DriverKind::ClaudeCli.usage_provider(), Some(CLAUDE)); + // A local model costs nothing to run, so it meters nothing. + assert_eq!(DriverKind::LlamaCpp.usage_provider(), None); + } + + /// The states the fixture exists to make reachable, and the one thing + /// it must not do: invent a reset time for a window that has none. + #[test] + fn the_fixture_says_each_state_the_screens_have_to_draw() { + let fixture = Fixture::new(); + assert!(fixture.read().is_none(), "unset until somebody sets it"); + + fixture.command("42 90"); + let (state, windows) = fixture.read().expect("set"); + assert_eq!(state, UsageState::Ok); + assert_eq!(windows[0].kind, "session"); + assert_eq!(windows[0].percent, 42.0); + assert!(windows[0].resets_at.is_some()); + + // Between blocks: no reset time, which the phone words as the + // window not running rather than as a time it could not read. + fixture.command("42 never"); + assert_eq!(fixture.read().expect("set").1[0].resets_at, None); + + fixture.command("unreachable no route to host"); + assert!(matches!( + fixture.read().expect("set").0, + UsageState::Unreachable { detail } if detail == "no route to host", + )); + + fixture.command("off"); + assert!(fixture.read().is_none()); + + // A word it does not know changes nothing and says what it takes. + fixture.command("42"); + let refused = fixture.command("sideways"); + assert!( + refused.contains("not one of this fixture's answers"), + "{refused}" + ); + assert_eq!(fixture.read().expect("still set").1[0].percent, 42.0); } #[test] From e4f0935f98996b0381478eba20b9233646ccc483 Mon Sep 17 00:00:00 2001 From: iris <2+iris@noreply.localhost> Date: Fri, 4 Sep 2026 17:58:11 -0400 Subject: [PATCH 02/12] Keep the second auth test under a subscriber, so the tripwire is not flaky `gates_every_route_and_never_logs_the_token` failed about one full-suite run in ten, on the assertion that a rejection *was* logged. Its sibling ends with an unauthenticated request of its own, made with no subscriber on that thread -- and tracing caches a callsite's interest process-wide the first time it is reached, so whichever test got there first decided whether the warning would ever be recorded. That is the rule already written at the top of "Things that have bitten", applied to one member of a set: the combined gating+logging test exists because of it, and the enrollment test added later did not get it. Twenty runs clean since. Co-Authored-By: Claude Opus 5 --- server/src/auth.rs | 14 ++++++++++++++ 1 file changed, 14 insertions(+) diff --git a/server/src/auth.rs b/server/src/auth.rs index 7719bce..b8dc6ac 100644 --- a/server/src/auth.rs +++ b/server/src/auth.rs @@ -207,6 +207,20 @@ mod tests { /// like any other. #[tokio::test] async fn a_spooled_enrollment_is_adopted_on_first_use() { + // Under a subscriber, like every other exercise of this middleware. + // `tracing` caches a callsite's interest process-wide the first time it + // is reached, so the refusal at the end of this test -- reached with no + // subscriber on this thread -- could cache the rejection warning as + // never-enabled and make the tripwire above see an empty log. That + // failed about one full-suite run in ten, in the test that exists to + // notice a credential leak, which is the worst place for a flake. + let _guard = tracing::subscriber::set_default( + tracing_subscriber::fmt() + .with_max_level(tracing::Level::TRACE) + .with_writer(std::io::sink) + .finish(), + ); + let dir = tempfile::tempdir().expect("tempdir"); let manager = manager_with_token(dir.path(), "first"); let spooled = generate_token(); From 1ff662c7c39d3d827c001737ed747ae89d408757 Mon Sep 17 00:00:00 2001 From: iris <2+iris@noreply.localhost> Date: Fri, 4 Sep 2026 21:23:58 -0400 Subject: [PATCH 03/12] Let a session choose how hard it thinks Output is about an eighth of what a session costs and thinking is nearly all of it -- prose is ~1.5% of output tokens, measured over 27,015 requests of this account's own transcripts -- so the level is the largest saving available short of shortening the conversation itself. Shaped like the working directory rather than like the model: the CLI's only two setting control requests are `set_model` and `set_permission_mode` (checked against the 2.1.258 binary), so `--effort` is read when the process launches and cannot be asked of a running one. `set_session_effort` records the level and stops the process; the next message or Start launches one that has it. That is also why the picker is in the session settings dialog beside Move, and not on the bar beside the model and the mode, which take effect mid-turn. `None` is a level in its own right -- the CLI's own default -- so the picker can return to it, and a blank is normalized to it at the boundary rather than stored as a level the CLI would reject. Offered only where it means something: `DriverKind::takes_effort` reports the capability and the phone leaves the row out entirely, rather than the session-type branch this app does not have anywhere else. A llama session would otherwise get a control whose only effect is stopping its process. Verified on the emulator against the sandbox's fake CLI: the picker sets it, the server reports it, and an echo session's dialog is unchanged. Co-Authored-By: Claude Opus 5 --- PLAN.md | 22 ++- .../src/main/kotlin/com/example/aiapp/Api.kt | 45 +++++++ .../kotlin/com/example/aiapp/SessionScreen.kt | 4 +- .../example/aiapp/SessionSettingsDialog.kt | 70 ++++++++++ server/src/config.rs | 30 +++++ server/src/routes.rs | 33 +++++ server/src/session/claude.rs | 6 + server/src/session/mod.rs | 125 +++++++++++++++++- 8 files changed, 329 insertions(+), 6 deletions(-) diff --git a/PLAN.md b/PLAN.md index d8185bf..a0c24a3 100644 --- a/PLAN.md +++ b/PLAN.md @@ -147,12 +147,26 @@ turn. Spawn: `claude -p --verbose --input-format stream-json --output-format stream-json --permission-mode ` in the chosen working directory, plus -`--model`. Wire-format notes are pinned against CLI 2.1.237 in -`session/claude.rs`'s module doc: permissions need the hidden -`--permission-prompt-tool stdio` flag, AskUserQuestion answers ride -`updatedInput.answers` keyed by question text, and `set_model`/`interrupt` +`--model` and, where one has been chosen, `--effort`. Wire-format notes are +pinned against CLI 2.1.237 in `session/claude.rs`'s module doc: permissions +need the hidden `--permission-prompt-tool stdio` flag, AskUserQuestion answers +ride `updatedInput.answers` keyed by question text, and `set_model`/`interrupt` are control requests. +**The thinking level is settled at launch** (added 2026-09-04, because it is +the largest saving available on a long session: output is about an eighth of +what a session costs and thinking is the bulk of output, against the ~1.5% that +is prose). The CLI's only two setting control requests are `set_model` and +`set_permission_mode` -- checked against the 2.1.258 binary -- so there is no +way to ask a running process to think differently. `set_session_effort` is +therefore shaped like `set_session_cwd` rather than like `set_session_model`: +it records the level and **stops the process**, and the next message or Start +launches one that has it. It lives in the session settings dialog beside the +working directory for that reason, not on the session bar beside the model and +the mode, which do take effect mid-turn. `None` is a level in its own right -- +the CLI's own default -- so the picker can return to it; a level this app named +as the default instead would be this app choosing one. + **`--resume` only ever runs when nothing else has that session open.** That is the rule behind the import refusal, the single `ClaudeDriver::launch` entry point, and the `Exited` correction below; two CLIs on one session file diff --git a/app/androidApp/src/main/kotlin/com/example/aiapp/Api.kt b/app/androidApp/src/main/kotlin/com/example/aiapp/Api.kt index 166a337..8dafacd 100644 --- a/app/androidApp/src/main/kotlin/com/example/aiapp/Api.kt +++ b/app/androidApp/src/main/kotlin/com/example/aiapp/Api.kt @@ -142,6 +142,20 @@ data class SessionSummary( val keepsOwnTranscript: Boolean, /** How much the session asks before acting; null when it was never set. */ val permissionMode: String?, + /** + * How hard the model thinks, or null for the CLI's own default. + * + * Null is a level somebody can choose, not only one to start in -- see [EFFORT_LEVELS]. It is + * reported rather than assumed for the same reason [permissionMode] is. + */ + val effort: String?, + /** + * Whether a thinking level does anything here -- a Claude CLI session, not a llama or echo one. + * + * Asked of the server rather than worked out from the provider's name, because this is a + * property of the driver's *kind* and the phone only has the name. + */ + val takesEffort: Boolean, /** * Whether this continues a session the machine already had, which changes what deleting means. */ @@ -203,6 +217,8 @@ private fun parseSession(session: JSONObject) = title = session.getString("title"), model = session.optString("model").ifEmpty { null }, permissionMode = session.optString("permissionMode").ifEmpty { null }, + effort = session.optString("effort").ifEmpty { null }, + takesEffort = session.optBoolean("takesEffort", false), imported = session.optBoolean("imported", false), notify = session.optBoolean("notify", true), cwd = session.optString("cwd").ifEmpty { null }, @@ -992,6 +1008,35 @@ fun setSessionModel(settings: ServerSettings, sessionId: String, model: String) */ val PERMISSION_MODES = listOf("manual", "acceptEdits", "auto", "bypassPermissions", "plan") +/** + * How hard the model thinks, as `claude --effort` takes them, cheapest first. + * + * Not offered alongside the model and the permission mode on the session's own bar, because it does + * not behave like them: the CLI has a control request for those two and none for this (checked + * against 2.1.258), so a level is settled when the process is launched. Changing it therefore stops + * the process, which is what the working directory beside it in this dialog does, and why it is + * here rather than on a bar whose other controls take effect mid-turn. + */ +val EFFORT_LEVELS = listOf("low", "medium", "high", "xhigh", "max") + +/** What the picker shows, and sends as null, for a session that has chosen no level. */ +const val DEFAULT_EFFORT = "default" + +/** + * Records how hard a session thinks and **stops its process**, since the level is read when the + * process is launched. The next message, or Start, runs one that has it. + * + * [level] is null for the CLI's own default. + */ +fun setSessionEffort(settings: ServerSettings, sessionId: String, level: String?) { + requestFromServer( + settings, + "/sessions/$sessionId/effort", + method = "POST", + jsonBody = JSONObject().put("effort", level ?: JSONObject.NULL).toString(), + ) {} +} + /** Switches how much a running session asks before acting, also in place. */ fun setSessionPermissionMode(settings: ServerSettings, sessionId: String, mode: String) { requestFromServer( diff --git a/app/androidApp/src/main/kotlin/com/example/aiapp/SessionScreen.kt b/app/androidApp/src/main/kotlin/com/example/aiapp/SessionScreen.kt index 6255810..f31cd0e 100644 --- a/app/androidApp/src/main/kotlin/com/example/aiapp/SessionScreen.kt +++ b/app/androidApp/src/main/kotlin/com/example/aiapp/SessionScreen.kt @@ -1846,6 +1846,8 @@ fun SessionScreen( settings = settings, sessionId = summary.id, title = title, + effort = summary.effort.takeIf { summary.takesEffort }, + takesEffort = summary.takesEffort, cachedBytes = cachedBytes, // The purge finishes before the epoch moves, because the relaunched opening effect // reads the same directory and would otherwise draw what is about to be deleted. The @@ -2249,7 +2251,7 @@ private const val ONE_TAP_MS = 250L * session is set to without spending a second line on saying it. */ @Composable -private fun PickerButton(current: String, options: List, onPick: (String) -> Unit) { +fun PickerButton(current: String, options: List, onPick: (String) -> Unit) { var open by remember { mutableStateOf(false) } // When an outside touch last closed the menu. // diff --git a/app/androidApp/src/main/kotlin/com/example/aiapp/SessionSettingsDialog.kt b/app/androidApp/src/main/kotlin/com/example/aiapp/SessionSettingsDialog.kt index 4e116b2..c756642 100644 --- a/app/androidApp/src/main/kotlin/com/example/aiapp/SessionSettingsDialog.kt +++ b/app/androidApp/src/main/kotlin/com/example/aiapp/SessionSettingsDialog.kt @@ -54,6 +54,16 @@ fun SessionSettingsDialog( */ title: String, onRenamed: (String) -> Unit, + /** + * How hard the model thinks, as the session reports it, or null for the CLI's own default. + * + * Taken from the row this dialog was opened over rather than fetched, because unlike the + * notification switch there is nothing else that changes it: the level is this app's to set and + * the server does not resolve it into something else. + */ + effort: String?, + /** Whether a level does anything here; the row is left out entirely where it does not. */ + takesEffort: Boolean, /** * What this phone is holding of the conversation, or null while that is being measured -- see * the Reload row below, which is what would discard it. @@ -69,6 +79,8 @@ fun SessionSettingsDialog( ) { val scope = rememberCoroutineScope() var name by remember(sessionId) { mutableStateOf(title) } + var level by remember(sessionId) { mutableStateOf(effort) } + var effortError by remember { mutableStateOf(null) } var saving by remember { mutableStateOf(false) } var error by remember { mutableStateOf(null) } // Null until the server has been asked. The row this dialog was opened over is a snapshot of @@ -125,6 +137,26 @@ fun SessionSettingsDialog( } } + /** + * Chooses a thinking level, which ends the process the old level was launched with. + * + * Put back if the request is refused, for the reason the notification switch below gives: a + * control that stays where it was put after a refusal is stating something untrue. + */ + fun setEffort(chosen: String?) { + val was = level + level = chosen + effortError = null + scope.launch { + try { + withContext(Dispatchers.IO) { setSessionEffort(settings, sessionId, chosen) } + } catch (e: ApiException) { + level = was + effortError = e.message + } + } + } + // Moved optimistically so the switch answers the finger that moved it, and put back if the // request is refused -- a switch that waits for a round trip reads as broken on a slow tunnel, // and one that stays moved after a refusal lies. @@ -257,6 +289,44 @@ fun SessionSettingsDialog( style = MaterialTheme.typography.bodySmall, ) } + // Left out rather than disabled, the one place this dialog does that: a disabled + // control teaches what the thing can do, and a llama session cannot do this at all + // -- the row would be teaching something false about it. + if (takesEffort) { + Spacer(Modifier.height(8.dp)) + Row( + verticalAlignment = Alignment.CenterVertically, + modifier = Modifier.fillMaxWidth(), + ) { + Text("Thinking", modifier = Modifier.weight(1f)) + PickerButton( + current = level ?: DEFAULT_EFFORT, + // The level the CLI picks for itself is in the list as well as in the + // button, so leaving a level is not a one-way trip -- the same + // correction the model picker carries. + options = listOf(DEFAULT_EFFORT) + EFFORT_LEVELS, + onPick = { chosen -> + setEffort(chosen.takeIf { it != DEFAULT_EFFORT }) + }, + ) + } + // What it costs, said where it is about to be pressed, like Move above: the + // CLI reads the level when it launches and has no control request for + // changing one. + Text( + "Changing this stops the session's process. It starts again with the " + + "next message, or with Start.", + style = MaterialTheme.typography.bodySmall, + color = MaterialTheme.colorScheme.onSurfaceVariant, + ) + effortError?.let { + Text( + it, + color = MaterialTheme.colorScheme.error, + style = MaterialTheme.typography.bodySmall, + ) + } + } Spacer(Modifier.height(8.dp)) Row( verticalAlignment = Alignment.CenterVertically, diff --git a/server/src/config.rs b/server/src/config.rs index 2eae1d8..02d7bc4 100644 --- a/server/src/config.rs +++ b/server/src/config.rs @@ -199,6 +199,24 @@ impl DriverKind { Self::Echo | Self::LlamaCpp => false, } } + + /// Whether a thinking level means anything to this kind, so the phone can + /// offer the control only where it does something. + /// + /// Reported from here rather than decided on the phone, and asked of the + /// *kind* rather than branched on: the alternative is the session-type + /// `if` this app does not have anywhere else. `--effort` is the Claude + /// CLI's; a llama session's sampling is `params`, and echo does not think. + /// + /// It matters more than a control that would simply do nothing, because + /// choosing a level stops the process -- so on a session that cannot use + /// one it is a button whose only effect is the cost. + pub fn takes_effort(self) -> bool { + match self { + Self::ClaudeCli => true, + Self::Echo | Self::LlamaCpp => false, + } + } } #[derive(Debug, Clone, Serialize, Deserialize)] @@ -233,6 +251,17 @@ pub struct SessionConfig { /// the CLI stays the one authority on which modes exist. #[serde(skip_serializing_if = "Option::is_none")] pub permission_mode: Option, + /// How hard the model thinks, passed straight to `--effort`. A string for + /// the same reason `permission_mode` is: the CLI owns which levels exist. + /// + /// Unlike the model and the mode, there is no control request that changes + /// one -- checked against 2.1.258, whose only two are `set_model` and + /// `set_permission_mode` -- so this is settled at launch and `None` means + /// whatever the CLI's own default is. That is a state the phone has to be + /// able to *choose*, not just start in, which is why it is an option + /// rather than a level with a default written here. + #[serde(skip_serializing_if = "Option::is_none")] + pub effort: Option, /// Settings the driver interprets, chosen at spawn. /// /// Deliberately untyped: what a temperature or a context size means is the @@ -441,6 +470,7 @@ mod tests { model: None, cwd: None, permission_mode: None, + effort: None, params: BTreeMap::new(), notify: true, throwaway: false, diff --git a/server/src/routes.rs b/server/src/routes.rs index 541f2f6..7ce44bd 100644 --- a/server/src/routes.rs +++ b/server/src/routes.rs @@ -42,6 +42,8 @@ //! which starts again in the new one //! POST /sessions/{id}/model {model} //! POST /sessions/{id}/permission-mode {permissionMode} +//! POST /sessions/{id}/effort {effort} -- null for the CLI's default; +//! settled at launch, so this stops the process //! POST /sessions/{id}/command {text} -- /compact, /clear, /rename x, or the dialect's own //! (starts the process first if it has exited) //! POST /sessions/{id}/compact @@ -134,6 +136,7 @@ pub fn router(manager: Arc) -> Router { .route("/sessions/{id}/cwd", post(set_cwd)) .route("/sessions/{id}/model", post(set_model)) .route("/sessions/{id}/permission-mode", post(set_permission_mode)) + .route("/sessions/{id}/effort", post(set_effort)) .route("/sessions/{id}/notify", post(set_notify)) .route("/notifications", get(notifications)) .route("/sessions/{id}/compact", post(compact)) @@ -644,6 +647,8 @@ struct SpawnRequest { cwd: Option, #[serde(default)] permission_mode: Option, + #[serde(default)] + effort: Option, /// Whatever the chosen driver understands -- llama.cpp's context size and /// sampling. Opaque here on purpose: see `SessionConfig::params`. #[serde(default)] @@ -831,6 +836,7 @@ async fn start_import( model: body.model.clone(), cwd: None, permission_mode: body.permission_mode.clone(), + effort: body.effort.clone(), params: std::collections::BTreeMap::new(), import: Some(session.clone()), }; @@ -863,6 +869,8 @@ struct ImportRequest { model: Option, #[serde(default)] permission_mode: Option, + #[serde(default)] + effort: Option, } /// Runs `work` on the server, marked as in flight for as long as it takes. @@ -1013,6 +1021,7 @@ async fn spawn(manager: &Arc, body: SpawnRequest) -> Result, +} + +/// Records how hard this session thinks, and stops the process so the next one +/// is launched with it -- `--effort` has no control request behind it. See +/// [`SessionManager::set_session_effort`]. +async fn set_effort( + State(manager): State>, + UrlPath(id): UrlPath, + axum::Json(body): axum::Json, +) -> Result { + manager + .set_session_effort(&id, body.effort.as_deref()) + .map_err(bad_request)?; + Ok(StatusCode::NO_CONTENT) +} + async fn set_permission_mode( State(manager): State>, UrlPath(id): UrlPath, diff --git a/server/src/session/claude.rs b/server/src/session/claude.rs index 5adfc93..5bffaaa 100644 --- a/server/src/session/claude.rs +++ b/server/src/session/claude.rs @@ -351,6 +351,12 @@ impl ClaudeDriver { if let Some(mode) = &meta.permission_mode { push("--permission-mode", mode); } + // Launch-only: see `SessionConfig::effort`. Omitted entirely when + // unset, so the CLI's own default is what an unchosen session gets + // rather than a level this app decided to call the default. + if let Some(effort) = &meta.effort { + push("--effort", effort); + } // Named at birth, so this session is the same session in the CLI's own // picker and in what other agents see. // diff --git a/server/src/session/mod.rs b/server/src/session/mod.rs index c7b017a..4c6cf81 100644 --- a/server/src/session/mod.rs +++ b/server/src/session/mod.rs @@ -63,6 +63,8 @@ pub struct SpawnSpec { pub model: Option, pub cwd: Option, pub permission_mode: Option, + /// See `SessionConfig::effort`. + pub effort: Option, /// Driver-interpreted settings; see `SessionConfig::params`. pub params: std::collections::BTreeMap, } @@ -118,6 +120,16 @@ pub struct SessionInfo { /// were confirming. #[serde(skip_serializing_if = "Option::is_none")] pub permission_mode: Option, + /// How hard it thinks; see `SessionConfig::effort`. Reported for the same + /// reason the mode is, and absent where nothing has been chosen -- which + /// the phone draws as the CLI's default rather than as a level. + #[serde(skip_serializing_if = "Option::is_none")] + pub effort: Option, + /// Whether a level means anything here; see `DriverKind::takes_effort`. + /// Reported beside the level because absent-and-irrelevant and + /// absent-and-unchosen are different answers, and only one of them is a + /// control worth drawing. + pub takes_effort: bool, /// Whether this session continues one the machine already had. /// Reported because it changes what deleting *means*: an imported /// session's real transcript belongs to the CLI and survives, so @@ -435,6 +447,7 @@ impl LiveSession { &self, setup_name: &str, cwd: Option<&Path>, + effort: Option<&str>, imported: bool, kind: Option, ) -> SessionInfo { @@ -446,6 +459,11 @@ impl LiveSession { title: self.shared.title.lock().unwrap().clone(), model: self.shared.model.lock().unwrap().clone(), permission_mode: self.shared.permission_mode.lock().unwrap().clone(), + // From the config rather than from `shared`, like the cwd beside + // it: neither can change under a running process, so there is no + // live value for one to disagree with. + effort: effort.map(str::to_string), + takes_effort: kind.is_some_and(DriverKind::takes_effort), context_tokens: *self.shared.context_tokens.lock().unwrap(), notify: *self.shared.notify.lock().unwrap(), max_image_edge: kind.and_then(DriverKind::max_image_edge), @@ -902,6 +920,7 @@ impl SessionManager { Some(session) => session.info( label_of(&inner.config, &meta.setup), meta.cwd.as_deref(), + meta.effort.as_deref(), import::read_cursor(&self.data_dir.join(&meta.id)).is_some(), kind_of(&inner.config, &meta.setup, &meta.provider), ), @@ -913,6 +932,9 @@ impl SessionManager { title: meta.title.clone(), model: meta.model.clone(), permission_mode: meta.permission_mode.clone(), + effort: meta.effort.clone(), + takes_effort: kind_of(&inner.config, &meta.setup, &meta.provider) + .is_some_and(DriverKind::takes_effort), context_tokens: None, max_image_edge: kind_of(&inner.config, &meta.setup, &meta.provider) .and_then(DriverKind::max_image_edge), @@ -1016,6 +1038,7 @@ impl SessionManager { model: spec.model, cwd: spec.cwd, permission_mode: spec.permission_mode, + effort: spec.effort, params: spec.params, // On by default. Not offered at spawn: a session's first turn // is exactly the one somebody is waiting for. @@ -1050,6 +1073,7 @@ impl SessionManager { let info = session.info( &setup.name, session.meta.cwd.as_deref(), + session.meta.effort.as_deref(), import::read_cursor(&self.data_dir.join(&id)).is_some(), Some(provider.kind), ); @@ -1247,6 +1271,47 @@ impl SessionManager { Ok(()) } + /// Sets how hard this session's model thinks. + /// + /// Shaped like [`SessionManager::set_session_cwd`] rather than like + /// [`SessionManager::set_session_model`], because `--effort` is a launch + /// flag with no control request behind it: the running process cannot be + /// asked, so the choice is recorded and the process ended, and the next + /// thing said to the session starts one that has it. Announcing it as a + /// settled change instead would put a level on the phone that the process + /// still running underneath was not using. + /// + /// `None` clears it, which is a level in its own right -- the CLI's own + /// default -- and the reason this takes an option rather than a string. + pub fn set_session_effort(&self, id: &str, effort: Option<&str>) -> Result<()> { + let effort = effort.map(str::trim).filter(|level| !level.is_empty()); + { + let mut inner = self.inner.write().unwrap(); + if !inner.config.sessions.iter().any(|meta| meta.id == id) { + bail!("no session {id}"); + } + let mut candidate = inner.config.clone(); + for meta in candidate.sessions.iter_mut().filter(|meta| meta.id == id) { + meta.effort = effort.map(str::to_string); + } + candidate.save(&self.config_path)?; + inner.config = candidate; + } + // Saved before the process is touched, for the reason `set_session_cwd` + // gives: a process that will not stop must not leave the session + // recorded as something nothing agrees with. + let dir = self.data_dir.join(id); + if let Some(record) = process::live(&dir) { + tracing::info!( + "session {id} effort now {} -- stopping pid {}", + effort.unwrap_or("default"), + record.pid + ); + process::stop(&record, process::STOP_GRACE); + } + Ok(()) + } + /// Ends this session's process, leaving the session -- its transcript, /// its place in the list, everything a phone is watching -- exactly /// where it is. [`SessionManager::start_session`] is the way back. @@ -2256,6 +2321,7 @@ mod tests { model: None, cwd: None, permission_mode: None, + effort: None, } } @@ -2558,7 +2624,10 @@ mod tests { assert_eq!(first.session_id, info.id); // The title travels with it, because the phone may have no screen // open to look one up on. - assert_eq!(first.title, session.info("m", None, false, None).title); + assert_eq!( + first.title, + session.info("m", None, None, false, None).title + ); manager.set_session_notify(&info.id, false).expect("off"); // Subscribed before the message, or the turn can finish in the gap @@ -3411,6 +3480,60 @@ mod tests { std::fs::write(path, rewritten).expect("write transcript"); } + /// A thinking level is stored and the process **ended**, because `--effort` + /// is read when the CLI launches and has no control request behind it. A + /// session left running would go on thinking at the old level underneath a + /// phone showing the new one -- the failure this app has already had once + /// with the model, and the one a stop makes impossible rather than + /// unlikely. + /// + /// Clearing it back to the CLI's own default is exercised too: that is a + /// level somebody can choose, not only one to start in, so a picker that + /// could not return to it would make leaving a level a one-way trip. + #[tokio::test] + async fn choosing_a_thinking_level_stores_it_and_ends_the_process() { + let dir = tempfile::tempdir().expect("tempdir"); + let config_path = dir.path().join("config.ron"); + let data_dir = dir.path().join("sessions"); + // The stand-in CLI rather than the echo driver: what is under test is + // that a *process* is ended, and an echo session has none to end. + let provider = seed_stand_in_cli(&config_path, dir.path()); + let manager = SessionManager::new( + config_path.clone(), + data_dir.clone(), + data_dir.join("models"), + ) + .expect("manager"); + let info = manager + .spawn_session(stand_in_spec(&provider)) + .expect("spawn"); + assert_eq!(info.effort, None, "nothing is chosen at spawn"); + let record = process::live(&data_dir.join(&info.id)).expect("the session has a process"); + + manager + .set_session_effort(&info.id, Some("low")) + .expect("store the level"); + assert_eq!( + manager.sessions()[0].effort.as_deref(), + Some("low"), + "stored, so the next start is launched with it" + ); + process::wait_gone(&[record], process::STOP_GRACE); + + // Blank is the same answer as unchosen; normalized here so a caller + // clearing the field cannot store a level the CLI would reject. + manager + .set_session_effort(&info.id, Some(" ")) + .expect("clear the level"); + assert_eq!( + manager.sessions()[0].effort, + None, + "the CLI's own default has to be reachable again" + ); + + manager.delete_session(&info.id).expect("delete"); + } + /// A setting changed on a session with nothing running is recorded as the /// session's own, rather than refused because there is no driver. The /// config already took it, so the refusal was about the driver while From 4821a02bd3d3db03b776ee4c54d9f32195ee9196 Mon Sep 17 00:00:00 2001 From: iris <2+iris@noreply.localhost> Date: Fri, 4 Sep 2026 21:42:05 -0400 Subject: [PATCH 04/12] Default thinking level for new sessions, and move the rigs out of AGENTS.md `Config::default_effort` is what a session starts at when nothing chose one, applied in `spawn_session` rather than filled in by the spawn screen so it holds for an import and a bare API call too. It is set by the spawn screen's own picker, whose label says so: one control, where new sessions are made, rather than a settings page for a single value. Not on a provider, because providers are discovered and the next rediscovery would erase it; not on the phone, because a second device would then spawn at a level nobody there chose. `GET`/`POST /defaults` carry it as a struct, so the permission mode -- still hardcoded to `auto` on the spawn screen -- can move there later without a second route. Only drivers that read a level are given one: an echo session was storing a `--effort` it never passes to anything, which is a config file answering a question about itself wrongly. Separately, `AGENTS.md` is 35 KB sent with every request in this repo, and 12 KB of it was rigs and reference measurements that only matter once you are running one. Those are the `ai-app-rigs` skill now -- the same text, still the only copy, read when the work touches it. 35,198 -> 20,813 chars. Verified on the emulator against the sandbox: the spawn screen pre-fills from the server, picking `low` spawned a session at `low` and left `/defaults` set to it, and an echo session spawned afterwards took no level at all. Co-Authored-By: Claude Opus 5 --- .claude/skills/ai-app-rigs/SKILL.md | 244 ++++++++++++++++++ AGENTS.md | 243 +---------------- PLAN.md | 11 + .../src/main/kotlin/com/example/aiapp/Api.kt | 24 ++ .../kotlin/com/example/aiapp/SpawnScreen.kt | 32 +++ server/src/config.rs | 13 + server/src/routes.rs | 33 +++ server/src/session/mod.rs | 103 +++++++- 8 files changed, 468 insertions(+), 235 deletions(-) create mode 100644 .claude/skills/ai-app-rigs/SKILL.md diff --git a/.claude/skills/ai-app-rigs/SKILL.md b/.claude/skills/ai-app-rigs/SKILL.md new file mode 100644 index 0000000..0c7ecbd --- /dev/null +++ b/.claude/skills/ai-app-rigs/SKILL.md @@ -0,0 +1,244 @@ +--- +name: ai-app-rigs +description: ai-app's test rigs, harness scripts and reference measurements - ui-sandbox.sh, debug-transcript.sh, transcript-bench.sh, stream-bench.sh, trace-draw.sh, the /usage fixture vocabulary, the fake CLI, the rule that no UI-driving script may tap a coordinate, how to test llama.cpp and ssh on this machine, how importing behaves, and the scroll/stream/explorer numbers not worth re-measuring. Read before running or writing a benchmark, driving the app's UI from a script, exercising the session lifecycle, testing a llama or remote session, or touching the import screen. +--- + +# ai-app: rigs, harnesses and measurements + +Moved out of `AGENTS.md` on 2026-09-04 so it is read when it is relevant +rather than sent with every request in this repo -- it was 12 KB of the 35 KB +that file cost on every one. Unchanged in the move, and still the only copy. + +## The rigs + +Each exists because something was invisible without it. + +- **`app/ui-sandbox.sh`** — a second `ai-server` with its own `$HOME`, config + and data directory, holding eight invented Claude Code transcripts and a + `claude` that is two lines of shell. **That isolation is the point**: the + import screen lists whatever is in `~/.claude/projects`, which in this VM is + real agent transcripts, so exercising *delete* against the ordinary server + deletes somebody's conversation and exercising *import* starts a real + `--resume` on the owner's account. + Its port and root derive from the checkout's name, so two checkouts' + sandboxes cannot reach each other, and its token is generated once into + `~/.config/ai-app/sandbox-token` and carried across restarts along with any + the enrolment flow appended — so the emulator app is enrolled **once** (the + start banner prints the command) and stays enrolled. It shares the real TLS + certificates, because the installed APK pins that CA. + Driving verbs, so none of this is re-derived per session: + `./ui-sandbox.sh spawn [title]` (an echo session, prints its id), + `./ui-sandbox.sh send SID text|@file`, and + `./ui-sandbox.sh api /path [curl args]`. + `./ui-sandbox.sh keep` restarts the server without wiping the sessions and + enrolment already there — for when the fixture under test was expensive to + build; plain `start` wipes them, which is right for the list-screen + fixtures and wrong for that. + It passes `--delay` by default, and `AI_SANDBOX_BIG_MB` puts one large + transcript among the small ones while `AI_SANDBOX_SPAWN_DELAY` makes the + fake CLI slow to start. Both exist because operations that finish in + milliseconds have states on the way that nothing can observe, and an + unobservable state is one where broken and working look identical. + It also builds a fixture tree at the sandbox home's `~/files` for the + explorer, holding the states otherwise only reachable by finding a real + machine in one: an empty directory, a name with a tab and one with an + apostrophe, a binary file, one over `FILE_LIMIT`, one `chmod 000`, a + symlink to a directory and a broken one, a source file per language, and + the three sizes the limits were measured against (`edit-32k.rs`, + `edit-128k.rs`, `big-source.rs`). Point a session at it with + `./ui-sandbox.sh api /sessions//cwd -X POST -H 'content-type: application/json' -d '{"cwd":"~/files"}'`. + The explorer's 409 is produced by editing the file on the machine + (`printf … > file`) between pressing the pencil and pressing save. +- **`app/debug-transcript.sh`** — a real conversation on the emulator. The + echo driver is the right rig for most things and the wrong one for anything + whose cost scales with what was actually written: a real reply is longer, + is real markdown, and carries tool calls whose input and output are + kilobytes. Two faults were invisible until a real transcript was loaded — a + page of history landing mid-fling threw the reader back to the newest end, + and parsing one real reply took 51ms against 4.6ms for a synthetic one. + `-b` takes the biggest conversation on the machine rather than the newest, + which is what a scrolling test wants; `--stop` takes it down. + It copies the transcript into `/tmp` and gives the server a `HOME` of its + own, so the import can only see the copy — importing spawns `claude + --resume`, and against the real file that is a second CLI writing to a + conversation somebody may still be in. **A transcript never goes in this + repository**: they hold whatever was said, read and written in that + session, and `~/repos` is shared with the host besides. +- **`/usage` in an echo session puts up an invented meter**, which is how the + rate-limit screens' states are reached without spending quota: `/usage 42`, + `/usage 95 20` (minutes left), `/usage 42 never` (the between-blocks window + with no reset time), `/usage 42 unreadable`, `/usage notloggedin`, + `/usage unreachable`, `/usage failed`, `/usage off`. The vocabulary is + `usage::Fixture`'s, since those are its states. With none set an echo + session meters nothing, which is the ordinary case and draws no bar. +- **A fake CLI exercises the process lifecycle without a token.** Point a + `claude_cli` provider's `command` at a two-line script — `#!/bin/sh` and + `cat > /dev/null` — and it behaves the way the lifecycle code cares about: + it holds the fifo open, records a real pid, writes nothing, and dies on a + signal. So adopt, stop, restart and start are all drivable without a real + `--resume` and without spending a turn on somebody's account. Reach for + this when what is under test is *whether a process is running*, and for + `debug-transcript.sh` when it is *what the transcript draws*. +- **`app/transcript-bench.sh`** is the standard scroll measurement: it opens + the first session (or `-k` keeps the current screen), scrolls a fixed + gesture loop, and prints the app's render report — the same one the in-app + copy button produces, whose `on screen:` line names what the viewport was + holding. Compare two runs with the same gestures; the emulator's absolute + frame times transfer nothing, the report's accounting does. Run it either + side of any change under `Markdown*.kt`, `Transcript*.kt` or + `SessionScreen.kt`'s list, and put the report in the commit. The numbers + that move first are the worst `record: one block`, the reparse mean while + streaming, and the draw phase's accounting line. +- **`app/stream-bench.sh [-k] FILE`** is that measurement for a reply still + arriving. It taps "Jump to latest" so the list is pinned to the newest end, + resets the report, sends FILE, waits for the transcript to stop growing, + and prints. Both of those are corrections to a first version that measured + nothing: a transcript parked further back never redraws while a reply + streams into it, and a session is idle at *both* ends of a turn, so polling + for idle answers before the turn has started. +- **`app/trace-draw.sh`** names what a scrolling frame spends inside the + framework, from `atrace` text output with no trace processor needed. It is + how the cost of a layout node per link was attributed to the framework + rather than guessed at. + +### Driving the UI + +**No script that drives this app's UI presses a coordinate.** Every control +is found by the name it already carries for assistive technology — +`ui-trace record --do "tap 'Session settings'"` — which resolves the label +against the screen at the moment of the gesture and fails the whole run when +it is not there. `app/bench-lib.sh` is what the bench scripts share for it. A +coordinate is a position measured once by hand, and anything that moves the +control makes the tap land on whatever now sits there — the bench then +reports a number that was never measured, which reads exactly like a result. +Both bench scripts pressed the render report at `tap 723 205` until that +button moved into the session settings dialog on 2026-09-03. The check that +none has crept back: + + grep -n "tap [0-9]" app/*.sh + +Swipes are still coordinates, deliberately: a gesture across a scrolling area +is a distance rather than a control. + +**Two traps in the emulator bench loop**, each of which cost a run. +`adb shell pm clear` removes the enrolment and the notification permission +along with the saved anchors, so the next run measures a permission dialog — +re-enrol with the command `ui-sandbox.sh` prints, and +`pm grant … POST_NOTIFICATIONS`. And a saved scroll anchor is per session id, +so the only way two builds start a scroll from the same place is a *fresh +session for each*. + +**The emulator is `~/repos/emulator-tools`' business, not this repo's.** +`emu up` creates and boots the AVD named after this checkout — whatever `emu +name` prints, never a name typed out here, since this file is the same in +every clone. `run-android.sh` is that plus a build and an install. The `adb` +on `PATH` after sourcing `android-env.sh` is that repo's wrapper, which fills +in `-s` from the same rule. Gradle does not go through it, so a Gradle init +script from `emulator-tools` runs `emu check` before `installDebug`, +`uninstallDebug` and `connectedAndroidTest` and fails rather than fanning out +to every attached device; when it refuses, say which device you mean at the +moment you use it — `ANDROID_SERIAL=$(emu serial) ./gradlew …`. + +### Testing llama.cpp and ssh here + +**Both are set up here as of 2026-09-04** and need nothing typed. The +prebuilt CPU llama.cpp lives outside the repo at `~/.local/opt/llama.cpp` +(the 15 MB `ubuntu-x64` release asset) and is symlinked as +`/usr/local/bin/llama-server`, which is what makes **discovery find it over +ssh**: `~/.local/bin` is not on the PATH a non-interactive ssh session gets. +It resolves its own libraries through `$ORIGIN`, so no `LD_LIBRARY_PATH` is +needed. One model is downloaded — `unsloth/Qwen3-0.6B-GGUF/Qwen3-0.6B-Q8_0.gguf`, +639 MB under `~/.local/share/ai-app/models` — and answers at usable speed on +this VM's 8 cores. **Do not test with a 2-bit quant**: the +IQ2_XXS of that model produces fluent nonsense, which reads exactly like a +broken driver — `llama-cli` produces the same from the file directly, which +is how to tell the two apart in a hurry. + +There is no second machine, so **ssh this VM to itself**. That is set up +too: the key is `~/.config/ai-app/ssh-self` (its public half is in +`~/.ssh/authorized_keys`, labelled removable), and the real config carries a +setup called **"this vm over ssh"** — `bob@127.0.0.1` with that +`identityFile` plus +`options: ["StrictHostKeyChecking=no", "UserKnownHostsFile=/tmp/ai-app-known-hosts"]` +so it touches nothing real — offering `claude-cli` and `llama-cpp`. It is the +whole rig for "does a remote llama session work", since the far machine is +this one and the model file is the same file. For a throwaway setup of your +own, point a provider's `command` at something harmless like `/bin/echo` +rather than at `claude`: the transport is what is under test, the process +exiting immediately is the signal, and it costs no tokens. The remote login +shell here is **fish**; the +remote script and `ssh.rs`'s POSIX quoting happen to mean the same thing in +both, but that is luck rather than design, and a shell that is neither is the +thing to suspect first if a remote spawn ever mangles an argument. + +## Importing + +The import list reports each session's **size as well as its line count**, +because the two disagree in the way that matters: these transcripts embed +screenshots as base64, so one line can be a megabyte. On this machine a 69 MB +session has 3,427 lines and a 44 MB one has 6,792 — nothing about a line +count tells you what continuing a session will cost. Shown, not warned about; +importing a large session is a choice somebody is entitled to make. + +**Never import a Claude Code session that is open in a terminal.** The app +refuses it — see PLAN.md for the incident that made that a refusal rather +than a warning. + +**One Claude Code session id can name two files, and the listing offers it +once.** Resuming from a different working directory makes the CLI write a +second transcript with the same id under that directory's project folder — an +ordinary state of a machine, not corruption. Everything downstream addresses +a session by id, and the phone keyed its list on it, so two rows sharing one +**closed the app** on a Compose duplicate-key throw. `parse_listing` keeps +the copy with the most lines, because the other is usually a few-hundred-byte +stub and is often the *newer* of the two, so recency is the wrong key. +Deleting removes every copy rather than the first, or the row came back after +a delete that reported success. The phone's half is `uniqueItems`, which +every list keyed on a server-chosen id goes through: a repeat there must +never be able to close the app, whatever produced it. + +**Deleting a session offers to take the machine's own transcript with it** — +`DELETE /sessions/{id}?deleteForeign=true`, behind a switch in the +confirmation, and only where the driver keeps a record of its own +(`keepsOwnTranscript`, which today means Claude Code). Off by default, +because leaving that copy is what makes an ordinary delete recoverable — and +the dialog's paragraph is rewritten when it is on rather than appended to, +since the sentence promising the conversation "should still be there to +import again" is exactly the one the switch makes false. The server deletes +the machine's copy *first*, so a machine it cannot reach leaves the session +where it was instead of half-deleted. + +## Measurements worth not re-taking + +- **What the transcript screen costs to scroll.** Taken 2026-08-30 on the GPU + emulator against a real imported transcript with the server at + `--delay 120`. Settled and flinging fast, both into fresh history and back + through rows already drawn: **5.2–5.9% janky frames, 99th percentile + 29–32ms, 0–2 slow UI-thread frames.** The stock Settings app on the same + device is 3.3% and 38ms, so this is at the platform floor. The number that + is *not* at the floor is the first few seconds after opening a session, + where every row on the way is being composed for the first time; that is + inherent to a lazy list and it is why a measurement taken before the screen + settles reads three times worse. **Settle first, then reset `gfxinfo`.** +- **The reset path is not reachable by reopening a session.** Measured + 2026-09-04 against a session streaming at 20 events a second: reopening one + with an anchor 1,800 events back connects **87–119 events behind**, well + under `CATCH_UP_LIMIT`'s 200, because the restore is two requests — the + opening page, then one span covering the whole distance. To exercise the + reset at all you have to lower `CATCH_UP_LIMIT` in a throwaway build; at 5 + the app takes the reset on a live connection, clears, refills and carries + on without reconnecting. +- **The session screen's stream survives backgrounding here** — 20 seconds at + the launcher while 415 events were produced brought no reconnect at all, + which is not what the comment above that loop expects, and is most likely + this emulator being headless rather than the phone's behaviour. +- **Reopening a cached session costs one request for one event** (the probe), + and scrolling the whole conversation back costs nothing more; a cold open + of the same 500-event session is two pages, 100 events. Measured + 2026-09-04 on the emulator against the sandbox. +- **Reading is cheap and editing is not.** The viewer handles a 1 MiB, + 28,000-line file because it draws one row per line; the editor is one + `BasicTextField`, which costs two seconds a frame at 128 kB and stops the + app at 1 MiB, so `EDIT_LIMIT` caps it at 32 kB with the reason said on + screen. If you make the editor faster, that number is what to move. + EXPLORER.md's "What the measurements said" has the rest. diff --git a/AGENTS.md b/AGENTS.md index cd01ade..f94676b 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -9,7 +9,15 @@ between them. its rationale, and what was rejected. Read it before changing anything structural, and update it in place when a decision changes rather than letting this file and the plan become two versions of the truth. This file is -the working notes layer: layout, commands, rigs, and things that have bitten. +the working notes layer: layout, commands, and things that have bitten. + +**The rigs are the `ai-app-rigs` skill** — the sandbox and bench scripts, the +rule that no UI-driving script may tap a coordinate, how to test llama.cpp and +ssh here, how importing behaves, and the measurements not worth re-taking. +They moved there on 2026-09-04 because they are 12 KB that only matter once +you are actually running one, and this file is sent with every request. Read +it before writing or running a benchmark, driving the UI from a script, or +touching the import screen. The central design point, worth not undoing by accident: **a session is a child process, translated into one common event model.** A new session type @@ -151,168 +159,6 @@ two icon buttons the same width without either being given one — and why genuine handshake against 10.66.0.1 with pinned TLS, no router or phone involved. That is how to verify the wg0-only posture. -## The rigs - -Each exists because something was invisible without it. - -- **`app/ui-sandbox.sh`** — a second `ai-server` with its own `$HOME`, config - and data directory, holding eight invented Claude Code transcripts and a - `claude` that is two lines of shell. **That isolation is the point**: the - import screen lists whatever is in `~/.claude/projects`, which in this VM is - real agent transcripts, so exercising *delete* against the ordinary server - deletes somebody's conversation and exercising *import* starts a real - `--resume` on the owner's account. - Its port and root derive from the checkout's name, so two checkouts' - sandboxes cannot reach each other, and its token is generated once into - `~/.config/ai-app/sandbox-token` and carried across restarts along with any - the enrolment flow appended — so the emulator app is enrolled **once** (the - start banner prints the command) and stays enrolled. It shares the real TLS - certificates, because the installed APK pins that CA. - Driving verbs, so none of this is re-derived per session: - `./ui-sandbox.sh spawn [title]` (an echo session, prints its id), - `./ui-sandbox.sh send SID text|@file`, and - `./ui-sandbox.sh api /path [curl args]`. - `./ui-sandbox.sh keep` restarts the server without wiping the sessions and - enrolment already there — for when the fixture under test was expensive to - build; plain `start` wipes them, which is right for the list-screen - fixtures and wrong for that. - It passes `--delay` by default, and `AI_SANDBOX_BIG_MB` puts one large - transcript among the small ones while `AI_SANDBOX_SPAWN_DELAY` makes the - fake CLI slow to start. Both exist because operations that finish in - milliseconds have states on the way that nothing can observe, and an - unobservable state is one where broken and working look identical. - It also builds a fixture tree at the sandbox home's `~/files` for the - explorer, holding the states otherwise only reachable by finding a real - machine in one: an empty directory, a name with a tab and one with an - apostrophe, a binary file, one over `FILE_LIMIT`, one `chmod 000`, a - symlink to a directory and a broken one, a source file per language, and - the three sizes the limits were measured against (`edit-32k.rs`, - `edit-128k.rs`, `big-source.rs`). Point a session at it with - `./ui-sandbox.sh api /sessions//cwd -X POST -H 'content-type: application/json' -d '{"cwd":"~/files"}'`. - The explorer's 409 is produced by editing the file on the machine - (`printf … > file`) between pressing the pencil and pressing save. -- **`app/debug-transcript.sh`** — a real conversation on the emulator. The - echo driver is the right rig for most things and the wrong one for anything - whose cost scales with what was actually written: a real reply is longer, - is real markdown, and carries tool calls whose input and output are - kilobytes. Two faults were invisible until a real transcript was loaded — a - page of history landing mid-fling threw the reader back to the newest end, - and parsing one real reply took 51ms against 4.6ms for a synthetic one. - `-b` takes the biggest conversation on the machine rather than the newest, - which is what a scrolling test wants; `--stop` takes it down. - It copies the transcript into `/tmp` and gives the server a `HOME` of its - own, so the import can only see the copy — importing spawns `claude - --resume`, and against the real file that is a second CLI writing to a - conversation somebody may still be in. **A transcript never goes in this - repository**: they hold whatever was said, read and written in that - session, and `~/repos` is shared with the host besides. -- **`/usage` in an echo session puts up an invented meter**, which is how the - rate-limit screens' states are reached without spending quota: `/usage 42`, - `/usage 95 20` (minutes left), `/usage 42 never` (the between-blocks window - with no reset time), `/usage 42 unreadable`, `/usage notloggedin`, - `/usage unreachable`, `/usage failed`, `/usage off`. The vocabulary is - `usage::Fixture`'s, since those are its states. With none set an echo - session meters nothing, which is the ordinary case and draws no bar. -- **A fake CLI exercises the process lifecycle without a token.** Point a - `claude_cli` provider's `command` at a two-line script — `#!/bin/sh` and - `cat > /dev/null` — and it behaves the way the lifecycle code cares about: - it holds the fifo open, records a real pid, writes nothing, and dies on a - signal. So adopt, stop, restart and start are all drivable without a real - `--resume` and without spending a turn on somebody's account. Reach for - this when what is under test is *whether a process is running*, and for - `debug-transcript.sh` when it is *what the transcript draws*. -- **`app/transcript-bench.sh`** is the standard scroll measurement: it opens - the first session (or `-k` keeps the current screen), scrolls a fixed - gesture loop, and prints the app's render report — the same one the in-app - copy button produces, whose `on screen:` line names what the viewport was - holding. Compare two runs with the same gestures; the emulator's absolute - frame times transfer nothing, the report's accounting does. Run it either - side of any change under `Markdown*.kt`, `Transcript*.kt` or - `SessionScreen.kt`'s list, and put the report in the commit. The numbers - that move first are the worst `record: one block`, the reparse mean while - streaming, and the draw phase's accounting line. -- **`app/stream-bench.sh [-k] FILE`** is that measurement for a reply still - arriving. It taps "Jump to latest" so the list is pinned to the newest end, - resets the report, sends FILE, waits for the transcript to stop growing, - and prints. Both of those are corrections to a first version that measured - nothing: a transcript parked further back never redraws while a reply - streams into it, and a session is idle at *both* ends of a turn, so polling - for idle answers before the turn has started. -- **`app/trace-draw.sh`** names what a scrolling frame spends inside the - framework, from `atrace` text output with no trace processor needed. It is - how the cost of a layout node per link was attributed to the framework - rather than guessed at. - -### Driving the UI - -**No script that drives this app's UI presses a coordinate.** Every control -is found by the name it already carries for assistive technology — -`ui-trace record --do "tap 'Session settings'"` — which resolves the label -against the screen at the moment of the gesture and fails the whole run when -it is not there. `app/bench-lib.sh` is what the bench scripts share for it. A -coordinate is a position measured once by hand, and anything that moves the -control makes the tap land on whatever now sits there — the bench then -reports a number that was never measured, which reads exactly like a result. -Both bench scripts pressed the render report at `tap 723 205` until that -button moved into the session settings dialog on 2026-09-03. The check that -none has crept back: - - grep -n "tap [0-9]" app/*.sh - -Swipes are still coordinates, deliberately: a gesture across a scrolling area -is a distance rather than a control. - -**Two traps in the emulator bench loop**, each of which cost a run. -`adb shell pm clear` removes the enrolment and the notification permission -along with the saved anchors, so the next run measures a permission dialog — -re-enrol with the command `ui-sandbox.sh` prints, and -`pm grant … POST_NOTIFICATIONS`. And a saved scroll anchor is per session id, -so the only way two builds start a scroll from the same place is a *fresh -session for each*. - -**The emulator is `~/repos/emulator-tools`' business, not this repo's.** -`emu up` creates and boots the AVD named after this checkout — whatever `emu -name` prints, never a name typed out here, since this file is the same in -every clone. `run-android.sh` is that plus a build and an install. The `adb` -on `PATH` after sourcing `android-env.sh` is that repo's wrapper, which fills -in `-s` from the same rule. Gradle does not go through it, so a Gradle init -script from `emulator-tools` runs `emu check` before `installDebug`, -`uninstallDebug` and `connectedAndroidTest` and fails rather than fanning out -to every attached device; when it refuses, say which device you mean at the -moment you use it — `ANDROID_SERIAL=$(emu serial) ./gradlew …`. - -### Testing llama.cpp and ssh here - -**Both are set up here as of 2026-09-04** and need nothing typed. The -prebuilt CPU llama.cpp lives outside the repo at `~/.local/opt/llama.cpp` -(the 15 MB `ubuntu-x64` release asset) and is symlinked as -`/usr/local/bin/llama-server`, which is what makes **discovery find it over -ssh**: `~/.local/bin` is not on the PATH a non-interactive ssh session gets. -It resolves its own libraries through `$ORIGIN`, so no `LD_LIBRARY_PATH` is -needed. One model is downloaded — `unsloth/Qwen3-0.6B-GGUF/Qwen3-0.6B-Q8_0.gguf`, -639 MB under `~/.local/share/ai-app/models` — and answers at usable speed on -this VM's 8 cores. **Do not test with a 2-bit quant**: the -IQ2_XXS of that model produces fluent nonsense, which reads exactly like a -broken driver — `llama-cli` produces the same from the file directly, which -is how to tell the two apart in a hurry. - -There is no second machine, so **ssh this VM to itself**. That is set up -too: the key is `~/.config/ai-app/ssh-self` (its public half is in -`~/.ssh/authorized_keys`, labelled removable), and the real config carries a -setup called **"this vm over ssh"** — `bob@127.0.0.1` with that -`identityFile` plus -`options: ["StrictHostKeyChecking=no", "UserKnownHostsFile=/tmp/ai-app-known-hosts"]` -so it touches nothing real — offering `claude-cli` and `llama-cpp`. It is the -whole rig for "does a remote llama session work", since the far machine is -this one and the model file is the same file. For a throwaway setup of your -own, point a provider's `command` at something harmless like `/bin/echo` -rather than at `claude`: the transport is what is under test, the process -exiting immediately is the signal, and it costs no tokens. The remote login -shell here is **fish**; the -remote script and `ssh.rs`'s POSIX quoting happen to mean the same thing in -both, but that is luck rather than design, and a shell that is neither is the -thing to suspect first if a remote spawn ever mangles an argument. - ## Where things run (host vs this VM) The machine itself — the two boxes, the shared `~/repos` mount, and why the @@ -359,43 +205,6 @@ day to day: in `process.json`; removing either by hand while the session is live loses output or replays it. -## Importing - -The import list reports each session's **size as well as its line count**, -because the two disagree in the way that matters: these transcripts embed -screenshots as base64, so one line can be a megabyte. On this machine a 69 MB -session has 3,427 lines and a 44 MB one has 6,792 — nothing about a line -count tells you what continuing a session will cost. Shown, not warned about; -importing a large session is a choice somebody is entitled to make. - -**Never import a Claude Code session that is open in a terminal.** The app -refuses it — see PLAN.md for the incident that made that a refusal rather -than a warning. - -**One Claude Code session id can name two files, and the listing offers it -once.** Resuming from a different working directory makes the CLI write a -second transcript with the same id under that directory's project folder — an -ordinary state of a machine, not corruption. Everything downstream addresses -a session by id, and the phone keyed its list on it, so two rows sharing one -**closed the app** on a Compose duplicate-key throw. `parse_listing` keeps -the copy with the most lines, because the other is usually a few-hundred-byte -stub and is often the *newer* of the two, so recency is the wrong key. -Deleting removes every copy rather than the first, or the row came back after -a delete that reported success. The phone's half is `uniqueItems`, which -every list keyed on a server-chosen id goes through: a repeat there must -never be able to close the app, whatever produced it. - -**Deleting a session offers to take the machine's own transcript with it** — -`DELETE /sessions/{id}?deleteForeign=true`, behind a switch in the -confirmation, and only where the driver keeps a record of its own -(`keepsOwnTranscript`, which today means Claude Code). Off by default, -because leaving that copy is what makes an ordinary delete recoverable — and -the dialog's paragraph is rewritten when it is on rather than appended to, -since the sentence promising the conversation "should still be there to -import again" is exactly the one the switch makes false. The server deletes -the machine's copy *first*, so a machine it cannot reach leaves the session -where it was instead of half-deleted. - ## Shared appearance - **A row something is happening to is dimmed, drained of colour, and says @@ -523,37 +332,3 @@ belongs in `~/.claude/TOOLCHAIN.md` or `~/.claude/MACHINE.md` instead. hop to `Dispatchers.Default`. The shape to watch for is a `withContext` that wraps the *fetch* and leaves the work done with the result outside it. -## Measurements worth not re-taking - -- **What the transcript screen costs to scroll.** Taken 2026-08-30 on the GPU - emulator against a real imported transcript with the server at - `--delay 120`. Settled and flinging fast, both into fresh history and back - through rows already drawn: **5.2–5.9% janky frames, 99th percentile - 29–32ms, 0–2 slow UI-thread frames.** The stock Settings app on the same - device is 3.3% and 38ms, so this is at the platform floor. The number that - is *not* at the floor is the first few seconds after opening a session, - where every row on the way is being composed for the first time; that is - inherent to a lazy list and it is why a measurement taken before the screen - settles reads three times worse. **Settle first, then reset `gfxinfo`.** -- **The reset path is not reachable by reopening a session.** Measured - 2026-09-04 against a session streaming at 20 events a second: reopening one - with an anchor 1,800 events back connects **87–119 events behind**, well - under `CATCH_UP_LIMIT`'s 200, because the restore is two requests — the - opening page, then one span covering the whole distance. To exercise the - reset at all you have to lower `CATCH_UP_LIMIT` in a throwaway build; at 5 - the app takes the reset on a live connection, clears, refills and carries - on without reconnecting. -- **The session screen's stream survives backgrounding here** — 20 seconds at - the launcher while 415 events were produced brought no reconnect at all, - which is not what the comment above that loop expects, and is most likely - this emulator being headless rather than the phone's behaviour. -- **Reopening a cached session costs one request for one event** (the probe), - and scrolling the whole conversation back costs nothing more; a cold open - of the same 500-event session is two pages, 100 events. Measured - 2026-09-04 on the emulator against the sandbox. -- **Reading is cheap and editing is not.** The viewer handles a 1 MiB, - 28,000-line file because it draws one row per line; the editor is one - `BasicTextField`, which costs two seconds a frame at 128 kB and stops the - app at 1 MiB, so `EDIT_LIMIT` caps it at 32 kB with the reason said on - screen. If you make the editor faster, that number is what to move. - EXPLORER.md's "What the measurements said" has the rest. diff --git a/PLAN.md b/PLAN.md index a0c24a3..2beee6d 100644 --- a/PLAN.md +++ b/PLAN.md @@ -167,6 +167,17 @@ the mode, which do take effect mid-turn. `None` is a level in its own right -- the CLI's own default -- so the picker can return to it; a level this app named as the default instead would be this app choosing one. +**What a new session starts at is `Config::default_effort`**, applied in +`spawn_session` rather than filled in by the spawn screen, so it holds for an +import and a bare API call as well. It is set by the spawn screen's own +picker, whose label says so: one control, where new sessions are made, rather +than a settings page for a single value. It is not on a provider, because +providers are discovered and the next rediscovery would erase it, and not on +the phone, because a second device would then spawn at a level nobody there +chose. `GET`/`POST /defaults` carry it, as a struct rather than a bare value +so the permission mode -- still hardcoded to `auto` on the spawn screen -- can +move there without a second route. + **`--resume` only ever runs when nothing else has that session open.** That is the rule behind the import refusal, the single `ClaudeDriver::launch` entry point, and the `Exited` correction below; two CLIs on one session file diff --git a/app/androidApp/src/main/kotlin/com/example/aiapp/Api.kt b/app/androidApp/src/main/kotlin/com/example/aiapp/Api.kt index 8dafacd..c7fb46f 100644 --- a/app/androidApp/src/main/kotlin/com/example/aiapp/Api.kt +++ b/app/androidApp/src/main/kotlin/com/example/aiapp/Api.kt @@ -485,6 +485,8 @@ fun spawnSession( model: String? = null, cwd: String? = null, permissionMode: String? = null, + /** Null for whatever the server's default is; see [fetchDefaultEffort]. */ + effort: String? = null, params: Map = emptyMap(), /** Continue this Claude Code session instead of starting an empty one. */ import: String? = null, @@ -502,6 +504,7 @@ fun spawnSession( if (!model.isNullOrBlank()) put("model", model) if (!cwd.isNullOrBlank()) put("cwd", cwd) if (!permissionMode.isNullOrBlank()) put("permissionMode", permissionMode) + if (!effort.isNullOrBlank()) put("effort", effort) if (!import.isNullOrBlank()) put("import", import) if (params.isNotEmpty()) { put("params", JSONObject(params.toMap())) @@ -1008,6 +1011,27 @@ fun setSessionModel(settings: ServerSettings, sessionId: String, model: String) */ val PERMISSION_MODES = listOf("manual", "acceptEdits", "auto", "bypassPermissions", "plan") +/** + * What a new session's thinking level is when nothing chose one, or null for the CLI's own. + * + * Held by the server rather than by this phone, because a second device would otherwise spawn + * sessions at a level the first one's owner never picked. + */ +fun fetchDefaultEffort(settings: ServerSettings): String? = + requestFromServer(settings, "/defaults") { + it.jsonObject().optString("effort").ifEmpty { null } + } + +/** Sets what new sessions start at. Nothing already running changes. */ +fun setDefaultEffort(settings: ServerSettings, level: String?) { + requestFromServer( + settings, + "/defaults", + method = "POST", + jsonBody = JSONObject().put("effort", level ?: JSONObject.NULL).toString(), + ) {} +} + /** * How hard the model thinks, as `claude --effort` takes them, cheapest first. * diff --git a/app/androidApp/src/main/kotlin/com/example/aiapp/SpawnScreen.kt b/app/androidApp/src/main/kotlin/com/example/aiapp/SpawnScreen.kt index 9807836..4077cfa 100644 --- a/app/androidApp/src/main/kotlin/com/example/aiapp/SpawnScreen.kt +++ b/app/androidApp/src/main/kotlin/com/example/aiapp/SpawnScreen.kt @@ -61,6 +61,11 @@ fun SpawnScreen( // "auto" rather than "manual": on a phone every ask is a round trip to a question card, and // answering "allow Bash?" dozens of times per task is what this app exists to avoid. var permissionMode by remember { mutableStateOf("auto") } + // Null until the server has been asked, and null again if it answers "no level chosen" -- the + // two are told apart by [defaultsAsked], because a picker that shows a level before the answer + // arrives is one you can spawn at without having chosen it. + var effort by remember { mutableStateOf(null) } + var defaultsAsked by remember { mutableStateOf(false) } var busy by remember { mutableStateOf(false) } // Only the spawn's own failure. The fetch's lives in `options`: this one leaves a filled-in // form worth keeping, and that one leaves nothing to fill in. @@ -75,6 +80,12 @@ fun SpawnScreen( var temperature by remember { mutableStateOf("") } LaunchedEffect(Unit) { + // Separate from the setups fetch below and deliberately not fatal: failing to learn the + // default must leave a screen you can still spawn from, so the picker stays on "default" + // and says so rather than the whole form refusing to draw. + runCatching { withContext(Dispatchers.IO) { fetchDefaultEffort(settings) } } + .onSuccess { effort = it } + defaultsAsked = true options = try { val fetched = withContext(Dispatchers.IO) { fetchSetups(settings) } @@ -262,6 +273,19 @@ fun SpawnScreen( selected = permissionMode, onSelect = { permissionMode = it }, ) + Spacer(Modifier.height(16.dp)) + + // Says what it does to *later* spawns as well, because it does: the level chosen here + // is stored as the default, which is the whole way that default is set. A picker that + // quietly changed a global would be the same control with the fact left out. + ChipGroup( + label = "Thinking (kept as the default for new sessions)", + options = listOf(DEFAULT_EFFORT) + EFFORT_LEVELS, + // The CLI's own default is a level in the list, so this cannot be a one-way trip. + // Disabled-looking until the server has answered, for the reason above. + selected = if (defaultsAsked) effort ?: DEFAULT_EFFORT else null, + onSelect = { chosen -> effort = chosen.takeIf { it != DEFAULT_EFFORT } }, + ) } Spacer(Modifier.height(24.dp)) @@ -279,6 +303,13 @@ fun SpawnScreen( try { val spawned = withContext(Dispatchers.IO) { + // Stored before the spawn and not after it: choosing a level is + // an intent about new sessions in general, so a spawn that then + // fails must not also lose the choice. Non-fatal for the same + // reason the fetch above is -- the session is what was asked for. + if (isClaude) { + runCatching { setDefaultEffort(settings, effort) } + } spawnSession( settings, // The id, not the label: labels are editable and the server @@ -291,6 +322,7 @@ fun SpawnScreen( if (isLlama) modelKey else model.trim().takeIf { isClaude }, cwd = cwd.trim().takeIf { isClaude }, permissionMode = permissionMode.takeIf { isClaude }, + effort = effort.takeIf { isClaude }, // Sent only when set, so blank means "whatever llama.cpp does // by default" rather than a zero. params = diff --git a/server/src/config.rs b/server/src/config.rs index 02d7bc4..d699cc8 100644 --- a/server/src/config.rs +++ b/server/src/config.rs @@ -30,6 +30,18 @@ pub struct Config { pub tokens: Vec, pub setups: Vec, pub sessions: Vec, + /// What a new session's thinking level is when nothing chose one. + /// + /// Here rather than on a provider because providers are *discovered*: a + /// default written onto one would be erased by the next rediscovery, which + /// is the kind of setting that looks like it stuck until the day it did + /// not. Here rather than on the phone because a second device would then + /// spawn sessions the first one's owner did not expect. + /// + /// `None` is the CLI's own default, and stays reachable: this is a level + /// somebody chose, not a level this app picked for them. + #[serde(default, skip_serializing_if = "Option::is_none")] + pub default_effort: Option, } /// A machine, and the things it can run. @@ -462,6 +474,7 @@ mod tests { }], }, ], + default_effort: Some("low".to_string()), sessions: vec![SessionConfig { id: "abc123".to_string(), setup: "vm".to_string(), diff --git a/server/src/routes.rs b/server/src/routes.rs index 7ce44bd..b369c92 100644 --- a/server/src/routes.rs +++ b/server/src/routes.rs @@ -54,6 +54,8 @@ //! POST /sessions/{id}/notify {notify} -- announce this one or not //! GET /notifications SSE: every session's attention-wanting //! moments, live only (see `notifications`) +//! GET /defaults {effort} -- what a new session starts at +//! POST /defaults {effort} -- null for the CLI's own default //! GET /usage cached usage windows per provider //! GET /models downloaded GGUFs, and what is being fetched //! GET /models/search?q=Q HuggingFace repositories matching Q @@ -137,6 +139,7 @@ pub fn router(manager: Arc) -> Router { .route("/sessions/{id}/model", post(set_model)) .route("/sessions/{id}/permission-mode", post(set_permission_mode)) .route("/sessions/{id}/effort", post(set_effort)) + .route("/defaults", get(defaults).post(set_defaults)) .route("/sessions/{id}/notify", post(set_notify)) .route("/notifications", get(notifications)) .route("/sessions/{id}/compact", post(compact)) @@ -1355,6 +1358,36 @@ struct PermissionModeRequest { mode: String, } +/// What new sessions start at. One field today; a struct rather than a bare +/// value because "the defaults" is the thing a phone asks for, and the next +/// one to move here -- the permission mode, which the spawn screen still +/// hardcodes -- must not need a second route. +#[derive(Serialize, Deserialize)] +#[serde(rename_all = "camelCase")] +#[serde(deny_unknown_fields)] +struct Defaults { + #[serde(default, skip_serializing_if = "Option::is_none")] + effort: Option, +} + +async fn defaults(State(manager): State>) -> axum::Json { + axum::Json(Defaults { + effort: manager.default_effort(), + }) +} + +/// Sets what a new session's thinking level is. Applied when a session is +/// spawned, so nothing already running changes underneath anybody. +async fn set_defaults( + State(manager): State>, + axum::Json(body): axum::Json, +) -> Result { + manager + .set_default_effort(body.effort.as_deref()) + .map_err(bad_request)?; + Ok(StatusCode::NO_CONTENT) +} + #[derive(Deserialize)] #[serde(rename_all = "camelCase")] #[serde(deny_unknown_fields)] diff --git a/server/src/session/mod.rs b/server/src/session/mod.rs index 4c6cf81..a13601e 100644 --- a/server/src/session/mod.rs +++ b/server/src/session/mod.rs @@ -1038,7 +1038,19 @@ impl SessionManager { model: spec.model, cwd: spec.cwd, permission_mode: spec.permission_mode, - effort: spec.effort, + // Applied here rather than on the spawn screen, so it holds + // however a session was made -- the phone, an import, or a bare + // API call -- instead of only where somebody remembered to fill it + // in. And only where the driver reads one: a llama session storing + // a level it never passes to anything is a config file that + // answers a question about itself wrongly. + effort: spec.effort.or_else(|| { + provider + .kind + .takes_effort() + .then(|| inner.config.default_effort.clone()) + .flatten() + }), params: spec.params, // On by default. Not offered at spawn: a session's first turn // is exactly the one somebody is waiting for. @@ -1283,6 +1295,27 @@ impl SessionManager { /// /// `None` clears it, which is a level in its own right -- the CLI's own /// default -- and the reason this takes an option rather than a string. + /// What a new session's thinking level is when nothing chose one, and the + /// setting of it. See `Config::default_effort`; `None` is the CLI's own. + /// + /// Only the default: a session already spawned keeps the level it was + /// given, because changing what running conversations do from a screen + /// about *new* ones is not something anybody asked for by setting a + /// default. + pub fn default_effort(&self) -> Option { + self.inner.read().unwrap().config.default_effort.clone() + } + + pub fn set_default_effort(&self, effort: Option<&str>) -> Result<()> { + let effort = effort.map(str::trim).filter(|level| !level.is_empty()); + let mut inner = self.inner.write().unwrap(); + let mut candidate = inner.config.clone(); + candidate.default_effort = effort.map(str::to_string); + candidate.save(&self.config_path)?; + inner.config = candidate; + Ok(()) + } + pub fn set_session_effort(&self, id: &str, effort: Option<&str>) -> Result<()> { let effort = effort.map(str::trim).filter(|level| !level.is_empty()); { @@ -3480,6 +3513,74 @@ mod tests { std::fs::write(path, rewritten).expect("write transcript"); } + /// A new session takes the stored default, and an explicit choice still + /// wins over it. + /// + /// Applied where the session is made rather than on the spawn screen, so + /// it holds for an import and a bare API call too -- a default that only + /// worked from one screen would be a default somebody had already set and + /// would reasonably believe was in force. + #[tokio::test] + async fn a_new_session_starts_at_the_stored_default_thinking_level() { + let dir = tempfile::tempdir().expect("tempdir"); + let config_path = dir.path().join("config.ron"); + let data_dir = dir.path().join("sessions"); + // Seeded with both kinds, because half of what this asks is that a + // driver which does not read a level is not given one. + let cli = seed_stand_in_cli(&config_path, dir.path()); + let manager = SessionManager::new( + config_path.clone(), + data_dir.clone(), + data_dir.join("models"), + ) + .expect("manager"); + assert_eq!( + manager.default_effort(), + None, + "nothing is set to begin with" + ); + + manager + .set_default_effort(Some("low")) + .expect("store the default"); + let took = manager.spawn_session(stand_in_spec(&cli)).expect("spawn"); + assert_eq!( + took.effort.as_deref(), + Some("low"), + "a new session takes it" + ); + + let chosen = manager + .spawn_session(SpawnSpec { + effort: Some("max".to_string()), + ..stand_in_spec(&cli) + }) + .expect("spawn"); + assert_eq!( + chosen.effort.as_deref(), + Some("max"), + "an explicit choice is not overwritten by the default" + ); + + // The case this change had no reason to touch: echo does not read a + // level, so storing one on it would be a config file describing a + // session in terms of something that never reaches it. + let echo = manager.spawn_session(echo_spec()).expect("spawn echo"); + assert_eq!( + echo.effort, None, + "a driver that does not take a level is not given the default" + ); + + // Clearing it is reachable, so the CLI's own default can be restored. + manager.set_default_effort(None).expect("clear the default"); + let cleared = manager.spawn_session(stand_in_spec(&cli)).expect("spawn"); + assert_eq!(cleared.effort, None, "and then new sessions choose nothing"); + + for id in [took.id, chosen.id, echo.id, cleared.id] { + manager.delete_session(&id).expect("delete"); + } + } + /// A thinking level is stored and the process **ended**, because `--effort` /// is read when the CLI launches and has no control request behind it. A /// session left running would go on thinking at the old level underneath a From 6bdec6e785e7b4f6167358cc991f4df697809c7b Mon Sep 17 00:00:00 2001 From: iris <2+iris@noreply.localhost> Date: Sat, 5 Sep 2026 05:21:43 -0400 Subject: [PATCH 05/12] Let a session resume itself when its usage limit lifts Off by default and per session: it spends quota the moment quota exists, with nobody watching, which is not a thing a default may decide. Switched on from the session settings dialog, with the message it sends editable ("continue" unless something else is typed). Running out of quota becomes a state rather than an error. The Claude driver recognises its dialect's sentence -- `Claude AI usage limit reached|1788546972` -- and reports `LimitReached` with the reset time it gave; nothing above a driver matches on a string. The transcript draws it as a divider, like a clear or a compaction. The schedule is a plan to *ask*, never a plan to send. Both reset times available are untrustworthy in the direction that matters -- the dialect's is written when the turn fails, the endpoint's moves when the window does -- so the wait ends in a question to the usage meter, and only `ok` with no window at 100% sends anything. A window still spent reschedules to its own reset time, which is what makes a limit that lifts late wait longer and one that lifts early resume sooner. A meter that cannot be asked is a longer wait too, never a send. A day after the limit was hit the wait gives up and says so in the transcript, so a machine that can never be asked is not retried for ever. The schedule is persisted on the session: a five-hour window outlasts a backend restart, and a wait forgotten across one never comes back. Driven end to end with echo, never a real account: `/limit [minutes]` reports the same event a real driver does and `/usage` sets what the meter answers, deliberately separate so the two can disagree. The wait moved from the dialect's two minutes to the meter's seven when the meter changed its mind, and the message went out on the first check after the meter came back under the limit. Also makes the settings dialog scrollable, which these two controls made necessary: at a 1.5x system font it clipped the last of them with nothing on screen to say so. Co-Authored-By: Claude Opus 5 --- AGENTS.md | 19 + PLAN.md | 48 ++ .../src/main/kotlin/com/example/aiapp/Api.kt | 54 +++ .../main/kotlin/com/example/aiapp/Dividers.kt | 41 ++ .../main/kotlin/com/example/aiapp/Events.kt | 16 + .../kotlin/com/example/aiapp/SessionScreen.kt | 1 + .../example/aiapp/SessionSettingsDialog.kt | 141 +++++- .../com/example/aiapp/TranscriptItems.kt | 13 + .../kotlin/com/example/aiapp/LimitRowTest.kt | 32 ++ server/src/config.rs | 49 ++ server/src/main.rs | 8 + server/src/resume.rs | 382 +++++++++++++++ server/src/routes.rs | 30 ++ server/src/session/claude/translate.rs | 92 +++- server/src/session/driver.rs | 19 + server/src/session/echo.rs | 40 ++ server/src/session/mod.rs | 440 +++++++++++++++++- 17 files changed, 1405 insertions(+), 20 deletions(-) create mode 100644 app/androidApp/src/test/kotlin/com/example/aiapp/LimitRowTest.kt create mode 100644 server/src/resume.rs diff --git a/AGENTS.md b/AGENTS.md index f94676b..9279365 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -205,6 +205,25 @@ day to day: in `process.json`; removing either by hand while the session is live loses output or replays it. +## Auto-resume + +**A session switched to it sends itself a message once the account's usage +limit lifts** — off by default, per session, in the session settings dialog. +PLAN.md's "Auto-resume" is the design; day to day: + +- **The schedule is a plan to ask.** `resume.rs` wakes at the scheduled time, + asks `GET /usage`'s meter for that machine and provider, and only sends when + it answers `ok` with nothing at 100%. Anything else — still spent, logged + out, unreachable — is a longer wait, and a still-spent window reschedules to + the reset time the *meter* now gives. +- **Test it with echo, never with a real account.** `/limit [minutes]` reports + the same `limitReached` event a real driver does, and `/usage 100 5` sets + what the meter answers. They are deliberately separate: the two disagreeing + is the case the design exists for. `/usage 20` is the limit lifting. +- The wait is on the session in `config.ron` (`resume`), so it survives a + backend restart. A day after the limit was hit it gives up and says so in + the transcript. + ## Shared appearance - **A row something is happening to is dimmed, drained of colour, and says diff --git a/PLAN.md b/PLAN.md index 2beee6d..3f1a0a0 100644 --- a/PLAN.md +++ b/PLAN.md @@ -596,6 +596,54 @@ always running. So absent means **not running**, and only a timestamp that arrives and cannot be parsed is unknown. `WindowEnd` in `ResetCountdown.kt` is the one rule both readers go through. +### Auto-resume (2026-09-05) + +**A session may pick itself back up when the account's usage limit lifts.** +Off unless somebody switched that session to it, because it spends quota the +moment quota exists and does so with nobody looking — that is not a thing a +default may decide. It sends one message, `continue` unless another was +typed, and then it is done; there is no retry loop around the conversation +itself. + +**Running out of quota is a state, not an error.** `Event::LimitReached` +carries the dialect's reset time where it gave one, and recognising it +belongs to the driver — the Claude CLI ends the turn with `is_error` and +`Claude AI usage limit reached|1788546972`, and nothing above the driver +matches on a string. The transcript draws it as a divider, like a clear or a +compaction: what a reader scrolling back wants from it is why the +conversation stops at that line. + +**The schedule is a plan to ask, never a plan to send.** Every reset time +available here is untrustworthy in the direction that matters: the dialect's +is written when the turn fails, and the endpoint's moves when the window +does. So the wait ends in a question to `usage.rs`, and only `ok` with no +window at 100% sends anything. A window still spent reschedules to *its own* +reset time — which is what makes a limit that lifts later than promised wait +longer, and one that lifts sooner resume sooner. A meter that cannot be +asked at all is a longer wait too, never a send: "we could not find out" +must not be able to produce the same action as "there is room". + +Bounded, because something has to be: a day after the limit was hit the wait +stops and says so in the session's own transcript. A machine that can never +be asked would otherwise be retried for ever with nothing on screen saying +so. + +The schedule is persisted on the session (`resume: Some(ScheduledResume)`), +not held in memory: a five-hour window routinely outlasts a backend restart, +and a wait forgotten across one is a session that silently never comes back. +`resume.rs` is the top layer — it holds the manager and the monitor and +neither holds it — which is what lets the decision be a pure function of a +snapshot and a clock. The pump reports limits downward on a broadcast, for +the reason `Shared` exists: the pump runs underneath the manager. + +**Exercised with echo, never with a real account.** `/limit [minutes]` in an +echo session reports the same event a real driver does, and `/usage` sets +what the meter answers — deliberately two commands, because the two +disagreeing is the state the whole design is about. The loop was driven end +to end that way on 2026-09-05: the wait moved from the dialect's two minutes +to the meter's seven when the meter changed its mind, and the message went +out on the first check after the meter came back under the limit. + ### HTTP surface **`routes.rs`'s module doc comment is the table.** REST for actions, one SSE diff --git a/app/androidApp/src/main/kotlin/com/example/aiapp/Api.kt b/app/androidApp/src/main/kotlin/com/example/aiapp/Api.kt index c7fb46f..d64df02 100644 --- a/app/androidApp/src/main/kotlin/com/example/aiapp/Api.kt +++ b/app/androidApp/src/main/kotlin/com/example/aiapp/Api.kt @@ -167,6 +167,24 @@ data class SessionSummary( * itself from a default is one you can turn off while believing you are reading it. */ val notify: Boolean, + /** + * Whether this session sends itself a message once its account's usage limit lifts, and what + * that message says. + * + * The message is what the server would actually send, with its own default already filled in, + * so the field shows the words rather than an empty box standing for them. + */ + val autoResume: Boolean, + val autoResumeMessage: String, + /** + * When the server next intends to check whether the limit has lifted, in epoch seconds, or null + * when nothing is waiting. + * + * A time to *ask*, not a time to resume: the server checks the meter at that moment and waits + * again if the limit is still on. Worded that way wherever it is shown, because a promise this + * app cannot keep is worse than no time at all. + */ + val resumeAt: Double?, /** * The directory the session works in, or null where it was never given one. * @@ -221,6 +239,12 @@ private fun parseSession(session: JSONObject) = takesEffort = session.optBoolean("takesEffort", false), imported = session.optBoolean("imported", false), notify = session.optBoolean("notify", true), + autoResume = session.optBoolean("autoResume", false), + // The server sends its own default rather than nothing, so an empty answer means an older + // server -- and this app's word for it is the same word. + autoResumeMessage = + session.optString("autoResumeMessage").ifEmpty { DEFAULT_RESUME_MESSAGE }, + resumeAt = if (session.has("resumeAt")) session.getDouble("resumeAt") else null, cwd = session.optString("cwd").ifEmpty { null }, contextTokens = if (session.has("contextTokens")) session.getLong("contextTokens") else null, @@ -1072,6 +1096,36 @@ fun setSessionPermissionMode(settings: ServerSettings, sessionId: String, mode: } /** Turns this session's notifications on or off. Stored on the backend -- see `SessionConfig`. */ +/** + * What an auto-resume says when nothing else was typed. Mirrors the server's own default, so a + * cleared field shows the word that would actually be sent instead of going blank. + */ +const val DEFAULT_RESUME_MESSAGE = "continue" + +/** + * Turns auto-resume on or off and sets what it would say, in one request because they are one + * decision -- see the server's `/sessions/{id}/auto-resume`. + */ +fun setSessionAutoResume( + settings: ServerSettings, + sessionId: String, + autoResume: Boolean, + message: String?, +) { + requestFromServer( + settings, + "/sessions/$sessionId/auto-resume", + method = "POST", + jsonBody = + JSONObject() + .put("autoResume", autoResume) + // Empty means the server's default rather than a session poked with nothing to + // read, which is the same rule the server applies to the field. + .put("message", message?.trim()?.ifEmpty { null } ?: JSONObject.NULL) + .toString(), + ) {} +} + fun setSessionNotify(settings: ServerSettings, sessionId: String, notify: Boolean) { requestFromServer( settings, diff --git a/app/androidApp/src/main/kotlin/com/example/aiapp/Dividers.kt b/app/androidApp/src/main/kotlin/com/example/aiapp/Dividers.kt index 01a13ab..57600e1 100644 --- a/app/androidApp/src/main/kotlin/com/example/aiapp/Dividers.kt +++ b/app/androidApp/src/main/kotlin/com/example/aiapp/Dividers.kt @@ -12,6 +12,10 @@ import androidx.compose.ui.Alignment import androidx.compose.ui.Modifier import androidx.compose.ui.graphics.Color import androidx.compose.ui.unit.dp +import java.time.Instant +import java.time.ZoneId +import java.time.format.DateTimeFormatter +import java.time.format.FormatStyle /** * A line across the transcript saying what left the session's context. @@ -47,3 +51,40 @@ fun TranscriptDivider(text: String, color: Color, modifier: Modifier = Modifier) fun ClearedRow(modifier: Modifier = Modifier) { TranscriptDivider("Context cleared", clearedColor, modifier) } + +/** + * The mark running out of quota leaves. + * + * The same red the usage bar takes when a window is spent, because it is the same fact in a second + * place: colour by consequence, so "there is nothing left to spend" is learned once. + * + * A time rather than a countdown. The row is folded once and never re-measured, so a span would go + * stale on screen the moment it was drawn; and this is when the *account* said it would reset, + * which is not a promise about when the session picks back up. A limit the session was told no + * reset time for says nothing about one -- that state has its own words rather than a plausible + * number. + */ +@Composable +fun LimitRow(item: TranscriptItem.LimitNote, modifier: Modifier = Modifier) { + TranscriptDivider(limitSummary(item.resetsAt, ZoneId.systemDefault()), overLimitColor, modifier) +} + +/** + * What the row says. Split out so the wording is testable without a screen, since the two states it + * has to keep apart -- a reset time that arrived and one that never did -- are exactly the pair + * that reads the same when it goes wrong. + * + * [zone] is a parameter rather than read here so a test says the same thing wherever it runs. + */ +fun limitSummary(resetsAt: Double?, zone: ZoneId): String { + val at = resetsAt?.let { + try { + DateTimeFormatter.ofLocalizedTime(FormatStyle.SHORT) + .withZone(zone) + .format(Instant.ofEpochSecond(it.toLong())) + } catch (_: Exception) { + null + } + } + return if (at == null) "Usage limit reached" else "Usage limit reached • resets $at" +} diff --git a/app/androidApp/src/main/kotlin/com/example/aiapp/Events.kt b/app/androidApp/src/main/kotlin/com/example/aiapp/Events.kt index f266f70..8580867 100644 --- a/app/androidApp/src/main/kotlin/com/example/aiapp/Events.kt +++ b/app/androidApp/src/main/kotlin/com/example/aiapp/Events.kt @@ -157,6 +157,18 @@ sealed class SessionEvent { */ data object Cleared : SessionEvent() + /** + * The session stopped because its account's usage limit was reached. + * + * Its own event rather than an [Error] carrying the CLI's sentence, because it is a state + * rather than something that went wrong -- and because the raw sentence is `Claude AI usage + * limit reached|1788546972`, which is not readable by the person it is shown to. + * + * [resetsAt] is epoch seconds and null where the session was told nothing. Only the server acts + * on it; what this draws it as is a time, not a countdown, because nothing here re-measures it. + */ + data class LimitReached(val resetsAt: Double?) : SessionEvent() + data class Error(val message: String) : SessionEvent() /** @@ -261,6 +273,10 @@ fun parseSeqEvent(json: String): SeqEvent { trigger = body.optString("trigger").ifEmpty { null }, ) "cleared" -> SessionEvent.Cleared + "limitReached" -> + SessionEvent.LimitReached( + if (body.has("resetsAt")) body.getDouble("resetsAt") else null + ) "error" -> SessionEvent.Error(body.getString("message")) else -> SessionEvent.Unknown(type) } diff --git a/app/androidApp/src/main/kotlin/com/example/aiapp/SessionScreen.kt b/app/androidApp/src/main/kotlin/com/example/aiapp/SessionScreen.kt index f31cd0e..64e031c 100644 --- a/app/androidApp/src/main/kotlin/com/example/aiapp/SessionScreen.kt +++ b/app/androidApp/src/main/kotlin/com/example/aiapp/SessionScreen.kt @@ -1545,6 +1545,7 @@ fun SessionScreen( is TranscriptItem.ClearedNote -> ClearedRow() is TranscriptItem.CompactedNote -> CompactedRow(item) + is TranscriptItem.LimitNote -> LimitRow(item) // Never reached: a peer message is flattened into // its own units. Here because a `when` over the // item kinds has to stay exhaustive. diff --git a/app/androidApp/src/main/kotlin/com/example/aiapp/SessionSettingsDialog.kt b/app/androidApp/src/main/kotlin/com/example/aiapp/SessionSettingsDialog.kt index c756642..0a68e41 100644 --- a/app/androidApp/src/main/kotlin/com/example/aiapp/SessionSettingsDialog.kt +++ b/app/androidApp/src/main/kotlin/com/example/aiapp/SessionSettingsDialog.kt @@ -6,8 +6,10 @@ import androidx.compose.foundation.layout.Spacer import androidx.compose.foundation.layout.fillMaxWidth import androidx.compose.foundation.layout.height import androidx.compose.foundation.layout.width +import androidx.compose.foundation.rememberScrollState import androidx.compose.foundation.text.KeyboardActions import androidx.compose.foundation.text.KeyboardOptions +import androidx.compose.foundation.verticalScroll import androidx.compose.material3.AlertDialog import androidx.compose.material3.CircularProgressIndicator import androidx.compose.material3.MaterialTheme @@ -26,6 +28,10 @@ import androidx.compose.ui.Alignment import androidx.compose.ui.Modifier import androidx.compose.ui.text.input.ImeAction import androidx.compose.ui.unit.dp +import java.time.Instant +import java.time.ZoneId +import java.time.format.DateTimeFormatter +import java.time.format.FormatStyle import kotlinx.coroutines.Dispatchers import kotlinx.coroutines.launch import kotlinx.coroutines.withContext @@ -89,6 +95,16 @@ fun SessionSettingsDialog( // and a spinner sits beside it, which is what not knowing looks like. var notify by remember(sessionId) { mutableStateOf(null) } var notifyError by remember { mutableStateOf(null) } + // The same three-state shape the notification switch has, for the same reason: until the + // server has answered, the switch is disabled rather than showing a position nothing confirmed. + var autoResume by remember(sessionId) { mutableStateOf(null) } + var resumeMessage by remember(sessionId) { mutableStateOf(DEFAULT_RESUME_MESSAGE) } + // When the server next intends to ask whether the limit has lifted, or null when nothing is + // waiting. Read once with everything else: it moves on the server's schedule, not this + // screen's, and a figure that redrew itself here would be this app re-measuring what it was + // told. + var resumeAt by remember(sessionId) { mutableStateOf(null) } + var resumeError by remember { mutableStateOf(null) } // Where the session works. Null until the server has been asked, for the same reason the switch // above is. An empty answer is a session that was never given a directory, which is not the // same as one whose directory is unknown -- the field is only enabled once one of those is @@ -102,6 +118,9 @@ fun SessionSettingsDialog( try { val fresh = withContext(Dispatchers.IO) { fetchSession(settings, sessionId) } notify = fresh.notify + autoResume = fresh.autoResume + resumeMessage = fresh.autoResumeMessage + resumeAt = fresh.resumeAt cwd = fresh.cwd.orEmpty() typedCwd = fresh.cwd.orEmpty() } catch (e: ApiException) { @@ -109,6 +128,8 @@ fun SessionSettingsDialog( // instead of offering a position nothing confirmed. notifyError = e.message notify = null + resumeError = e.message + autoResume = null } } @@ -174,6 +195,39 @@ fun SessionSettingsDialog( } } + /** + * Turns auto-resume on or off, or changes what it would say. + * + * One request for both, because the server takes one: switching it on and typing the message + * are two halves of the same decision, and sending them separately would leave a moment where + * the session is armed with the old words. + * + * Put back if refused, like the notification switch. Turning it off also clears what was + * scheduled -- said here rather than only on the server, or the row would go on naming a time + * that no longer exists. + */ + fun setAutoResume(on: Boolean, message: String) { + val wasOn = autoResume + val wasMessage = resumeMessage + val wasAt = resumeAt + autoResume = on + resumeMessage = message + if (!on) resumeAt = null + resumeError = null + scope.launch { + try { + withContext(Dispatchers.IO) { + setSessionAutoResume(settings, sessionId, on, message) + } + } catch (e: ApiException) { + autoResume = wasOn + resumeMessage = wasMessage + resumeAt = wasAt + resumeError = e.message + } + } + } + // Nothing to do when the name has not changed, so the button says so rather than sending a // request whose success would look exactly like the failure of having typed nothing. val changed = name.trim().isNotEmpty() && name.trim() != title @@ -200,7 +254,10 @@ fun SessionSettingsDialog( onDismissRequest = onDismiss, title = { Text("Session settings") }, text = { - Column { + // Scrollable, because this dialog grew past a screenful: a Material dialog constrains + // its own height and clips what does not fit, so the last control on the list is one + // large system font away from being unreachable with nothing on screen to say so. + Column(Modifier.verticalScroll(rememberScrollState())) { OutlinedTextField( value = name, onValueChange = { name = it }, @@ -244,6 +301,70 @@ fun SessionSettingsDialog( ) } Spacer(Modifier.height(8.dp)) + Row( + verticalAlignment = Alignment.CenterVertically, + modifier = Modifier.fillMaxWidth(), + ) { + Text("Resume after a usage limit", modifier = Modifier.weight(1f)) + if (autoResume == null && resumeError == null) { + CircularProgressIndicator( + modifier = Modifier.width(16.dp).height(16.dp), + strokeWidth = 2.dp, + ) + Spacer(Modifier.width(8.dp)) + } + Switch( + checked = autoResume == true, + onCheckedChange = { setAutoResume(it, resumeMessage) }, + enabled = autoResume != null, + ) + } + // Disabled rather than hidden while the switch is off: a field that comes and goes + // makes its own presence the signal, and a visible one teaches what the switch will + // do. Committed on the keyboard's Done rather than on every keystroke, so typing a + // sentence is one request instead of one per letter. + OutlinedTextField( + value = resumeMessage, + onValueChange = { resumeMessage = it }, + label = { Text("Message to send") }, + // What an empty field means, in the field: the server's own word rather than a + // session poked with nothing to read. + placeholder = { Text(DEFAULT_RESUME_MESSAGE) }, + singleLine = true, + enabled = autoResume == true, + modifier = Modifier.fillMaxWidth(), + keyboardOptions = KeyboardOptions(imeAction = ImeAction.Done), + keyboardActions = + KeyboardActions(onDone = { setAutoResume(true, resumeMessage) }), + ) + // What it does and what it costs, in the order it happens. The last sentence is the + // one that matters: the time below is when the server will *ask*, not a promise + // about when the session speaks. + Text( + "When this session stops because the account is out of quota, the server " + + "checks the limit and sends this message once it has lifted. It checks " + + "again if the limit is still on.", + style = MaterialTheme.typography.bodySmall, + color = MaterialTheme.colorScheme.onSurfaceVariant, + ) + // Only where something is actually waiting. Absent is not a state worth a row: a + // session that has not hit a limit has nothing scheduled, which the reader can see + // from the switch. + resumeAt?.let { at -> + Text( + "Waiting now -- next check ${formatCheckTime(at)}.", + style = MaterialTheme.typography.bodySmall, + color = MaterialTheme.colorScheme.onSurfaceVariant, + ) + } + resumeError?.let { + Text( + it, + color = MaterialTheme.colorScheme.error, + style = MaterialTheme.typography.bodySmall, + ) + } + Spacer(Modifier.height(8.dp)) Row( verticalAlignment = Alignment.CenterVertically, modifier = Modifier.fillMaxWidth(), @@ -399,3 +520,21 @@ fun SessionSettingsDialog( dismissButton = { TextButton(onClick = onDismiss) { Text("Close") } }, ) } + +/** + * When the server will next look, as a local time. + * + * A time rather than a countdown, for the reason the transcript's own limit row gives: this screen + * reads the figure once, and a span drawn from a value nothing refreshes goes stale while somebody + * is looking at it. + */ +private fun formatCheckTime(epochSeconds: Double): String = + try { + DateTimeFormatter.ofLocalizedTime(FormatStyle.SHORT) + .withZone(ZoneId.systemDefault()) + .format(Instant.ofEpochSecond(epochSeconds.toLong())) + } catch (_: Exception) { + // A time that cannot be read is not a time to show: the sentence above still says a check + // is coming, which is the part the reader can act on. + "soon" + } diff --git a/app/androidApp/src/main/kotlin/com/example/aiapp/TranscriptItems.kt b/app/androidApp/src/main/kotlin/com/example/aiapp/TranscriptItems.kt index 5089ca2..96eefa1 100644 --- a/app/androidApp/src/main/kotlin/com/example/aiapp/TranscriptItems.kt +++ b/app/androidApp/src/main/kotlin/com/example/aiapp/TranscriptItems.kt @@ -175,6 +175,18 @@ sealed class TranscriptItem { val preTokens: Long?, val postTokens: Long?, ) : TranscriptItem() + + /** + * The account ran out of quota, so the turn stopped here. + * + * A divider rather than an error: nothing failed, and what a reader scrolling back needs from + * it is the same thing a clear or a compaction gives them -- why the conversation stops at this + * line. + * + * [resetsAt] is epoch seconds and null where the session was told nothing, which is a state the + * row has words for rather than a time it invents. + */ + data class LimitNote(override val seq: Long, val resetsAt: Double?) : TranscriptItem() } /** @@ -458,6 +470,7 @@ fun foldEvent(items: List, entry: SeqEvent): List items + TranscriptItem.LimitNote(entry.seq, event.resetsAt) is SessionEvent.Cleared -> items + TranscriptItem.ClearedNote(entry.seq) is SessionEvent.Compacted -> items + TranscriptItem.CompactedNote(entry.seq, event.preTokens, event.postTokens) diff --git a/app/androidApp/src/test/kotlin/com/example/aiapp/LimitRowTest.kt b/app/androidApp/src/test/kotlin/com/example/aiapp/LimitRowTest.kt new file mode 100644 index 0000000..e4c030f --- /dev/null +++ b/app/androidApp/src/test/kotlin/com/example/aiapp/LimitRowTest.kt @@ -0,0 +1,32 @@ +package com.example.aiapp + +import java.time.ZoneId +import kotlin.test.Test +import kotlin.test.assertEquals +import kotlin.test.assertTrue + +/** + * What the transcript says where a session ran out of quota. + * + * The pair worth a test is the one that reads the same when it goes wrong: a reset time that + * arrived and one that never did. The second must not turn into a plausible-looking time, because a + * reader has no way of telling an invented one from a reported one. + */ +class LimitRowTest { + private val utc = ZoneId.of("UTC") + + @Test + fun `a reported reset time is shown as a time`() { + // 2026-09-05T12:00:00Z. Asserted as a prefix and the clock reading rather than as the + // whole string: the platform's own short-time format is what this asks for, and it + // differs by JDK and locale down to which space character separates the meridiem. + val summary = limitSummary(1_788_609_600.0, utc) + assertTrue(summary.startsWith("Usage limit reached • resets "), summary) + assertTrue(summary.contains("12:00"), summary) + } + + @Test + fun `a limit with no reset time says only what is known`() { + assertEquals("Usage limit reached", limitSummary(null, utc)) + } +} diff --git a/server/src/config.rs b/server/src/config.rs index d699cc8..475a4cb 100644 --- a/server/src/config.rs +++ b/server/src/config.rs @@ -293,6 +293,30 @@ pub struct SessionConfig { /// turned off in one tap where one that never arrived is not diagnosable. #[serde(default = "notify_default")] pub notify: bool, + /// Whether a session stopped by the account's usage limit sends itself a + /// message once the limit lifts, instead of waiting for a person. + /// + /// Off unless somebody asked for it. It spends quota the moment it becomes + /// available and it does so while nobody is looking, which is exactly the + /// kind of thing that must not happen because a default said so. + #[serde(default, skip_serializing_if = "not_set")] + pub auto_resume: bool, + /// What that message says. `None` is [`DEFAULT_RESUME_MESSAGE`], and stays + /// reachable: it is this app's word, not one somebody chose, so clearing + /// the field goes back to it rather than sending an empty message. + #[serde(default, skip_serializing_if = "Option::is_none")] + pub auto_resume_message: Option, + /// The message this session owes itself once the limit lifts, and when to + /// try. Written when a limit is hit, moved when the wait turns out to be + /// wrong, and cleared when the message goes out or auto-resume is turned + /// off -- see [`ScheduledResume`]. + /// + /// Persisted rather than held in memory because the wait outlives the + /// process doing it: a five-hour window and a weekly one both routinely + /// outlast a backend restart, and a resume forgotten across one is a + /// session that silently never comes back. + #[serde(default, skip_serializing_if = "Option::is_none")] + pub resume: Option, /// Whether this session's process is stopped when the server exits, instead /// of being left running for the next start to adopt. /// @@ -309,6 +333,28 @@ pub struct SessionConfig { pub created: f64, } +/// A message owed to a session whose account ran out, and when to try sending +/// it. +/// +/// `since` is the whole reason this is a struct: the wait is rescheduled every +/// time the meter is asked and still says no, so `at` alone cannot say how long +/// this has been going on -- and something has to, or a machine that can never +/// be asked is retried until somebody notices. See `crate::resume`. +#[derive(Debug, Clone, Copy, Serialize, Deserialize)] +#[serde(rename_all = "camelCase")] +pub struct ScheduledResume { + /// Epoch seconds: when the limit is next worth checking. Never a promise + /// that the message goes out then -- the meter is asked first. + pub at: f64, + /// Epoch seconds the limit was hit. + pub since: f64, +} + +/// What an auto-resume says when nothing else was chosen. One word, because +/// the session already knows what it was doing and this is only the nudge that +/// lets it carry on. +pub const DEFAULT_RESUME_MESSAGE: &str = "continue"; + fn notify_default() -> bool { true } @@ -486,6 +532,9 @@ mod tests { effort: None, params: BTreeMap::new(), notify: true, + auto_resume: false, + auto_resume_message: None, + resume: None, throwaway: false, created: 1234.5, }], diff --git a/server/src/main.rs b/server/src/main.rs index 9fde3b7..3571b7b 100644 --- a/server/src/main.rs +++ b/server/src/main.rs @@ -16,6 +16,7 @@ mod config; mod files; mod media; mod models; +mod resume; mod routes; mod session; mod setups; @@ -268,6 +269,13 @@ async fn main() -> Result<()> { // that sets it is typed; the monitor is what serves it. let monitor = Arc::new(usage::UsageMonitor::new(manager.usage_fixture())); + // The one thing in here that acts without a request behind it: a session + // switched to auto-resume waits out its account's usage limit and picks + // itself back up. Started whether or not any session has it on, because + // the setting is per session and changes from the phone -- see + // `resume::run`. + tokio::spawn(resume::run(Arc::clone(&manager), Arc::clone(&monitor))); + // The bearer-token middleware wraps the entire router -- routes and fallback // alike -- here and only here, so a new route can't forget auth. let app = routes::router(Arc::clone(&manager)) diff --git a/server/src/resume.rs b/server/src/resume.rs new file mode 100644 index 0000000..2ef56d4 --- /dev/null +++ b/server/src/resume.rs @@ -0,0 +1,382 @@ +//! Auto-resume: picking a session back up when its account's usage limit +//! lifts. +//! +//! Off unless a session was switched to it, because this spends quota the +//! moment quota exists and does it while nobody is watching. What it does is +//! narrow on purpose: it sends one message -- "continue" unless something else +//! was typed -- to a session that stopped because the account ran out, and +//! then it is done. There is no retry loop around the conversation itself. +//! +//! **The schedule is a plan to ask, never a plan to send.** A reset time is +//! the one thing here that cannot be trusted: the dialect's is a hint written +//! when the turn failed, the endpoint's moves when the window moves, and both +//! are wrong across the case this exists for -- a limit that lifts later than +//! it said. So the wait ends in a *question* to [`crate::usage`], and only an +//! answer that says the limits no longer apply sends anything. Every other +//! answer, including one that cannot be got at all, becomes a new wait. +//! +//! This is the top layer: it holds the session manager and the usage monitor +//! and neither holds it. That is what lets the decision below be a pure +//! function of a snapshot and a clock, which is the whole of what is worth +//! testing here. + +use std::sync::Arc; +use std::time::Duration; + +use crate::session::{LimitHit, OwedResume, SessionManager, now}; +use crate::usage::{UsageMonitor, UsageSnapshot, UsageState}; + +/// How often to look at the schedule. Coarse deliberately: a wait measured in +/// hours does not deserve a fine-grained clock, and the meter behind it is +/// cached for three minutes anyway. +const TICK: Duration = Duration::from_secs(60); + +/// How close to a scheduled check is close enough to ask the meter. Anything +/// further out is left alone, so a session waiting five hours costs nothing +/// until the last few minutes of it. +const NEARLY: f64 = 300.0; + +/// How long to wait after an answer that decided nothing -- the machine could +/// not be asked, or it says the limit is still on with no reset time. +const BACKOFF: f64 = 300.0; + +/// The least time to wait before asking again, whatever a reset time says. A +/// window that claims to reset in the past would otherwise be asked about on +/// every tick. +const AT_LEAST: f64 = 60.0; + +/// How long after the limit was hit to stop waiting. +/// +/// Something has to bound it, or a machine that can never be asked -- an +/// unplugged laptop, a setup somebody edited away -- is retried for ever with +/// nothing on screen saying so. A day is past the longest window Claude +/// reports, so reaching this means the wait was never going to end on its own. +const GIVE_UP: f64 = 24.0 * 60.0 * 60.0; + +/// The percentage at which a window is spent. The API counts up to 100, so +/// this is an equality in all but name; written as a threshold because a +/// figure arriving slightly over is a full window, not a corrupt one. +const SPENT: f64 = 100.0; + +/// What to do about one owed resume, having asked the meter. +#[derive(Debug, Clone, Copy, PartialEq)] +pub enum Step { + /// The limits no longer apply: send the message. + Send, + /// Ask again at this epoch second. + WaitUntil(f64), + /// This has been waiting longer than anything real would take. + GiveUp, +} + +/// Runs the schedule until the server stops. +/// +/// Two things wake it: the tick, and a session reporting that it has just run +/// out. The second is not an optimisation -- a limit hit is what *creates* a +/// schedule, and a tick that happened a moment before it would otherwise leave +/// the session unrecorded until the next one. +pub async fn run(manager: Arc, monitor: Arc) { + let mut limits = manager.subscribe_limits(); + loop { + tokio::select! { + _ = tokio::time::sleep(TICK) => {} + hit = limits.recv() => match hit { + Ok(LimitHit { session_id, resets_at }) => note(&manager, &session_id, resets_at), + // Lagged: some reports were dropped, and a session that hit a + // limit while this was busy has no schedule. Nothing is lost + // for good -- the sweep below reads the config, and the + // session will report again the next time it is poked -- but + // it is worth saying, because until then that session waits + // for a person. + Err(tokio::sync::broadcast::error::RecvError::Lagged(missed)) => { + tracing::warn!("auto-resume missed {missed} limit reports"); + } + Err(tokio::sync::broadcast::error::RecvError::Closed) => return, + }, + } + sweep(&manager, &monitor).await; + } +} + +/// Records a limit against the session that hit it, if it is one that resumes. +pub(crate) fn note(manager: &SessionManager, session_id: &str, resets_at: Option) { + match manager.note_limit(session_id, resets_at) { + Ok(true) => tracing::info!("session {session_id} hit its usage limit; auto-resume is on"), + Ok(false) => {} + Err(err) => tracing::error!("couldn't schedule a resume for {session_id}: {err:#}"), + } +} + +/// One pass over everything owed a message. +async fn sweep(manager: &SessionManager, monitor: &Arc) { + let at = now(); + for owed in manager.owed_resumes() { + if owed.scheduled.at - at > NEARLY { + continue; + } + // Asked per session rather than once for the whole sweep: the answer + // is cached per machine and per meter, so several sessions on one + // account share one fetch, and a machine nobody is waiting on is not + // dialled at all. + let snapshot = snapshot_for(Arc::clone(monitor), manager, &owed).await; + match decide(snapshot.as_ref(), &owed, now()) { + Step::Send => match manager.resume_now(&owed.session_id) { + Ok(message) => tracing::info!( + "the limit on {} has lifted; sent \"{message}\" to {}", + owed.setup, + owed.session_id + ), + Err(err) => { + tracing::error!("couldn't resume {}: {err:#}", owed.session_id) + } + }, + Step::WaitUntil(next) => { + if let Err(err) = manager.reschedule_resume(&owed.session_id, next) { + tracing::error!( + "couldn't move {}'s resume to {next}: {err:#}", + owed.session_id + ); + } + } + Step::GiveUp => { + // About the machine rather than in the state's own words: the + // detail is in the log, and what lands in the transcript has + // to read on a phone. + let why = match snapshot.as_ref().map(|snapshot| &snapshot.state) { + Some(UsageState::Ok) => "the limit has not lifted in a day".to_string(), + _ => format!("{} could not be asked for a day", owed.setup), + }; + if let Err(err) = manager.abandon_resume(&owed.session_id, &why) { + tracing::error!("couldn't clear {}'s resume: {err:#}", owed.session_id); + } + } + } + } +} + +/// The numbers for the machine and the meter this session is billed against, +/// and `None` when nothing reports on it. +/// +/// Blocking work, so it goes to a blocking thread: the fetch behind it reads a +/// credential file over ssh and then makes an HTTP call. +async fn snapshot_for( + monitor: Arc, + manager: &SessionManager, + owed: &OwedResume, +) -> Option { + let setups: Vec<_> = manager + .setups() + .into_iter() + .filter(|setup| setup.id == owed.setup) + .collect(); + if setups.is_empty() { + return None; + } + let provider = owed.provider; + tokio::task::spawn_blocking(move || { + monitor + .snapshots(&setups) + .into_iter() + .find(|snapshot| snapshot.provider == provider) + }) + .await + .unwrap_or_default() +} + +/// What one owed resume should do, given what the meter said and the time. +/// +/// A pure function of the two, which is what makes the rule inspectable: every +/// answer that is not "the limits no longer apply" is a longer wait, and the +/// only thing that ends the waiting other than success is the clock. +/// +/// The reset time comes from the *snapshot* rather than from the schedule, so +/// a window that turns out to reset later than the dialect said pushes the +/// check back, and one that resets sooner pulls it forward. That is the case +/// the whole design is about: the first answer was a guess, this one is a +/// measurement. +pub fn decide(snapshot: Option<&UsageSnapshot>, owed: &OwedResume, at: f64) -> Step { + let step = match snapshot { + // The meter answered with numbers, which is the only answer that can + // send anything. + Some(snapshot) if snapshot.state == UsageState::Ok => { + let spent: Vec<&crate::usage::UsageWindow> = snapshot + .windows + .iter() + .filter(|window| window.percent >= SPENT) + .collect(); + if spent.is_empty() { + Step::Send + } else { + // The earliest of the spent windows: it is the first moment + // the situation can have changed, and if the others are still + // full this comes straight back here. + match spent + .iter() + .filter_map(|window| epoch_of(window.resets_at.as_deref())) + .min_by(f64::total_cmp) + { + Some(resets) => Step::WaitUntil(resets), + // Spent with no reset time anybody could read. Not a + // reason to send: what is known is that the limit is on. + None => Step::WaitUntil(at + BACKOFF), + } + } + } + // Logged out, unreachable, or the endpoint refused us -- and nothing + // at all, which is a session whose machine or provider has gone. None + // of them says the limit has lifted, and sending on any of them is + // exactly the "inferred value presented as a measured one" this is + // built to avoid. + _ => Step::WaitUntil(at + BACKOFF), + }; + match step { + // Waiting past the point where a real window would have reset means + // whatever is wrong is not going to fix itself. + Step::WaitUntil(_) if at - owed.scheduled.since > GIVE_UP => Step::GiveUp, + Step::WaitUntil(next) => Step::WaitUntil(next.max(at + AT_LEAST)), + other => other, + } +} + +/// An RFC-3339 timestamp as epoch seconds, and `None` for one that is absent +/// or unreadable -- the same two answers the phone's countdown makes, kept +/// apart from each other nowhere here because both mean "this cannot decide +/// when to ask". +fn epoch_of(resets_at: Option<&str>) -> Option { + let text = resets_at?; + time::OffsetDateTime::parse(text, &time::format_description::well_known::Rfc3339) + .ok() + .map(|at| at.unix_timestamp() as f64) +} + +#[cfg(test)] +mod tests { + use super::*; + use crate::config::ScheduledResume; + use crate::usage::UsageWindow; + + fn owed(since: f64) -> OwedResume { + OwedResume { + session_id: "s1".to_string(), + setup: "local".to_string(), + provider: crate::usage::CLAUDE, + scheduled: ScheduledResume { at: since, since }, + } + } + + fn snapshot(state: UsageState, windows: Vec) -> UsageSnapshot { + UsageSnapshot { + provider: crate::usage::CLAUDE.to_string(), + setup: "local".to_string(), + setup_name: "this machine".to_string(), + state, + windows, + fetched_at: 0.0, + } + } + + fn window(percent: f64, resets_at: Option<&str>) -> UsageWindow { + UsageWindow { + kind: "session".to_string(), + label: "5-hour window".to_string(), + percent, + resets_at: resets_at.map(str::to_string), + active: true, + } + } + + #[test] + fn a_meter_with_room_in_it_is_the_only_thing_that_sends() { + let clear = snapshot(UsageState::Ok, vec![window(41.0, None)]); + assert_eq!(decide(Some(&clear), &owed(0.0), 100.0), Step::Send); + } + + #[test] + fn a_window_still_spent_moves_the_check_to_its_own_reset_time() { + // The case the feature exists for: the wait was scheduled for one + // time, the limit is still on, and the endpoint now names another. + let at = 1_788_546_972.0; + let later = "2026-09-05T12:00:00+00:00"; + let spent = snapshot(UsageState::Ok, vec![window(100.0, Some(later))]); + assert_eq!( + decide(Some(&spent), &owed(at - 60.0), at), + Step::WaitUntil(epoch_of(Some(later)).expect("parses")) + ); + } + + #[test] + fn a_reset_time_already_past_still_waits_a_little() { + let at = 1_788_546_972.0; + let spent = snapshot( + UsageState::Ok, + vec![window(100.0, Some("2020-01-01T00:00:00+00:00"))], + ); + assert_eq!( + decide(Some(&spent), &owed(at - 60.0), at), + Step::WaitUntil(at + AT_LEAST) + ); + } + + #[test] + fn the_earliest_spent_window_is_the_one_worth_waiting_on() { + let at = 1_788_546_972.0; + let soon = "2026-09-05T12:00:00+00:00"; + let far = "2026-09-09T12:00:00+00:00"; + let mut weekly = window(100.0, Some(far)); + weekly.kind = "weekly_all".to_string(); + let spent = snapshot(UsageState::Ok, vec![window(100.0, Some(soon)), weekly]); + assert_eq!( + decide(Some(&spent), &owed(at - 60.0), at), + Step::WaitUntil(epoch_of(Some(soon)).expect("parses")) + ); + } + + #[test] + fn a_meter_that_could_not_be_asked_never_sends() { + let at = 1_788_546_972.0; + for state in [ + UsageState::NotLoggedIn, + UsageState::Unreachable { + detail: "no route".to_string(), + }, + UsageState::Failed { + detail: "429".to_string(), + }, + ] { + let broken = snapshot(state.clone(), Vec::new()); + assert_eq!( + decide(Some(&broken), &owed(at - 60.0), at), + Step::WaitUntil(at + BACKOFF), + "{state:?}" + ); + } + // And no snapshot at all -- a machine or provider edited away under a + // session that was waiting on it. + assert_eq!( + decide(None, &owed(at - 60.0), at), + Step::WaitUntil(at + BACKOFF) + ); + } + + #[test] + fn waiting_longer_than_any_real_window_gives_up_rather_than_retrying_for_ever() { + let at = 1_788_546_972.0; + let broken = snapshot( + UsageState::Unreachable { + detail: "no route".to_string(), + }, + Vec::new(), + ); + assert_eq!( + decide(Some(&broken), &owed(at - GIVE_UP - 1.0), at), + Step::GiveUp + ); + // A meter that answers is still allowed to send on the same tick: the + // ceiling bounds waiting, not resuming. + let clear = snapshot(UsageState::Ok, vec![window(3.0, None)]); + assert_eq!( + decide(Some(&clear), &owed(at - GIVE_UP - 1.0), at), + Step::Send + ); + } +} diff --git a/server/src/routes.rs b/server/src/routes.rs index b369c92..b42ff11 100644 --- a/server/src/routes.rs +++ b/server/src/routes.rs @@ -52,6 +52,8 @@ //! DELETE /sessions/{id} kill process, delete transcript + files //! (?deleteForeign=true removes the machine's own copy too) //! POST /sessions/{id}/notify {notify} -- announce this one or not +//! POST /sessions/{id}/auto-resume {autoResume, message?} -- carry on by itself +//! once the account's usage limit lifts //! GET /notifications SSE: every session's attention-wanting //! moments, live only (see `notifications`) //! GET /defaults {effort} -- what a new session starts at @@ -141,6 +143,7 @@ pub fn router(manager: Arc) -> Router { .route("/sessions/{id}/effort", post(set_effort)) .route("/defaults", get(defaults).post(set_defaults)) .route("/sessions/{id}/notify", post(set_notify)) + .route("/sessions/{id}/auto-resume", post(set_auto_resume)) .route("/notifications", get(notifications)) .route("/sessions/{id}/compact", post(compact)) .route("/sessions/{id}/command", post(command)) @@ -1440,6 +1443,33 @@ async fn set_notify( Ok(StatusCode::NO_CONTENT) } +#[derive(Deserialize)] +#[serde(rename_all = "camelCase", deny_unknown_fields)] +struct AutoResumeRequest { + auto_resume: bool, + /// What to send when the limit lifts. Absent -- and empty, which is what a + /// cleared field sends -- means this app's own default word, which is a + /// choice a caller has to be able to make rather than only start in. + #[serde(default)] + message: Option, +} + +/// Turns auto-resume on or off, and sets what it would say. +/// +/// One request for both, because they are one decision: switching it on +/// without saying what to send is the ordinary case, and changing the words +/// while it is off is how somebody sets it up before it is needed. +async fn set_auto_resume( + State(manager): State>, + UrlPath(id): UrlPath, + axum::Json(body): axum::Json, +) -> Result { + manager + .set_session_auto_resume(&id, body.auto_resume, body.message.as_deref()) + .map_err(bad_request)?; + Ok(StatusCode::NO_CONTENT) +} + #[derive(Deserialize)] #[serde(deny_unknown_fields)] struct CommandRequest { diff --git a/server/src/session/claude/translate.rs b/server/src/session/claude/translate.rs index 3baae30..39c1650 100644 --- a/server/src/session/claude/translate.rs +++ b/server/src/session/claude/translate.rs @@ -224,12 +224,12 @@ impl Translator { .and_then(Value::as_bool) .unwrap_or(false) { - events.push(Event::Error { - message: message - .get("result") - .and_then(Value::as_str) - .unwrap_or("the turn ended with an error") - .to_string(), + let said = message.get("result").and_then(Value::as_str); + events.push(match said.and_then(usage_limit) { + Some(resets_at) => Event::LimitReached { resets_at }, + None => Event::Error { + message: said.unwrap_or("the turn ended with an error").to_string(), + }, }); } let context = self.context.take(); @@ -603,6 +603,37 @@ impl Translator { } } +/// Whether a failed turn failed because the account is out of quota, and when +/// the CLI said the limit lifts. +/// +/// The wording is the CLI's: a turn stopped by the limit ends with `is_error` +/// and a result of `Claude AI usage limit reached|1788546972`, the reset being +/// epoch seconds after a pipe. Matched on the sentence rather than on a code +/// because the CLI sends none, so this is deliberately loose about everything +/// but the four words. +/// +/// The two `None`s mean different things and both are real. The outer one is +/// "some other failure". The inner one is "the limit is reached and the CLI did +/// not say until when" -- which is not a reason to invent a time: `crate::resume` +/// asks the usage endpoint before sending anything, and that answer is the one +/// that decides. +/// +/// Milliseconds are accepted as well as seconds and told apart by magnitude, +/// since a wrong guess would schedule a resume tens of thousands of years out +/// and look exactly like auto-resume being broken. +fn usage_limit(result: &str) -> Option> { + if !result.to_ascii_lowercase().contains("usage limit reached") { + return None; + } + let stamp = result + .rsplit('|') + .next() + .and_then(|tail| tail.trim().parse::().ok()) + .filter(|stamp| *stamp > 0.0) + .map(|stamp| if stamp > 1e11 { stamp / 1000.0 } else { stamp }); + Some(stamp) +} + /// A string field that is there and not empty, or `None`. The CLI omits these /// rather than sending them empty, but a caller that sends `""` means the same /// thing and should not produce a description that draws as a blank line. @@ -1275,6 +1306,55 @@ mod tests { ); } + /// Running out of quota is a state, not a failure of the work. + /// + /// The naive reading -- an error result like any other -- is what shipped + /// before this: the transcript said "Claude AI usage limit reached|…" in + /// red, which is neither readable nor actionable, and nothing above the + /// driver could tell it apart from a broken tool call. + #[test] + fn a_turn_stopped_by_the_usage_limit_says_so_and_carries_the_reset() { + let dir = tempfile::tempdir().expect("tempdir"); + let mut translator = Translator::new(dir.path().to_path_buf()); + let events = translate_lines( + &mut translator, + &[ + r#"{"type":"result","subtype":"error_during_execution","is_error":true,"result":"Claude AI usage limit reached|1788546972","usage":{}}"#, + ], + ); + assert_eq!( + events[0], + Event::LimitReached { + resets_at: Some(1_788_546_972.0) + } + ); + } + + #[test] + fn a_limit_the_cli_gave_no_reset_for_is_reported_without_one() { + let dir = tempfile::tempdir().expect("tempdir"); + let mut translator = Translator::new(dir.path().to_path_buf()); + let events = translate_lines( + &mut translator, + &[ + r#"{"type":"result","subtype":"error_during_execution","is_error":true,"result":"Claude AI usage limit reached","usage":{}}"#, + ], + ); + // Not a time this side invented: the meter is asked before anything is + // sent, and a made-up reset would only decide when to ask. + assert_eq!(events[0], Event::LimitReached { resets_at: None }); + } + + #[test] + fn a_reset_in_milliseconds_is_not_read_as_the_year_58000() { + assert_eq!( + usage_limit("Claude AI usage limit reached|1788546972000"), + Some(Some(1_788_546_972.0)) + ); + // And anything that is not the limit stays an ordinary failure. + assert_eq!(usage_limit("something broke"), None); + } + /// Pressing Stop is not a failure, and the CLI cannot tell you which it was. /// /// An interrupted turn arrives as exactly the same shape a broken one does, diff --git a/server/src/session/driver.rs b/server/src/session/driver.rs index c0e53ad..3af525b 100644 --- a/server/src/session/driver.rs +++ b/server/src/session/driver.rs @@ -302,6 +302,25 @@ pub enum Event { /// it, which is why this is written down rather than left to be inferred /// from a second example that does not exist. Cleared, + /// The account behind this session has no quota left, so the turn stopped + /// without finishing. + /// + /// Its own event rather than an [`Event::Error`] carrying the dialect's + /// sentence, because two things act on it that cannot read English: the + /// transcript draws it as a state the session is in rather than as a + /// failure of something it did, and `crate::resume` schedules the message + /// that picks the work back up. Recognising it belongs to the driver, which + /// is the only layer that knows its dialect's wording -- above here nothing + /// matches on strings. + /// + /// `resets_at` is epoch seconds, and `None` is a real state: the dialect + /// said the limit was hit without saying when it lifts. Nothing here + /// invents one -- what the wait is actually decided against is the usage + /// endpoint, and this is the hint that starts the waiting. + LimitReached { + #[serde(default, skip_serializing_if = "Option::is_none")] + resets_at: Option, + }, Error { message: String, }, diff --git a/server/src/session/echo.rs b/server/src/session/echo.rs index 454ee1d..20c2c3f 100644 --- a/server/src/session/echo.rs +++ b/server/src/session/echo.rs @@ -33,6 +33,13 @@ //! `/usage 42 never`, `/usage notloggedin`, `/usage unreachable`, //! `/usage failed`. The vocabulary is `usage::Fixture`'s, where the states //! live. +//! - `/limit [minutes]` -- a turn that stops because the account is out of +//! quota, saying the limit lifts in `minutes` (default 5, and `never` for a +//! limit with no stated reset). What it exists for is auto-resume, which is +//! otherwise reachable only by actually exhausting somebody's account: pair +//! it with `/usage 100 5` for a meter that agrees, and then `/usage 20` for +//! the moment the limit lifts. The wait itself is decided by the meter, so +//! those two commands are the whole rig. //! - `/compact` -- a compaction, start to finish. //! - `/stream N` -- one long answer in N small pieces, 50ms apart: the shape a //! real model's reply arrives in, and the one where the row a reader is @@ -344,6 +351,39 @@ impl EchoDriver { return; } + // A turn that ends the way a real one does when the account runs out: + // the same event a real driver reports, so what acts on it -- the + // transcript row and `crate::resume` -- is exercised rather than + // imitated. The meter it should agree with is `/usage`'s fixture, + // deliberately separate: the two disagreeing is a state worth being + // able to produce, since it is what a stale reset time looks like. + if let Some(rest) = text.strip_prefix("/limit") { + if announce { + self.emit(Event::MessageTaken { + id: None, + text: text.clone(), + attachments, + }); + } + let rest = rest.trim(); + let resets_at = match rest { + "never" | "none" => None, + "" => Some(super::now() + 5.0 * 60.0), + minutes => Some(super::now() + minutes.parse::().unwrap_or(5.0) * 60.0), + }; + self.emit(Event::Status { + state: SessionStatus::Running, + }); + self.emit(Event::AssistantText { + delta: "Working on it".to_string(), + }); + self.emit(Event::LimitReached { resets_at }); + self.emit(Event::Status { + state: SessionStatus::Idle, + }); + return; + } + // The same word the real CLI takes, so a phone drives both the same way. // `Driver::compact` is what the manager's route calls; this is the typed // path onto it. diff --git a/server/src/session/mod.rs b/server/src/session/mod.rs index a13601e..463c658 100644 --- a/server/src/session/mod.rs +++ b/server/src/session/mod.rs @@ -28,7 +28,8 @@ use serde::Serialize; use tokio::sync::{broadcast, mpsc}; use crate::config::{ - Config, DriverKind, ProviderConfig, SessionConfig, SetupConfig, SshConfig, TokenEntry, + Config, DEFAULT_RESUME_MESSAGE, DriverKind, ProviderConfig, ScheduledResume, SessionConfig, + SetupConfig, SshConfig, TokenEntry, }; use claude::ClaudeDriver; use driver::{ @@ -49,6 +50,44 @@ const EVENT_BUFFER: usize = 256; /// -- the newest "your turn" is the one still true. const NOTIFICATION_BUFFER: usize = 64; +/// Fan-out buffer for limit reports. One per session per rate-limit window, +/// so a handful a day across everything -- but sized like the notifications +/// above rather than at 1, because the only subscriber is a task that may be +/// mid-tick when several arrive. +const LIMIT_BUFFER: usize = 64; + +/// How long after a limit with no stated reset to ask the meter about it. +/// Short, because the meter is the authority and this is only how soon it is +/// worth the first question. +const FIRST_CHECK: f64 = 60.0; + +/// A session that stopped because its account is out of quota, as the pump +/// saw it. +/// +/// Broadcast downward rather than acted on here, for the reason `Shared` +/// gives: the pump runs underneath the manager and reaching back up would +/// invert that. `crate::resume` is the one subscriber, and what it does with +/// this is decided by the session's own `auto_resume`. +/// The two channels a pump reports on, which carry what this layer records +/// but does not act on: what a phone should be told, and what +/// `crate::resume` should schedule. +/// +/// One struct because they travel together through every launch and every +/// pump, and a second one arriving should not be a third parameter on both. +#[derive(Clone)] +pub struct Announcements { + notifications: broadcast::Sender, + limits: broadcast::Sender, +} + +#[derive(Debug, Clone)] +pub struct LimitHit { + pub session_id: String, + /// Epoch seconds the dialect said the limit lifts, and `None` where it + /// said nothing. Only ever a hint -- see [`Event::LimitReached`]. + pub resets_at: Option, +} + pub fn now() -> f64 { SystemTime::now() .duration_since(UNIX_EPOCH) @@ -96,6 +135,54 @@ pub enum NotificationKind { Finished, } +/// What a session's auto-resume setting looks like from outside: on or off, +/// what it would say, and when it next intends to check. +/// +/// One struct rather than three parameters on [`LiveSession::info`], and read +/// from the config rather than from the launch snapshot beside it, for the +/// reason `cwd` is: all three change under a running session. +#[derive(Debug, Clone)] +pub struct AutoResumeView { + pub on: bool, + pub message: String, + pub at: Option, +} + +impl AutoResumeView { + fn of(meta: &SessionConfig) -> Self { + Self { + on: meta.auto_resume, + message: resume_message(meta), + at: meta.resume.map(|scheduled| scheduled.at), + } + } +} + +/// A session with a message owed to it once its account has quota again -- +/// see [`SessionManager::owed_resumes`]. +/// +/// Carries no message: what to send is read under the lock at the moment it is +/// sent (see [`SessionManager::resume_now`]), because a wait lasts hours and +/// the words can be edited from the phone inside one. +#[derive(Debug, Clone)] +pub struct OwedResume { + pub session_id: String, + /// The machine whose account ran out, which is the one to ask. + pub setup: String, + /// Which meter reports on it -- a `crate::usage::UsageProvider::name`, the + /// same pairing `SessionInfo::usage_provider` uses. + pub provider: &'static str, + pub scheduled: ScheduledResume, +} + +/// What a session's auto-resume says, with the default filled in. One place, +/// so the phone is shown the words that would actually be sent. +fn resume_message(meta: &SessionConfig) -> String { + meta.auto_resume_message + .clone() + .unwrap_or_else(|| DEFAULT_RESUME_MESSAGE.to_string()) +} + /// One row of `GET /sessions`. #[derive(Debug, Clone, Serialize)] #[serde(rename_all = "camelCase")] @@ -162,6 +249,19 @@ pub struct SessionInfo { /// reason `permission_mode` is: a switch that guesses its own position /// is how you turn something off while believing you are reading it. pub notify: bool, + /// Whether this session sends itself a message when its account's usage + /// limit lifts, and what that message says. Reported for the same reason + /// `notify` is. + pub auto_resume: bool, + /// The words that would be sent, with the default already filled in -- + /// the phone shows what would actually happen rather than an empty field + /// meaning "something". + pub auto_resume_message: String, + /// Epoch seconds this session next intends to check whether the limit has + /// lifted, and absent when nothing is waiting. A measurement rather than + /// a promise: what decides is the meter, asked at that moment. + #[serde(skip_serializing_if = "Option::is_none")] + pub resume_at: Option, pub status: SessionStatus, pub last_activity: f64, pub created: f64, @@ -450,6 +550,7 @@ impl LiveSession { effort: Option<&str>, imported: bool, kind: Option, + resume: AutoResumeView, ) -> SessionInfo { SessionInfo { id: self.meta.id.clone(), @@ -466,6 +567,9 @@ impl LiveSession { takes_effort: kind.is_some_and(DriverKind::takes_effort), context_tokens: *self.shared.context_tokens.lock().unwrap(), notify: *self.shared.notify.lock().unwrap(), + auto_resume: resume.on, + auto_resume_message: resume.message, + resume_at: resume.at, max_image_edge: kind.and_then(DriverKind::max_image_edge), usage_provider: kind.and_then(DriverKind::usage_provider), imported, @@ -491,8 +595,9 @@ pub struct SessionManager { /// Downloaded GGUF models, shared by every session that names one, /// which is why they sit beside the session directories. models_dir: PathBuf, - /// Where every session's pump sends what a phone should be told about. - notifications: broadcast::Sender, + /// Where every session's pump reports what this layer does not act on -- + /// see [`Announcements`]. + announce: Announcements, /// Imports and deletes running against a machine's Claude Code /// sessions: like the notifications, state the phone reads but does not /// own. @@ -520,6 +625,11 @@ impl SessionManager { wg_app_link::private::create_dir(&data_dir)?; let (notifications, _) = broadcast::channel(NOTIFICATION_BUFFER); + let (limits, _) = broadcast::channel(LIMIT_BUFFER); + let announce = Announcements { + notifications, + limits, + }; // Made here rather than passed in, and handed *out* to the usage // monitor by whoever wires the two together: every echo driver // this manager builds gets a clone, including the ones built @@ -540,7 +650,7 @@ impl SessionManager { models_dir: &models_dir, usage: &usage_fixture, }, - notifications.clone(), + announce.clone(), // Nothing is started here; see `Launching`. Launching::Restart, ) @@ -557,7 +667,7 @@ impl SessionManager { config_path, data_dir, models_dir, - notifications, + announce, pending: Arc::new(pending::Registry::default()), spawn_throwaway: false, usage_fixture, @@ -923,6 +1033,7 @@ impl SessionManager { meta.effort.as_deref(), import::read_cursor(&self.data_dir.join(&meta.id)).is_some(), kind_of(&inner.config, &meta.setup, &meta.provider), + AutoResumeView::of(meta), ), None => SessionInfo { id: meta.id.clone(), @@ -941,6 +1052,9 @@ impl SessionManager { usage_provider: kind_of(&inner.config, &meta.setup, &meta.provider) .and_then(DriverKind::usage_provider), notify: meta.notify, + auto_resume: meta.auto_resume, + auto_resume_message: resume_message(meta), + resume_at: meta.resume.map(|scheduled| scheduled.at), imported: import::read_cursor(&self.data_dir.join(&meta.id)).is_some(), keeps_own_transcript: keeps_own_transcript( &inner.config, @@ -959,8 +1073,15 @@ impl SessionManager { /// Every session's attention-wanting moments, on one stream. One /// connection for the whole backend rather than one per session: the /// phone subscribes while showing no session at all. + /// Every session running out of quota, on one stream -- the other half of + /// [`SessionManager::owed_resumes`]. Subscribed to by `crate::resume`, so + /// a limit hit is acted on when it happens rather than at the next tick. + pub fn subscribe_limits(&self) -> broadcast::Receiver { + self.announce.limits.subscribe() + } + pub fn subscribe_notifications(&self) -> broadcast::Receiver { - self.notifications.subscribe() + self.announce.notifications.subscribe() } /// Imports and deletes running against importable sessions -- see @@ -1055,6 +1176,12 @@ impl SessionManager { // On by default. Not offered at spawn: a session's first turn // is exactly the one somebody is waiting for. notify: true, + // Off, and not offered at spawn either -- for the opposite + // reason: this one spends quota with nobody watching, so it is + // asked for on a session somebody already has, never inherited. + auto_resume: false, + auto_resume_message: None, + resume: None, // Recorded on the session rather than remembered here, so // whichever server is running when the time comes knows what to // do with it -- see `SessionConfig::throwaway`. @@ -1067,7 +1194,7 @@ impl SessionManager { &setup, &provider, self.env(), - self.notifications.clone(), + self.announce.clone(), Launching::Asked(seed), )?; let mut candidate = inner.config.clone(); @@ -1088,6 +1215,7 @@ impl SessionManager { session.meta.effort.as_deref(), import::read_cursor(&self.data_dir.join(&id)).is_some(), Some(provider.kind), + AutoResumeView::of(&session.meta), ); inner.live.insert(id, session); Ok(info) @@ -1150,6 +1278,176 @@ impl SessionManager { Ok(()) } + /// Turns auto-resume on or off for one session, and sets what it will + /// say. + /// + /// Turning it off cancels anything already scheduled, which is the path + /// out of the state the previous call put the session in: a message left + /// owed by a switch somebody has since turned off would arrive hours + /// later with nothing on screen to explain it. + /// + /// An empty message is not a message -- it is what a cleared field sends + /// -- so it means [`DEFAULT_RESUME_MESSAGE`] rather than a session poked + /// with nothing to read. + pub fn set_session_auto_resume( + &self, + id: &str, + auto_resume: bool, + message: Option<&str>, + ) -> Result<()> { + let mut inner = self.inner.write().unwrap(); + if !inner.config.sessions.iter().any(|meta| meta.id == id) { + bail!("no session {id}"); + } + let mut candidate = inner.config.clone(); + for meta in candidate.sessions.iter_mut().filter(|meta| meta.id == id) { + meta.auto_resume = auto_resume; + meta.auto_resume_message = message + .map(str::trim) + .filter(|message| !message.is_empty()) + .map(str::to_string); + if !auto_resume { + meta.resume = None; + } + } + candidate.save(&self.config_path)?; + inner.config = candidate; + Ok(()) + } + + /// Records that a session ran out of quota, and when to look again. + /// + /// Does nothing for a session that does not auto-resume, and nothing for + /// one already waiting: a turn that fails twice against the same window + /// is the same wait, and taking the second report would push the check + /// back every time the session was poked. + /// + /// `resets_at` is the dialect's hint and is used only to decide when to + /// *ask*; [`crate::resume`] asks the meter before anything is sent. A + /// session told nothing is checked shortly, since the meter is the + /// authority either way. + pub fn note_limit(&self, id: &str, resets_at: Option) -> Result { + let mut inner = self.inner.write().unwrap(); + let meta = inner + .config + .sessions + .iter() + .find(|meta| meta.id == id) + .with_context(|| format!("no session {id}"))?; + if !meta.auto_resume || meta.resume.is_some() { + return Ok(false); + } + let at = now(); + let scheduled = ScheduledResume { + at: resets_at.unwrap_or(at + FIRST_CHECK), + since: at, + }; + let mut candidate = inner.config.clone(); + for meta in candidate.sessions.iter_mut().filter(|meta| meta.id == id) { + meta.resume = Some(scheduled); + } + candidate.save(&self.config_path)?; + inner.config = candidate; + Ok(true) + } + + /// Every session with a message owed to it, oldest schedule first. + /// + /// Carries what deciding needs rather than a session id to look things up + /// by, so the scheduler holds no lock while it makes a network call: the + /// machine and the meter to ask, and the words to send. + pub fn owed_resumes(&self) -> Vec { + let inner = self.inner.read().unwrap(); + let mut owed: Vec = inner + .config + .sessions + .iter() + .filter_map(|meta| { + let scheduled = meta.resume?; + Some(OwedResume { + session_id: meta.id.clone(), + setup: meta.setup.clone(), + provider: kind_of(&inner.config, &meta.setup, &meta.provider)? + .usage_provider()?, + scheduled, + }) + }) + .collect(); + owed.sort_by(|a, b| a.scheduled.at.total_cmp(&b.scheduled.at)); + owed + } + + /// Moves a scheduled check later (or earlier), leaving everything else + /// about it alone -- including when the limit was hit, which is what + /// bounds the retrying. + pub fn reschedule_resume(&self, id: &str, at: f64) -> Result<()> { + let mut inner = self.inner.write().unwrap(); + let mut candidate = inner.config.clone(); + for meta in candidate.sessions.iter_mut().filter(|meta| meta.id == id) { + if let Some(scheduled) = meta.resume.as_mut() { + scheduled.at = at; + } + } + candidate.save(&self.config_path)?; + inner.config = candidate; + Ok(()) + } + + /// Sends the message this session is owed and clears the schedule. + /// + /// Cleared first, and saved before the message goes out: a send that + /// fails leaves nothing owed, where a schedule left standing by a failed + /// send is one that fires again on the next tick and every tick after. + /// The session is started if it has none, exactly as any other message + /// does. + pub fn resume_now(&self, id: &str) -> Result { + let message = { + let mut inner = self.inner.write().unwrap(); + let meta = inner + .config + .sessions + .iter() + .find(|meta| meta.id == id) + .with_context(|| format!("no session {id}"))?; + let message = resume_message(meta); + let mut candidate = inner.config.clone(); + for meta in candidate.sessions.iter_mut().filter(|meta| meta.id == id) { + meta.resume = None; + } + candidate.save(&self.config_path)?; + inner.config = candidate; + message + }; + self.send_message(id, message.clone(), Vec::new())?; + Ok(message) + } + + /// Gives up on a scheduled resume, and says so in the transcript. + /// + /// In the transcript because that is where somebody looking at this + /// session will be: a wait that quietly stopped waiting is + /// indistinguishable from one still going, and the session is sitting + /// there having said nothing since the limit was hit. + pub fn abandon_resume(&self, id: &str, why: &str) -> Result<()> { + { + let mut inner = self.inner.write().unwrap(); + let mut candidate = inner.config.clone(); + for meta in candidate.sessions.iter_mut().filter(|meta| meta.id == id) { + meta.resume = None; + } + candidate.save(&self.config_path)?; + inner.config = candidate; + } + if let Some(session) = self.session(id) { + let _ = session.sink.send(Event::Error { + message: format!( + "auto-resume gave up on this session: {why}. Send it something to carry on." + ), + }); + } + Ok(()) + } + /// Renames a session: persisted, shown, and passed on to whatever is /// running it. /// @@ -1540,7 +1838,7 @@ impl SessionManager { &setup, &provider, self.env(), - self.notifications.clone(), + self.announce.clone(), Launching::Asked(None), )?; inner.live.insert(id.to_string(), session); @@ -1935,7 +2233,7 @@ fn launch( setup: &SetupConfig, provider: &ProviderConfig, env: Env<'_>, - notifications: broadcast::Sender, + announce: Announcements, why: Launching, ) -> Result> { let dir = env.data_dir.join(&meta.id); @@ -2069,7 +2367,7 @@ fn launch( Arc::clone(&shared), events.clone(), Arc::clone(&commands), - notifications, + announce, )); Ok(Arc::new(LiveSession { @@ -2191,7 +2489,7 @@ async fn pump( shared: Arc, events: broadcast::Sender, commands: Arc, - notifications: broadcast::Sender, + announce: Announcements, ) { // Messages the session has been given and not started reading, which is // what makes a turn ending not the same thing as the work ending. @@ -2270,7 +2568,7 @@ async fn pump( { // No subscribers is the ordinary case -- nobody has // the app open -- and it is not an error. - let _ = notifications.send(Notification { + let _ = announce.notifications.send(Notification { session_id: id.clone(), title: shared.title.lock().unwrap().clone(), kind, @@ -2278,6 +2576,16 @@ async fn pump( }); } } + if let Event::LimitReached { resets_at } = &entry.event { + // Sent whether or not this session auto-resumes: whether + // to act is the manager's decision, and it is the one + // holding the setting. No subscribers is the ordinary + // case -- nothing waits on this in the tests. + let _ = announce.limits.send(LimitHit { + session_id: id.clone(), + resets_at: *resets_at, + }); + } *shared.last_activity.lock().unwrap() = ts; *shared.written.lock().unwrap() += 1; // The turn's own first line, kept for whatever arrives at the @@ -2631,6 +2939,99 @@ mod tests { ); } + /// The whole of the server's half of auto-resume, driven by echo: a + /// limit is reported, the session that asked for it is scheduled, and the + /// one that did not is left alone. + /// + /// Echo rather than the Claude CLI on purpose -- reaching this state for + /// real means exhausting an account, and the event both drivers report is + /// the same one. + #[tokio::test] + async fn a_limit_schedules_a_resume_only_where_one_was_asked_for() { + let dir = tempfile::tempdir().expect("tempdir"); + let config_path = dir.path().join("config.ron"); + let data_dir = dir.path().join("sessions"); + seed_echo_only(&config_path); + let manager = SessionManager::new(config_path, data_dir.clone(), data_dir.join("models")) + .expect("manager"); + let quiet = manager.spawn_session(echo_spec()).expect("spawn"); + let resuming = manager.spawn_session(echo_spec()).expect("spawn"); + manager + .set_session_auto_resume(&resuming.id, true, Some("carry on")) + .expect("on"); + + let mut limits = manager.subscribe_limits(); + for id in [&quiet.id, &resuming.id] { + manager + .session(id) + .expect("live") + .send_message("/limit 10".to_string(), Vec::new()); + } + // Both report; only one is owed anything. Drained rather than slept + // through, so the assertions below cannot run before the events they + // are about. + for _ in 0..2 { + let hit = tokio::time::timeout(Duration::from_secs(5), limits.recv()) + .await + .expect("a limit within five seconds") + .expect("channel open"); + crate::resume::note(&manager, &hit.session_id, hit.resets_at); + } + + let owed = manager.owed_resumes(); + assert_eq!( + owed.iter().map(|owed| &owed.session_id).collect::>(), + vec![&resuming.id], + "a session nobody switched on was scheduled anyway" + ); + // The dialect's hint decides when to *ask*, so it is what was written + // down -- ten minutes out, not the minute a session told nothing gets. + assert!( + owed[0].scheduled.at - now() > FIRST_CHECK, + "the reset time the session reported was ignored" + ); + + // Turning it off is the way out of the state turning it on created. + manager + .set_session_auto_resume(&resuming.id, false, None) + .expect("off"); + assert!( + manager.owed_resumes().is_empty(), + "a message stayed owed after auto-resume was switched off" + ); + } + + /// What the phone reads back, which is what its switch and its text field + /// are drawn from. + #[tokio::test] + async fn a_session_reports_its_auto_resume_setting_and_its_default_words() { + let dir = tempfile::tempdir().expect("tempdir"); + let config_path = dir.path().join("config.ron"); + let data_dir = dir.path().join("sessions"); + seed_echo_only(&config_path); + let manager = SessionManager::new(config_path, data_dir.clone(), data_dir.join("models")) + .expect("manager"); + let info = manager.spawn_session(echo_spec()).expect("spawn"); + assert!(!info.auto_resume); + // The default is reported rather than left empty: the field shows + // what would actually be sent. + assert_eq!(info.auto_resume_message, DEFAULT_RESUME_MESSAGE); + assert_eq!(info.resume_at, None); + + // An empty message is what a cleared field sends, and means the + // default rather than a session poked with nothing to read. + manager + .set_session_auto_resume(&info.id, true, Some(" ")) + .expect("on"); + let fresh = manager + .sessions() + .into_iter() + .find(|session| session.id == info.id) + .expect("listed"); + assert!(fresh.auto_resume); + assert_eq!(fresh.auto_resume_message, DEFAULT_RESUME_MESSAGE); + } + /// The switch reaches the running pump, not just the config file. The /// failure is silent in the direction that matters: a /// `set_session_notify(false)` writing only the config looks correct on @@ -2659,7 +3060,20 @@ mod tests { // open to look one up on. assert_eq!( first.title, - session.info("m", None, None, false, None).title + session + .info( + "m", + None, + None, + false, + None, + AutoResumeView { + on: false, + message: DEFAULT_RESUME_MESSAGE.to_string(), + at: None, + }, + ) + .title ); manager.set_session_notify(&info.id, false).expect("off"); From 7b63330aaad94520860c79342276af0631a9b453 Mon Sep 17 00:00:00 2001 From: iris <2+iris@noreply.localhost> Date: Sat, 5 Sep 2026 11:58:17 -0400 Subject: [PATCH 06/12] Say when a usage 401 is an expired login, not an unreachable endpoint A 401 is the endpoint answering and refusing the stored OAuth token, which Claude Code refreshes as it runs -- so a machine whose CLI has been idle hands us a stale one. Reporting it as "usage endpoint unreachable" pointed at the network instead of at the one thing that fixes it. --- server/src/usage.rs | 39 ++++++++++++++++++++++++++++++--------- 1 file changed, 30 insertions(+), 9 deletions(-) diff --git a/server/src/usage.rs b/server/src/usage.rs index 214e7a1..78d20ea 100644 --- a/server/src/usage.rs +++ b/server/src/usage.rs @@ -206,15 +206,8 @@ impl UsageProvider for ClaudeUsage { .and_then(|mut response| response.body_mut().read_to_string()) { Ok(text) => text, - Err(err) => { - // The error string can embed the URL but never the token. - return self.snapshot( - UsageState::Failed { - detail: format!("usage endpoint unreachable: {err}"), - }, - Vec::new(), - ); - } + // The error string can embed the URL but never the token. + Err(err) => return self.snapshot(UsageState::Failed { detail: why(&err) }, Vec::new()), }; let body: Value = match serde_json::from_str(&text) { Ok(body) => body, @@ -231,6 +224,25 @@ impl UsageProvider for ClaudeUsage { } } +/// What a failed call to the usage endpoint should say. +/// +/// A status is not a network fault and must not be reported as one: the +/// endpoint answered, and 401 in particular says the stored token has expired +/// -- Claude Code refreshes it as it runs, so a machine whose CLI has been +/// idle long enough hands us a stale one. That is fixable, and the message is +/// the only place anybody finds out how. +fn why(err: &ureq::Error) -> String { + match err { + ureq::Error::StatusCode(401) => { + format!( + "the Claude login on this machine has expired (401); run `claude` there, or re-run `/login`, to refresh {CREDENTIALS}" + ) + } + ureq::Error::StatusCode(code) => format!("usage endpoint refused the request: HTTP {code}"), + other => format!("usage endpoint unreachable: {other}"), + } +} + /// Which kind of "no credentials" a failed read was. /// /// The distinction is the point of having both states. `cat` failing because @@ -631,6 +643,15 @@ mod tests { use super::*; use crate::config::DriverKind; + #[test] + fn an_expired_login_is_not_reported_as_an_unreachable_endpoint() { + let stale = why(&ureq::Error::StatusCode(401)); + assert!(stale.contains("expired"), "{stale}"); + assert!(!stale.contains("unreachable"), "{stale}"); + assert!(why(&ureq::Error::StatusCode(500)).contains("HTTP 500")); + assert!(why(&ureq::Error::HostNotFound).contains("unreachable")); + } + #[test] fn parses_the_limits_array_defensively() { // Trimmed from a live 2026-08-24 response. From eff5c8b0c090bcfeac205e6e82b9c3be7cebaa27 Mon Sep 17 00:00:00 2001 From: iris <2+iris@noreply.localhost> Date: Sat, 5 Sep 2026 12:07:34 -0400 Subject: [PATCH 07/12] Let the machine's own CLI refresh an expired token, and retry once A 401 from the usage endpoint means the stored access token has expired. Refreshing it here is not an option: Anthropic's OAuth rotates the refresh token, so a second refresher invalidates the CLI's copy and forces a re-login on a machine that usually has a live session on it. So run the CLI there instead and re-read what it wrote. `doctor` rather than `auth status`: probed against 2.1.258 with an invalid token, `auth status` answers loggedIn:true from the file alone and never reaches the network. The same probe showed a failed refresh blanks both tokens, which is why this stays on the 401 path. Also gives ProviderConfig one program() so the CLI's default path is not written down twice. --- server/src/config.rs | 25 ++++++ server/src/session/claude.rs | 2 +- server/src/session/llama.rs | 2 +- server/src/usage.rs | 157 ++++++++++++++++++++++++++++------- 4 files changed, 154 insertions(+), 32 deletions(-) diff --git a/server/src/config.rs b/server/src/config.rs index 475a4cb..7e5fb8f 100644 --- a/server/src/config.rs +++ b/server/src/config.rs @@ -90,6 +90,16 @@ pub struct ProviderConfig { pub models: Vec, } +impl ProviderConfig { + /// The executable to run for this provider: its override, or its kind's + /// default. + pub fn program(&self) -> &str { + self.command + .as_deref() + .unwrap_or(self.kind.default_program()) + } +} + /// How to reach a setup that isn't this machine, with the system `ssh` client /// -- so `~/.ssh/config`, agents and jump hosts all keep working, and there is /// one place to configure connections. A remote session is the identical @@ -194,6 +204,21 @@ impl DriverKind { } } + /// The executable a provider of this kind runs when it names none. + /// + /// Here rather than at each spawn site because it is not only the spawn + /// that runs it: `usage` runs the Claude CLI too, to have it refresh its + /// own OAuth token, and a default that disagreed with the driver's would + /// ask the wrong binary on a machine with two installs. + pub fn default_program(self) -> &'static str { + match self { + Self::ClaudeCli => "claude", + Self::LlamaCpp => "llama-server", + // Echo is translated in-process; nothing is spawned for it. + Self::Echo => "echo", + } + } + /// Whether the conversation exists outside this app, so that deleting the /// session here does not end it. /// diff --git a/server/src/session/claude.rs b/server/src/session/claude.rs index 5bffaaa..69c4d54 100644 --- a/server/src/session/claude.rs +++ b/server/src/session/claude.rs @@ -392,7 +392,7 @@ impl ClaudeDriver { let stdout = create_log(&session_dir.join(STDOUT_LOG))?; let stderr = create_log(&session_dir.join(STDERR_LOG))?; - let program = provider.command.as_deref().unwrap_or("claude"); + let program = provider.program(); let launch = Launch::new(program, args, meta.cwd.as_deref()); let child = transport.spawn( &launch, diff --git a/server/src/session/llama.rs b/server/src/session/llama.rs index 108467d..c9ce819 100644 --- a/server/src/session/llama.rs +++ b/server/src/session/llama.rs @@ -145,7 +145,7 @@ impl LlamaDriver { } } - let program = provider.command.as_deref().unwrap_or("llama-server"); + let program = provider.program(); let launch = Launch::new(program, args, meta.cwd.as_deref()).reaching(forward); // Its output goes to files, not pipes. Not only so the process can // outlive this server: nothing ever read those pipes, so a chatty diff --git a/server/src/usage.rs b/server/src/usage.rs index 78d20ea..1f14a9e 100644 --- a/server/src/usage.rs +++ b/server/src/usage.rs @@ -141,6 +141,10 @@ pub struct ClaudeUsage { pub setup_name: String, /// How to reach that machine. `Here` for the backend's own. pub transport: Transport, + /// The CLI to run there, for the one thing this asks of it: refreshing its + /// own expired token. The provider's, so a machine with the CLI somewhere + /// odd is asked at the same path its sessions run. + pub program: String, } /// Where Claude Code keeps its credentials, as a shell word rather than a path: @@ -198,46 +202,109 @@ impl UsageProvider for ClaudeUsage { Ok(token) => token, Err(state) => return self.snapshot(state, Vec::new()), }; - let text = match ureq::get(USAGE_URL) + let body = match self.call(&token) { + Ok(body) => body, + Err(Refused::Other(detail)) => { + return self.snapshot(UsageState::Failed { detail }, Vec::new()); + } + Err(Refused::Unauthorized) => match self.after_cli_refresh(&token) { + Ok(body) => body, + Err(state) => return self.snapshot(state, Vec::new()), + }, + }; + self.snapshot(UsageState::Ok, parse_windows(&body)) + } +} + +/// Why one call to the usage endpoint did not produce numbers. +/// +/// 401 is apart from the rest because it is the only one with a way out: the +/// endpoint answered, and it means the access token has expired rather than +/// that anything is broken. +enum Refused { + Unauthorized, + Other(String), +} + +impl ClaudeUsage { + /// One call to the endpoint with one token. + fn call(&self, token: &str) -> Result { + let text = ureq::get(USAGE_URL) .header("Authorization", &format!("Bearer {token}")) .header("anthropic-beta", "oauth-2025-04-20") .header("User-Agent", USER_AGENT) .call() .and_then(|mut response| response.body_mut().read_to_string()) - { - Ok(text) => text, // The error string can embed the URL but never the token. - Err(err) => return self.snapshot(UsageState::Failed { detail: why(&err) }, Vec::new()), - }; - let body: Value = match serde_json::from_str(&text) { - Ok(body) => body, - Err(err) => { - return self.snapshot( - UsageState::Failed { - detail: format!("usage endpoint sent non-JSON: {err}"), - }, - Vec::new(), - ); - } - }; - self.snapshot(UsageState::Ok, parse_windows(&body)) + .map_err(|err| match err { + ureq::Error::StatusCode(401) => Refused::Unauthorized, + other => Refused::Other(why(&other)), + })?; + serde_json::from_str(&text) + .map_err(|err| Refused::Other(format!("usage endpoint sent non-JSON: {err}"))) + } + + /// Have the machine's own CLI refresh its token, then ask once more. + /// + /// **The CLI does the refresh, never this.** Anthropic's OAuth rotates the + /// refresh token, so whoever refreshes second presents a dead one and the + /// machine is logged out until somebody runs `/login` on it -- and the + /// machine we would be refreshing on is usually one with a live session of + /// its own. Running the CLI keeps it the only writer of + /// `.credentials.json`. + /// + /// `doctor` rather than the `auth status` it reads like, measured against + /// this CLI (2.1.258) on 2026-09-05 with a deliberately invalid token: + /// `auth status` reports `loggedIn: true` off the file alone and never + /// touches the network, so it would have refreshed nothing while looking + /// like it had. `doctor` resolves the account, which is what makes it + /// refresh, and it spends no quota. The same probe showed what a *failed* + /// refresh does -- the CLI blanks both tokens -- so this must stay on the + /// 401 path, where the access token is already dead, and never be used to + /// refresh speculatively. + /// + /// Only a token that actually changed is retried, so a CLI that refreshed + /// nothing costs one call rather than two, and this cannot become a loop. + fn after_cli_refresh(&self, stale: &str) -> Result { + let launch = Launch::new(&self.program, vec!["doctor".to_string()], None); + if let Err(err) = self.transport.capture_blocking(&launch) { + return Err(UsageState::Failed { + detail: format!( + "the Claude login on {} has expired, and `{} doctor` couldn't be run there to refresh it: {err:#}", + self.setup_name, self.program + ), + }); + } + let fresh = self.access_token()?; + if fresh == stale { + return Err(self.still_expired()); + } + self.call(&fresh).map_err(|err| match err { + Refused::Unauthorized => self.still_expired(), + Refused::Other(detail) => UsageState::Failed { detail }, + }) + } + + /// A login the CLI could not renew: the one state here somebody has to act + /// on, so it says where and what to run. + fn still_expired(&self) -> UsageState { + UsageState::Failed { + detail: format!( + "the Claude login on {} has expired and could not be refreshed; run `{} /login` there", + self.setup_name, self.program + ), + } } } /// What a failed call to the usage endpoint should say. /// /// A status is not a network fault and must not be reported as one: the -/// endpoint answered, and 401 in particular says the stored token has expired -/// -- Claude Code refreshes it as it runs, so a machine whose CLI has been -/// idle long enough hands us a stale one. That is fixable, and the message is -/// the only place anybody finds out how. +/// endpoint answered. 401 never reaches here -- it has its own way out in +/// [`ClaudeUsage::after_cli_refresh`] -- so what is left is a refusal nobody +/// on this side can fix. fn why(err: &ureq::Error) -> String { match err { - ureq::Error::StatusCode(401) => { - format!( - "the Claude login on this machine has expired (401); run `claude` there, or re-run `/login`, to refresh {CREDENTIALS}" - ) - } ureq::Error::StatusCode(code) => format!("usage endpoint refused the request: HTTP {code}"), other => format!("usage endpoint unreachable: {other}"), } @@ -552,6 +619,7 @@ fn providers_for(setup: &SetupConfig, fixture: &Fixture) -> Vec Date: Sat, 5 Sep 2026 13:41:15 -0400 Subject: [PATCH 08/12] Show a session's subagents as subcards, each with a read-only transcript A subagent is a second transcript owned by a session, in the same event model, with no process and no controls. The claude translator routes lines carrying parent_tool_use_id to a per-subagent translator and transcript under /subagents/; three routes expose the list, a transcript page and the SSE stream. Echo grows /subagent [n] as the rig. On the phone a card with subagents ends in a chevron expander, collapsed by default, opening to outlined subcards styled like dev-updater's components; a subcard opens SessionScreen in read-only form, addressed through TranscriptAddress so paging, cache and stream are shared. Design in SUBAGENTS.md; choices awaiting review in DECISIONS.md. Co-Authored-By: Claude Fable 5.1 --- AGENTS.md | 4 + DECISIONS.md | 44 ++ PLAN.md | 14 + SUBAGENTS.md | 123 ++++ .../src/main/kotlin/com/example/aiapp/Api.kt | 42 +- .../main/kotlin/com/example/aiapp/AppRoot.kt | 52 +- .../kotlin/com/example/aiapp/EventStream.kt | 4 +- .../kotlin/com/example/aiapp/MainScreen.kt | 3 + .../com/example/aiapp/SessionListScreen.kt | 149 ++++- .../kotlin/com/example/aiapp/SessionScreen.kt | 601 ++++++++++-------- .../com/example/aiapp/TranscriptAddress.kt | 27 + .../com/example/aiapp/TranscriptCache.kt | 11 +- .../com/example/aiapp/TranscriptSource.kt | 10 +- .../com/example/aiapp/TranscriptCacheTest.kt | 2 +- server/src/routes.rs | 133 +++- server/src/session/claude.rs | 27 +- server/src/session/claude/translate.rs | 304 +++++++-- server/src/session/echo.rs | 111 +++- server/src/session/llama.rs | 5 + server/src/session/mod.rs | 129 +++- server/src/session/subagent.rs | 490 ++++++++++++++ 21 files changed, 1953 insertions(+), 332 deletions(-) create mode 100644 DECISIONS.md create mode 100644 SUBAGENTS.md create mode 100644 app/androidApp/src/main/kotlin/com/example/aiapp/TranscriptAddress.kt create mode 100644 server/src/session/subagent.rs diff --git a/AGENTS.md b/AGENTS.md index 9279365..58cc91e 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -61,6 +61,10 @@ Module-by-module intent is in PLAN.md's "Backend layout". projects version-locked to the commit this repo pins. What deliberately did **not** move is the API surface and the config *schema*: routes, drivers, sessions and setups are what makes this project itself. +- `SUBAGENTS.md` — a session's subagents as transcripts of their own + (`server/src/session/subagent.rs`, the subcards in `SessionListScreen.kt` + and the read-only form of `SessionScreen.kt`); `DECISIONS.md` holds the + choices made there that are still awaiting review. - `EXPLORER.md` — the file explorer's design (`server/src/files.rs` and `FilesScreen.kt` / `FileViewer.kt` / `FileEditor.kt`). - `TRANSCRIPT_CACHE.md` — the phone's copy of what it has been sent. Read it diff --git a/DECISIONS.md b/DECISIONS.md new file mode 100644 index 0000000..65ff286 --- /dev/null +++ b/DECISIONS.md @@ -0,0 +1,44 @@ +# Decisions awaiting review + +Choices made while working autonomously, for Bryan to keep or change. Each +says what was picked and why; the detail is in the design doc it names. +Delete an entry once it has been looked at. + +## Subagent views (2026-09-05, `SUBAGENTS.md`) + +Made on my own judgement, limited blast radius: + +1. **A subagent is a transcript, not a session.** It has no process, + controls or settings; it is addressed as `/sessions/{id}/subagents/{sub}` + and stored under the session's directory, so deleting the session takes + it. Alternative rejected: registering it as a session of its own, which + would give it a card in the main list and a driver that can do nothing. +2. **Read-only view is the session screen minus its controls**, rather than + a second, simpler transcript screen. Keeps paging, caching, selection + and rendering in one place. Cost: a `readOnly` mode threaded through + `SessionScreen`. +3. **The list only carries a count.** Each session row says how many + subagents it has; their titles and statuses are fetched when the card is + expanded. Keeps `GET /sessions` from reading every subagent transcript. + Consequence: an expanded card's statuses refresh with the list, not live. +4. **Expanded/collapsed is remembered per session on the phone**, not on + the server. Collapsed by default, per the transcript convention that new + things arrive collapsed. +5. **Subagents of imported sessions are not shown.** The import path still + skips `isSidechain` records; the CLI's own `subagents/agent-*.jsonl` files + are not read. Only subagents run while this backend was watching exist. +6. **Echo grows `/subagent [n]`** as the test rig, so nothing here needs a + paid turn to exercise. + +Deferred, because they reach further than this feature: + +- **Live status on the list.** Whether the session list should follow a + stream at all (it refreshes on demand today) decides whether subagent + status can ever be live there. Not changed. +- **Nested subagents.** A subagent's own Task calls are shown as tool calls + in its transcript and are not given transcripts of their own. Supporting + that is the same mechanism one level down, but the UI would need nested + expanders. + +- **The subagent status row says "context unknown".** Nothing measures a + subagent's context; the row could leave it out rather than admit it. diff --git a/PLAN.md b/PLAN.md index 3f1a0a0..c28c664 100644 --- a/PLAN.md +++ b/PLAN.md @@ -644,6 +644,20 @@ to end that way on 2026-09-05: the wait moved from the dialect's two minutes to the meter's seven when the meter changed its mind, and the message went out on the first check after the meter came back under the limit. +### Subagents (2026-09-05) + +**A subagent is a second transcript owned by a session, in the same event +model, with no process and no controls of its own.** Full design and wire +shape in `SUBAGENTS.md`, kept separate because the app half is being built +against it in parallel and it is the shared contract between the two. The +one-paragraph reason: a session's Task-tool helpers already speak the common +event model on the parent's own stdout (each line carrying +`parent_tool_use_id`), so giving each one its own small transcript — same +file format, same paging routes, same SSE stream, reused by addressing rather +than by copying — costs a routing step in the translator and a registry +(`session/subagent.rs`) rather than a second session type with a driver, a +process and a config entry it does not need. + ### HTTP surface **`routes.rs`'s module doc comment is the table.** REST for actions, one SSE diff --git a/SUBAGENTS.md b/SUBAGENTS.md new file mode 100644 index 0000000..b173aa7 --- /dev/null +++ b/SUBAGENTS.md @@ -0,0 +1,123 @@ +# Subagents + +A session's subagents -- the helpers a Claude Code session starts through its +Task tool -- each get a transcript of their own, listed under the session's +card and readable in the same transcript view the session has. Designed +2026-09-05; the decisions Bryan has not yet reviewed are in `DECISIONS.md`. + +## What a subagent is here + +**A subagent is a second transcript owned by a session, in the same event +model, with no process and no controls.** It is not a session: it cannot be +messaged, stopped or started, and it has no setup, model or usage of its +own. Everything it shares with a session -- the transcript file format, the +paging routes, the SSE stream, the phone's cache and rendering -- is reused +by addressing, not by copying. + +The CLI reports a subagent's messages on the parent's own stream-json +output, each carrying `parent_tool_use_id` = the id of the Task `tool_use` +that started it. Before this the translator dropped those lines +(`subagent_events_are_not_duplicated_into_the_transcript`); now it routes +them to that subagent's own translator and transcript. The parent's +transcript still shows only the Task call itself. + +## Storage + +Under the session directory: + +``` +/subagents//meta.json {title, created} +/subagents//transcript.jsonl same SeqEvent lines as the session's +``` + +The id is the Task tool_use id (`toolu_…`), which is unique, stable across a +backend restart, and already the key everything on the parent side uses. +Only ids matching `[A-Za-z0-9_-]+` are ever created or looked up, since the +id becomes a path. + +The transcript's sequence numbers are its own, starting at 1. `Transcript`, +`read_window`, `catch_up` and `read_after` work on it unchanged. + +Its path out: deleting the session deletes its directory, subagents included. +There is no separate delete. + +## Lifecycle, as events in the subagent's transcript + +1. Created on the first child line for an unseen parent id (or, when the + parent Task call was seen, at that call). First lines written: + `Status Running`, then `UserMessage { text: }` when + the prompt is known -- it genuinely is the subagent's first user turn. +2. Every child line is translated by that subagent's own `Translator` + (one per subagent: tool ids are unique but streaming deltas are by + content-block index, and parallel subagents interleave). +3. When the parent's `tool_result` for the Task id arrives, the parent gets + its `ToolEnd` as before, and the subagent gets `Status Exited`. +4. When the parent session's process exits (`Status Exited` on the + session), every subagent still `Running` gets `Status Exited` too: its + process was the parent's. + +A subagent that was mid-flight when the backend restarted keeps working: +the registry reopens the existing transcript on the next child line, and +the file continues its sequence. If its Task call finished while the backend +was down nothing ever closes it -- its last status stays `Running`, which +the list reports as **unknown** rather than as running (see the wire shape). + +Title: the Task call's `description` input, then ` ()` when +one is given; falling back to the tool's name when the child arrives before +(or without) the parent call being seen. + +## Server layout + +- `session/subagent.rs` -- the registry: `Subagents` (per session, in + `Shared`), `Subagent` (its `Transcript` behind a mutex plus a + `broadcast::Sender`), `record(id, event)`, `start(id, title, + prompt)`, `finish(id)`, `finish_all()`, `list()` from disk. Drivers get an + `Arc` beside their `EventSink`; llama ignores it. +- `session/claude/translate.rs` -- routes child lines by parent id, holds + one child `Translator` per subagent, remembers pending Task calls' + description/prompt/subagent_type. +- `session/echo.rs` -- `/subagent [n]`: the test rig. Starts *n* (default 1) + subagents at once, each named "helper k". Each writes the prompt as its + user message, streams a few words of text, runs one `Bash` tool call, then + finishes about three seconds after starting, and the parent's Task calls + end when their subagent does. Three seconds so the running state can be + seen on the phone. +- `routes.rs` -- three routes, in the doc table. + +## Wire shape + +``` +GET /sessions/{id} SessionInfo gains `subagents: N` (count, 0 when none) +GET /sessions same field on each row +GET /sessions/{id}/subagents [{id, title, status, created, lastActivity}], oldest first +GET /sessions/{id}/subagents/{sub}/transcript exactly the session transcript's query and answer +GET /sessions/{id}/subagents/{sub}/events?after=N exactly the session events stream +``` + +`status` is the transcript's last `Status` event, serialised like a session's +(`running`, `exited`), except that a subagent whose session is not itself +running cannot be running: the list answers `unknown` for that one. The +phone words these as *running*, *finished* and *unknown* on the subcard. + +The count on `SessionInfo` is a directory listing, so the list stays cheap. +The per-subagent status is only read when the list route is asked for. + +## Phone + +- `SessionSummary.subagents: Int`. A card with a non-zero count ends in an + expander row -- a full-width `Chevron(Pointing.Down)` row that flips to + `Pointing.Up` -- collapsed by default. Expanding fetches + `/sessions/{id}/subagents` and draws one `OutlinedCard` per subagent, + indented inside the session card, the way dev-updater draws a project's + components: title, then the status word and a relative time. The + expansion state is per session id and survives a refresh of the list. +- Tapping a subcard opens `Screen.Subagent`, which is `SessionScreen` in + **read-only** form: the same transcript, paging, cache, selection, + images and status row, with the composer, the process button, the model + picker, the files button, the settings cog and the usage bar left out. + The header shows the subagent's title with the session's title beneath + it. Back returns to the list. +- Addressing: `fetchTranscript`, `EventStream`, `TranscriptSource` and the + cache take a transcript address rather than a session id -- + `sessions/{id}` or `sessions/{id}/subagents/{sub}` -- so the cache nests a + subagent's copy under its session's and the same code serves both. diff --git a/app/androidApp/src/main/kotlin/com/example/aiapp/Api.kt b/app/androidApp/src/main/kotlin/com/example/aiapp/Api.kt index d64df02..832f89b 100644 --- a/app/androidApp/src/main/kotlin/com/example/aiapp/Api.kt +++ b/app/androidApp/src/main/kotlin/com/example/aiapp/Api.kt @@ -223,6 +223,14 @@ data class SessionSummary( val usageProvider: String?, val status: String, val lastActivity: Double, + /** + * How many subagents this session has, however their own status now reads. + * + * A directory listing on the server rather than a status read per subagent, so the list stays + * cheap; the per-subagent state is only fetched when the card is expanded. Zero on a server + * that predates subagents, so this app still opens against one. + */ + val subagents: Int, ) private fun parseSession(session: JSONObject) = @@ -252,6 +260,7 @@ private fun parseSession(session: JSONObject) = usageProvider = session.optString("usageProvider").ifEmpty { null }, status = session.getString("status"), lastActivity = session.getDouble("lastActivity"), + subagents = session.optInt("subagents", 0), ) fun fetchSessions(settings: ServerSettings): List = @@ -267,6 +276,35 @@ fun fetchSessions(settings: ServerSettings): List = fun fetchSession(settings: ServerSettings, sessionId: String): SessionSummary = requestFromServer(settings, "/sessions/$sessionId") { parseSession(it.jsonObject()) } +/** + * One row of `GET /sessions/{id}/subagents`, oldest first. + * + * A subagent is a second transcript owned by a session -- no process, no controls of its own -- so + * this carries only what a card needs to draw and to open it; see SUBAGENTS.md. [status] is + * "running", "exited" or "unknown": a subagent whose session is not itself running cannot be + * running, and the list says so rather than reporting a state that cannot hold. + */ +data class SubagentSummary( + val id: String, + val title: String, + val status: String, + val created: Double, + val lastActivity: Double, +) + +fun fetchSubagents(settings: ServerSettings, sessionId: String): List = + requestFromServer(settings, "/sessions/$sessionId/subagents") { + it.jsonObjects { row -> + SubagentSummary( + id = row.getString("id"), + title = row.getString("title"), + status = row.getString("status"), + created = row.getDouble("created"), + lastActivity = row.getDouble("lastActivity"), + ) + } + } + // What the server offers, so the spawn screen has no hardcoded lists: a setup added to the server's // config.ron appears here with no app rebuild. // @@ -968,7 +1006,7 @@ fun startImport( */ fun fetchTranscript( settings: ServerSettings, - sessionId: String, + address: TranscriptAddress, before: Long? = null, limit: Int = 80, // Count [limit] in rows, not events, joining a reply's streamed deltas into one -- so a page of @@ -987,7 +1025,7 @@ fun fetchTranscript( if (coalesce) append("&coalesce=true") if (after != null) append("&after=").append(after) } - return requestFromServer(settings, "/sessions/$sessionId/transcript$query") { connection -> + return requestFromServer(settings, "/${address.urlPath}/transcript$query") { connection -> val body = JSONArray(connection.inputStream.bufferedReader().readText()) // The text as well as the event: the transcript cache stores the one and the fold needs the // other, and they have to be the same line. diff --git a/app/androidApp/src/main/kotlin/com/example/aiapp/AppRoot.kt b/app/androidApp/src/main/kotlin/com/example/aiapp/AppRoot.kt index c1b000a..c8efd06 100644 --- a/app/androidApp/src/main/kotlin/com/example/aiapp/AppRoot.kt +++ b/app/androidApp/src/main/kotlin/com/example/aiapp/AppRoot.kt @@ -1,7 +1,9 @@ package com.example.aiapp import androidx.activity.compose.BackHandler +import androidx.compose.foundation.background import androidx.compose.foundation.layout.Box +import androidx.compose.foundation.layout.fillMaxSize import androidx.compose.foundation.layout.imePadding import androidx.compose.foundation.layout.padding import androidx.compose.material3.AlertDialog @@ -34,7 +36,23 @@ import kotlinx.coroutines.withContext * session, spawning one, and settings. */ private sealed class Screen { - data object Main : Screen() + /** + * The session list, with a subagent's own transcript over it when [subagent] is set. + * + * A layer on this screen rather than a screen of its own, for the same reason [Session.files] + * is: [SessionListScreen] owns which cards are expanded and what each expansion fetched, kept + * in `remember`, and a subagent is opened from a card's expander. As a sibling `Screen` it was + * disposed and recreated on every return, which lost that state -- an expanded card collapsed + * itself the moment its own subagent's view was closed. + */ + data class Main(val subagent: SubagentTarget? = null) : Screen() + + /** + * One subagent's own transcript, read-only. See [SessionScreen]'s `subagent` parameter and + * SUBAGENTS.md's "Phone". Closing it returns to [Main] under it, not to [Session]: a subagent + * is opened from the session list's card rather than from inside the session it belongs to. + */ + data class SubagentTarget(val summary: SessionSummary, val subagent: SubagentSummary) /** * One session, with the file explorer over it when [files] is set. @@ -81,7 +99,7 @@ fun AppRoot( val context = LocalContext.current val scope = rememberCoroutineScope() var settings by remember(settingsVersion) { mutableStateOf(loadServerSettings(context)) } - var screen by remember { mutableStateOf(Screen.Main) } + var screen by remember { mutableStateOf(Screen.Main()) } // A notification tap this could not follow, and why. Null both before one is asked for and // after one succeeds, since success is a screen rather than a message. var failedOpen by remember { mutableStateOf(null) } @@ -96,7 +114,7 @@ fun AppRoot( share = shareRequest // A session already open takes it. Otherwise the list is where the choice is made, // whatever screen was showing: Spawn and Settings have nowhere to put a file. - if (screen !is Screen.Session) screen = Screen.Main + if (screen !is Screen.Session) screen = Screen.Main() } } @@ -123,7 +141,7 @@ fun AppRoot( existing = null, onSaved = { saved -> settings = saved - screen = Screen.Main + screen = Screen.Main() }, onBack = null, ) @@ -136,7 +154,7 @@ fun AppRoot( // shows, so it always refetches. val goToMain = { reloadToken++ - screen = Screen.Main + screen = Screen.Main() } if (screen !is Screen.Main) { BackHandler(onBack = goToMain) @@ -185,6 +203,9 @@ fun AppRoot( reloadToken = reloadToken, share = share, onOpen = { screen = Screen.Session(it) }, + onOpenSubagent = { summary, subagent -> + screen = here.copy(subagent = Screen.SubagentTarget(summary, subagent)) + }, onSpawn = { screen = Screen.Spawn }, onImported = { imported -> reloadToken++ @@ -192,6 +213,27 @@ fun AppRoot( }, onSettings = { screen = Screen.Settings }, ) + // Its own back handler is registered after MainScreen's, so it is the one the + // platform asks first while a subagent is open -- the same rule the files + // explorer's handler follows over its session, below. + here.subagent?.let { target -> + BackHandler { screen = here.copy(subagent = null) } + // Its own opaque background: this screen was always the sole content under + // the theme's own Surface before, so it never had to paint one -- stacked over + // the list here, the space between its own cards let the list underneath show + // through without this. The same fix FilesScreen needed over its session. + Box(Modifier.fillMaxSize().background(MaterialTheme.colorScheme.background)) { + key(target.summary.id, target.subagent.id) { + SessionScreen( + settings = current, + summary = target.summary, + onBack = { screen = here.copy(subagent = null) }, + onFiles = {}, + subagent = target.subagent, + ) + } + } + } } is Screen.Session -> // Keyed on the id, because a different session is a different screen rather than this diff --git a/app/androidApp/src/main/kotlin/com/example/aiapp/EventStream.kt b/app/androidApp/src/main/kotlin/com/example/aiapp/EventStream.kt index 748b3fc..6f30cb2 100644 --- a/app/androidApp/src/main/kotlin/com/example/aiapp/EventStream.kt +++ b/app/androidApp/src/main/kotlin/com/example/aiapp/EventStream.kt @@ -14,7 +14,7 @@ private const val RESET_EVENT = "reset" * mean. [close] from any thread ends it, and the caller owns reconnecting -- with the last seq it * saw as the new cursor. */ -class EventStream(settings: ServerSettings, private val sessionId: String) { +class EventStream(settings: ServerSettings, private val address: TranscriptAddress) { private val stream = Sse(settings) fun close() = stream.close() @@ -35,7 +35,7 @@ class EventStream(settings: ServerSettings, private val sessionId: String) { // one and the screen folds the other, and they have to be the same line. onEvent: (raw: String, event: SeqEvent) -> Unit, ) { - stream.run("/sessions/$sessionId/events?after=$after", onOpen) { name, data -> + stream.run("/${address.urlPath}/events?after=$after", onOpen) { name, data -> // A named frame carries no payload and a data frame has no name. if (name == RESET_EVENT) onReset() else if (data.isNotEmpty()) onEvent(data, parseSeqEvent(data)) diff --git a/app/androidApp/src/main/kotlin/com/example/aiapp/MainScreen.kt b/app/androidApp/src/main/kotlin/com/example/aiapp/MainScreen.kt index f18ef22..4252ac4 100644 --- a/app/androidApp/src/main/kotlin/com/example/aiapp/MainScreen.kt +++ b/app/androidApp/src/main/kotlin/com/example/aiapp/MainScreen.kt @@ -48,6 +48,8 @@ fun MainScreen( /** What another app shared in and no session has taken yet; see [ShareRequest]. */ share: ShareRequest? = null, onOpen: (SessionSummary) -> Unit, + /** Opens one session's subagent, from the expander under its card. */ + onOpenSubagent: (SessionSummary, SubagentSummary) -> Unit, onSpawn: () -> Unit, onImported: (SessionSummary) -> Unit, onSettings: () -> Unit, @@ -139,6 +141,7 @@ fun MainScreen( settings = settings, reloadToken = token, onOpen = onOpen, + onOpenSubagent = onOpenSubagent, onSpawn = onSpawn, ) MainTab.Import -> diff --git a/app/androidApp/src/main/kotlin/com/example/aiapp/SessionListScreen.kt b/app/androidApp/src/main/kotlin/com/example/aiapp/SessionListScreen.kt index 1c3cc2b..425d26b 100644 --- a/app/androidApp/src/main/kotlin/com/example/aiapp/SessionListScreen.kt +++ b/app/androidApp/src/main/kotlin/com/example/aiapp/SessionListScreen.kt @@ -1,7 +1,9 @@ package com.example.aiapp import androidx.compose.foundation.ExperimentalFoundationApi +import androidx.compose.foundation.clickable import androidx.compose.foundation.combinedClickable +import androidx.compose.foundation.layout.Arrangement import androidx.compose.foundation.layout.Box import androidx.compose.foundation.layout.Column import androidx.compose.foundation.layout.Row @@ -17,6 +19,7 @@ import androidx.compose.material3.Card import androidx.compose.material3.CircularProgressIndicator import androidx.compose.material3.FloatingActionButton import androidx.compose.material3.MaterialTheme +import androidx.compose.material3.OutlinedCard import androidx.compose.material3.Switch import androidx.compose.material3.Text import androidx.compose.material3.TextButton @@ -30,6 +33,8 @@ import androidx.compose.runtime.setValue import androidx.compose.ui.Alignment import androidx.compose.ui.Modifier import androidx.compose.ui.platform.LocalContext +import androidx.compose.ui.semantics.contentDescription +import androidx.compose.ui.semantics.semantics import androidx.compose.ui.unit.dp import kotlinx.coroutines.Dispatchers import kotlinx.coroutines.launch @@ -47,12 +52,40 @@ fun SessionListScreen( settings: ServerSettings, reloadToken: Int, onOpen: (SessionSummary) -> Unit, + /** Opens one session's subagent, from the expander under its card. */ + onOpenSubagent: (SessionSummary, SubagentSummary) -> Unit, onSpawn: () -> Unit, ) { val scope = rememberCoroutineScope() var listState by remember { mutableStateOf>>(LoadState.Loading) } var confirmingDelete by remember { mutableStateOf(null) } + // Which session cards are expanded to show their subagents, and what each expansion fetched. + // Ids rather than a flag on the row for the same reason `deleting` is: the rows are rebuilt + // from + // whatever the server last said, and this belongs to the reader's own choice, which survives a + // refresh. + var expandedSessions by remember { mutableStateOf(setOf()) } + var subagentLoads by remember { + mutableStateOf(mapOf>>()) + } + + fun loadSubagents(sessionId: String) { + subagentLoads = subagentLoads + (sessionId to LoadState.Loading) + scope.launch { + subagentLoads = + subagentLoads + + (sessionId to + try { + LoadState.Loaded( + withContext(Dispatchers.IO) { fetchSubagents(settings, sessionId) } + ) + } catch (e: ApiException) { + LoadState.failed(e) + }) + } + } + // Failures that belong to one session rather than to the list, keyed by its id and shown on its // own card. The two scopes are decided by whether the server answered: it answered and refused, // so this says nothing about the other rows. @@ -84,6 +117,14 @@ fun SessionListScreen( withContext(Dispatchers.IO) { transcriptCache.retainOnly(loaded.value.map { it.id }.toSet()) } + // A session gone from this answer cannot still be expanded, and an expanded one + // that is still here asks again -- its subagents may have changed since the + // last + // fetch. + val ids = loaded.value.map { it.id }.toSet() + expandedSessions = expandedSessions intersect ids + subagentLoads = subagentLoads.filterKeys { it in ids } + expandedSessions.forEach(::loadSubagents) loaded } catch (e: ApiException) { LoadState.failed(e) @@ -127,6 +168,17 @@ fun SessionListScreen( deleting = session.id in deleting, onOpen = { onOpen(session) }, onLongPress = { confirmingDelete = session }, + expanded = session.id in expandedSessions, + subagents = subagentLoads[session.id], + onToggleSubagents = { + if (session.id in expandedSessions) { + expandedSessions = expandedSessions - session.id + } else { + expandedSessions = expandedSessions + session.id + loadSubagents(session.id) + } + }, + onOpenSubagent = { subagent -> onOpenSubagent(session, subagent) }, ) Spacer(Modifier.height(12.dp)) } @@ -225,7 +277,7 @@ fun SessionListScreen( deleteSession(settings, session.id, alsoDeleteForeign) // After it succeeded, not before: a refused delete leaves the // session exactly as it was, and its transcript with it. - transcriptCache.session(session.id).purge() + transcriptCache.session(TranscriptAddress(session.id)).purge() } // Only this row, and only what changed. Refetching the list instead // put every other session back through loading and handed the @@ -276,6 +328,12 @@ private fun SessionCard( deleting: Boolean, onOpen: () -> Unit, onLongPress: () -> Unit, + /** Whether the expander below is open. Collapsed by default; see [SessionListScreen]. */ + expanded: Boolean, + /** What the expander's own fetch answered, or null before it has been asked. */ + subagents: LoadState>?, + onToggleSubagents: () -> Unit, + onOpenSubagent: (SubagentSummary) -> Unit, ) { BusyItem(label = if (deleting) "deleting" else null) { Card( @@ -332,11 +390,100 @@ private fun SessionCard( color = MaterialTheme.colorScheme.error, ) } + // Nothing at all for a card with no subagents: a disabled expander here would be + // noise on every ordinary session's card. Its own row at the bottom rather than + // beside the title or the machine line, so opening it never displaces text that was + // already on screen -- see UI_RULES on a control not displacing the text beside it. + if (session.subagents > 0) { + Spacer(Modifier.height(8.dp)) + Row( + horizontalArrangement = Arrangement.Center, + modifier = + Modifier.fillMaxWidth() + .clickable(enabled = !deleting, onClick = onToggleSubagents) + .semantics { + contentDescription = + if (expanded) "Collapse subagents" else "Expand subagents" + }, + ) { + Chevron(if (expanded) Pointing.Up else Pointing.Down) + } + if (expanded) { + Spacer(Modifier.height(4.dp)) + Column(verticalArrangement = Arrangement.spacedBy(8.dp)) { + when (subagents) { + null, + is LoadState.Loading -> + CircularProgressIndicator( + modifier = Modifier.width(20.dp).height(20.dp), + strokeWidth = 2.dp, + ) + is LoadState.Error -> + // Said here rather than left silent: a fetch that failed and an + // expander that simply found nothing must not look the same -- + // see UI_RULES on designing the unknown state first. + Text( + subagents.message, + style = MaterialTheme.typography.bodySmall, + color = MaterialTheme.colorScheme.error, + ) + is LoadState.Loaded -> + subagents.value.forEach { subagent -> + SubagentCard( + subagent, + onClick = { onOpenSubagent(subagent) }, + ) + } + } + } + } + } } } } } +/** + * One subagent, indented inside its session's card -- the way dev-updater draws a project's + * components (`ComponentCard`, `UpdaterScreen.kt`): an outlined card, not the session card's own + * filled one, so the nesting reads as one step rather than as another session. + */ +@Composable +private fun SubagentCard(subagent: SubagentSummary, onClick: () -> Unit) { + OutlinedCard(Modifier.fillMaxWidth().clickable(onClick = onClick)) { + Column(Modifier.padding(horizontal = 12.dp, vertical = 8.dp)) { + Text(subagent.title, style = MaterialTheme.typography.titleSmall) + Spacer(Modifier.height(2.dp)) + Row(modifier = Modifier.fillMaxWidth()) { + Text( + subagentStatusLabel(subagent.status), + style = MaterialTheme.typography.bodySmall, + color = MaterialTheme.colorScheme.onSurfaceVariant, + modifier = Modifier.weight(1f), + ) + Text( + relativeTime(subagent.lastActivity), + style = MaterialTheme.typography.bodySmall, + color = MaterialTheme.colorScheme.onSurfaceVariant, + ) + } + } + } +} + +/** + * The subcard's word for a subagent's status -- see SUBAGENTS.md's "Wire shape". Its own function + * rather than a branch inside [StatusText], because a subagent's three states are not that + * composable's five: "exited" reads as "finished" here, since its process was always its parent's + * and never something of its own to have merely stopped. + */ +private fun subagentStatusLabel(status: String) = + when (status) { + "running" -> "running" + "exited" -> "finished" + else -> "unknown" + } + @Composable fun StatusText(status: String) { val (label, color) = diff --git a/app/androidApp/src/main/kotlin/com/example/aiapp/SessionScreen.kt b/app/androidApp/src/main/kotlin/com/example/aiapp/SessionScreen.kt index 64e031c..2b54b6f 100644 --- a/app/androidApp/src/main/kotlin/com/example/aiapp/SessionScreen.kt +++ b/app/androidApp/src/main/kotlin/com/example/aiapp/SessionScreen.kt @@ -216,16 +216,33 @@ fun SessionScreen( share: ShareRequest? = null, /** Said once [share] has been attached here, so it is not attached again. */ onShareTaken: () -> Unit = {}, + /** + * Draws this screen read-only, on a subagent's own transcript instead of the session's. + * + * A subagent has no process and no controls of its own -- see SUBAGENTS.md's "Phone" -- so + * every gate below keyed on this switches off the composer, the files button, the settings cog, + * the usage bar and notifications, while everything that draws a transcript (paging, cache, + * selection, images, the status row, stream reconnects) is reused unchanged, pointed at + * [address] instead of the session's own. + */ + subagent: SubagentSummary? = null, ) { DebugStats.count("session screen recomposed") + val isSubagent = subagent != null + val address = TranscriptAddress(summary.id, subagent?.id) val scope = rememberCoroutineScope() val topEdgeHeld = remember { TopEdgeHold() } var items by remember { mutableStateOf(listOf()) } - var status by remember { mutableStateOf(summary.status) } + var status by remember { mutableStateOf(subagent?.status ?: summary.status) } // Seeded from the row this screen was opened from, so a conversation already under way says how // much it is holding before any turn happens here. Null is "nobody has measured it", which is a // different answer from an empty context and is drawn differently. - var contextTokens by remember(summary.id) { mutableStateOf(summary.contextTokens) } + // + // A subagent has no context measurement of its own, so it always starts unmeasured rather than + // borrowing the parent session's figure -- see UI_RULES on not showing an inferred value as one + // that was measured. + var contextTokens by + remember(address) { mutableStateOf(if (isSubagent) null else summary.contextTokens) } // When the current compaction started. The moment comes off the `compacting` status event // itself -- the server timestamps every transcript line -- rather than off this device noticing // one, which is what makes it survive leaving the session and reopening it. @@ -241,7 +258,13 @@ fun SessionScreen( val context = LocalContext.current // Seeded from what was left in the box last time and written back on every keystroke, so // leaving the screen does not throw away a half-typed message. See `Drafts.kt`. - var input by remember(summary.id) { mutableStateOf(atEnd(loadDraft(context, summary.id))) } + // + // A subagent has no box to type into, so it never touches a draft at all -- not this session's, + // which is what reading one keyed only by `summary.id` would do here. + var input by + remember(summary.id) { + mutableStateOf(if (isSubagent) atEnd("") else atEnd(loadDraft(context, summary.id))) + } // A model the reader has chosen and not yet confirmed. See [ModelSwitchWarning]: switching // makes the session re-read the whole conversation. var pendingModel by remember { mutableStateOf(null) } @@ -294,27 +317,26 @@ fun SessionScreen( // Reload throws away what it was reading from. val cache = remember(settings) { TranscriptCache(cacheRoot(context, settings)) } val source = - remember(summary.id, epoch) { - TranscriptSource(settings, summary.id, cache.session(summary.id)) - } + remember(address, epoch) { TranscriptSource(settings, address, cache.session(address)) } // Whether the cached tail has been shown to still be the server's own line. Nothing is resumed // from a cached cursor until it has, and a probe that could not be made leaves this false for // the stream loop to try again. - var probePassed by remember(summary.id, epoch) { mutableStateOf(false) } + var probePassed by remember(address, epoch) { mutableStateOf(false) } // Whether the opening effect is still settling that question. It draws the cached rows and // lifts [ready] before the answer arrives, which is the point of the cache -- so the stream // below waits for this rather than for `ready`, or it asks the same question twice. - var probing by remember(summary.id, epoch) { mutableStateOf(true) } + var probing by remember(address, epoch) { mutableStateOf(true) } // The oldest sequence number loaded, and whether there is more behind it. Paging backwards is // what keeps opening a long session cheap. var oldestSeq by remember { mutableLongStateOf(0L) } - // Where this session was last being read, from this device's own store. Read once, because the - // answer stops being interesting the moment the list is on screen. - val savedAnchor = remember(summary.id, epoch) { loadScrollAnchor(context, summary.id) } + // Where this transcript was last being read, from this device's own store, keyed by the address + // rather than the session id so a subagent's saved position cannot collide with its session's. + // Read once, because the answer stops being interesting the moment the list is on screen. + val savedAnchor = remember(address, epoch) { loadScrollAnchor(context, address.cachePath) } // Whether the saved position is still being put back. Nothing is drawn while it is: opening at // the newest end and then travelling to the anchor is exactly the journey a reader must never // see. - var restoring by remember(summary.id, epoch) { mutableStateOf(savedAnchor != null) } + var restoring by remember(address, epoch) { mutableStateOf(savedAnchor != null) } // Messages the server has taken and the session has not read yet, by the id that will resolve // them. From the event stream rather than from what this screen sent, so they survive leaving // the session -- and a message sent from another device is drawn waiting on this one too. @@ -327,11 +349,11 @@ fun SessionScreen( var loadingHistory by remember { mutableStateOf(false) } var ready by remember { mutableStateOf(false) } // Replies parsed ahead of the rows that draw them; see [ParsedReplies]. - val replies = remember(summary.id) { ParsedReplies() } - // Keyed like everything else describing one session's transcript. `rememberLazyListState` saves - // through `rememberSaveable`, and this screen restores by its own anchor instead -- two - // restores would fight over the first frame. - val listState = remember(summary.id) { LazyListState() } + val replies = remember(address) { ParsedReplies() } + // Keyed like everything else describing one transcript. `rememberLazyListState` saves through + // `rememberSaveable`, and this screen restores by its own anchor instead -- two restores would + // fight over the first frame. + val listState = remember(address) { LazyListState() } // Whether the newest message is on screen right now. The list is reversed, so the newest end is // the scrolling start: nothing behind you is exactly being at the bottom. Asked of the scroll // state rather than of item indices, because a zero-height first item makes an index ambiguous. @@ -637,7 +659,7 @@ fun SessionScreen( // ended and carries live events only. The window comes from this phone's own copy when there is // one, and then costs a single request to check that the server's transcript is still the one // it came from. See TRANSCRIPT_CACHE.md. - LaunchedEffect(summary.id, epoch) { + LaunchedEffect(address, epoch) { /** * One opening window onto the screen, whichever side it came from. * @@ -667,11 +689,16 @@ fun SessionScreen( // A replay is as old as the last visit; the row this screen was opened from was // fetched moments ago. So the transcript comes from the cache and everything that // is not the transcript comes from the summary -- otherwise a session that finished - // an hour ago opens saying "working" until the stream connects. - status = summary.status - model = summary.model - permissionMode = summary.permissionMode ?: "auto" - if (summary.status != "compacting") compactingSince = null + // an hour ago opens saying "working" until the stream connects. A subagent's status + // comes from its own summary, never the parent session's: they are two different + // things running or not, and the parent's model and permission mode do not apply to + // it at all. + status = subagent?.status ?: summary.status + if (!isSubagent) { + model = summary.model + permissionMode = summary.permissionMode ?: "auto" + } + if (status != "compacting") compactingSince = null // Nothing to put back, so these rows are the screen and the probe can return under // them. A restore still has history to fetch and is gated below. if (savedAnchor == null) ready = true @@ -798,7 +825,7 @@ fun SessionScreen( // at the top on their return. Switching apps is a choice somebody made, not a fault to report. // Stopping the stream deliberately makes the drop a close rather than an error, and resuming // reconnects from the same cursor. - LaunchedEffect(summary.id, ready, epoch, lifecycleOwner) { + LaunchedEffect(address, ready, epoch, lifecycleOwner) { if (!ready) return@LaunchedEffect // The opening effect draws cached rows and lifts `ready` *before* it has checked that the // cursor under them is still the server's, so `ready` is no longer the whole gate. Without @@ -868,17 +895,22 @@ fun SessionScreen( // The screen going away entirely, which the lifecycle scope above does not cover: a composable // can leave the composition while the activity stays started. Keyed on the epoch as well, so // Reload's replacement source is the one a later disposal closes. - DisposableEffect(summary.id, epoch) { onDispose { source.close() } } + DisposableEffect(address, epoch) { onDispose { source.close() } } // Nothing gets announced about the session somebody is reading; see NotificationService. // RESUMED rather than STARTED because "looking at it" means the foreground. - LaunchedEffect(summary.id, lifecycleOwner) { - lifecycleOwner.repeatOnLifecycle(Lifecycle.State.RESUMED) { - NotificationService.showing(context, summary.id) - try { - awaitCancellation() - } finally { - NotificationService.stoppedShowing(summary.id) + // + // Not for a subagent: it has no notifications of its own, and it is not the session this would + // otherwise mark as being read. + if (!isSubagent) { + LaunchedEffect(summary.id, lifecycleOwner) { + lifecycleOwner.repeatOnLifecycle(Lifecycle.State.RESUMED) { + NotificationService.showing(context, summary.id) + try { + awaitCancellation() + } finally { + NotificationService.stoppedShowing(summary.id) + } } } } @@ -924,7 +956,7 @@ fun SessionScreen( val (index, offset, awayFromNewest) = settled saveScrollAnchor( context, - summary.id, + address.cachePath, // Nothing to restore at the newest end, which is where a session with no anchor // opens anyway. One *before* the index, because item zero is the "below" slot. if (!awayFromNewest) null @@ -947,7 +979,7 @@ fun SessionScreen( // // There is no correction beside this one. Following the newest message is not an effect: the // list is reversed, so an arriving message extends the end the viewport is pinned to. - val unitSizes = remember(summary.id) { HashMap() } + val unitSizes = remember(address) { HashMap() } LaunchedEffect(listState, moreHistory) { snapshotFlow { listState.layoutInfo } .collect { info -> @@ -983,21 +1015,25 @@ fun SessionScreen( } } - LaunchedEffect(summary.setupName, summary.provider) { - offeredModels = - try { - withContext(Dispatchers.IO) { - fetchSetups(settings) - .firstOrNull { it.name == summary.setupName } - ?.providers - ?.firstOrNull { it.name == summary.provider } - ?.models - .orEmpty() + // Only for the model picker, which a subagent does not have. + if (!isSubagent) { + LaunchedEffect(summary.setupName, summary.provider) { + offeredModels = + try { + withContext(Dispatchers.IO) { + fetchSetups(settings) + .firstOrNull { it.name == summary.setupName } + ?.providers + ?.firstOrNull { it.name == summary.provider } + ?.models + .orEmpty() + } + } catch (_: Exception) { + // Not worth reporting: the picker simply has nothing to offer, which is + // visible. + emptyList() } - } catch (_: Exception) { - // Not worth reporting: the picker simply has nothing to offer, which is visible. - emptyList() - } + } } /** @@ -1148,8 +1184,9 @@ fun SessionScreen( } // One poll for the machines' limits, read by everything on this screen that reports them. - val usageFeed = rememberUsageFeed(settings) - val usage = usageFeed.forSession(summary) + // Nothing meters a subagent -- it has no account of its own -- so it never starts this poll. + val usageFeed = if (isSubagent) null else rememberUsageFeed(settings) + val usage = usageFeed?.forSession(summary) ?: SessionUsage.NotMetered RecordFrames() var usageOpen by remember { mutableStateOf(false) } var settingsOpen by remember { mutableStateOf(false) } @@ -1236,20 +1273,35 @@ fun SessionScreen( // A ring's worth, which is what the arrow already keeps on its other three sides. Spacer(Modifier.width(GLYPH_BUTTON_MARGIN)) Column(Modifier.weight(1f)) { - Text(title, style = MaterialTheme.typography.titleMedium) - // Machine first, then what runs on it -- the same order and the same wording - // everywhere this pair appears, so it reads as one fact rather than two - // sentences with different grammar. - // - // No model. The picker in the footer already shows what this session is set to, - // and showing it twice means two things to keep in step -- they disagreed for a - // moment on every model change. - Text( - "${summary.setupName} · ${summary.provider}", - style = MaterialTheme.typography.bodySmall, - color = MaterialTheme.colorScheme.onSurfaceVariant, - ) + // A subagent's own title, with the session's beneath it in a smaller style -- + // the header says whose conversation this is as well as what it is. Otherwise + // just the session's title, as before. + if (subagent != null) { + Text(subagent.title, style = MaterialTheme.typography.titleMedium) + Text( + title, + style = MaterialTheme.typography.bodySmall, + color = MaterialTheme.colorScheme.onSurfaceVariant, + ) + } else { + Text(title, style = MaterialTheme.typography.titleMedium) + // Machine first, then what runs on it -- the same order and the same + // wording everywhere this pair appears, so it reads as one fact rather than + // two sentences with different grammar. + // + // No model. The picker in the footer already shows what this session is set + // to, and showing it twice means two things to keep in step -- they + // disagreed for a moment on every model change. + Text( + "${summary.setupName} · ${summary.provider}", + style = MaterialTheme.typography.bodySmall, + color = MaterialTheme.colorScheme.onSurfaceVariant, + ) + } } + // None of this is a subagent's: it has no files of its own to browse, no settings, + // and nothing meters it -- see SUBAGENTS.md's "Phone". + // // Beside the provider it reports on, which is the line directly to its left. Its // real home is this provider's settings, which do not exist yet. A session on a // provider with no such service gets an honest "unavailable" rather than a hidden @@ -1264,42 +1316,47 @@ fun SessionScreen( // Usage, files, settings -- widest scope first, narrowing to the right, so the cog // stays at the end where every other screen keeps it. Asked for in this order by // Iris on 2026-09-03. - Row { - GlyphButton( - USAGE_GLYPH, - "Usage", - { usageOpen = true }, - colour = usageGlyphColour(usage), - ) - // The machine's files, which is where the answer to "what did it actually - // change" is. It opens *over* this screen rather than replacing it. - GlyphButton( - FOLDER_GLYPH, - "Files", - onClick = { - onFiles( - FilesTarget( - setup = summary.setup, - setupName = summary.setupName, - // Where this session works, and the machine's own home when it - // was never given a directory -- resolved there rather than - // guessed at here, since this app does not know that home. - start = summary.cwd?.takeIf { it.isNotBlank() } ?: "~", + if (!isSubagent) { + Row { + GlyphButton( + USAGE_GLYPH, + "Usage", + { usageOpen = true }, + colour = usageGlyphColour(usage), + ) + // The machine's files, which is where the answer to "what did it actually + // change" is. It opens *over* this screen rather than replacing it. + GlyphButton( + FOLDER_GLYPH, + "Files", + onClick = { + onFiles( + FilesTarget( + setup = summary.setup, + setupName = summary.setupName, + // Where this session works, and the machine's own home when + // it was never given a directory -- resolved there rather + // than guessed at here, since this app does not know that + // home. + start = summary.cwd?.takeIf { it.isNotBlank() } ?: "~", + ) ) - ) - }, - ) - // What it opens is about this session, so it sits at the end of the session's - // own row. A cog and not a word because there will be more, and a bar of words - // has nowhere to put it. - GlyphButton(SETTINGS_GLYPH, "Session settings", { settingsOpen = true }) + }, + ) + // What it opens is about this session, so it sits at the end of the + // session's own row. A cog and not a word because there will be more, and a + // bar of words has nowhere to put it. + GlyphButton(SETTINGS_GLYPH, "Session settings", { settingsOpen = true }) + } } } // Under the header, above everything the session itself says: it is a fact about the // machine rather than a turn in the conversation, and it is the number that decides - // whether to keep going. - SessionUsageBar(usage) + // whether to keep going. Nothing meters a subagent. + if (!isSubagent) { + SessionUsageBar(usage) + } (streamError ?: actionError)?.let { message -> Text( @@ -1644,185 +1701,208 @@ fun SessionScreen( ) } + // Kept for a subagent -- see SUBAGENTS.md's "Phone" -- with the wording that turns + // "exited" into "finished" for one, since it has no process to leave running or stop. SessionStatusRow( status = status, compactingFor = compactingFor, contextTokens = contextTokens, + subagent = isSubagent, ) - // Between the transcript and the box: above what is being typed, so the list does not - // cover the thing the command is about, and below everything that explains it. - CommandSuggestions( - // Nothing to suggest about a suggestion that was just taken. `/compact` is a whole - // command *and* a prefix of itself, so picking it left the list standing there with - // the one row already chosen. Held by what was picked rather than by a flag, so - // typing anything else brings the list back without a second thing to reset. - commands = if (input.text == picked) emptyList() else suggestedCommands(input.text), - onPick = { command -> - // At the end of what was inserted, which is where the reader carries on typing: - // a command with an argument is put in the box half-written, and a cursor left - // at the front makes the next keystroke the first character of "/rename". - input = atEnd(command.typed()) - picked = command.typed() - }, - ) - - // Always enabled -- a send while the session is running becomes a steering message - // injected at the next tool boundary, which is the point of the whole app. - // - // The field gets a row of its own, above the buttons: sharing one put the full width - // behind three controls, so the thing being typed into was the narrowest on the row. - Column(Modifier.fillMaxWidth().padding(8.dp)) { - // Directly above the box they will be sent from, so what is attached is visible - // rather than counted: the "+2" on the button below said how many and never which. - PendingAttachments( - settings = settings, - sessionId = summary.id, - refs = pendingAttachments, - onRemove = { pendingAttachments = pendingAttachments - it }, - ) - OutlinedTextField( - value = input, - onValueChange = { - input = it - saveDraft(context, summary.id, it.text) + // Everything from here down is the composer: a subagent cannot be messaged, so none of + // it applies -- see SUBAGENTS.md's "Phone". + if (!isSubagent) { + // Between the transcript and the box: above what is being typed, so the list does + // not cover the thing the command is about, and below everything that explains it. + CommandSuggestions( + // Nothing to suggest about a suggestion that was just taken. `/compact` is a + // whole command *and* a prefix of itself, so picking it left the list standing + // there with the one row already chosen. Held by what was picked rather than by + // a flag, so typing anything else brings the list back without a second thing + // to + // reset. + commands = + if (input.text == picked) emptyList() else suggestedCommands(input.text), + onPick = { command -> + // At the end of what was inserted, which is where the reader carries on + // typing: a command with an argument is put in the box half-written, and a + // cursor left at the front makes the next keystroke the first character of + // "/rename". + input = atEnd(command.typed()) + picked = command.typed() }, - modifier = Modifier.fillMaxWidth(), - // No longer "(+image)": the images are on screen above this, and a placeholder - // saying so said it in words beside the thing itself. - placeholder = { Text("Message") }, - maxLines = 4, ) - Row( - verticalAlignment = Alignment.CenterVertically, - modifier = Modifier.fillMaxWidth(), - ) { - // Photo or file, asked here rather than by two buttons: the row is full, and - // attaching is one action whichever picker answers it. - var attaching by remember { mutableStateOf(false) } - Box { - // Just "+". The count it used to carry was standing in for showing them. - BubbleButton(onClick = { attaching = true }) { Text("+") } - DropdownMenu( - expanded = attaching, - onDismissRequest = { attaching = false }, - // See PickerButton: without this the menu opens a status bar's height - // away from the button in an edge-to-edge activity. - properties = PopupProperties(clippingEnabled = false), - shape = BubbleMenuShape, - ) { - DropdownMenuItem( - text = { Text("Photo") }, - onClick = { - attaching = false - pickImage.launch( - PickVisualMediaRequest( - ActivityResultContracts.PickVisualMedia.ImageOnly - ) - ) - }, - ) - DropdownMenuItem( - text = { Text("File") }, - onClick = { - attaching = false - pickFile.launch(arrayOf("*/*")) - }, - ) - } - } - // The settings share what is left after the actions have taken what they need. - // A Row hands out intrinsic widths in order and clips whatever runs past the - // edge, so with these laid out first the arrival of Stop pushed Send off the - // screen entirely -- the app's central control, gone at the moment it is most - // in use. + + // Always enabled -- a send while the session is running becomes a steering message + // injected at the next tool boundary, which is the point of the whole app. + // + // The field gets a row of its own, above the buttons: sharing one put the full + // width + // behind three controls, so the thing being typed into was the narrowest on the + // row. + Column(Modifier.fillMaxWidth().padding(8.dp)) { + // Directly above the box they will be sent from, so what is attached is visible + // rather than counted: the "+2" on the button below said how many and never + // which. + PendingAttachments( + settings = settings, + sessionId = summary.id, + refs = pendingAttachments, + onRemove = { pendingAttachments = pendingAttachments - it }, + ) + OutlinedTextField( + value = input, + onValueChange = { + input = it + saveDraft(context, summary.id, it.text) + }, + modifier = Modifier.fillMaxWidth(), + // No longer "(+image)": the images are on screen above this, and a + // placeholder saying so said it in words beside the thing itself. + placeholder = { Text("Message") }, + maxLines = 4, + ) Row( verticalAlignment = Alignment.CenterVertically, - modifier = Modifier.weight(1f), + modifier = Modifier.fillMaxWidth(), ) { - if (offeredModels.isNotEmpty()) { + // Photo or file, asked here rather than by two buttons: the row is full, + // and + // attaching is one action whichever picker answers it. + var attaching by remember { mutableStateOf(false) } + Box { + // Just "+". The count it used to carry was standing in for showing + // them. + BubbleButton(onClick = { attaching = true }) { Text("+") } + DropdownMenu( + expanded = attaching, + onDismissRequest = { attaching = false }, + // See PickerButton: without this the menu opens a status bar's + // height away from the button in an edge-to-edge activity. + properties = PopupProperties(clippingEnabled = false), + shape = BubbleMenuShape, + ) { + DropdownMenuItem( + text = { Text("Photo") }, + onClick = { + attaching = false + pickImage.launch( + PickVisualMediaRequest( + ActivityResultContracts.PickVisualMedia.ImageOnly + ) + ) + }, + ) + DropdownMenuItem( + text = { Text("File") }, + onClick = { + attaching = false + pickFile.launch(arrayOf("*/*")) + }, + ) + } + } + // The settings share what is left after the actions have taken what they + // need. A Row hands out intrinsic widths in order and clips whatever runs + // past the edge, so with these laid out first the arrival of Stop pushed + // Send off the screen entirely -- the app's central control, gone at the + // moment it is most in use. + Row( + verticalAlignment = Alignment.CenterVertically, + modifier = Modifier.weight(1f), + ) { + if (offeredModels.isNotEmpty()) { + PickerButton( + current = modelLabel(model), + // What the machine offers, plus the state a session is in when + // it has chosen none of them. The button has always been able + // to + // say "default"; until this the list could not, so leaving it + // was a one-way trip. + options = listOf(DEFAULT_MODEL) + offeredModels, + // Not set here. The button follows what the session reports it + // is set to, which arrives a moment later and is sometimes a + // different answer -- a name the CLI resolved, or no change at + // all on a provider whose model is fixed. Asked about first, + // unless there is nothing to lose by it -- see + // [ModelSwitchWarning]. + onPick = { chosen -> + if ( + modelLabel(chosen) == modelLabel(model) || + !worthWarningAbout(status, contextTokens, items) + ) { + act { setSessionModel(settings, summary.id, chosen) } + } else { + pendingModel = chosen + } + }, + ) + } PickerButton( - current = modelLabel(model), - // What the machine offers, plus the state a session is in when it - // has chosen none of them. The button has always been able to say - // "default"; until this the list could not, so leaving it was a - // one-way trip. - options = listOf(DEFAULT_MODEL) + offeredModels, - // Not set here. The button follows what the session reports it is - // set to, which arrives a moment later and is sometimes a different - // answer -- a name the CLI resolved, or no change at all on a - // provider whose model is fixed. Asked about first, unless there is - // nothing to lose by it -- see [ModelSwitchWarning]. + current = permissionMode, + options = PERMISSION_MODES, onPick = { chosen -> - if ( - modelLabel(chosen) == modelLabel(model) || - !worthWarningAbout(status, contextTokens, items) - ) { - act { setSessionModel(settings, summary.id, chosen) } - } else { - pendingModel = chosen - } + act { setSessionPermissionMode(settings, summary.id, chosen) } }, ) } - PickerButton( - current = permissionMode, - options = PERMISSION_MODES, - onPick = { chosen -> - act { setSessionPermissionMode(settings, summary.id, chosen) } - }, - ) - } - // The same filled shape as the button beside it, not an outlined one: these are - // two things you can do about the session, and weighting one as secondary said - // they were a primary action and its qualifier. What separates them is the - // colour and the mark, which is what they mean. - // - // Always here, rather than arriving with the turn as it used to. A control that - // comes and goes makes its own presence the signal, and a button always in the - // same place also cannot push Send off the end of the row by turning up. - val process = - when { - running -> ProcessAction.Pause - status == "exited" -> ProcessAction.Start - else -> ProcessAction.Stop - } - Button( - onClick = { - processInFlight = true - act(onDone = { processInFlight = false }) { - process.perform(settings, summary.id) + // The same filled shape as the button beside it, not an outlined one: these + // are two things you can do about the session, and weighting one as + // secondary said they were a primary action and its qualifier. What + // separates them is the colour and the mark, which is what they mean. + // + // Always here, rather than arriving with the turn as it used to. A control + // that comes and goes makes its own presence the signal, and a button + // always + // in the same place also cannot push Send off the end of the row by turning + // up. + val process = + when { + running -> ProcessAction.Pause + status == "exited" -> ProcessAction.Start + else -> ProcessAction.Stop } - }, - enabled = !processInFlight, - colors = actionButtonColors(process.colour()), - ) { - Glyph( - process.glyph, - colour = LocalContentColor.current, - modifier = Modifier.semantics { contentDescription = process.label }, - ) - } - Spacer(Modifier.width(8.dp)) - // The paper plane, with a clock on it while a turn is in flight: sending then - // queues the message for the next tool boundary rather than starting a turn of - // its own, and the two have to be told apart at a glance. The label says the - // same thing to a screen reader. - // - // Disabled while there is nothing to send, rather than pressable and silent: - // `send` has always returned early on an empty composer, so the button promised - // something it would not do. Disabled and not hidden, for the reason above. - Button( - onClick = { send() }, - enabled = input.text.isNotBlank() || pendingAttachments.isNotEmpty(), - colors = actionButtonColors(if (running) queueColor else sendColor), - ) { - Glyph( - if (running) QUEUE_GLYPH else SEND_GLYPH, - colour = LocalContentColor.current, - modifier = - Modifier.semantics { contentDescription = sendLabel(running) }, - ) + Button( + onClick = { + processInFlight = true + act(onDone = { processInFlight = false }) { + process.perform(settings, summary.id) + } + }, + enabled = !processInFlight, + colors = actionButtonColors(process.colour()), + ) { + Glyph( + process.glyph, + colour = LocalContentColor.current, + modifier = + Modifier.semantics { contentDescription = process.label }, + ) + } + Spacer(Modifier.width(8.dp)) + // The paper plane, with a clock on it while a turn is in flight: sending + // then queues the message for the next tool boundary rather than starting a + // turn of its own, and the two have to be told apart at a glance. The label + // says the same thing to a screen reader. + // + // Disabled while there is nothing to send, rather than pressable and + // silent: + // `send` has always returned early on an empty composer, so the button + // promised something it would not do. Disabled and not hidden, for the + // reason above. + Button( + onClick = { send() }, + enabled = input.text.isNotBlank() || pendingAttachments.isNotEmpty(), + colors = actionButtonColors(if (running) queueColor else sendColor), + ) { + Glyph( + if (running) QUEUE_GLYPH else SEND_GLYPH, + colour = LocalContentColor.current, + modifier = + Modifier.semantics { contentDescription = sendLabel(running) }, + ) + } } } } @@ -1833,7 +1913,7 @@ fun SessionScreen( // is the screen's business rather than any row's. See [SessionImageViewer]. fullImage?.let { ref -> SessionImageViewer(settings, summary.id, ref) { fullImage = null } } if (usageOpen) { - UsageDialog(feed = usageFeed, onDismiss = { usageOpen = false }) + usageFeed?.let { UsageDialog(feed = it, onDismiss = { usageOpen = false }) } } if (settingsOpen) { // Measured when the dialog opens rather than kept up to date: what the reader is being told @@ -2125,6 +2205,13 @@ private fun SessionStatusRow( /** Context the session is holding, or null where nothing has measured it. */ contextTokens: Long?, modifier: Modifier = Modifier, + /** + * Whether this row is for a subagent rather than a session, which changes only one word: + * "exited" reads as "finished" there too, the same as the subagent list's own card -- a + * subagent's process was always its parent's, so "exited" would read as a fault rather than the + * ordinary way one of these ends. + */ + subagent: Boolean = false, ) { DebugStats.count("status row recomposed") Row( @@ -2181,7 +2268,7 @@ private fun SessionStatusRow( Text( when (status) { "idle" -> "idle" - "exited" -> "exited" + "exited" -> if (subagent) "finished" else "exited" "awaitingInput" -> "your turn" "unknown" -> "can't tell" else -> status diff --git a/app/androidApp/src/main/kotlin/com/example/aiapp/TranscriptAddress.kt b/app/androidApp/src/main/kotlin/com/example/aiapp/TranscriptAddress.kt new file mode 100644 index 0000000..ae2a97f --- /dev/null +++ b/app/androidApp/src/main/kotlin/com/example/aiapp/TranscriptAddress.kt @@ -0,0 +1,27 @@ +package com.example.aiapp + +/** + * Where one transcript lives: a session's own, or one of its subagents'. + * + * The single mechanism [fetchTranscript], [EventStream], [TranscriptSource] and + * [TranscriptCache.session] all take, rather than each growing its own branch between a session and + * a subagent -- see SUBAGENTS.md's "Phone" and "Wire shape". A caller that has only a session id + * builds one with the one-argument constructor; a subagent's screen supplies both ids. + */ +data class TranscriptAddress(val sessionId: String, val subagentId: String? = null) { + /** The URL segment naming this transcript, before `/transcript` or `/events`. */ + val urlPath: String + get() = + if (subagentId == null) "sessions/$sessionId" + else "sessions/$sessionId/subagents/$subagentId" + + /** + * Where this transcript's cache lives on the phone, relative to the cache root. + * + * A subagent's nests under its session's directory rather than sitting beside it, so deleting a + * session's cache directory takes its subagents' with it -- the same one-way door the server's + * own storage describes. + */ + val cachePath: String + get() = if (subagentId == null) sessionId else "$sessionId/subagents/$subagentId" +} diff --git a/app/androidApp/src/main/kotlin/com/example/aiapp/TranscriptCache.kt b/app/androidApp/src/main/kotlin/com/example/aiapp/TranscriptCache.kt index 606628e..d489f3a 100644 --- a/app/androidApp/src/main/kotlin/com/example/aiapp/TranscriptCache.kt +++ b/app/androidApp/src/main/kotlin/com/example/aiapp/TranscriptCache.kt @@ -35,8 +35,15 @@ class TranscriptCache( private val root: File, private val warn: (String) -> Unit = { Log.w("ai-app", it) }, ) { - /** The cache for one session, whether or not anything has been stored for it yet. */ - fun session(id: String): SessionCache = SessionCache(File(root, id), warn) + /** + * The cache for one transcript, whether or not anything has been stored for it yet. + * + * A subagent's [TranscriptAddress.cachePath] nests it under its session's directory, so + * deleting the session (below) takes its subagents' caches with it -- there is no separate + * purge for one. + */ + fun session(address: TranscriptAddress): SessionCache = + SessionCache(File(root, address.cachePath), warn) /** * Deletes every session directory not in [ids], called after a successful list fetch. The path diff --git a/app/androidApp/src/main/kotlin/com/example/aiapp/TranscriptSource.kt b/app/androidApp/src/main/kotlin/com/example/aiapp/TranscriptSource.kt index 8a7bdcd..10c7c32 100644 --- a/app/androidApp/src/main/kotlin/com/example/aiapp/TranscriptSource.kt +++ b/app/androidApp/src/main/kotlin/com/example/aiapp/TranscriptSource.kt @@ -18,7 +18,7 @@ import java.util.concurrent.atomic.AtomicReference */ class TranscriptSource( private val settings: ServerSettings, - private val sessionId: String, + private val address: TranscriptAddress, val cache: SessionCache, ) { private val stream = AtomicReference(null) @@ -65,7 +65,7 @@ class TranscriptSource( val tail = cache.tail() ?: return false // `before = seq + 1` is the newest event with seq <= the cursor, which is the event *at* // the cursor when the server still has one there. - val answer = fetchTranscript(settings, sessionId, before = tail.seq + 1, limit = 1) + val answer = fetchTranscript(settings, address, before = tail.seq + 1, limit = 1) val matches = answer.size == 1 && try { @@ -83,7 +83,7 @@ class TranscriptSource( */ suspend fun fetchOpening(): List { DebugStats.count("transcript page from server") - val page = fetchTranscript(settings, sessionId, limit = OPENING_WINDOW) + val page = fetchTranscript(settings, address, limit = OPENING_WINDOW) page.forEach { (line, entry) -> cache.append(line, entry.seq) } cache.flush() return page.map { it.second } @@ -108,7 +108,7 @@ class TranscriptSource( val page = fetchTranscript( settings, - sessionId, + address, before = before, limit = limit, coalesce = coalesce, @@ -131,7 +131,7 @@ class TranscriptSource( * well lose. */ fun follow(after: Long, onOpen: () -> Unit, onReset: () -> Unit, onEvent: (SeqEvent) -> Unit) { - val opened = EventStream(settings, sessionId) + val opened = EventStream(settings, address) stream.getAndSet(opened)?.close() try { opened.run(after, onOpen, onReset) { raw, entry -> diff --git a/app/androidApp/src/test/kotlin/com/example/aiapp/TranscriptCacheTest.kt b/app/androidApp/src/test/kotlin/com/example/aiapp/TranscriptCacheTest.kt index 04b4db9..e25341f 100644 --- a/app/androidApp/src/test/kotlin/com/example/aiapp/TranscriptCacheTest.kt +++ b/app/androidApp/src/test/kotlin/com/example/aiapp/TranscriptCacheTest.kt @@ -23,7 +23,7 @@ class TranscriptCacheTest { private fun cache() = TranscriptCache(File(temp, "v1/host_8443")) { said += it } - private fun session(id: String = "s") = cache().session(id) + private fun session(id: String = "s") = cache().session(TranscriptAddress(id)) private fun line(seq: Long, type: String = "toolStart") = """{"seq":$seq,"ts":1.5,"type":"$type","id":"x"}""" diff --git a/server/src/routes.rs b/server/src/routes.rs index b42ff11..aac707e 100644 --- a/server/src/routes.rs +++ b/server/src/routes.rs @@ -29,6 +29,12 @@ //! GET /sessions/{id}/transcript a page of history: ?before=N (newest when absent), //! ?limit=N, ?coalesce=true to count rows not deltas, //! ?after=N to floor it at what the caller already holds +//! GET /sessions/{id}/subagents [{id, title, status, created, lastActivity}], oldest +//! first -- see SUBAGENTS.md +//! GET /sessions/{id}/subagents/{sub}/transcript exactly the transcript route above, +//! against that subagent's own transcript +//! GET /sessions/{id}/subagents/{sub}/events?after=N exactly the events route above, +//! against that subagent's own stream //! POST /sessions/{id}/message {text, attachmentIds?} //! (starts the process first if it has exited) //! POST /sessions/{id}/unqueue {messageId} -- take back one not read yet @@ -99,6 +105,7 @@ use tokio_stream::wrappers::{BroadcastStream, ReceiverStream}; use crate::session::driver::{SessionCommand, Unqueued}; use crate::session::pending::Operation; +use crate::session::subagent::{Subagent, SubagentInfo}; use crate::session::transcript::{CATCH_UP_LIMIT, CatchUp, SeqEvent, catch_up}; use crate::session::{LiveSession, SessionInfo, SessionManager, SpawnSpec}; @@ -130,6 +137,15 @@ pub fn router(manager: Arc) -> Router { .route("/sessions/{id}", get(read_session).delete(delete_session)) .route("/sessions/{id}/events", get(events)) .route("/sessions/{id}/transcript", get(transcript)) + .route("/sessions/{id}/subagents", get(list_subagents)) + .route( + "/sessions/{id}/subagents/{sub}/transcript", + get(subagent_transcript), + ) + .route( + "/sessions/{id}/subagents/{sub}/events", + get(subagent_events), + ) .route("/sessions/{id}/message", post(message)) .route("/sessions/{id}/unqueue", post(unqueue)) .route("/sessions/{id}/answer", post(answer)) @@ -209,6 +225,18 @@ fn lookup(manager: &SessionManager, id: &str) -> Result, ApiErr .ok_or_else(|| ApiError::NotFound(format!("no session {id}"))) } +/// A session's subagent by id -- the second half of the lookup every +/// `/sessions/{id}/subagents/{sub}/...` route needs. `Arc` because reopening +/// one from disk (a subagent this process has not touched yet) inserts it +/// into the registry, and a route holding a borrow across that would be +/// holding the registry's lock the whole request. +fn lookup_subagent(session: &LiveSession, sub: &str) -> Result, ApiError> { + session + .subagents() + .get(sub) + .ok_or_else(|| ApiError::NotFound(format!("no subagent {sub}"))) +} + async fn list_sessions(State(manager): State>) -> axum::Json> { axum::Json(manager.sessions()) } @@ -1747,8 +1775,36 @@ async fn transcript( Query(query): Query, ) -> Result>, ApiError> { let session = lookup(&manager, &id)?; + transcript_page(session.transcript_path(), &id, query) +} + +/// Exactly [`transcript`]'s route and answer, against one subagent's own +/// transcript instead of its session's -- see `SUBAGENTS.md`'s wire shape. +async fn subagent_transcript( + State(manager): State>, + UrlPath((id, sub)): UrlPath<(String, String)>, + Query(query): Query, +) -> Result>, ApiError> { + let session = lookup(&manager, &id)?; + let subagent = lookup_subagent(&session, &sub)?; + transcript_page( + &subagent.transcript_path(), + &format!("{id}/subagents/{sub}"), + query, + ) +} + +/// A page of history at `path`, newest first to open with -- the one +/// implementation [`transcript`] and [`subagent_transcript`] share, since a +/// subagent's transcript is read exactly the way a session's is. `label` is +/// only for the debug line below. +fn transcript_page( + path: &Path, + label: &str, + query: TranscriptQuery, +) -> Result>, ApiError> { let events = crate::session::transcript::read_window( - session.transcript_path(), + path, query.before, query.after, query.limit, @@ -1760,7 +1816,7 @@ async fn transcript( // for events and draws rows, and the ratio between them is a property of // the conversation. `RUST_LOG=ai_server=debug`. tracing::debug!( - session = %id, + session = %label, before = ?query.before, after = ?query.after, limit = query.limit, @@ -1778,23 +1834,74 @@ async fn events( headers: HeaderMap, ) -> Result>>, ApiError> { let session = lookup(&manager, &id)?; - let cursor = headers - .get("last-event-id") - .and_then(|value| value.to_str().ok()) - .and_then(|value| value.parse().ok()) - .unwrap_or(query.after); - + let cursor = cursor_of(&headers, query.after); // Subscribe before reading the file so nothing can land in the gap // between replay and live; overlap is deduplicated by seq. let live = session.subscribe(); - let (tx, stream) = mpsc::channel(64); - tokio::spawn(stream_session( + Ok(sse_stream( session.transcript_path().to_path_buf(), cursor, live, - tx, - )); - Ok(Sse::new(ReceiverStream::new(stream).map(Ok)).keep_alive(KeepAlive::default())) + )) +} + +/// Exactly [`events`]'s route and answer, against one subagent's own stream +/// instead of its session's -- see `SUBAGENTS.md`'s wire shape. +async fn subagent_events( + State(manager): State>, + UrlPath((id, sub)): UrlPath<(String, String)>, + Query(query): Query, + headers: HeaderMap, +) -> Result>>, ApiError> { + let session = lookup(&manager, &id)?; + let subagent = lookup_subagent(&session, &sub)?; + let cursor = cursor_of(&headers, query.after); + let live = subagent.subscribe(); + Ok(sse_stream(subagent.transcript_path(), cursor, live)) +} + +/// The cursor an SSE reconnect resumes from: the native `Last-Event-ID` +/// takes precedence over the query parameter, same cursor either way. +fn cursor_of(headers: &HeaderMap, query_after: u64) -> u64 { + headers + .get("last-event-id") + .and_then(|value| value.to_str().ok()) + .and_then(|value| value.parse().ok()) + .unwrap_or(query_after) +} + +/// Spawns the backlog-then-live task and wraps it as the response, the one +/// piece [`events`] and [`subagent_events`] share. +fn sse_stream( + transcript: PathBuf, + cursor: u64, + live: broadcast::Receiver, +) -> Sse>> { + let (tx, stream) = mpsc::channel(64); + tokio::spawn(stream_session(transcript, cursor, live, tx)); + Sse::new(ReceiverStream::new(stream).map(Ok)).keep_alive(KeepAlive::default()) +} + +/// `GET /sessions/{id}/subagents`: every subagent this session has started, +/// oldest first, with a status read from its own transcript -- see +/// `SUBAGENTS.md`'s wire shape. A subagent whose last status is `Running` is +/// reported `Unknown` instead when the session itself is not running: its +/// process was the session's, and a session with none has nothing left to +/// ask. +async fn list_subagents( + State(manager): State>, + UrlPath(id): UrlPath, +) -> Result>, ApiError> { + let session = lookup(&manager, &id)?; + // Anything but `Exited` or `Unknown` has a process behind it, which is + // what decides whether a subagent still reading `Running` from its own + // transcript can be believed -- see `SUBAGENTS.md`'s wire shape. + let running = !matches!( + session.status(), + crate::session::driver::SessionStatus::Exited + | crate::session::driver::SessionStatus::Unknown + ); + Ok(axum::Json(session.subagents().list(running))) } /// Every session's attention-wanting moments, on one stream. diff --git a/server/src/session/claude.rs b/server/src/session/claude.rs index 69c4d54..ac88f6b 100644 --- a/server/src/session/claude.rs +++ b/server/src/session/claude.rs @@ -59,6 +59,7 @@ use tokio::sync::mpsc; use super::driver::{AttachmentRef, Driver, Event, EventSink, SessionStatus, Unqueued}; use super::process; +use super::subagent::Subagents; use super::transport::{Launch, Streams, Transport}; use crate::config::{ProviderConfig, SessionConfig}; use translate::{AnswerOutcome, Setting, Translator, starts_a_model_call}; @@ -213,8 +214,12 @@ impl ClaudeDriver { transport: &Transport, session_dir: &Path, sink: EventSink, + subagents: Arc, ) -> Result { - let state = Arc::new(Mutex::new(Translator::new(session_dir.to_path_buf()))); + let state = Arc::new(Mutex::new(Translator::new( + session_dir.to_path_buf(), + subagents, + ))); let queue = Arc::new(Mutex::new(Queue::default())); let reading = Arc::new(AtomicBool::new(true)); @@ -1169,7 +1174,10 @@ mod tests { /// this" and "the transcript records that". fn events_from_lines(lines: &[&str]) -> Vec { let dir = tempfile::tempdir().expect("temp dir"); - let state = Arc::new(Mutex::new(Translator::new(dir.path().to_path_buf()))); + let state = Arc::new(Mutex::new(Translator::new( + dir.path().to_path_buf(), + Arc::new(Subagents::new(dir.path().to_path_buf())), + ))); let queue = Arc::new(Mutex::new(Queue::default())); let (sink, mut out) = mpsc::unbounded_channel::(); for line in lines { @@ -1193,7 +1201,10 @@ mod tests { interject: impl FnOnce(&Arc>), ) -> Vec { let dir = tempfile::tempdir().expect("temp dir"); - let state = Arc::new(Mutex::new(Translator::new(dir.path().to_path_buf()))); + let state = Arc::new(Mutex::new(Translator::new( + dir.path().to_path_buf(), + Arc::new(Subagents::new(dir.path().to_path_buf())), + ))); let queue = Arc::new(Mutex::new(Queue::default())); let (sink, mut out) = mpsc::unbounded_channel::(); let mut interject = Some(interject); @@ -1487,7 +1498,10 @@ mod tests { // doing. let dir = tempfile::tempdir().expect("tempdir"); let (sink, mut received) = mpsc::unbounded_channel(); - let state = Arc::new(Mutex::new(Translator::new(dir.path().to_path_buf()))); + let state = Arc::new(Mutex::new(Translator::new( + dir.path().to_path_buf(), + Arc::new(Subagents::new(dir.path().to_path_buf())), + ))); let queue = Arc::new(Mutex::new(Queue::default())); let text = r#"{"type":"stream_event","event":{"type":"content_block_delta","index":0,"delta":{"type":"text_delta","text":"working"}},"parent_tool_use_id":null}"#; @@ -1532,7 +1546,10 @@ mod tests { // session back to work. let dir = tempfile::tempdir().expect("tempdir"); let (sink, mut received) = mpsc::unbounded_channel(); - let state = Arc::new(Mutex::new(Translator::new(dir.path().to_path_buf()))); + let state = Arc::new(Mutex::new(Translator::new( + dir.path().to_path_buf(), + Arc::new(Subagents::new(dir.path().to_path_buf())), + ))); let queue = Arc::new(Mutex::new(Queue::default())); queue.lock().unwrap().close(&sink, "the session ended"); diff --git a/server/src/session/claude/translate.rs b/server/src/session/claude/translate.rs index 39c1650..6a4bd09 100644 --- a/server/src/session/claude/translate.rs +++ b/server/src/session/claude/translate.rs @@ -11,10 +11,12 @@ use std::collections::HashMap; use std::path::{Path, PathBuf}; +use std::sync::{Arc, Mutex}; use serde_json::{Value, json}; use super::super::driver::{Event, QuestionOption, SessionStatus, context_tokens}; +use super::super::subagent::Subagents; /// Whether this line is the CLI opening a fresh model call. /// @@ -92,10 +94,20 @@ pub(super) struct Translator { /// reports none rather than repeating the previous turn's. context: Option, session_dir: PathBuf, + /// This session's subagents, shared with every child translator below -- + /// see `SUBAGENTS.md`. One registry per session, so a subagent started + /// through this translator or any of its children lands in the same + /// place a route reads it back from. + subagents: Arc, + /// One translator per subagent id, holding *its* streaming and + /// tool-tracking state -- separate from the parent's because tool ids + /// are unique but a `stream_event`'s content-block index is not, and + /// parallel subagents interleave their deltas on one stdout. + children: HashMap>>, } impl Translator { - pub(super) fn new(session_dir: PathBuf) -> Self { + pub(super) fn new(session_dir: PathBuf, subagents: Arc) -> Self { Self { session_id: None, pending: HashMap::new(), @@ -103,6 +115,8 @@ impl Translator { interrupting: false, context: None, session_dir, + subagents, + children: HashMap::new(), } } @@ -122,13 +136,54 @@ impl Translator { pub(super) fn translate(&mut self, message: &Value) -> Vec { // Events from subagents (Task tool internals) carry a // parent_tool_use_id; the transcript shows the Task tool's own - // start/end instead of every nested step. - if message - .get("parent_tool_use_id") - .is_some_and(|id| !id.is_null()) - { - return Vec::new(); + // start/end instead of every nested step. Routed into that + // subagent's own transcript rather than dropped -- see + // `SUBAGENTS.md`. + if let Some(parent_id) = message.get("parent_tool_use_id").and_then(Value::as_str) { + return self.translate_child(parent_id, message); } + self.dispatch(message) + } + + /// A line belonging to a subagent rather than to this translator's own + /// session. Always returns nothing to the *caller*: everything it + /// produces goes into the subagent's own transcript instead. + fn translate_child(&mut self, id: &str, message: &Value) -> Vec { + match self.subagents.get(id) { + Some(subagent) if !subagent.is_open() => { + // The Task call already ended (or this line is stale from a + // resumed conversation) -- see `SUBAGENTS.md`'s lifecycle. + tracing::debug!("dropping a line for subagent {id}, which has already finished"); + return Vec::new(); + } + Some(_) => {} + None => { + // Nobody has heard of this id yet: the Task call itself + // either has not been seen or never will be. Started here + // with the best title available -- the tool name of this + // first line -- since SUBAGENTS.md's real title only + // arrives with the Task call. + self.subagents.start(id, &fallback_title(message), None); + } + } + let child = self + .children + .entry(id.to_string()) + .or_insert_with(|| { + Arc::new(Mutex::new(Translator::new( + self.session_dir.clone(), + Arc::clone(&self.subagents), + ))) + }) + .clone(); + let events = child.lock().unwrap().dispatch(message); + for event in events { + self.subagents.record(id, event); + } + Vec::new() + } + + fn dispatch(&mut self, message: &Value) -> Vec { match message.get("type").and_then(Value::as_str) { Some("system") => self.translate_system(message), // The CLI's own announcement that `/clear` took effect, sent just @@ -379,22 +434,47 @@ impl Translator { content .iter() .filter(|block| block.get("type").and_then(Value::as_str) == Some("tool_use")) - .map(|block| Event::ToolStart { - id: block + .map(|block| { + let id = block .get("id") .and_then(Value::as_str) .unwrap_or_default() - .to_string(), - tool: block + .to_string(); + let tool = block .get("name") .and_then(Value::as_str) .unwrap_or_default() - .to_string(), - input: block.get("input").cloned().unwrap_or(Value::Null), + .to_string(); + let input = block.get("input").cloned().unwrap_or(Value::Null); + // A subagent this call is about to start -- see + // `SUBAGENTS.md`'s lifecycle #1. The parent's own transcript + // still shows only the Task call itself, below. + if tool == "Task" || tool == "Agent" { + self.start_subagent_from_task(&id, &input); + } + Event::ToolStart { id, tool, input } }) .collect() } + /// Starts the subagent a Task call names, with the title and prompt + /// SUBAGENTS.md describes: the call's `description`, then + /// `()` when one is given, falling back to the tool's own + /// name when there is no description to build one from. + fn start_subagent_from_task(&self, id: &str, input: &Value) { + let description = text_field(input, "description"); + let subagent_type = text_field(input, "subagent_type"); + let prompt = input.get("prompt").and_then(Value::as_str); + let title = match (description, subagent_type) { + (Some(description), Some(subagent_type)) => { + format!("{description} ({subagent_type})") + } + (Some(description), None) => description, + (None, _) => "Task".to_string(), + }; + self.subagents.start(id, &title, prompt); + } + fn translate_control_request(&mut self, message: &Value) -> Vec { let request = &message["request"]; if request.get("subtype").and_then(Value::as_str) != Some("can_use_tool") { @@ -598,11 +678,33 @@ impl Translator { id: about.clone(), output: texts.join("\n"), }); + // A no-op unless `about` is a subagent's own id -- see + // `SUBAGENTS.md`'s lifecycle #3: the parent gets this `ToolEnd` + // like any other tool result, and the subagent it names (if it + // names one) gets its `Status::Exited`. + self.subagents.finish(&about); } events } } +/// The title to start a subagent under when its own first line arrives +/// before (or without) its Task call ever being seen: the tool name of that +/// first line, which is the only thing known about it yet. `"subagent"` for +/// a line this cannot even find a tool name in, such as one that opens with +/// something other than a tool call. +fn fallback_title(message: &Value) -> String { + message["message"]["content"] + .as_array() + .into_iter() + .flatten() + .find(|block| block.get("type").and_then(Value::as_str) == Some("tool_use")) + .and_then(|block| block.get("name")) + .and_then(Value::as_str) + .unwrap_or("subagent") + .to_string() +} + /// Whether a failed turn failed because the account is out of quota, and when /// the CLI said the limit lifts. /// @@ -696,10 +798,18 @@ mod tests { .collect() } + /// A fresh, empty subagent registry over the same temp dir a test's + /// translator writes into -- every test here is about the parent's own + /// events, so what a registry does with a subagent is `subagent.rs`'s + /// tests to make, not these. + fn test_subagents(dir: &tempfile::TempDir) -> Arc { + Arc::new(Subagents::new(dir.path().to_path_buf())) + } + #[test] fn captures_the_resume_token_and_the_settings_from_init() { let dir = tempfile::tempdir().expect("tempdir"); - let mut translator = Translator::new(dir.path().to_path_buf()); + let mut translator = Translator::new(dir.path().to_path_buf(), test_subagents(&dir)); let events = translate_lines( &mut translator, &[ @@ -721,7 +831,7 @@ mod tests { #[test] fn a_setting_is_reported_when_the_cli_accepts_it_and_not_before() { let dir = tempfile::tempdir().expect("tempdir"); - let mut translator = Translator::new(dir.path().to_path_buf()); + let mut translator = Translator::new(dir.path().to_path_buf(), test_subagents(&dir)); // What `set_model` does: remember, send, and say nothing yet. translator.expect_setting("req-a".to_string(), Setting::Model("sonnet".to_string())); @@ -798,7 +908,7 @@ mod tests { // The line it sends just after answering `set_permission_mode`, which is // also how a mode changed from the terminal arrives. let dir = tempfile::tempdir().expect("tempdir"); - let mut translator = Translator::new(dir.path().to_path_buf()); + let mut translator = Translator::new(dir.path().to_path_buf(), test_subagents(&dir)); let events = translate_lines( &mut translator, &[ @@ -818,7 +928,7 @@ mod tests { fn streams_text_deltas_and_skips_the_consolidated_copy() { // Real lines (trimmed) from the 2.1.237 probe. let dir = tempfile::tempdir().expect("tempdir"); - let mut translator = Translator::new(dir.path().to_path_buf()); + let mut translator = Translator::new(dir.path().to_path_buf(), test_subagents(&dir)); let events = translate_lines( &mut translator, &[ @@ -838,7 +948,7 @@ mod tests { #[test] fn tool_use_and_result_become_tool_events() { let dir = tempfile::tempdir().expect("tempdir"); - let mut translator = Translator::new(dir.path().to_path_buf()); + let mut translator = Translator::new(dir.path().to_path_buf(), test_subagents(&dir)); let events = translate_lines( &mut translator, &[ @@ -865,7 +975,7 @@ mod tests { #[test] fn subagent_events_are_not_duplicated_into_the_transcript() { let dir = tempfile::tempdir().expect("tempdir"); - let mut translator = Translator::new(dir.path().to_path_buf()); + let mut translator = Translator::new(dir.path().to_path_buf(), test_subagents(&dir)); let events = translate_lines( &mut translator, &[ @@ -875,10 +985,132 @@ mod tests { assert!(events.is_empty()); } + /// A child line does not just vanish from the parent -- it lands in its + /// own subagent's transcript, with that transcript's own sequence + /// numbers, starting at 1 like any other. + #[test] + fn a_child_line_lands_in_its_own_subagents_transcript() { + let dir = tempfile::tempdir().expect("tempdir"); + let subagents = test_subagents(&dir); + let mut translator = Translator::new(dir.path().to_path_buf(), Arc::clone(&subagents)); + translate_lines( + &mut translator, + &[ + r#"{"type":"assistant","message":{"content":[{"type":"tool_use","id":"toolu_c1","name":"Bash","input":{"command":"echo hi"}}]},"parent_tool_use_id":"toolu_parent"}"#, + ], + ); + let subagent = subagents.get("toolu_parent").expect("subagent started"); + let lines = crate::session::transcript::read_after(&subagent.transcript_path(), 0) + .expect("read subagent transcript"); + assert_eq!(lines[0].seq, 1); + assert_eq!( + lines[0].event, + Event::Status { + state: SessionStatus::Running + } + ); + assert!( + lines.iter().any( + |entry| matches!(&entry.event, Event::ToolStart { tool, .. } if tool == "Bash") + ) + ); + } + + /// The title and prompt shown for a subagent come from the Task call + /// that started it, not from anything guessed at its first line. + #[test] + fn the_subagent_takes_its_title_and_prompt_from_the_task_call() { + let dir = tempfile::tempdir().expect("tempdir"); + let subagents = test_subagents(&dir); + let mut translator = Translator::new(dir.path().to_path_buf(), Arc::clone(&subagents)); + translate_lines( + &mut translator, + &[ + r#"{"type":"assistant","message":{"content":[{"type":"tool_use","id":"toolu_task","name":"Task","input":{"description":"Investigate the bug","prompt":"Find why X fails","subagent_type":"general-purpose"}}]},"parent_tool_use_id":null}"#, + ], + ); + let rows = subagents.list(true); + assert_eq!(rows.len(), 1); + assert_eq!(rows[0].title, "Investigate the bug (general-purpose)"); + let subagent = subagents.get(&rows[0].id).expect("subagent"); + let lines = crate::session::transcript::read_after(&subagent.transcript_path(), 0) + .expect("read subagent transcript"); + assert!(lines.iter().any( + |entry| matches!(&entry.event, Event::UserMessage { text, .. } if text == "Find why X fails") + )); + } + + /// The parent's `tool_result` for the Task id is what ends the + /// subagent -- SUBAGENTS.md's lifecycle #3 -- and nothing else does. + #[test] + fn the_parents_tool_result_finishes_the_subagent() { + let dir = tempfile::tempdir().expect("tempdir"); + let subagents = test_subagents(&dir); + let mut translator = Translator::new(dir.path().to_path_buf(), Arc::clone(&subagents)); + translate_lines( + &mut translator, + &[ + r#"{"type":"assistant","message":{"content":[{"type":"tool_use","id":"toolu_task2","name":"Task","input":{"description":"helper"}}]},"parent_tool_use_id":null}"#, + ], + ); + let subagent = subagents.get("toolu_task2").expect("subagent started"); + assert!(subagent.is_open()); + translate_lines( + &mut translator, + &[ + r#"{"type":"user","message":{"role":"user","content":[{"type":"tool_result","tool_use_id":"toolu_task2","content":"done","is_error":false}]},"parent_tool_use_id":null}"#, + ], + ); + assert!(!subagent.is_open()); + } + + /// Two subagents running at once keep two separate transcripts: tool ids + /// are unique but a `stream_event`'s content-block index is not, so + /// sharing translation state between them would cross their streams. + #[test] + fn two_parallel_subagents_keep_separate_transcripts() { + let dir = tempfile::tempdir().expect("tempdir"); + let subagents = test_subagents(&dir); + let mut translator = Translator::new(dir.path().to_path_buf(), Arc::clone(&subagents)); + translate_lines( + &mut translator, + &[ + r#"{"type":"assistant","message":{"content":[{"type":"tool_use","id":"toolu_a","name":"Bash","input":{}}]},"parent_tool_use_id":"toolu_task_a"}"#, + r#"{"type":"assistant","message":{"content":[{"type":"tool_use","id":"toolu_b","name":"Read","input":{}}]},"parent_tool_use_id":"toolu_task_b"}"#, + ], + ); + let a = subagents.get("toolu_task_a").expect("subagent a"); + let b = subagents.get("toolu_task_b").expect("subagent b"); + let a_events = crate::session::transcript::read_after(&a.transcript_path(), 0) + .expect("read a's transcript"); + let b_events = crate::session::transcript::read_after(&b.transcript_path(), 0) + .expect("read b's transcript"); + assert!( + a_events.iter().any( + |entry| matches!(&entry.event, Event::ToolStart { tool, .. } if tool == "Bash") + ) + ); + assert!( + b_events.iter().any( + |entry| matches!(&entry.event, Event::ToolStart { tool, .. } if tool == "Read") + ) + ); + assert!( + !a_events.iter().any( + |entry| matches!(&entry.event, Event::ToolStart { tool, .. } if tool == "Read") + ) + ); + assert!( + !b_events.iter().any( + |entry| matches!(&entry.event, Event::ToolStart { tool, .. } if tool == "Bash") + ) + ); + } + #[test] fn a_permission_request_becomes_an_allow_deny_question() { let dir = tempfile::tempdir().expect("tempdir"); - let mut translator = Translator::new(dir.path().to_path_buf()); + let mut translator = Translator::new(dir.path().to_path_buf(), test_subagents(&dir)); let events = translate_lines( &mut translator, &[ @@ -927,7 +1159,7 @@ mod tests { #[test] fn denying_a_permission_sends_deny() { let dir = tempfile::tempdir().expect("tempdir"); - let mut translator = Translator::new(dir.path().to_path_buf()); + let mut translator = Translator::new(dir.path().to_path_buf(), test_subagents(&dir)); translate_lines( &mut translator, &[ @@ -945,7 +1177,7 @@ mod tests { // The real 2.1.237 shape, verified live: answers go back inside // updatedInput, keyed by the question text. let dir = tempfile::tempdir().expect("tempdir"); - let mut translator = Translator::new(dir.path().to_path_buf()); + let mut translator = Translator::new(dir.path().to_path_buf(), test_subagents(&dir)); let events = translate_lines( &mut translator, &[ @@ -997,7 +1229,7 @@ mod tests { // in the event: a phone that had to read this dialect's tool input to // find them would be the only place that knew how. let dir = tempfile::tempdir().expect("tempdir"); - let mut translator = Translator::new(dir.path().to_path_buf()); + let mut translator = Translator::new(dir.path().to_path_buf(), test_subagents(&dir)); let events = translate_lines( &mut translator, &[ @@ -1045,7 +1277,7 @@ mod tests { #[test] fn images_in_tool_results_are_saved_and_referenced() { let dir = tempfile::tempdir().expect("tempdir"); - let mut translator = Translator::new(dir.path().to_path_buf()); + let mut translator = Translator::new(dir.path().to_path_buf(), test_subagents(&dir)); // A 1x1 PNG, the smallest real payload worth round-tripping. let png = "iVBORw0KGgoAAAANSUhEUgAAAAEAAAABCAYAAAAfFcSJAAAADUlEQVR42mNk+M9QDwADhgGAWjR9awAAAABJRU5ErkJggg=="; let line = format!( @@ -1074,7 +1306,7 @@ mod tests { #[test] fn a_turn_result_reports_usage_and_returns_to_idle() { let dir = tempfile::tempdir().expect("tempdir"); - let mut translator = Translator::new(dir.path().to_path_buf()); + let mut translator = Translator::new(dir.path().to_path_buf(), test_subagents(&dir)); let events = translate_lines( &mut translator, &[ @@ -1109,7 +1341,7 @@ mod tests { #[test] fn a_turn_started_by_another_agent_records_who_and_what() { let dir = tempfile::tempdir().expect("tempdir"); - let mut translator = Translator::new(dir.path().to_path_buf()); + let mut translator = Translator::new(dir.path().to_path_buf(), test_subagents(&dir)); let events = translate_lines( &mut translator, &[ @@ -1143,7 +1375,7 @@ mod tests { #[test] fn an_ordinary_turn_carries_no_peer_note() { let dir = tempfile::tempdir().expect("tempdir"); - let mut translator = Translator::new(dir.path().to_path_buf()); + let mut translator = Translator::new(dir.path().to_path_buf(), test_subagents(&dir)); let events = translate_lines( &mut translator, &[ @@ -1168,7 +1400,7 @@ mod tests { #[test] fn the_context_is_what_the_last_message_held_not_the_turn_added_up() { let dir = tempfile::tempdir().expect("tempdir"); - let mut translator = Translator::new(dir.path().to_path_buf()); + let mut translator = Translator::new(dir.path().to_path_buf(), test_subagents(&dir)); let events = translate_lines( &mut translator, &[ @@ -1208,7 +1440,7 @@ mod tests { // Note the snake_case keys -- the CLI's transcript file writes the same // records in camelCase. let dir = tempfile::tempdir().expect("tempdir"); - let mut translator = Translator::new(dir.path().to_path_buf()); + let mut translator = Translator::new(dir.path().to_path_buf(), test_subagents(&dir)); let events = translate_lines( &mut translator, &[ @@ -1238,7 +1470,7 @@ mod tests { #[test] fn a_failed_compaction_says_why_and_leaves_the_turn_running() { let dir = tempfile::tempdir().expect("tempdir"); - let mut translator = Translator::new(dir.path().to_path_buf()); + let mut translator = Translator::new(dir.path().to_path_buf(), test_subagents(&dir)); let events = translate_lines( &mut translator, &[ @@ -1265,7 +1497,7 @@ mod tests { #[test] fn a_boundary_without_counts_says_so_rather_than_inventing_them() { let dir = tempfile::tempdir().expect("tempdir"); - let mut translator = Translator::new(dir.path().to_path_buf()); + let mut translator = Translator::new(dir.path().to_path_buf(), test_subagents(&dir)); let events = translate_lines( &mut translator, &[ @@ -1285,7 +1517,7 @@ mod tests { #[test] fn an_error_result_surfaces_the_message() { let dir = tempfile::tempdir().expect("tempdir"); - let mut translator = Translator::new(dir.path().to_path_buf()); + let mut translator = Translator::new(dir.path().to_path_buf(), test_subagents(&dir)); let events = translate_lines( &mut translator, &[ @@ -1315,7 +1547,7 @@ mod tests { #[test] fn a_turn_stopped_by_the_usage_limit_says_so_and_carries_the_reset() { let dir = tempfile::tempdir().expect("tempdir"); - let mut translator = Translator::new(dir.path().to_path_buf()); + let mut translator = Translator::new(dir.path().to_path_buf(), test_subagents(&dir)); let events = translate_lines( &mut translator, &[ @@ -1333,7 +1565,7 @@ mod tests { #[test] fn a_limit_the_cli_gave_no_reset_for_is_reported_without_one() { let dir = tempfile::tempdir().expect("tempdir"); - let mut translator = Translator::new(dir.path().to_path_buf()); + let mut translator = Translator::new(dir.path().to_path_buf(), test_subagents(&dir)); let events = translate_lines( &mut translator, &[ @@ -1366,7 +1598,7 @@ mod tests { #[test] fn a_turn_stopped_on_purpose_is_not_an_error() { let dir = tempfile::tempdir().expect("tempdir"); - let mut translator = Translator::new(dir.path().to_path_buf()); + let mut translator = Translator::new(dir.path().to_path_buf(), test_subagents(&dir)); let stopped_result = r#"{"type":"result","subtype":"error_during_execution","is_error":true,"result":"Interrupted by user","usage":{}}"#; translator.expect_interrupt(); @@ -1398,7 +1630,7 @@ mod tests { #[test] fn replayed_and_synthetic_user_text_is_skipped() { let dir = tempfile::tempdir().expect("tempdir"); - let mut translator = Translator::new(dir.path().to_path_buf()); + let mut translator = Translator::new(dir.path().to_path_buf(), test_subagents(&dir)); let events = translate_lines( &mut translator, &[ diff --git a/server/src/session/echo.rs b/server/src/session/echo.rs index 20c2c3f..c899524 100644 --- a/server/src/session/echo.rs +++ b/server/src/session/echo.rs @@ -48,6 +48,10 @@ //! and height the app draws, in one session, which is what a scrolling //! problem needs in order to be reproduced twice the same way. //! - `/table [columns]` -- a markdown table with cells too long for one line. +//! - `/subagent [n]` -- n subagents at once (default 1), each named +//! "helper k", its prompt recorded as its own first user message: a +//! streamed reply, one Bash call, then it finishes about three seconds +//! later, the same lifecycle a real Task call has -- see `SUBAGENTS.md`. //! //! `/slow` earns its place: a queued message, a Stop button and a spinner are //! states that only exist mid-turn, and the obvious way to get one -- ask a @@ -62,6 +66,7 @@ use std::time::Duration; use super::driver::{ AttachmentRef, Driver, Event, EventSink, QuestionOption, SessionStatus, Unqueued, }; +use super::subagent::Subagents; /// Delay between streamed deltas -- long enough that streaming is visibly /// streaming, short enough that tests waiting on a full turn stay fast. @@ -108,6 +113,10 @@ pub struct EchoDriver { /// says it recovered, and a clear leaves it unmeasured. What is real is /// which way the numbers move. context: Arc, + /// This session's subagents -- see `SUBAGENTS.md`. `/subagent` is the + /// test rig for the same registry the claude driver routes real Task + /// calls into. + subagents: Arc, } impl EchoDriver { @@ -384,6 +393,55 @@ impl EchoDriver { return; } + // `n` subagents at once, each with its own transcript in the + // registry a real Task call routes into -- see `SUBAGENTS.md`. The + // parent's own Task calls end when their subagent does, three + // seconds later, which is long enough to see the running state on + // the phone before it finishes. + if let Some(rest) = text.strip_prefix("/subagent") { + let n = rest.trim().parse::().unwrap_or(1).clamp(1, 8); + if announce { + self.emit(Event::MessageTaken { + id: None, + text: text.clone(), + attachments, + }); + } + self.emit(Event::Status { + state: SessionStatus::Running, + }); + let sink = self.sink.clone(); + let subagents = Arc::clone(&self.subagents); + tokio::spawn(async move { + let mut helpers = Vec::new(); + for k in 1..=n { + let id = format!("echo-subagent-{k}-{}", super::random_hex()); + let title = format!("helper {k}"); + let prompt = format!( + "You are helper {k} of {n}. Say a few words, run a command, then stop." + ); + let _ = sink.send(Event::ToolStart { + id: id.clone(), + tool: "Task".to_string(), + input: serde_json::json!({ + "description": title, + "prompt": prompt, + "subagent_type": "general-purpose", + }), + }); + subagents.start(&id, &title, Some(&prompt)); + helpers.push((id, sink.clone(), Arc::clone(&subagents))); + } + for (id, sink, subagents) in helpers { + tokio::spawn(run_helper(id, sink, subagents)); + } + let _ = sink.send(Event::Status { + state: SessionStatus::Idle, + }); + }); + return; + } + // The same word the real CLI takes, so a phone drives both the same way. // `Driver::compact` is what the manager's route calls; this is the typed // path onto it. @@ -679,7 +737,12 @@ impl EchoDriver { }); } - pub fn new(sink: EventSink, session_dir: PathBuf, usage: crate::usage::Fixture) -> Self { + pub fn new( + sink: EventSink, + session_dir: PathBuf, + usage: crate::usage::Fixture, + subagents: Arc, + ) -> Self { let driver = Self { sink, pending_questions: Mutex::new(Vec::new()), @@ -688,6 +751,7 @@ impl EchoDriver { queued: Arc::new(Mutex::new(Vec::new())), session_dir, usage, + subagents, }; driver.emit(Event::Status { state: SessionStatus::Idle, @@ -809,6 +873,51 @@ async fn write_beat(sink: &EventSink, session_dir: &Path, beat: usize) { tokio::time::sleep(Duration::from_millis(120)).await; } +/// One `/subagent` helper: a few streamed words, one Bash call, then +/// `Status::Exited` about three seconds after it started -- long enough that +/// its `Running` state can be seen on the phone before it finishes. The +/// parent's own Task call for it ends at the same moment, the same way a +/// real Task's `tool_result` ends it. +async fn run_helper(id: String, sink: EventSink, subagents: Arc) { + let start = tokio::time::Instant::now(); + for word in "Working on it now.".split_inclusive(' ') { + subagents.record( + &id, + Event::AssistantText { + delta: word.to_string(), + }, + ); + tokio::time::sleep(DELTA_DELAY).await; + } + let tool_id = format!("{id}-bash"); + subagents.record( + &id, + Event::ToolStart { + id: tool_id.clone(), + tool: "Bash".to_string(), + input: serde_json::json!({ "command": "echo helper done" }), + }, + ); + tokio::time::sleep(DELTA_DELAY).await; + subagents.record( + &id, + Event::ToolEnd { + id: tool_id, + output: "helper done".to_string(), + }, + ); + let target = Duration::from_secs(3); + let elapsed = start.elapsed(); + if elapsed < target { + tokio::time::sleep(target - elapsed).await; + } + subagents.finish(&id); + let _ = sink.send(Event::ToolEnd { + id, + output: "subagent finished".to_string(), + }); +} + /// A message written during a turn and waiting for it to end: the id of the /// `MessageQueued` that announced it, what it said, and what was attached. All /// three, because all three are what the `MessageTaken` at the other end owes. diff --git a/server/src/session/llama.rs b/server/src/session/llama.rs index c9ce819..e70e041 100644 --- a/server/src/session/llama.rs +++ b/server/src/session/llama.rs @@ -77,6 +77,7 @@ impl LlamaDriver { /// in a different currency: two servers holding the same model is twice the /// memory, and the second would bind a different port while the phone kept /// talking to the first. + #[allow(clippy::too_many_arguments)] pub fn launch( meta: &SessionConfig, provider: &ProviderConfig, @@ -85,6 +86,10 @@ impl LlamaDriver { transcript: &Path, session_dir: &Path, sink: EventSink, + // llama.cpp has no notion of a Task call, so this is accepted only + // to keep one shape across every driver's launch -- see + // `SUBAGENTS.md`'s "Server layout". + _subagents: Arc, ) -> Result { let model = meta.model.as_deref().context( "a llama.cpp session needs a model -- one of the downloaded ones, by its key", diff --git a/server/src/session/mod.rs b/server/src/session/mod.rs index 463c658..8d1adc5 100644 --- a/server/src/session/mod.rs +++ b/server/src/session/mod.rs @@ -15,6 +15,7 @@ pub mod import; pub mod llama; pub mod pending; pub mod process; +pub mod subagent; pub mod transcript; pub mod transport; @@ -37,6 +38,7 @@ use driver::{ }; use echo::EchoDriver; use llama::LlamaDriver; +use subagent::Subagents; use transcript::{SeqEvent, Transcript}; use transport::Transport; @@ -265,6 +267,11 @@ pub struct SessionInfo { pub status: SessionStatus, pub last_activity: f64, pub created: f64, + /// How many subagents this session has started, from a directory + /// listing rather than reading each one's status -- see + /// `GET /sessions/{id}/subagents` for that. 0 when it has none, not + /// absent: every session can say this without asking anything. + pub subagents: usize, } /// What is running a session at this moment, and `None` when nothing is. @@ -293,6 +300,10 @@ pub struct LiveSession { events: broadcast::Sender, transcript_path: PathBuf, shared: Arc, + /// This session's subagents -- see `SUBAGENTS.md`. Built once at launch + /// and handed to whichever driver replaces it across a stop/start, so a + /// subagent started before a Stop is still there to read after a Start. + subagents: Arc, } /// Commands waiting for the session to be between turns. @@ -502,6 +513,19 @@ impl LiveSession { &self.transcript_path } + pub fn subagents(&self) -> &Arc { + &self.subagents + } + + /// What this session is doing right now, as the pump last recorded it -- + /// the same word `SessionInfo::status` reports. Read here rather than + /// only through `SessionManager::sessions` for + /// `GET /sessions/{id}/subagents`, which needs exactly this and nothing + /// else `SessionInfo` carries. + pub fn status(&self) -> SessionStatus { + *self.shared.status.lock().unwrap() + } + /// The session's directory (attachments in, produced files out live in /// `attachments/` and `files/` under it). pub fn dir(&self) -> &Path { @@ -578,6 +602,7 @@ impl LiveSession { status: *self.shared.status.lock().unwrap(), last_activity: *self.shared.last_activity.lock().unwrap(), created: self.meta.created, + subagents: subagent::count(self.dir()), } } } @@ -1065,6 +1090,7 @@ impl SessionManager { status: status_of_unlaunched(&self.data_dir.join(&meta.id)), last_activity: meta.created, created: meta.created, + subagents: subagent::count(&self.data_dir.join(&meta.id)), }, }) .collect() @@ -1828,6 +1854,7 @@ impl SessionManager { session.dir(), session.transcript_path(), &session.sink, + session.subagents(), )?); } // Nothing is live for this one -- a session whose launch failed @@ -2289,6 +2316,11 @@ fn launch( let (sink, source) = mpsc::unbounded_channel(); let (events, _) = broadcast::channel(EVENT_BUFFER); + // Built once per session, here, rather than per driver: a subagent + // started before a Stop has to still be there to read after a Start, + // and only `launch` runs once across that boundary -- `start_if_exited` + // replaces the driver alone. + let subagents = Arc::new(subagent::Subagents::new(dir.clone())); let shared = Arc::new(Shared { // What it was last known to be doing, not an assumption. A driver // that has something to say corrects this within its first poll. @@ -2350,7 +2382,18 @@ fn launch( let driver = Arc::new(Mutex::new( driving - .then(|| make_driver(&meta, setup, provider, env, &dir, &transcript_path, &sink)) + .then(|| { + make_driver( + &meta, + setup, + provider, + env, + &dir, + &transcript_path, + &sink, + &subagents, + ) + }) .transpose()?, )); @@ -2368,6 +2411,7 @@ fn launch( events.clone(), Arc::clone(&commands), announce, + Arc::clone(&subagents), )); Ok(Arc::new(LiveSession { @@ -2378,6 +2422,7 @@ fn launch( events, transcript_path, shared, + subagents, })) } @@ -2388,6 +2433,7 @@ fn launch( /// what [`SessionManager::start_session`] builds. That path replaces the /// driver and nothing else, so it has to construct one the same way rather /// than becoming a second answer to "what runs this". +#[allow(clippy::too_many_arguments)] fn make_driver( meta: &SessionConfig, setup: &SetupConfig, @@ -2396,13 +2442,17 @@ fn make_driver( dir: &Path, transcript_path: &Path, sink: &EventSink, + subagents: &Arc, ) -> Result> { Ok(match provider.kind { DriverKind::Echo => Arc::new(EchoDriver::new( sink.clone(), dir.to_path_buf(), env.usage.clone(), + Arc::clone(subagents), )), + // llama.cpp has no notion of a Task call, so it takes the registry + // and never touches it -- see `SUBAGENTS.md`'s "Server layout". DriverKind::LlamaCpp => Arc::new(LlamaDriver::launch( meta, provider, @@ -2411,6 +2461,7 @@ fn make_driver( transcript_path, dir, sink.clone(), + Arc::clone(subagents), )?), DriverKind::ClaudeCli => Arc::new(ClaudeDriver::launch( meta, @@ -2418,6 +2469,7 @@ fn make_driver( &Transport::for_setup(setup), dir, sink.clone(), + Arc::clone(subagents), )?), }) } @@ -2482,6 +2534,7 @@ fn notification_for( } } +#[allow(clippy::too_many_arguments)] async fn pump( id: String, mut transcript: Transcript, @@ -2490,6 +2543,7 @@ async fn pump( events: broadcast::Sender, commands: Arc, announce: Announcements, + subagents: Arc, ) { // Messages the session has been given and not started reading, which is // what makes a turn ending not the same thing as the work ending. @@ -2611,7 +2665,12 @@ async fn pump( } => commands.take_one(), Event::Status { state: SessionStatus::Exited, - } => commands.abandon("this session's process has exited"), + } => { + commands.abandon("this session's process has exited"); + // The process behind every open subagent was this + // session's own -- see `SUBAGENTS.md`'s lifecycle #4. + subagents.finish_all(); + } // The two ends of a message's wait. A `UserMessage` with // no id never waited -- it was sent between turns, and // counting it would take the total below zero. @@ -2752,6 +2811,7 @@ mod tests { sink.clone(), dir.path().to_path_buf(), crate::usage::Fixture::new(), + Arc::new(subagent::Subagents::new(dir.path().to_path_buf())), ))))), sink, waiting: Mutex::new(VecDeque::new()), @@ -4399,4 +4459,69 @@ mod tests { let seen = collect_turn(&mut rx).await; assert!(seen.first().expect("events").seq > last_seq); } + + /// `/subagent 2` is the test rig for `SUBAGENTS.md`'s whole feature: + /// each helper gets its own transcript with its prompt as its first + /// user message, `SessionInfo::subagents` counts them from the + /// directory, and each finishes on its own a few seconds later. + #[tokio::test] + async fn subagent_helpers_get_their_own_transcripts_and_finish() { + let dir = tempfile::tempdir().expect("tempdir"); + let config_path = dir.path().join("config.ron"); + let data_dir = dir.path().join("sessions"); + seed_echo_only(&config_path); + let manager = SessionManager::new(config_path, data_dir.clone(), data_dir.join("models")) + .expect("manager"); + let info = manager.spawn_session(echo_spec()).expect("spawn"); + let session = manager.session(&info.id).expect("live"); + + session.send_message("/subagent 2".to_string(), Vec::new()); + // Both helpers exist as soon as their Task calls go out, well before + // either finishes. + let deadline = tokio::time::Instant::now() + Duration::from_secs(2); + loop { + if session.subagents().list(true).len() == 2 { + break; + } + assert!( + tokio::time::Instant::now() < deadline, + "both helpers should have started by now" + ); + tokio::time::sleep(Duration::from_millis(20)).await; + } + let rows = session.subagents().list(true); + let mut titles: Vec<&str> = rows.iter().map(|row| row.title.as_str()).collect(); + titles.sort_unstable(); + assert_eq!(titles, ["helper 1", "helper 2"]); + assert!(rows.iter().all(|row| row.status == SessionStatus::Running)); + assert_eq!(manager.sessions()[0].subagents, 2); + + // Each subagent's own transcript opens with its prompt. + let first = session.subagents().get(&rows[0].id).expect("subagent"); + let events = + transcript::read_after(&first.transcript_path(), 0).expect("read subagent transcript"); + assert!( + events + .iter() + .any(|entry| matches!(&entry.event, Event::UserMessage { text, .. } if text.contains("helper"))) + ); + + // Each finishes on its own about three seconds after it started. + let deadline = tokio::time::Instant::now() + Duration::from_secs(5); + loop { + if session + .subagents() + .list(true) + .iter() + .all(|row| row.status == SessionStatus::Exited) + { + break; + } + assert!( + tokio::time::Instant::now() < deadline, + "both helpers should have finished by now" + ); + tokio::time::sleep(Duration::from_millis(50)).await; + } + } } diff --git a/server/src/session/subagent.rs b/server/src/session/subagent.rs new file mode 100644 index 0000000..c825e9f --- /dev/null +++ b/server/src/session/subagent.rs @@ -0,0 +1,490 @@ +//! A session's subagents -- see `SUBAGENTS.md`. +//! +//! **A subagent is a second transcript owned by a session, in the same event +//! model, with no process and no controls.** It shares the transcript file +//! format, the paging routes, and the SSE stream with a session by +//! addressing, not by copying: `Transcript`, `read_window` and `catch_up` +//! work on a subagent's file unchanged. +//! +//! Storage is `/subagents//{meta.json,transcript.jsonl}`, +//! where `` is the Task tool_use id that started it -- unique, stable +//! across a backend restart, and already the key the parent side uses. Only +//! ids matching [`is_subagent_id`] are ever turned into a path. + +use std::collections::HashMap; +use std::fs; +use std::path::{Path, PathBuf}; +use std::sync::{Arc, Mutex}; + +use anyhow::{Context, Result}; +use serde::{Deserialize, Serialize}; +use tokio::sync::broadcast; + +use super::driver::{Event, SessionStatus}; +use super::transcript::{SeqEvent, Transcript}; + +/// Fan-out buffer for one subagent's SSE subscribers. Smaller than a +/// session's: a subagent's whole conversation is usually a handful of tool +/// calls, not an hours-long session. +const EVENT_BUFFER: usize = 64; + +/// Whether `id` is safe to become a path segment under a session's +/// `subagents/` directory. Mirrors `import::is_session_id`'s reasoning: the +/// id arrives as a value inside JSON the CLI sent, and it becomes a +/// directory name, so a `/` or `..` in it must never be trusted. +fn is_subagent_id(id: &str) -> bool { + !id.is_empty() + && id.len() <= 200 + && id + .bytes() + .all(|b| b.is_ascii_alphanumeric() || b == b'_' || b == b'-') +} + +/// What a subagent's directory holds beside its transcript. Small and +/// separate from `Subagent` itself because this is exactly what survives a +/// backend restart on disk, and nothing else does. +#[derive(Debug, Clone, Serialize, Deserialize)] +struct Meta { + title: String, + /// Epoch seconds. Absent from `SubagentInfo`'s sort key deliberately: + /// `list` sorts by this rather than by directory order, which a + /// filesystem does not promise. + created: f64, +} + +/// One row of `GET /sessions/{id}/subagents`. +#[derive(Debug, Clone, Serialize)] +#[serde(rename_all = "camelCase")] +pub struct SubagentInfo { + pub id: String, + pub title: String, + pub status: SessionStatus, + pub created: f64, + pub last_activity: f64, +} + +/// One subagent: its own transcript and broadcast, same shape as a +/// session's but with no driver behind it. +pub struct Subagent { + dir: PathBuf, + transcript: Mutex, + events: broadcast::Sender, + /// Mirrors the transcript's last `Status` event, kept live rather than + /// read back from `Transcript::last_status` -- that answers "as of + /// opening" (see its own doc comment) and never moves for an append made + /// through *this* object, which is every append a live subagent ever + /// makes. Without this, `finish` immediately after `start` in the same + /// process read the file's stale opening status and reported itself + /// still open. + status: Mutex, +} + +impl Subagent { + pub fn transcript_path(&self) -> PathBuf { + self.dir.join("transcript.jsonl") + } + + pub fn subscribe(&self) -> broadcast::Receiver { + self.events.subscribe() + } + + /// Whether this subagent's last recorded status is not `Exited` -- + /// what decides whether a further child line still belongs in its + /// transcript. See `SUBAGENTS.md`'s lifecycle: "a child line whose + /// subagent finished already... is ignored". + pub fn is_open(&self) -> bool { + *self.status.lock().unwrap() != SessionStatus::Exited + } + + fn append(&self, event: Event) { + let mut transcript = self.transcript.lock().unwrap(); + match transcript.append(event, super::now()) { + Ok(entry) => { + if let Event::Status { state } = &entry.event { + *self.status.lock().unwrap() = *state; + } + // No subscribers is fine; the transcript already has it, + // same as a session's pump. + let _ = self.events.send(entry); + } + Err(err) => tracing::error!("subagent transcript append failed: {err:#}"), + } + } +} + +/// Every subagent one session has started, keyed by the Task tool_use id +/// that names it. +/// +/// Lives beside a session's driver rather than inside it: a claude driver +/// holds an `Arc` to this and routes child lines into it; echo uses it for +/// its `/subagent` rig; llama ignores it, since it has no notion of a Task +/// call. One instance per live session, built at launch and handed to +/// whichever driver replaces it across a stop/start. +pub struct Subagents { + /// The session's own directory; subagents live under `/subagents`. + dir: PathBuf, + live: Mutex>>, +} + +impl Subagents { + pub fn new(session_dir: PathBuf) -> Self { + Self { + dir: session_dir, + live: Mutex::new(HashMap::new()), + } + } + + fn subagents_dir(&self) -> PathBuf { + self.dir.join("subagents") + } + + /// Opens the subagent named `id`, creating it if this is the first + /// anyone has heard of it -- on disk as well as in memory, so a + /// subagent from before a backend restart is reopened rather than + /// recreated. `title`/`prompt` are used only at creation: reopening an + /// existing one keeps its original title and never repeats the prompt + /// into its transcript a second time. + fn open_or_create(&self, id: &str, title: &str, prompt: Option<&str>) -> Result> { + let dir = self.subagents_dir().join(id); + let meta_path = dir.join("meta.json"); + let existed = meta_path.is_file(); + let meta = if existed { + let text = fs::read_to_string(&meta_path) + .with_context(|| format!("read {}", meta_path.display()))?; + serde_json::from_str::(&text).context("parse subagent meta")? + } else { + wg_app_link::private::create_dir(&dir)?; + let meta = Meta { + title: title.to_string(), + created: super::now(), + }; + wg_app_link::private::write_file( + &meta_path, + serde_json::to_string(&meta) + .context("serialize subagent meta")? + .as_bytes(), + )?; + meta + }; + let mut transcript = Transcript::open(&dir.join("transcript.jsonl"))?; + if !existed { + // First lines, in order: the subagent is running the moment it + // exists, and its prompt -- when known -- is genuinely its first + // user turn. Written once, here, so a reopen never repeats them. + transcript.append( + Event::Status { + state: SessionStatus::Running, + }, + meta.created, + )?; + if let Some(prompt) = prompt { + transcript.append( + Event::UserMessage { + id: None, + text: prompt.to_string(), + attachments: Vec::new(), + }, + meta.created, + )?; + } + } + // A freshly created subagent is running by construction (its only + // lines so far are `Status::Running` and maybe its prompt); a + // reopened one takes whatever the file last said, since this + // `Transcript` has not been appended to yet in this process. + let status = if existed { + transcript.last_status().unwrap_or(SessionStatus::Running) + } else { + SessionStatus::Running + }; + let (events, _) = broadcast::channel(EVENT_BUFFER); + Ok(Arc::new(Subagent { + dir, + transcript: Mutex::new(transcript), + events, + status: Mutex::new(status), + })) + } + + /// Starts a subagent unless one is already known by this id -- see + /// `SUBAGENTS.md`'s lifecycle: created at the Task call or at the first + /// child line, whichever comes first, and never twice. A bad id is + /// refused rather than turned into a path. + pub fn start(&self, id: &str, title: &str, prompt: Option<&str>) { + if !is_subagent_id(id) { + tracing::debug!("refusing to start a subagent with a bad id {id:?}"); + return; + } + let mut live = self.live.lock().unwrap(); + if live.contains_key(id) { + return; + } + match self.open_or_create(id, title, prompt) { + Ok(subagent) => { + live.insert(id.to_string(), subagent); + } + Err(err) => tracing::error!("couldn't start subagent {id}: {err:#}"), + } + } + + /// The subagent named `id`, reopening it from disk on first use in this + /// process if one is there. `None` for an id nothing has ever started -- + /// deliberately not created here, since a route or a routing decision is + /// not the Task call that is supposed to be the only way one begins. + pub fn get(&self, id: &str) -> Option> { + if !is_subagent_id(id) { + return None; + } + if let Some(existing) = self.live.lock().unwrap().get(id).cloned() { + return Some(existing); + } + if !self.subagents_dir().join(id).join("meta.json").is_file() { + return None; + } + // Title and prompt are ignored: the directory already exists, so + // `open_or_create` reads its own meta rather than using either. + match self.open_or_create(id, "", None) { + Ok(subagent) => { + self.live + .lock() + .unwrap() + .insert(id.to_string(), Arc::clone(&subagent)); + Some(subagent) + } + Err(err) => { + tracing::error!("couldn't reopen subagent {id}: {err:#}"); + None + } + } + } + + /// Appends one event to a subagent's own transcript. A no-op, with a + /// debug log, for an id nothing was started under -- a child line for a + /// subagent this registry never opened is dropped rather than guessed + /// at. + pub fn record(&self, id: &str, event: Event) { + match self.live.lock().unwrap().get(id).cloned() { + Some(subagent) => subagent.append(event), + None => tracing::debug!("dropping an event for unknown subagent {id}"), + } + } + + /// The parent's `tool_result` for this Task id arrived: the subagent's + /// own `Status::Exited`. A no-op for an id that is not a subagent's, so + /// callers can call this for every `tool_result` without first checking + /// whether it belongs to one. + pub fn finish(&self, id: &str) { + if let Some(subagent) = self.live.lock().unwrap().get(id).cloned() + && subagent.is_open() + { + subagent.append(Event::Status { + state: SessionStatus::Exited, + }); + } + } + + /// The parent session's process is gone, so nothing still open here has + /// a process behind it either -- see `SUBAGENTS.md`'s lifecycle #4. + pub fn finish_all(&self) { + let subagents: Vec> = self.live.lock().unwrap().values().cloned().collect(); + for subagent in subagents { + if subagent.is_open() { + subagent.append(Event::Status { + state: SessionStatus::Exited, + }); + } + } + } + + /// Every subagent under this session's directory, oldest first -- + /// `GET /sessions/{id}/subagents`. Read straight from disk rather than + /// from `live`, so a subagent from before this process started (or one + /// this run has not yet touched) still shows up; one file read per + /// subagent, which is fine at the handful a session usually has. + /// + /// `session_running` is what turns a subagent whose last status is + /// `Running` into `Unknown`: its process was the session's, and the + /// session has none. + pub fn list(&self, session_running: bool) -> Vec { + let mut rows: Vec = match fs::read_dir(self.subagents_dir()) { + Ok(entries) => entries + .filter_map(Result::ok) + .filter_map(|entry| info_of(&entry.path(), session_running)) + .collect(), + // No directory is no subagents, not a fault worth reporting. + Err(_) => Vec::new(), + }; + rows.sort_by(|a, b| { + a.created + .partial_cmp(&b.created) + .unwrap_or(std::cmp::Ordering::Equal) + }); + rows + } +} + +fn info_of(subagent_dir: &Path, session_running: bool) -> Option { + let id = subagent_dir.file_name()?.to_str()?.to_string(); + let meta_path = subagent_dir.join("meta.json"); + let text = fs::read_to_string(&meta_path).ok()?; + let meta: Meta = serde_json::from_str(&text).ok()?; + let transcript = Transcript::open(&subagent_dir.join("transcript.jsonl")).ok()?; + // The subagent's first line is always `Status::Running`, written before + // this directory is discoverable at all, so `None` here is not a state a + // reader can actually observe -- but it is not this function's place to + // invent one, so a status this build does not expect to see falls back + // to the word the lifecycle promises it started in. + let last_status = transcript.last_status().unwrap_or(SessionStatus::Running); + let status = if last_status == SessionStatus::Running && !session_running { + SessionStatus::Unknown + } else { + last_status + }; + Some(SubagentInfo { + id, + title: meta.title, + status, + created: meta.created, + last_activity: transcript.last_activity().unwrap_or(meta.created), + }) +} + +/// How many subagents a session has, for `SessionInfo::subagents`: a +/// directory listing, so the session list stays cheap and only the +/// dedicated route pays for reading a status out of each one. +pub fn count(session_dir: &Path) -> usize { + fs::read_dir(session_dir.join("subagents")) + .map(|entries| entries.filter_map(Result::ok).count()) + .unwrap_or(0) +} + +#[cfg(test)] +mod tests { + use super::*; + + #[test] + fn a_bad_id_is_refused_rather_than_turned_into_a_path() { + let dir = tempfile::tempdir().expect("tempdir"); + let subagents = Subagents::new(dir.path().to_path_buf()); + subagents.start("../../etc", "escape", None); + assert!(subagents.get("../../etc").is_none()); + assert!(!dir.path().join("subagents").exists()); + } + + #[test] + fn starting_twice_keeps_the_first_title_and_does_not_repeat_the_prompt() { + let dir = tempfile::tempdir().expect("tempdir"); + let subagents = Subagents::new(dir.path().to_path_buf()); + subagents.start("toolu_1", "first title", Some("do the thing")); + subagents.start("toolu_1", "second title", Some("do the thing")); + + let rows = subagents.list(true); + assert_eq!(rows.len(), 1); + assert_eq!(rows[0].title, "first title"); + + let events = crate::session::transcript::read_after( + &subagents.get("toolu_1").unwrap().transcript_path(), + 0, + ) + .expect("read"); + assert_eq!( + events + .iter() + .filter(|e| matches!(e.event, Event::UserMessage { .. })) + .count(), + 1 + ); + } + + #[test] + fn a_reopened_subagent_continues_its_own_transcript() { + let dir = tempfile::tempdir().expect("tempdir"); + { + let subagents = Subagents::new(dir.path().to_path_buf()); + subagents.start("toolu_2", "helper", Some("go")); + subagents.record( + "toolu_2", + Event::AssistantText { + delta: "working".to_string(), + }, + ); + } + // A fresh registry, the way a backend restart builds one. + let subagents = Subagents::new(dir.path().to_path_buf()); + let subagent = subagents.get("toolu_2").expect("reopened"); + assert!(subagent.is_open()); + subagents.record( + "toolu_2", + Event::AssistantText { + delta: " more".to_string(), + }, + ); + let events = + crate::session::transcript::read_after(&subagent.transcript_path(), 0).expect("read"); + // Status, UserMessage, two AssistantText deltas, seq continuing. + assert_eq!(events.len(), 4); + assert_eq!(events.last().unwrap().seq, 4); + } + + #[test] + fn finishing_appends_exited_and_further_lines_are_droppable() { + let dir = tempfile::tempdir().expect("tempdir"); + let subagents = Subagents::new(dir.path().to_path_buf()); + subagents.start("toolu_3", "helper", None); + subagents.finish("toolu_3"); + let subagent = subagents.get("toolu_3").unwrap(); + assert!(!subagent.is_open()); + // On disk too, not only in the live cache `is_open` reads. + assert_eq!( + Transcript::open(&subagent.transcript_path()) + .expect("reopen") + .last_status(), + Some(SessionStatus::Exited) + ); + + // Finishing an id that was never a subagent is a no-op, not a panic. + subagents.finish("never-started"); + } + + #[test] + fn finish_all_closes_only_what_is_still_open() { + let dir = tempfile::tempdir().expect("tempdir"); + let subagents = Subagents::new(dir.path().to_path_buf()); + subagents.start("toolu_4", "one", None); + subagents.start("toolu_5", "two", None); + subagents.finish("toolu_4"); + subagents.finish_all(); + + let rows = subagents.list(false); + assert_eq!(rows.len(), 2); + for row in rows { + assert_eq!(row.status, SessionStatus::Exited); + } + } + + #[test] + fn a_subagent_still_running_when_the_session_is_not_reports_unknown() { + let dir = tempfile::tempdir().expect("tempdir"); + let subagents = Subagents::new(dir.path().to_path_buf()); + subagents.start("toolu_6", "helper", None); + + assert_eq!(subagents.list(true)[0].status, SessionStatus::Running); + assert_eq!(subagents.list(false)[0].status, SessionStatus::Unknown); + } + + #[test] + fn list_is_oldest_first_and_the_count_matches_the_directory() { + let dir = tempfile::tempdir().expect("tempdir"); + let subagents = Subagents::new(dir.path().to_path_buf()); + assert_eq!(count(dir.path()), 0); + subagents.start("toolu_a", "a", None); + std::thread::sleep(std::time::Duration::from_millis(2)); + subagents.start("toolu_b", "b", None); + let rows = subagents.list(true); + assert_eq!( + rows.iter().map(|r| r.id.as_str()).collect::>(), + ["toolu_a", "toolu_b"] + ); + assert_eq!(count(dir.path()), 2); + } +} From cf10b17c5b9e1ba27f390f856eb37e667b0a19e1 Mon Sep 17 00:00:00 2001 From: iris <2+iris@noreply.localhost> Date: Sat, 5 Sep 2026 14:44:00 -0400 Subject: [PATCH 09/12] End a subagent on its own end_turn, not the parent's tool_result, and give the expander a touch-sized row The Agent tool runs subagents in the background, so the parent's result arrives at launch while the subagent works on for minutes; finishing on it read a running agent as finished with a transcript cut off at launch. A subagent now ends on its own message_delta end_turn, and a later line for a finished one reopens it, since a background agent can be messaged again. The card's expander row was only the chevron's height, so a tap for it landed on the first subcard; it is the platform's 48dp minimum now. Co-Authored-By: Claude Fable 5.1 --- SUBAGENTS.md | 32 ++- .../com/example/aiapp/SessionListScreen.kt | 6 + server/src/session/claude/translate.rs | 184 ++++++++++++++++-- server/src/session/subagent.rs | 31 ++- 4 files changed, 227 insertions(+), 26 deletions(-) diff --git a/SUBAGENTS.md b/SUBAGENTS.md index b173aa7..0d8a1c7 100644 --- a/SUBAGENTS.md +++ b/SUBAGENTS.md @@ -50,17 +50,35 @@ There is no separate delete. 2. Every child line is translated by that subagent's own `Translator` (one per subagent: tool ids are unique but streaming deltas are by content-block index, and parallel subagents interleave). -3. When the parent's `tool_result` for the Task id arrives, the parent gets - its `ToolEnd` as before, and the subagent gets `Status Exited`. -4. When the parent session's process exits (`Status Exited` on the +3. **The parent's `tool_result` never finishes a subagent.** The Task tool + runs in the background by default: the `tool_result` -- "Async agent + launched..." -- arrives the moment it *starts*, while the subagent goes + on working for however long its own turn takes, sometimes minutes. What + ends it is its own turn ending: the raw API's `message_delta` on its + stream carrying `stop_reason: "end_turn"` (a `stop_reason` of `tool_use` + is the model about to call one, not an end), or a `result` line for its + own turn if a future CLI version ever sends one. Either maps to + `Status Exited`; the subagent's vocabulary has no `Idle`, so the + equivalent event `dispatch` produces for an ordinary session is dropped + rather than written. A shipped version of this finished on the + `tool_result` instead, which read a running background agent as + "finished" with its transcript truncated at the moment it launched. +4. **A child line for a subagent that already finished reopens it** + (`Status Running`) rather than being dropped: a background Task can be + sent another message long after its first turn ended, and that is + exactly what a further line for it means. Same transcript, same child + `Translator`, just picking back up. +5. When the parent session's process exits (`Status Exited` on the session), every subagent still `Running` gets `Status Exited` too: its process was the parent's. A subagent that was mid-flight when the backend restarted keeps working: the registry reopens the existing transcript on the next child line, and -the file continues its sequence. If its Task call finished while the backend -was down nothing ever closes it -- its last status stays `Running`, which -the list reports as **unknown** rather than as running (see the wire shape). +the file continues its sequence -- the same reopening #4 describes, whether +what closed it was a restart or its own `end_turn`. If its turn ended while +the backend was down nothing recorded that until the next line arrives, so +its last status stays `Running`, which the list reports as **unknown** +rather than as running (see the wire shape) until then. Title: the Task call's `description` input, then ` ()` when one is given; falling back to the tool's name when the child arrives before @@ -71,7 +89,7 @@ one is given; falling back to the tool's name when the child arrives before - `session/subagent.rs` -- the registry: `Subagents` (per session, in `Shared`), `Subagent` (its `Transcript` behind a mutex plus a `broadcast::Sender`), `record(id, event)`, `start(id, title, - prompt)`, `finish(id)`, `finish_all()`, `list()` from disk. Drivers get an + prompt)`, `finish(id)`, `reopen(id)`, `finish_all()`, `list()` from disk. Drivers get an `Arc` beside their `EventSink`; llama ignores it. - `session/claude/translate.rs` -- routes child lines by parent id, holds one child `Translator` per subagent, remembers pending Task calls' diff --git a/app/androidApp/src/main/kotlin/com/example/aiapp/SessionListScreen.kt b/app/androidApp/src/main/kotlin/com/example/aiapp/SessionListScreen.kt index 425d26b..27faf3f 100644 --- a/app/androidApp/src/main/kotlin/com/example/aiapp/SessionListScreen.kt +++ b/app/androidApp/src/main/kotlin/com/example/aiapp/SessionListScreen.kt @@ -11,6 +11,7 @@ import androidx.compose.foundation.layout.Spacer import androidx.compose.foundation.layout.fillMaxSize import androidx.compose.foundation.layout.fillMaxWidth import androidx.compose.foundation.layout.height +import androidx.compose.foundation.layout.heightIn import androidx.compose.foundation.layout.padding import androidx.compose.foundation.layout.width import androidx.compose.foundation.lazy.LazyColumn @@ -396,10 +397,15 @@ private fun SessionCard( // already on screen -- see UI_RULES on a control not displacing the text beside it. if (session.subagents > 0) { Spacer(Modifier.height(8.dp)) + // The platform's minimum touch height, not the chevron's own ten or so dp: + // at the chevron's height a tap meant for it landed on the first subcard + // beneath and opened a subagent instead. Row( horizontalArrangement = Arrangement.Center, + verticalAlignment = Alignment.CenterVertically, modifier = Modifier.fillMaxWidth() + .heightIn(min = 48.dp) .clickable(enabled = !deleting, onClick = onToggleSubagents) .semantics { contentDescription = diff --git a/server/src/session/claude/translate.rs b/server/src/session/claude/translate.rs index 6a4bd09..72964e4 100644 --- a/server/src/session/claude/translate.rs +++ b/server/src/session/claude/translate.rs @@ -151,10 +151,12 @@ impl Translator { fn translate_child(&mut self, id: &str, message: &Value) -> Vec { match self.subagents.get(id) { Some(subagent) if !subagent.is_open() => { - // The Task call already ended (or this line is stale from a - // resumed conversation) -- see `SUBAGENTS.md`'s lifecycle. - tracing::debug!("dropping a line for subagent {id}, which has already finished"); - return Vec::new(); + // Not stale: the Task tool runs in the background by + // default, so a finished subagent can still be sent another + // message later (SendMessage) and start working again. A + // line arriving after `finish` means exactly that, not a + // conversation that is over -- see `SUBAGENTS.md`. + self.subagents.reopen(id); } Some(_) => {} None => { @@ -178,7 +180,28 @@ impl Translator { .clone(); let events = child.lock().unwrap().dispatch(message); for event in events { - self.subagents.record(id, event); + // The subagent's own vocabulary is Running/Exited/Unknown, never + // Idle -- a background Task is either working or it has ended, + // never merely "between turns" the way a session is. Dropped + // here rather than never produced, so a `result` line's own + // `Idle` (dispatch's ordinary end-of-turn event, for a subagent + // dialect that ever sends one) is caught the same way a + // `message_delta` would be. + if !matches!( + event, + Event::Status { + state: SessionStatus::Idle + } + ) { + self.subagents.record(id, event); + } + } + // What actually ends a subagent's turn: not the parent's + // `tool_result`, which for a background Task arrives at launch + // ("Async agent launched...") long before the work is done -- see + // `SUBAGENTS.md`. + if ends_a_turn(message) { + self.subagents.finish(id); } Vec::new() } @@ -678,11 +701,11 @@ impl Translator { id: about.clone(), output: texts.join("\n"), }); - // A no-op unless `about` is a subagent's own id -- see - // `SUBAGENTS.md`'s lifecycle #3: the parent gets this `ToolEnd` - // like any other tool result, and the subagent it names (if it - // names one) gets its `Status::Exited`. - self.subagents.finish(&about); + // Deliberately does *not* finish a subagent `about` might name: + // the Task tool runs in the background by default, so this + // `tool_result` -- "Async agent launched..." -- arrives at + // launch, long before the subagent's own work is done. What + // ends it is its own turn ending, handled in `translate_child`. } events } @@ -705,6 +728,32 @@ fn fallback_title(message: &Value) -> String { .to_string() } +/// Whether this line is a subagent's *own* turn ending -- the only thing +/// that does, per `SUBAGENTS.md`: not the parent's `tool_result`, which for +/// a background Task arrives at launch rather than at completion. +/// +/// Checked on the raw line rather than on what `dispatch` returns, so this +/// never has to touch the shared `translate_stream_event`/`dispatch` code a +/// top-level session's own turn-ending also goes through -- a subagent's +/// idea of "ended" must not change when a real session's does. +/// +/// `message_delta` is the raw API's own signal, carrying the stop reason: +/// `end_turn` is genuinely done, `tool_use` means the model is about to call +/// one and there is more coming. A `result` line is the CLI's own shape for +/// a top-level turn; a subagent has not been observed to send one, but +/// SUBAGENTS.md counts it too in case a future CLI version does. +fn ends_a_turn(message: &Value) -> bool { + match message.get("type").and_then(Value::as_str) { + Some("stream_event") => { + let event = &message["event"]; + event.get("type").and_then(Value::as_str) == Some("message_delta") + && event["delta"].get("stop_reason").and_then(Value::as_str) == Some("end_turn") + } + Some("result") => true, + _ => false, + } +} + /// Whether a failed turn failed because the account is out of quota, and when /// the CLI said the limit lifts. /// @@ -1043,7 +1092,13 @@ mod tests { /// The parent's `tool_result` for the Task id is what ends the /// subagent -- SUBAGENTS.md's lifecycle #3 -- and nothing else does. #[test] - fn the_parents_tool_result_finishes_the_subagent() { + fn the_parents_tool_result_does_not_finish_the_subagent() { + // The Task tool runs in the background by default: this + // `tool_result` is "Async agent launched...", arriving the moment + // the subagent *starts*, while it goes on working for however long + // its own turn takes. Finishing it here was the bug -- a running + // background agent read as "finished" with its transcript truncated + // at launch. let dir = tempfile::tempdir().expect("tempdir"); let subagents = test_subagents(&dir); let mut translator = Translator::new(dir.path().to_path_buf(), Arc::clone(&subagents)); @@ -1058,10 +1113,115 @@ mod tests { translate_lines( &mut translator, &[ - r#"{"type":"user","message":{"role":"user","content":[{"type":"tool_result","tool_use_id":"toolu_task2","content":"done","is_error":false}]},"parent_tool_use_id":null}"#, + r#"{"type":"user","message":{"role":"user","content":[{"type":"tool_result","tool_use_id":"toolu_task2","content":"Async agent launched","is_error":false}]},"parent_tool_use_id":null}"#, ], ); + assert!(subagent.is_open()); + } + + /// What actually ends a subagent: the raw API's own `message_delta` + /// saying its turn stopped with `end_turn`. Never written into the + /// subagent's own transcript as `Idle` -- its vocabulary has no such + /// state. + #[test] + fn the_subagents_own_end_turn_finishes_it() { + let dir = tempfile::tempdir().expect("tempdir"); + let subagents = test_subagents(&dir); + let mut translator = Translator::new(dir.path().to_path_buf(), Arc::clone(&subagents)); + translate_lines( + &mut translator, + &[ + r#"{"type":"assistant","message":{"content":[{"type":"tool_use","id":"toolu_task3","name":"Task","input":{"description":"helper"}}]},"parent_tool_use_id":null}"#, + r#"{"type":"stream_event","event":{"type":"message_delta","delta":{"stop_reason":"end_turn"}},"parent_tool_use_id":"toolu_task3"}"#, + ], + ); + let subagent = subagents.get("toolu_task3").expect("subagent started"); assert!(!subagent.is_open()); + let lines = crate::session::transcript::read_after(&subagent.transcript_path(), 0) + .expect("read subagent transcript"); + assert!( + !lines + .iter() + .any(|entry| matches!(&entry.event, Event::Status { state } if *state == SessionStatus::Idle)), + "a subagent's transcript must never carry Idle: {lines:?}" + ); + assert_eq!( + lines.last().unwrap().event, + Event::Status { + state: SessionStatus::Exited + } + ); + } + + /// `stop_reason: "tool_use"` is the model about to call a tool, with + /// more of the turn still coming -- not an end. + #[test] + fn a_stop_reason_of_tool_use_does_not_finish_the_subagent() { + let dir = tempfile::tempdir().expect("tempdir"); + let subagents = test_subagents(&dir); + let mut translator = Translator::new(dir.path().to_path_buf(), Arc::clone(&subagents)); + translate_lines( + &mut translator, + &[ + r#"{"type":"assistant","message":{"content":[{"type":"tool_use","id":"toolu_task4","name":"Task","input":{"description":"helper"}}]},"parent_tool_use_id":null}"#, + r#"{"type":"stream_event","event":{"type":"message_delta","delta":{"stop_reason":"tool_use"}},"parent_tool_use_id":"toolu_task4"}"#, + ], + ); + assert!( + subagents + .get("toolu_task4") + .expect("subagent started") + .is_open() + ); + } + + /// A background Task can be sent another message long after its first + /// turn ended -- a further child line for it reopens rather than being + /// dropped, and the same transcript and child translator carry on. + #[test] + fn a_line_after_finish_reopens_the_subagent_rather_than_being_dropped() { + let dir = tempfile::tempdir().expect("tempdir"); + let subagents = test_subagents(&dir); + let mut translator = Translator::new(dir.path().to_path_buf(), Arc::clone(&subagents)); + translate_lines( + &mut translator, + &[ + r#"{"type":"assistant","message":{"content":[{"type":"tool_use","id":"toolu_task5","name":"Task","input":{"description":"helper"}}]},"parent_tool_use_id":null}"#, + r#"{"type":"stream_event","event":{"type":"message_delta","delta":{"stop_reason":"end_turn"}},"parent_tool_use_id":"toolu_task5"}"#, + ], + ); + let subagent = subagents.get("toolu_task5").expect("subagent started"); + assert!(!subagent.is_open()); + + translate_lines( + &mut translator, + &[ + r#"{"type":"assistant","message":{"content":[{"type":"tool_use","id":"toolu_more","name":"Bash","input":{}}]},"parent_tool_use_id":"toolu_task5"}"#, + ], + ); + assert!(subagent.is_open()); + let lines = crate::session::transcript::read_after(&subagent.transcript_path(), 0) + .expect("read subagent transcript"); + // Running, [prompt], Exited, Running (reopened), then the new line's + // own ToolStart -- the same transcript throughout, not a new one. + assert!( + lines.iter().any( + |entry| matches!(&entry.event, Event::ToolStart { tool, .. } if tool == "Bash") + ) + ); + assert_eq!( + lines + .iter() + .filter(|entry| matches!( + &entry.event, + Event::Status { + state: SessionStatus::Running + } + )) + .count(), + 2, + "expected one Running at creation and one at the reopen: {lines:?}" + ); } /// Two subagents running at once keep two separate transcripts: tool ids diff --git a/server/src/session/subagent.rs b/server/src/session/subagent.rs index c825e9f..a6a23e5 100644 --- a/server/src/session/subagent.rs +++ b/server/src/session/subagent.rs @@ -89,9 +89,9 @@ impl Subagent { } /// Whether this subagent's last recorded status is not `Exited` -- - /// what decides whether a further child line still belongs in its - /// transcript. See `SUBAGENTS.md`'s lifecycle: "a child line whose - /// subagent finished already... is ignored". + /// what decides whether a further child line reopens it (see + /// `Subagents::reopen`) rather than continuing straight through. See + /// `SUBAGENTS.md`'s lifecycle. pub fn is_open(&self) -> bool { *self.status.lock().unwrap() != SessionStatus::Exited } @@ -269,10 +269,12 @@ impl Subagents { } } - /// The parent's `tool_result` for this Task id arrived: the subagent's - /// own `Status::Exited`. A no-op for an id that is not a subagent's, so - /// callers can call this for every `tool_result` without first checking - /// whether it belongs to one. + /// The subagent's own turn ended: its `Status::Exited`. Called from + /// `translate_child` on the subagent's own `end_turn`, never on the + /// parent's `tool_result` -- a background Task's `tool_result` arrives + /// at launch, not at completion, so it says nothing about whether this + /// is over. A no-op for an id that is not a subagent's or is already + /// closed. pub fn finish(&self, id: &str) { if let Some(subagent) = self.live.lock().unwrap().get(id).cloned() && subagent.is_open() @@ -283,6 +285,21 @@ impl Subagents { } } + /// A line arrived for a subagent that had already finished: it is + /// working again, not stale -- a background Task can be sent another + /// message long after its first turn ended. Appends `Status::Running` + /// so the list stops reporting it as finished; a no-op if it was not + /// actually closed, so a caller need not check first. + pub fn reopen(&self, id: &str) { + if let Some(subagent) = self.live.lock().unwrap().get(id).cloned() + && !subagent.is_open() + { + subagent.append(Event::Status { + state: SessionStatus::Running, + }); + } + } + /// The parent session's process is gone, so nothing still open here has /// a process behind it either -- see `SUBAGENTS.md`'s lifecycle #4. pub fn finish_all(&self) { From 8d23a207921fe77115bc1da68625bdc0eb09a67d Mon Sep 17 00:00:00 2001 From: iris <2+iris@noreply.localhost> Date: Sat, 5 Sep 2026 21:35:33 -0400 Subject: [PATCH 10/12] iris: AndroidAppState::platform_ready, a JavaVM+View handle for later JNI calls Default no-op lifecycle hook, called once from new_peer right after new. P0's bench build needs to call BatteryManager/ClipboardManager through the view's own Context from a background thread as well as the UI thread, and neither a JavaVM nor a GlobalRef to the view was reachable from AndroidAppState::new before this. Existing implementors (Client, TranscriptClient) are unaffected. Co-Authored-By: Claude Fable 5.1 --- iris/src/android/view.rs | 20 ++++++++++++++++++-- 1 file changed, 18 insertions(+), 2 deletions(-) diff --git a/iris/src/android/view.rs b/iris/src/android/view.rs index f735c21..ed3128b 100644 --- a/iris/src/android/view.rs +++ b/iris/src/android/view.rs @@ -4,7 +4,7 @@ use accesskit_android::Adapter as AccessAdapter; use android_view::{ AccessibilityNodeInfo, AccessibilityNodeProvider, Bundle, CallbackCtx, Context, InputConnection, KeyEvent, MotionEvent, Rect, View, ViewPeer, - jni::{JNIEnv, sys::jint}, + jni::{JNIEnv, JavaVM, objects::GlobalRef, sys::jint}, ndk::event::{Keycode, MotionAction}, }; // `marker::Sized` explicitly: `crate::prelude::*` below also brings in the @@ -107,6 +107,19 @@ pub trait AndroidAppState: HasAndroidUiState { fn back_pressed(&mut self, rsc: &mut AndroidRsc, render: &mut UiRenderState) -> bool { false } + /// Called once, right after `new`, with a fresh `JavaVM` handle and a + /// global reference to this app's own `View` -- for a caller that + /// needs to call into Java itself beyond what a [`RequestRedraw`] + /// handle already covers (P0's bench build calling + /// `BatteryManager`/`ClipboardManager` through the view's `Context`, + /// docs/RUST.md). Not folded into `new` itself: most implementors need + /// nothing here, and `new`'s job is building the widget tree, not + /// holding a platform handle -- the default does nothing. `vm`/`view` + /// are independent handles from the ones `new_peer` keeps for its own + /// `RequestRedraw` (a fresh `get_java_vm`/`new_global_ref` each), so + /// storing them has no effect on that mechanism. + #[allow(unused_variables)] + fn platform_ready(&mut self, rsc: &mut AndroidRsc, vm: JavaVM, view: GlobalRef) {} } /// The android-view analogue of `default::DefaultRsc` -- identical in @@ -561,7 +574,10 @@ pub fn new_peer<'local, State: AndroidAppState>( }; let shared = Rc::new(RefCell::new(Shared::default())); let ui_state = AndroidUiState::new(shared.clone()); - let state = State::new(ui_state, &mut rsc); + let mut state = State::new(ui_state, &mut rsc); + let platform_vm = env.get_java_vm().unwrap(); + let platform_view = env.new_global_ref(&view.0).unwrap(); + state.platform_ready(&mut rsc, platform_vm, platform_view); let peer = IrisViewPeer { rsc, render: UiRenderState::new(), From 683db4908a7224e1228fc668ca14b6248870f4a1 Mon Sep 17 00:00:00 2001 From: iris <2+iris@noreply.localhost> Date: Sat, 5 Sep 2026 21:35:44 -0400 Subject: [PATCH 11/12] iris-android-app: a `bench` feature, P0's iris half A third AndroidAppState (BenchClient) on top of transcript-screen: embeds app/bench-fixture/assets/transcript.jsonl with include_str! (no server, no enrollment), folds the first 3,200 lines through client_core's real fold_page as the opening backlog, and holds the rest back as a streaming tail. "Run benchmark" resets FrameReport, animates the same 24-swipe/ 6-cycle scroll BenchRun.kt drives (List::scroll in ~60Hz steps, since iris has no built-in tween), then replays the tail at 20/s through fold_event -- the same fold path a live SSE reply takes -- and shows a report in a selectable TextEdit. "Copy report" puts it on the clipboard. The report adds process CPU time (libc::getrusage), peak RSS (/proc/self/ status's VmHWM) and battery current (BatteryManager.getIntProperty via direct JNI, bench_jni.rs's PlatformHandle) to FrameStats's existing frames/janky%/percentiles/CPU-GPU-split line -- "unavailable" rather than a fabricated number wherever the platform can't answer. build.rs now exits early under the bench feature before requiring a live server's host/port/token/CA: BenchClient never calls build_transport(). app/build.gradle gains a signed `release` build type (previously only debug) so the cdylib cargo ndk builds can be packaged for a phone, the same key app/build-apk.sh generates. Co-Authored-By: Claude Fable 5.1 --- iris/android-app/Cargo.lock | 2 + iris/android-app/Cargo.toml | 25 ++ iris/android-app/app/build.gradle | 30 ++ iris/android-app/build.rs | 8 + iris/android-app/src/bench_client.rs | 418 +++++++++++++++++++++++++++ iris/android-app/src/bench_jni.rs | 134 +++++++++ iris/android-app/src/lib.rs | 22 +- 7 files changed, 637 insertions(+), 2 deletions(-) create mode 100644 iris/android-app/src/bench_client.rs create mode 100644 iris/android-app/src/bench_jni.rs diff --git a/iris/android-app/Cargo.lock b/iris/android-app/Cargo.lock index 4f1a070..592dbfd 100644 --- a/iris/android-app/Cargo.lock +++ b/iris/android-app/Cargo.lock @@ -1768,9 +1768,11 @@ dependencies = [ "client-core", "event-model", "iris", + "libc", "log", "serde_json", "tabs-ui", + "tokio", "transcript-ui", ] diff --git a/iris/android-app/Cargo.toml b/iris/android-app/Cargo.toml index 776cbb4..d0551bd 100644 --- a/iris/android-app/Cargo.toml +++ b/iris/android-app/Cargo.toml @@ -32,6 +32,23 @@ transcript-ui = { path = "../transcript-ui", optional = true } client-core = { path = "../../client-core", optional = true } event-model = { path = "../../event-model", optional = true } serde_json = { version = "1", features = ["float_roundtrip"], optional = true } +# P0's bench build only (docs/RUST.md): `getrusage(RUSAGE_SELF)` for +# process CPU time, matching `libc::getrusage`'s mention in that box over +# parsing `/proc/self/stat` by hand and assuming `USER_HZ`. Already in the +# workspace's own dependency tree transitively (`iris/Cargo.lock`, pinned +# at 0.2.179) -- this makes it a direct dependency at the same version +# rather than a second, possibly-drifting resolution. +libc = { version = "0.2.179", optional = true } +# P0's bench build only: the scroll animation and the streaming phase are +# both a sequence of `sleep`s inside the async task `rsc.spawn_task` already +# runs on iris's own tokio runtime (`iris/src/task.rs`'s `Tasks::init`), and +# the battery sampler is a second, concurrent task on that same runtime +# (`tokio::spawn`) -- so this crate needs `tokio` directly rather than only +# through `iris`. `rt`+`time` only: no I/O, no macros, nothing this crate +# doesn't call. Version matches the one `iris`'s own dependency tree already +# resolves to (`iris/Cargo.lock`), so there is one copy of the runtime, not +# two. +tokio = { version = "1.53.1", features = ["rt", "time"], optional = true } [features] default = ["tabs-screen"] @@ -41,6 +58,14 @@ transcript-screen = ["dep:transcript-ui", "dep:client-core", "dep:event-model", # instead of SwiftShader's software Vulkan. See `iris/Cargo.toml`'s own doc # on the feature this forwards to. force-gles = ["iris/force-gles"] +# P0's iris half (docs/RUST.md, docs/AGENTS.md's "The rigs"): the same +# checked-in fixture, scroll loop and streaming phase the Compose `bench` +# build type drives, run here against `transcript-ui`'s real screen with no +# server. Depends on `transcript-screen` for `transcript-ui`/`client-core`/ +# `event-model` -- `lib.rs`'s `ActiveClient` selection gives this feature +# priority over `transcript-screen`'s own `TranscriptClient` when both are +# listed, which is how this crate's build command names both explicitly. +bench = ["transcript-screen", "dep:libc", "dep:tokio"] [profile.release] panic = "abort" diff --git a/iris/android-app/app/build.gradle b/iris/android-app/app/build.gradle index 4af67aa..b1f18e9 100644 --- a/iris/android-app/app/build.gradle +++ b/iris/android-app/app/build.gradle @@ -19,9 +19,39 @@ android { versionName = "1.0" } + // A release build must be signed, and the key is per machine rather than per repo -- same + // reasoning and the same key as `app/build-apk.sh` (the Compose app): it is what a phone + // recognises the app by, and a secret never lives in a checkout (the mount is shared with an + // untrusted VM). `build-apk.sh` generates this key once and points at it through the + // environment; without it a release build here is unsigned, which is fine for everything + // except installing. + def keystore = System.getenv("AI_APP_KEYSTORE") + signingConfigs { + if (keystore != null) { + release { + storeFile = file(keystore) + storePassword = System.getenv("AI_APP_KEYSTORE_PASSWORD") + keyAlias = "ai-app" + keyPassword = storePassword + } + } + } + buildTypes { debug { } + // P0's iris half (docs/RUST.md's P0 box): the build a phone actually runs. The `.so` + // itself is built separately with `cargo ndk --release --features "transcript-screen + // force-gles bench"` straight into src/main/jniLibs/ (this crate's own Cargo.toml) -- + // Gradle here only packages and signs whatever is already there, the same division as the + // debug/tabs-screen build this project started with. `applicationIdSuffix` keeps it + // installable beside a debug build of the tabs demo rather than replacing it. + release { + applicationIdSuffix ".bench" + if (keystore != null) { + signingConfig = signingConfigs.release + } + } } compileOptions { diff --git a/iris/android-app/build.rs b/iris/android-app/build.rs index 18704fe..f6cb3fb 100644 --- a/iris/android-app/build.rs +++ b/iris/android-app/build.rs @@ -22,6 +22,14 @@ fn main() { if std::env::var_os("CARGO_FEATURE_TRANSCRIPT_SCREEN").is_none() { return; } + // P0's bench build (docs/RUST.md) opens the checked-in fixture with no + // server at all -- `bench_client.rs` never references the `pinned` + // module this generates, so requiring a live server's host/port/token/ + // CA to build it (as plain `transcript-screen` does, below) would be a + // pointless requirement for a build that talks to nothing. + if std::env::var_os("CARGO_FEATURE_BENCH").is_some() { + return; + } println!("cargo:rerun-if-env-changed=AI_APP_TRANSCRIPT_HOST"); println!("cargo:rerun-if-env-changed=AI_APP_TRANSCRIPT_PORT"); println!("cargo:rerun-if-env-changed=AI_APP_TRANSCRIPT_TOKEN"); diff --git a/iris/android-app/src/bench_client.rs b/iris/android-app/src/bench_client.rs new file mode 100644 index 0000000..2b8c755 --- /dev/null +++ b/iris/android-app/src/bench_client.rs @@ -0,0 +1,418 @@ +//! P0's iris half (docs/RUST.md's P0 box, docs/AGENTS.md's "The rigs"): +//! the same fixture, scroll loop and streaming phase the Compose `bench` +//! build type's `BenchRun.kt`/`BenchFixture.kt` drive, run here against +//! `transcript-ui`'s real screen with no server -- a frame-time comparison +//! that measures the renderer rather than the data or the network. +//! +//! **Reuses `transcript_client.rs`'s shape** (folded items, a full +//! `transcript_ui::build_tree` rebuild per event) with the network half +//! replaced by the checked-in fixture, embedded with `include_str!` -- +//! `app/bench-fixture/assets/transcript.jsonl`, 1,915,760 bytes, generated +//! by `app/bench-fixture/generate.py` and never a real transcript (that +//! file's own README). The first 3,200 lines are the opening backlog, +//! folded once through `client_core::transcript_fold::fold_page` exactly +//! as a real `/transcript` page would be; the remaining ~400 are the +//! streaming tail, replayed one at a time through `fold_event` -- the same +//! fold path a live SSE reply arrives on -- by the "Run benchmark" +//! control below. + +use crate::bench_jni::PlatformHandle; +use android_view::jni::{JavaVM, objects::GlobalRef}; +use client_core::transcript_fold::{TranscriptItem, fold_event, fold_page, group_tool_runs}; +use event_model::SeqEvent; +use iris::android::{AndroidAppState, AndroidRsc, AndroidUiState, HasAndroidUiState}; +use iris::prelude::*; +use std::sync::Arc; +use std::sync::atomic::{AtomicBool, Ordering}; +use std::time::Duration; + +/// bench-fixture/README.md: the first `BACKLOG_COUNT` non-blank lines are +/// the opening window; the rest are the streaming tail. Kept in sync with +/// `BenchFixture.kt`'s identical constant by hand -- both read the same +/// checked-in file, so a mismatch would only mean the two apps' bench +/// builds open a different split of it, not a wrong-vs-right answer. +const BACKLOG_COUNT: usize = 3200; + +/// `BenchRun.kt`'s own constants -- kept identical so the two apps' bench +/// runs are the same gesture and the same load, which is the entire point +/// of a shared fixture and a shared scripted loop (P0's pass condition). +const CYCLES: usize = 6; +const SWIPE_PX: f32 = 900.0; +const SWIPE_MS: u64 = 200; +const SWIPE_PAUSE_MS: u64 = 500; +const STREAM_EVENTS_PER_SEC: u64 = 20; +const STREAM_SECONDS: u64 = 20; +/// One animation step's target cadence -- close enough to 60Hz that a +/// `List::scroll` swipe is many small moves rather than one jump, so +/// frames are actually rendered along the way (the point of animating it +/// at all rather than calling `scroll` once per swipe). +const ANIM_STEP_MS: u64 = 16; + +const FIXTURE_JSONL: &str = include_str!("../../../app/bench-fixture/assets/transcript.jsonl"); + +pub struct BenchClient { + ui_state: AndroidUiState, + content: WeakWidget, + report_display: WeakWidget, + screen: Option, + items: Vec, + /// The events not yet streamed -- consumed by `start_benchmark`'s own + /// clone, kept here only as the source a second run would need (the + /// button can be pressed more than once; `running` just stops overlap, + /// not repeat). + stream_tail: Vec, + platform: Option>, + last_report: Option, + running: bool, +} + +impl HasAndroidUiState for BenchClient { + fn android_state(&self) -> &AndroidUiState { + &self.ui_state + } + fn android_state_mut(&mut self) -> &mut AndroidUiState { + &mut self.ui_state + } +} + +/// Parses the fixture once: `serde_json::Value`s for the backlog +/// (`fold_page` takes a page of raw wire JSON, same as a real +/// `/transcript` response) and folded `SeqEvent`s for the tail (`fold_event` +/// takes one live wire event at a time, same as a real SSE frame). +fn parse_fixture() -> (Vec, Vec) { + let lines: Vec<&str> = FIXTURE_JSONL + .lines() + .filter(|line| !line.trim().is_empty()) + .collect(); + let mut backlog = Vec::with_capacity(BACKLOG_COUNT.min(lines.len())); + let mut stream_tail = Vec::new(); + for (i, line) in lines.iter().enumerate() { + let value: serde_json::Value = + serde_json::from_str(line).expect("bench fixture is generated JSON, always valid"); + if i < BACKLOG_COUNT { + backlog.push(value); + } else { + let event: SeqEvent = serde_json::from_value(value) + .expect("bench fixture event matches event-model's SeqEvent"); + stream_tail.push(event); + } + } + (backlog, stream_tail) +} + +fn placeholder(rsc: &mut Rsc, message: &str) -> StrongWidget { + wtext(message.to_string()) + .color(Color::WHITE) + .wrap(true) + .pad(16) + .add_strong(rsc) + .any() +} + +/// `getrusage(RUSAGE_SELF)`'s user+system time, in ms -- `None` only if +/// the syscall itself fails, which UI_RULES.md's "never present an +/// inferred value as a measured one" says to keep apart from a real (and +/// here, impossible) zero. +fn process_cpu_ms() -> Option { + // SAFETY: `rusage` is a plain-old-data struct `getrusage` fully + // initialises on success; on failure it is never read. + unsafe { + let mut usage: libc::rusage = std::mem::zeroed(); + if libc::getrusage(libc::RUSAGE_SELF, &mut usage) != 0 { + return None; + } + let user_ms = usage.ru_utime.tv_sec as u64 * 1000 + usage.ru_utime.tv_usec as u64 / 1000; + let sys_ms = usage.ru_stime.tv_sec as u64 * 1000 + usage.ru_stime.tv_usec as u64 / 1000; + Some(user_ms + sys_ms) + } +} + +/// `VmHWM` from `/proc/self/status` -- the process's peak RSS since it +/// started, in kB. Same source `BenchRun.kt`'s `peakRssLine` reads, so the +/// two reports' numbers mean the same thing. +fn peak_rss_kb() -> Option { + std::fs::read_to_string("/proc/self/status") + .ok()? + .lines() + .find_map(|line| line.strip_prefix("VmHWM:")) + .and_then(|rest| rest.trim().strip_suffix("kB")) + .and_then(|n| n.trim().parse().ok()) +} + +fn battery_line(samples: &[i32]) -> String { + if samples.is_empty() { + return " battery current: unavailable on this device".to_string(); + } + let mean = samples.iter().map(|&v| v as i64).sum::() / samples.len() as i64; + let min = samples.iter().min().unwrap(); + let max = samples.iter().max().unwrap(); + format!( + " battery current: mean {mean}\u{b5}A over {} samples (min {min}, max {max})", + samples.len() + ) +} + +impl AndroidAppState for BenchClient { + fn new(mut ui_state: AndroidUiState, rsc: &mut AndroidRsc) -> Self { + let content = WidgetPtr::new().add(rsc); + let loading = placeholder(rsc, "Loading fixture..."); + content(rsc).set(loading); + + let report_display = wtext("") + .editable(EditMode::MultiLine) + .text_align(Align::LEFT) + .wrap(true) + .size(14) + .color(Color::WHITE) + .attr::(()) + .label("Benchmark report") + .add(rsc); + + let controls = bench_controls(rsc); + let tree = ( + controls, + content.height(rest(2)), + report_display.height(rest(1)).pad(8), + ) + .span(Dir::DOWN) + .add_strong(rsc) + .any(); + ui_state.set_root(tree); + + let mut client = Self { + ui_state, + content, + report_display, + screen: None, + items: Vec::new(), + stream_tail: Vec::new(), + platform: None, + last_report: None, + running: false, + }; + + let (backlog, stream_tail) = parse_fixture(); + client.stream_tail = stream_tail; + match fold_page(&backlog) { + Ok(items) => { + client.items = items; + client.rebuild_transcript(rsc); + } + Err(message) => { + client.show_message(rsc, &format!("Couldn't fold the bench fixture: {message}")) + } + } + client + } + + fn platform_ready(&mut self, _rsc: &mut AndroidRsc, vm: JavaVM, view: GlobalRef) { + self.platform = Some(Arc::new(PlatformHandle::new(vm, view))); + } + + fn back_pressed(&mut self, _rsc: &mut AndroidRsc, _render: &mut UiRenderState) -> bool { + false + } +} + +type Rsc = AndroidRsc; + +fn bench_controls(rsc: &mut Rsc) -> WeakWidget { + let run_rect = rect(Color::rgb(40, 70, 40)) + .on( + CursorSense::click(), + |ctx: EventIdCtx<'_, Rsc, _, _>, rsc: &mut Rsc| { + ctx.state.start_benchmark(rsc); + }, + ) + .label("Run benchmark"); + let run = ( + run_rect, + wtext("Run benchmark").size(18).text_align(Align::CENTER), + ) + .stack() + .pad(8) + .add(rsc); + + let copy_rect = rect(Color::rgb(50, 50, 60)) + .on( + CursorSense::click(), + |ctx: EventIdCtx<'_, Rsc, _, _>, _rsc: &mut Rsc| { + ctx.state.copy_report(); + }, + ) + .label("Copy report"); + let copy = ( + copy_rect, + wtext("Copy report").size(18).text_align(Align::CENTER), + ) + .stack() + .pad(8) + .add(rsc); + + (run, copy).span(Dir::RIGHT).height(56).add(rsc) +} + +impl BenchClient { + fn show_message(&mut self, rsc: &mut Rsc, message: &str) { + let widget = placeholder(rsc, message); + (self.content)(rsc).set(widget); + self.screen = None; + } + + fn rebuild_transcript(&mut self, rsc: &mut Rsc) { + let rows = group_tool_runs(&self.items); + let (screen, tree) = transcript_ui::build_tree(rsc, rows); + (self.content)(rsc).set(tree); + self.screen = Some(screen); + } + + fn copy_report(&mut self) { + let Some(report) = &self.last_report else { + log::info!("iris bench report: nothing to copy -- run the benchmark first"); + return; + }; + let Some(platform) = &self.platform else { + log::info!("iris bench report: no platform handle, can't reach the clipboard"); + return; + }; + if platform.copy_to_clipboard("iris bench report", report) { + log::info!("iris bench report: copied to clipboard"); + } else { + log::info!("iris bench report: clipboard copy failed"); + } + } + + /// P0's scripted run: `BenchRun.kt`'s scroll loop, then its streaming + /// phase, then the report -- run in-process for the same reason that + /// file's own doc gives (no usable system tracing on a real phone, no + /// agent that can drive one). + fn start_benchmark(&mut self, rsc: &mut Rsc) { + if self.running { + log::info!("iris bench report: already running"); + return; + } + self.running = true; + self.android_state_mut().frame_report.reset(); + self.report_display.edit(rsc).set("Running benchmark..."); + + let redraw = rsc.tasks.redraw_handle(); + let platform = self.platform.clone(); + let stream_tail = self.stream_tail.clone(); + let cpu_start = process_cpu_ms(); + + rsc.spawn_task(async move |mut ctx| { + // The swipe loop: two drags toward newer content, two back -- + // a cycle returns to where it started, so the whole loop + // measures steady-state scrolling. `BenchRun.kt`'s own + // comment on this shape. + for _ in 0..CYCLES { + for delta in [SWIPE_PX, SWIPE_PX, -SWIPE_PX, -SWIPE_PX] { + animate_scroll(&mut ctx, &redraw, delta, SWIPE_MS).await; + tokio::time::sleep(Duration::from_millis(SWIPE_PAUSE_MS)).await; + } + } + + // Pinned to the newest end before streaming starts, matching + // `stream-bench.sh`'s "Jump to latest" tap. + ctx.update(|state: &mut BenchClient, rsc| { + if let Some(screen) = &state.screen { + (screen.list)(rsc).jump_to_end(); + } + }); + redraw.request_redraw(); + + // The battery sampler runs concurrently with the streaming + // phase, once a second, the same cadence `BatterySampler` uses + // on the Compose side -- via its own JNI-attached thread, not + // `ctx.update`, since a sample needs no widget-tree access. + let sampler_done = Arc::new(AtomicBool::new(false)); + let samples = Arc::new(std::sync::Mutex::new(Vec::::new())); + let sampler = platform.clone().map(|platform| { + let done = sampler_done.clone(); + let samples = samples.clone(); + tokio::spawn(async move { + while !done.load(Ordering::Relaxed) { + if let Some(value) = platform.battery_current_ua() { + samples.lock().unwrap().push(value); + } + tokio::time::sleep(Duration::from_secs(1)).await; + } + }) + }); + + let total = (STREAM_EVENTS_PER_SEC * STREAM_SECONDS) as usize; + let mut sent = 0usize; + for event in stream_tail.into_iter().take(total) { + ctx.update(move |state: &mut BenchClient, rsc| { + state.items = fold_event(&state.items, &event); + state.rebuild_transcript(rsc); + }); + redraw.request_redraw(); + sent += 1; + tokio::time::sleep(Duration::from_millis(1000 / STREAM_EVENTS_PER_SEC)).await; + } + // Lets the last few deltas land and draw before the report is + // read -- `BenchRun.kt`'s own closing delay. + tokio::time::sleep(Duration::from_millis(300)).await; + + sampler_done.store(true, Ordering::Relaxed); + if let Some(sampler) = sampler { + let _ = sampler.await; + } + let battery = battery_line(&samples.lock().unwrap()); + let cpu_line = match (cpu_start, process_cpu_ms()) { + (Some(start), Some(end)) => { + format!(" process CPU time over this run: {}ms", end.saturating_sub(start)) + } + _ => " process CPU time over this run: unavailable".to_string(), + }; + let rss_line = match peak_rss_kb() { + Some(kb) => format!(" peak RSS: {kb}kB"), + None => " peak RSS: unavailable (/proc/self/status unreadable)".to_string(), + }; + + ctx.update(move |state: &mut BenchClient, rsc| { + state.running = false; + let scroll_line = format!( + " scroll: {CYCLES} cycles ({} swipes), streamed {sent}/{total} fixture events", + CYCLES * 4 + ); + let frames_line = match state.android_state().frame_report.report() { + Some(stats) => format!("{stats}"), + None => "no frames recorded".to_string(), + }; + let report = format!( + "iris bench report\n{frames_line}\n{scroll_line}\n{cpu_line}\n{rss_line}\n{battery}" + ); + log::info!("iris bench report: {report}"); + state.report_display.edit(rsc).set(&report); + state.last_report = Some(report); + }); + redraw.request_redraw(); + }); + } +} + +/// Moves `List::scroll` by `total_px` over `duration_ms`, in ~60Hz steps, +/// so the swipe is many rendered frames rather than one jump -- the same +/// shape `animateScrollBy(SWIPE_PX, tween(SWIPE_MS))` gives on the Compose +/// side, in the one place the two backends have to differ (iris's `List` +/// has no built-in tween, so this drives it by hand). +async fn animate_scroll( + ctx: &mut iris::task::TaskCtx, + redraw: &Arc, + total_px: f32, + duration_ms: u64, +) { + let steps = (duration_ms / ANIM_STEP_MS).max(1); + let step_px = total_px / steps as f32; + for _ in 0..steps { + ctx.update(move |state: &mut BenchClient, rsc| { + if let Some(screen) = &state.screen { + (screen.list)(rsc).scroll(step_px); + } + }); + redraw.request_redraw(); + tokio::time::sleep(Duration::from_millis(ANIM_STEP_MS)).await; + } +} diff --git a/iris/android-app/src/bench_jni.rs b/iris/android-app/src/bench_jni.rs new file mode 100644 index 0000000..0a5ae4b --- /dev/null +++ b/iris/android-app/src/bench_jni.rs @@ -0,0 +1,134 @@ +//! JNI calls the `bench` feature needs that go through the shell's own +//! Java side rather than anything `iris`/`android-view` already wraps: +//! `BatteryManager.getIntProperty(BATTERY_PROPERTY_CURRENT_NOW)` for the +//! per-second battery sample, and `ClipboardManager.setPrimaryClip` for +//! the "Copy report" control (P0's iris half, docs/RUST.md). Neither is +//! part of `android_view::context`'s own `Context`/`Resources` wrappers +//! (that file's own `// TODO: more methods?`), so this calls them +//! directly rather than growing that crate's wrapper for two one-off +//! calls this crate alone needs. +//! +//! Holds its own `JavaVM` + `GlobalRef` to the view (handed in through +//! [`iris::android::AndroidAppState::platform_ready`]) so it can attach +//! whichever thread calls it -- the battery sampler runs on a background +//! tokio task, not the UI thread the rest of `IrisViewPeer`'s JNI calls +//! run on. `JavaVM::attach_current_thread` is safe to call from a thread +//! already attached (the `jni` crate detects it and does not double +//! attach), so no caller here needs to know or care which thread it is. + +use android_view::jni::{ + JNIEnv, JavaVM, + objects::{GlobalRef, JObject, JValue}, +}; + +/// `android.os.BatteryManager.BATTERY_PROPERTY_CURRENT_NOW` -- not exposed +/// as a constant anywhere reachable without the Android SDK jar, so named +/// here with its source rather than left as a bare `2`. +const BATTERY_PROPERTY_CURRENT_NOW: i32 = 2; + +pub struct PlatformHandle { + vm: JavaVM, + view: GlobalRef, +} + +impl PlatformHandle { + pub fn new(vm: JavaVM, view: GlobalRef) -> Self { + Self { vm, view } + } + + fn context<'e>(&self, env: &mut JNIEnv<'e>) -> Option> { + env.call_method( + self.view.as_obj(), + "getContext", + "()Landroid/content/Context;", + &[], + ) + .ok()? + .l() + .ok() + } + + fn system_service<'e>( + &self, + env: &mut JNIEnv<'e>, + context: &JObject<'e>, + name: &str, + ) -> Option> { + let jname = env.new_string(name).ok()?; + env.call_method( + context, + "getSystemService", + "(Ljava/lang/String;)Ljava/lang/Object;", + &[JValue::Object(jname.as_ref())], + ) + .ok()? + .l() + .ok() + } + + /// One sample of `BATTERY_PROPERTY_CURRENT_NOW`, in microamps. `None` + /// on any JNI failure, on a device with no `BatteryManager` service, + /// or when the platform itself answers "not supported" -- `0` or + /// `Integer.MIN_VALUE` are both documented SDK answers for that, and + /// both would read as a real (and wrong) measurement if folded into an + /// average rather than named apart. UI_RULES.md: never present an + /// inferred value as a measured one. + pub fn battery_current_ua(&self) -> Option { + let mut guard = self.vm.attach_current_thread().ok()?; + let env: &mut JNIEnv = &mut guard; + let context = self.context(env)?; + let battery_manager = self.system_service(env, &context, "batterymanager")?; + let value = env + .call_method( + &battery_manager, + "getIntProperty", + "(I)I", + &[JValue::Int(BATTERY_PROPERTY_CURRENT_NOW)], + ) + .ok()? + .i() + .ok()?; + if value == 0 || value == i32::MIN { + None + } else { + Some(value) + } + } + + /// Puts `text` on the system clipboard through `ClipboardManager` -- + /// `true` only if the whole JNI chain (service lookup, `ClipData`, + /// `setPrimaryClip`) succeeded. + pub fn copy_to_clipboard(&self, label: &str, text: &str) -> bool { + self.try_copy_to_clipboard(label, text).is_some() + } + + fn try_copy_to_clipboard(&self, label: &str, text: &str) -> Option<()> { + let mut guard = self.vm.attach_current_thread().ok()?; + let env: &mut JNIEnv = &mut guard; + let context = self.context(env)?; + let clipboard = self.system_service(env, &context, "clipboard")?; + let jlabel = env.new_string(label).ok()?; + let jtext = env.new_string(text).ok()?; + let clip = env + .call_static_method( + "android/content/ClipData", + "newPlainText", + "(Ljava/lang/CharSequence;Ljava/lang/CharSequence;)Landroid/content/ClipData;", + &[ + JValue::Object(jlabel.as_ref()), + JValue::Object(jtext.as_ref()), + ], + ) + .ok()? + .l() + .ok()?; + env.call_method( + &clipboard, + "setPrimaryClip", + "(Landroid/content/ClipData;)V", + &[JValue::Object(&clip)], + ) + .ok()?; + Some(()) + } +} diff --git a/iris/android-app/src/lib.rs b/iris/android-app/src/lib.rs index 6a8cc32..8825688 100644 --- a/iris/android-app/src/lib.rs +++ b/iris/android-app/src/lib.rs @@ -23,6 +23,18 @@ //! A build picks one screen or the other, never both, so `Client` and //! `TranscriptClient` are cfg-gated apart rather than switched at runtime -- //! there is no in-app navigation to switch *to* on either side yet. +//! +//! **`bench` feature (P0's iris half, docs/RUST.md):** a third +//! `AndroidAppState`, `bench_client::BenchClient`, on the same axis -- +//! `transcript_ui::build_tree` again, this time against the checked-in +//! fixture (`app/bench-fixture/assets/transcript.jsonl`) instead of a real +//! server, with a "Run benchmark" control that drives the same scroll loop +//! and streaming phase the Compose `bench` build type's `BenchRun.kt` +//! does. `bench` depends on `transcript-screen` (Cargo.toml) for +//! `transcript-ui`/`client-core`/`event-model`, so both features end up +//! enabled together -- `ActiveClient` below gives `bench` priority in that +//! case, the same way `transcript-screen` already takes priority over the +//! default `tabs-screen`. use android_view::{ Context, View, @@ -39,7 +51,11 @@ use iris::prelude::*; use log::LevelFilter; use std::ffi::c_void; -#[cfg(feature = "transcript-screen")] +#[cfg(feature = "bench")] +mod bench_client; +#[cfg(feature = "bench")] +mod bench_jni; +#[cfg(all(feature = "transcript-screen", not(feature = "bench")))] mod transcript_client; /// The app's `View` subclass, matching the Java side's package -- @@ -85,8 +101,10 @@ impl AndroidAppState for Client { #[cfg(not(feature = "transcript-screen"))] type ActiveClient = Client; -#[cfg(feature = "transcript-screen")] +#[cfg(all(feature = "transcript-screen", not(feature = "bench")))] type ActiveClient = transcript_client::TranscriptClient; +#[cfg(feature = "bench")] +type ActiveClient = bench_client::BenchClient; extern "system" fn new_view_peer<'local>( env: JNIEnv<'local>, From 00767eed4dd7597a0366edd00a5755047a0ca15e Mon Sep 17 00:00:00 2001 From: iris <2+iris@noreply.localhost> Date: Sat, 5 Sep 2026 21:35:51 -0400 Subject: [PATCH 12/12] docs: P0's iris half done -- bench feature, emulator smoke run, APK RUST.md's P0 box gets the iris-half account: the fixture, the scroll/stream mechanism, the report fields, build commands (all clean), packaging (no cargo xtask apk yet, so a new Gradle release build type on top of cargo ndk), and the emulator smoke run's report next to Compose's own. Used a second, differently-named AVD rather than contend with the session already on this checkout's own emulator. DECISIONS.md's P0 entry gets a matching summary bullet. IRIS.md records AndroidAppState::platform_ready. IRIS_TODO.md notes the one gap found: no read-only selectable text primitive, so the bench report's TextEdit picks up a keyboard on tap it has nothing to type into. Co-Authored-By: Claude Fable 5.1 --- docs/DECISIONS.md | 19 +++++ docs/IRIS.md | 23 ++++++ docs/IRIS_TODO.md | 12 ++++ docs/RUST.md | 174 ++++++++++++++++++++++++++++++++++++++++++++++ 4 files changed, 228 insertions(+) diff --git a/docs/DECISIONS.md b/docs/DECISIONS.md index 8e4f056..45a2ade 100644 --- a/docs/DECISIONS.md +++ b/docs/DECISIONS.md @@ -18,6 +18,25 @@ marked **DEFERRED** are ones the agent chose not to decide alone. flagged here because it is the first half of something Iris explicitly asked to see before P1. +- **P0's iris half is also built and smoke-tested on the emulator, + 2026-09-05.** A new `bench` Cargo feature on `iris-android-app`, on top + of `transcript-screen`: the same checked-in fixture (`include_str!`, no + asset pipeline needed), the same 24-swipe scroll loop animated through + `List::scroll` and the same 400-event/20s streaming phase through + `fold_event`, "Run benchmark"/"Copy report" as named accessible + controls, and the same three added report fields (process CPU time, + peak RSS, battery current) via direct JNI calls + (`bench_jni.rs::PlatformHandle`) since `android_view` has no + `BatteryManager`/`ClipboardManager` wrapper of its own. One small public + API addition to get there: `AndroidAppState::platform_ready` (`IRIS.md`), + a default-no-op lifecycle hook handing an implementor a `JavaVM` + + `GlobalRef` it can call Java through from any thread. Packaged with a + new `release` build type on `iris-android-app`'s own Gradle project + (there was previously only `debug`), signed with the same key + `app/build-apk.sh` generates. Smoke run and the full report are in + RUST.md's P0 box; not attempted this pass: the real on-phone runs and + Iris's pass/fail call, which is the actual gate. + - **The intermittent touch-scroll dropout is root-caused and fixed: a missed `ACTION_DOWN` hit-test, not the previously-suspected coalesced first `ACTION_MOVE`.** Diagnosed by temporary logcat tracing of every diff --git a/docs/IRIS.md b/docs/IRIS.md index 94bac63..a4fd1eb 100644 --- a/docs/IRIS.md +++ b/docs/IRIS.md @@ -8,6 +8,29 @@ capability that moved. Small and trivial changes do not go here. An entry gives the date, what changed, why, and a short before/after where it helps judge the change without the session that made it. Newest first. +## 2026-09-05: `AndroidAppState::platform_ready` (RUST.md's P0 box, iris half) + +Added a second, optional lifecycle method to `iris::android::AndroidAppState` +(`iris/src/android/view.rs`), called once from `new_peer` right after `new`: + +```rust +fn platform_ready(&mut self, rsc: &mut AndroidRsc, vm: JavaVM, view: GlobalRef) {} +``` + +Default does nothing, so every existing implementor (`Client`, +`TranscriptClient`) is unaffected. It exists for a caller that needs to call +into Java itself beyond what a `RequestRedraw` handle already covers -- +P0's bench build (`iris-android-app`'s new `bench` feature, +`bench_client.rs`/`bench_jni.rs`) uses it to hold a `JavaVM` + `GlobalRef` +to the view so its "Copy report" control and once-a-second battery sampler +can call `BatteryManager`/`ClipboardManager` through the view's own +`Context`, from a background tokio task as well as the UI thread. `new` +itself was not extended with these two parameters: most implementors need +nothing here, and `new`'s job is building the widget tree, not holding a +platform handle. `vm`/`view` are independent handles from the ones +`new_peer` keeps for its own `RequestRedraw` (a fresh `get_java_vm`/ +`new_global_ref` each), so storing them has no effect on that mechanism. + ## 2026-09-05 (later the same day): `iris_core::FrameReport` (RUST.md's I5 box) New public type, `iris_core::FrameReport` (re-exported from `iris_core`'s diff --git a/docs/IRIS_TODO.md b/docs/IRIS_TODO.md index d4edcc6..13222be 100644 --- a/docs/IRIS_TODO.md +++ b/docs/IRIS_TODO.md @@ -82,6 +82,18 @@ order and what "done" looks like. Tick and date them in place. were exactly the same root cause measured two different ways. Frame 2 now reports 0 (see the numbers above); not a separate fix. +- [ ] **A read-only text display has no widget of its own — P0's bench + report area is a `TextEdit` standing in for one (2026-09-05).** The only + way to get selectable text on screen today is `.editable(...)` plus + `.attr::(())` (`Selectable` is only implemented for + `TextEdit`, `iris/src/attr.rs`), which also makes the field focusable — + tapping the bench report opens the soft keyboard over text nothing lets + you type into. Harmless for a bench-only debug screen (not fixed this + pass), but a real "selectable, not editable" text primitive would + remove the keyboard side effect and is worth having before another + screen wants the same thing (P1's own transcript rows already read + their content from a `TextEdit` for the same reason). + ## Build - [x] **Benchmarks**, not unit tests, run on demand (2026-09-05; a diff --git a/docs/RUST.md b/docs/RUST.md index 089a19a..9839e7d 100644 --- a/docs/RUST.md +++ b/docs/RUST.md @@ -3487,6 +3487,180 @@ device. this session was told not to touch `iris/`), and anything past the emulator — the actual on-phone runs and Iris's pass/fail call. + **iris half: done, 2026-09-05.** A `bench` Cargo feature on + `iris-android-app`, built on top of `transcript-screen` + (`bench = ["transcript-screen", "dep:libc", "dep:tokio"]`, + `iris/android-app/Cargo.toml`), gives `lib.rs`'s `ActiveClient` + priority a third `AndroidAppState` (`bench_client::BenchClient`) + over `TranscriptClient` when both features are listed together -- + matching the exact build command below, which lists both. + + **Fixture.** `include_str!("../../../app/bench-fixture/assets/ + transcript.jsonl")` (1,915,760 bytes) at compile time -- no asset + pipeline needed the way the Compose half's Gradle source set does. + `bench_client::parse_fixture` splits the same way `BenchFixture.kt` + does: the first 3,200 non-blank lines parsed as `serde_json::Value`s + and folded once through `client_core::transcript_fold::fold_page` + (the real fold a `/transcript` page goes through), the rest parsed + as `event_model::SeqEvent`s and held back as the streaming tail. + `build.rs` (transcript-screen's own) now exits early under `bench` + before requiring a live server's host/port/token/CA -- `BenchClient` + never calls `build_transport()`, so that requirement made no sense + for a build that talks to nothing. + + **"Run benchmark" (`.label("Run benchmark")`) and "Copy report" + (`.label("Copy report")`)** sit in a fixed bar above the transcript; + a selectable `TextEdit` (`.attr::(())`, the same + attribute the composer field uses) below it shows the report text. + Pressing "Run benchmark" resets `FrameReport`, then drives + `List::scroll` in ~60Hz steps (`ANIM_STEP_MS = 16`) to animate each + 900px/200ms swipe rather than jumping it -- iris's `List` has no + built-in tween the way `animateScrollBy(tween(...))` gives Compose, + so this is the one place the two backends' bench code has to differ + in shape rather than only in numbers -- through the same + `rsc.tasks.redraw_handle()` + manual `request_redraw()` per step + `transcript_client.rs` already established (a `Tasks::spawn`d + future's *automatic* redraw fires once, after the whole future + completes, which would show nothing moving until the run ends). + After the scroll loop, `List::jump_to_end()` pins to the newest + content (matching `stream-bench.sh`'s "Jump to latest" tap), then + 400 fixture events replay at 20/s through `fold_event` -- the same + fold path a live SSE frame takes in `transcript_client.rs`'s own + `apply_event` -- each one triggering `rebuild_transcript`'s full + `transcript_ui::build_tree` rebuild, same tradeoff as + `TranscriptClient`/`desktop-app`. A battery sampler runs + concurrently on its own `tokio::spawn`d task (not through + `ctx.update`, since a JNI battery read needs no widget-tree access), + attaching whichever thread it runs on via a stored `JavaVM` -- + `AndroidAppState::platform_ready` (new, `IRIS.md`) is what hands + `bench_client.rs` that `JavaVM` + a `GlobalRef` to the view, since + neither was reachable from `AndroidAppState::new` before this box. + + **Report fields.** `FrameStats`'s existing `Display` (frames, janky + %, p50/p90/p99, worst, and I5's own `cpu_p50`/`gpu_wait_p50` CPU/GPU + split) plus a `bench:`-shaped tail this box added: process CPU time + via `libc::getrusage(RUSAGE_SELF)` (user+system time; chosen over + parsing `/proc/self/stat` by hand to avoid assuming `USER_HZ`), peak + RSS from `/proc/self/status`'s `VmHWM` (same source `BenchRun.kt` + reads), and battery current sampled once a second via + `BatteryManager.getIntProperty(BATTERY_PROPERTY_CURRENT_NOW)` + through direct JNI calls (`bench_jni.rs`'s `PlatformHandle` -- + `android_view::context`'s own `Context`/`Resources` wrappers have no + `getSystemService`, so this calls it directly rather than growing + that crate's wrapper for two one-off calls). `0`/`Integer.MIN_VALUE` + read as "unavailable" rather than folded into the average, matching + `BatterySampler`'s own rule and UI_RULES.md's "never present an + inferred value as a measured one." The report is logged under the + existing `iris-android-app` logcat tag on a line starting `iris + bench report:` (grep-able the same way `transcript_client.rs`'s + "Frame report" control already is), shown in the on-screen + `TextEdit`, and copied to the system clipboard by "Copy report" + through `ClipboardManager.setPrimaryClip` (`bench_jni.rs`, same + `PlatformHandle`). + + **Build commands, all clean this pass:** + - `cargo fmt --all -- --check` (iris workspace) and + `cd iris/android-app && cargo fmt --all -- --check`: clean. + - `cargo clippy --workspace --all-targets` (iris workspace): clean + (only the pre-existing `wgpu`/`winit`/`naga` future-incompat + notice). + - `cargo test --workspace` (iris workspace): 39 + 8 + 10 = the same + pre-existing counts, all passing, unaffected by this box (it + touched no logic under test there beyond `AndroidAppState`'s new + default no-op method). + - `cargo ndk -t x86_64 -P 26 clippy --features "transcript-screen + force-gles bench" --lib -- -D warnings` (`iris/android-app`): + clean. + - `cargo ndk -t arm64-v8a -P 26 -o app/src/main/jniLibs/ build + --release --features "transcript-screen force-gles bench"`: + clean, `arm64-v8a/libmain.so` produced. The pre-existing "unused + dependency `tabs-ui`" Cargo advisory also appears on a plain + `--features transcript-screen` build with no `bench` (confirmed + by building that combination alone with fake env vars) -- not + something this box introduced, and not a clippy/rustc warning + (AGENTS.md's "keep the build clean" gate is `cargo clippy`, which + stays silent on it). + + **Packaging.** No `cargo xtask apk` exists for `iris/android-app` + yet (I2's own Gradle project is the only pipeline), so this reused + that split rather than inventing one: `cargo ndk --release` above + builds the cdylib straight into `app/src/main/jniLibs/`, then a new + `release` build type in `app/build.gradle` (there was previously + only `debug`) packages and signs it -- + `AI_APP_KEYSTORE=~/.config/ai-app/release.jks` + + `AI_APP_KEYSTORE_PASSWORD` (the same key `app/build-apk.sh` + generates for the Compose app) via `gradle :app:assembleRelease`, + with `applicationIdSuffix ".bench"` so it installs beside the plain + tabs demo rather than replacing it. `aapt2 dump badging` on the + result: `package: name='dev.iris.android.demo.bench'`, one native + library, `lib/arm64-v8a/libmain.so`. `apksigner verify + --print-certs` shows the same `CN=ai-app` certificate + `compose-bench-arm64.apk` is signed with. + + **Emulator smoke run, 2026-09-05.** This checkout's own AVD + (`ai-app-2`) was in use by the session recording I5's clean-scroll + comparison in this same file (its Compose app was in the + foreground, confirmed via `dumpsys window`/`dumpsys activity + processes` before touching anything) -- rather than contend for it + (AGENTS.md's "coordinate with peer agents"), a second, + differently-named AVD was created (`AVD_NAME=ai-app-2-bench emu + up`, `pixel_10`/`android-36`/`google_apis`/`x86_64`, cold boot, host + GPU, no `EMU_GPU=software`), with 12GB of the VM's memory still + available after both were up (this-machine-android's "two are + comfortable" guidance). Installed via `adb -s emulator-5556 install + -r`, launched, driven by `ui-trace record -s emulator-5556 --do + "tap 'Run benchmark'"` (the control resolved by its accessibility + label, per AGENTS.md's "no coordinate" rule), then read back over + `adb logcat`: + + iris bench report + frames=372 janky%=56.99 p50=19.5ms p90=219.5ms p99=284.5ms worst=369.3ms (measures redraw-start to after present() is called, not GPU/compositor completion) cpu_p50=0.4ms gpu_wait_p50=13.9ms (redraw-start-to-submit vs. submit-to-after-present) + scroll: 6 cycles (24 swipes), streamed 400/400 fixture events + process CPU time over this run: 24665ms + peak RSS: 224600kB + battery current: mean 900000µA over 21 samples (min 900000, max 900000) + + "Copy report" was pressed immediately after and logged `iris bench + report: copied to clipboard` (`ClipboardManager.setPrimaryClip` + succeeded). No crash (`adb logcat`'s `FATAL`/`AndroidRuntime` lines + checked -- only `ui-trace`'s own runtime, unrelated), process alive + throughout (`dumpsys activity processes`), 400/400 stream events + confirmed sent. + + Read this the same way the Compose half's own box already asks to + read its number: this is software-rasterised (well, GLES-over-virgl + under `force-gles`, per I5's "Where iris's frame time goes") + emulator output, "the harness runs end to end and produces every + field P0 asked for," not a phone number -- and the battery current + is again the emulator's fixed 900000µA mocked charger reporting a + constant, exactly what the Compose box's own run found, not a real + battery answering. `cpu_p50=0.4ms` (iris's own per-frame CPU work) + against a much larger `gpu_wait_p50`/`p50` again matches I5's "Where + iris's frame time goes" finding under real GPU rendering (`-gpu + host`, `force-gles`) -- the frame-time budget here is dominated by + the driver/compositor wait, not by iris's layout or primitive + building, though this run's `janky%`/`p90`/`p99` are considerably + worse than that earlier isolated pass, most likely the cost of this + AVD's very first cold boot plus running two emulators on this VM at + once (a fair comparison against Compose would need both apps run + back-to-back on the same freshly-booted device, not attempted this + pass since the second AVD was torn down immediately after per + AGENTS.md's "stop yours when you are done with it"). + + Copied to `~/host/bench/iris-bench-arm64.apk` (15,445,468 bytes) and + `~/host/bench/README.md`'s "iris" section filled in (install, open, + tap "Run benchmark", read the report from the on-screen text or + logcat, tap "Copy report", paste back). + + **Not done this pass**: the actual on-phone runs and Iris's + pass/fail call between the two reports (P0's own pass condition) -- + that needs Iris's phone, which this session has no access to. + `iris/src/android/view.rs` was touched (`AndroidAppState:: + platform_ready`, `new_peer`'s wiring) -- confirmed to not be one of + the three files the concurrent `device_limits()` work on this + branch was using (`iris/core/src/render/mod.rs`, + `iris/src/android/render.rs`, `iris/src/default/render.rs`). + - [ ] **P1 — session screen parity.** History paging backward (with the page-boundary healing `client-core` does not have yet, below), `TranscriptSource`-backed cache/server stitching, jump-to-latest,