Files
ai-app/PLAN.md
T
irisandClaude Opus 5 127b25e60a Meter a session by its provider, and let llama.cpp run over ssh
The rate-limit bar answered a question about an account, and picked the
answer by machine. One machine runs echo, the Claude CLI and a local
model side by side, so every echo session on it drew the CLI's five-hour
window: a quota that session cannot spend and could never run down. A
session now names its meter (`usageProvider`, from
`DriverKind::usage_provider`, which `usage::providers_for` reads too so
the two lists cannot disagree), and the phone matches on machine *and*
provider. Nothing meters echo or llama, and nothing at all is drawn --
including while the first fetch is out, since "checking" under a session
that turns out to meter nothing is a row the screen then withdraws.

Echo gets a meter it can be *told* about instead: `/usage 42`,
`/usage 95 20`, `/usage 42 never`, `/usage notloggedin`,
`/usage unreachable`, `/usage failed`, `/usage off`. Those states cost
real quota to arrange, which is why none of them had been looked at.

And llama.cpp runs wherever a setup says, which was the last of phase 5.
`Transport::reserve_port` is the second half of what a transport is --
"run this" plus "reach this port" -- returning the port the server binds
there and the port that reaches it here, and `Launch::reaching` puts the
`-L` tunnel on the connection that already carries the command. Three
things that came out of building it:

- A forwarded launch gets a pty and every other one keeps `-T`. Killing
  the ssh client ends a CLI by closing the stdin it reads; llama-server
  never reads its stdin, so the same kill left it running on the far
  machine with the model loaded -- one orphan per stopped session.
- The model is looked for on the machine that will serve it, at that
  machine's own models directory, so `GET /setups/{id}/models` is what
  the spawn screen offers rather than the backend's own downloads.
- The readiness poll watches the process, not only the port: a model
  that will not load exits in a second and would otherwise have been
  reported as "gave up after 300s". The failure carries the log's tail.

Exercised end to end against this VM over ssh to itself: spawn, load,
answer, outlive a backend restart, be adopted, answer again, and stop --
with both the ssh client and the far llama-server gone afterwards. The
local path, the Claude bar and the spawn screen checked on the emulator.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-04 17:45:32 -04:00

1251 lines
74 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# ai-app — plan
A phone interface to AI coding sessions — Claude Code and llama.cpp for now —
built to replace the Claude app for day-to-day use. Two motivations: local
models need a front end at all, and owning the client means fixing the things
the official app gets wrong (e.g. it won't deliver a typed message until the
session fully finishes its turn, where the TUI injects it at the next tool
boundary).
Same shape as `../dev-updater`: a Rust (Axum) backend on the desktop, a
Kotlin/Compose Android app, pinned self-signed TLS between them.
## The one idea everything hangs off
Both session types are **a child process speaking JSONL over stdio**:
- Claude Code: `claude -p --input-format stream-json --output-format stream-json`
— bidirectional streaming JSON. User messages sent while a turn is running
are injected at the next opportunity (the TUI behavior we want), a control
protocol carries interrupts and permission requests, `--resume <id>` picks a
session back up after a backend restart.
- llama.cpp: **pi in RPC mode** (`pi --mode rpc`), pointed at a llama-server
endpoint. Same deal: JSONL on stdio, `prompt` (with images), `steer` for
mid-run injection, `abort`, `set_model`, `compact` / `set_auto_compaction`,
session files that survive restarts, structured events for streaming text
and tool executions.
So the backend has one abstraction — spawn a process, translate its dialect to
a common event stream, keep an append-only transcript — and two translators.
SSH support falls out of the same shape: a remote session is the identical
command run as `ssh <host> <cmd>`; stdio doesn't care.
Decisions already made (2026-08-24):
- llama.cpp harness: **pi RPC now**, with the session abstraction kept clean
enough that a custom Rust agent loop can be added as a third driver later.
- The backend **manages llama-server itself** (start with a chosen GGUF, stop,
swap models), locally and over SSH.
- Claude permission prompts are **interactive in the app**, with a per-session
permission mode chosen at spawn.
- **One backend** on the main machine; the phone talks only to it, and it
reaches other hosts via SSH. Remote hosts need the CLIs installed but no
backend.
## Architecture
### Setups and providers (decided and built 2026-08-28, superseding the below)
**A setup is a machine, and it carries the providers that machine has.**
Optional ssh details, plus the list of what can be run there. Spawning is
then two choices in order: pick a setup, then pick one of its providers.
This replaces the independent providers × hosts model recorded below,
which is what the code does today. What went wrong with it: the two axes
are not actually independent. A provider is only real on a machine where
that CLI is installed, so a free cross-product offers combinations that
cannot work — `claude-cli` on a machine with no `claude`, and every
provider paired with a host the driver ignores entirely (`EchoDriver`
takes no host, so "Run on" is a control that silently does nothing for
it). Grouping providers under the machine they exist on makes the picker
show only what is true.
Settled while building it:
- **Echo is seeded, not implicit.** It lives in the setup with no ssh,
because it runs in-process and has no transport to cross. It is written
into `config.ron` on first run rather than conjured at read time — a
provider nobody can see in the file is one nobody can edit from the
phone, which is the opposite of what this app is for.
- **Migrated once, then the migration was deleted** (2026-08-28). Unknown
fields default away, so a `providers:`/`hosts:` file would have loaded as
an empty config and then been seeded over, losing everything silently.
The first answer to that was to *refuse* such a file, which was the wrong
trade and proved it: this process is how a phone reaches the backend at
all, so refusing to start stranded the person who would have to fix it,
as a crash loop with nothing reachable to explain it. It was replaced by
a migration that kept the token hashes, backed the old file up, and
rebuilt the rest — which is discoverable now anyway.
That migration has since run on the one host there is, so it is gone
again, per the standing rule that migration code is deleted once the
update carrying it has been received. With one backend and one phone,
nothing is left on the old shape, and a second parsing path nothing
exercises only constrains later changes to the schema. A file in the old
shape now fails to parse, which is correct because no such file exists.
- **Still to do: editing setups from the phone.** `GET /setups` exists;
writing them is not built, so a new machine is still a hand edit on the
backend. That is the remaining gap against the standing preference that
configuration be reachable from the app. Key material is the honest
exception — a setup names an identity file that must already exist on
the backend machine, because a private key must not travel.
The superseded model, for the reasoning it recorded:
Two independent axes, configured separately and chosen per session:
- A **provider** is *what* runs: a driver kind, the command to invoke, and
the models worth offering. `claude-cli` is the first — named for the CLI
specifically, since bare "claude" would suggest the credit-billed API,
which this is not. llama.cpp becomes a second provider later.
- A **host** is *where* it runs: an ssh target. Absent means the backend
machine itself.
Sessions name both. Keeping them independent is what the motivating setup
requires: the backend runs on the machine the phone can reach (where
WireGuard terminates), which is not necessarily where a CLI is installed —
here the Claude CLI lives only in a VM on that machine, while llama.cpp
will be on the host itself. Pinning a host into a provider would make "the
Claude CLI" and "the Claude CLI over there" two things to configure and
choose between, and would stop the same provider from being sent somewhere
else for one session.
```
Android app (Compose)
│ HTTPS (pinned CA) — REST for actions, SSE for live events
backend (Rust/Axum, desktop)
├─ SessionManager ── Session ── Driver (trait)
│ ├─ ClaudeDriver (claude stream-json)
│ └─ PiDriver (pi --mode rpc)
│ each driver's process is spawned locally or as `ssh host …`,
│ decided per session by the host it names
├─ LlamaServerManager (llama-server lifecycle, local + SSH)
├─ UsageMonitor (Anthropic OAuth usage endpoint)
└─ config.ron + per-session transcript files
```
### Backend layout (`server/`)
Mirroring dev-updater's stack: axum 0.8, axum-server + rustls, tokio, serde,
clap, tracing. Rust edition 2024, warning-clean, clippy in CI habit.
- `main.rs` — bootstrap, TLS listener.
- `routes.rs` — the whole HTTP table in one module doc comment (as in
dev-updater).
- `session/mod.rs``SessionManager`: the live session registry, every
mutation funnels through it (the `registry.rs` pattern: in-memory and
on-disk state can't come apart).
- `session/driver.rs` — the `Driver` trait and the common event model.
- `session/claude.rs`, `session/pi.rs` — the two translators.
- `session/transcript.rs` — append-only JSONL event log per session, with
monotonically increasing sequence numbers (the phone's resume cursor).
- `llama.rs``LlamaServerManager`.
- `ssh.rs` — the ssh command builder (host configs ended up in `config.rs`
with the rest of the schema, so this module is only the wrapping; named
for what it does rather than `hosts.rs` as first sketched).
- `usage.rs` — Anthropic usage polling.
- `config.rs` — persisted schema.
- `certs.rs` — the TLS certificates, generated in process on first start
(added 2026-08-25, replacing a `gen-dev-cert.sh` that shelled out to
openssl).
- `private.rs` — creating files and directories owner-only. One module
owns the modes so "nothing this server writes is readable by anyone
else" is checkable in one place instead of re-argued at each `create`
(added 2026-08-25; config, certs, and session dirs had three copies).
- `media.rs` — the image media-type/extension table, shared by the four
places that have to agree on it: storing an upload, serving it back,
handing one to a driver's dialect, and saving one a tool produced.
`session/pi.rs` and `llama.rs` are phase 4 and not built yet; everything
else above exists.
### The common event model
Driver output, whatever the dialect, is normalized into one event enum before
it touches the transcript or the phone:
- `UserMessage { text }` — what the user sent, echoed into the transcript
by the manager (not by drivers) so every device renders the conversation
from the one stream. (Added 2026-08-24 during phase 1: without it,
reconnects and second devices would lose the user's side.)
- `AssistantText { delta }` — streaming text (rendered as markdown).
- `ToolStart / ToolUpdate / ToolEnd { tool, input, output }` — the "view tools
it's running" screen is just these.
- `Image { ref }` — images in output (screenshots from tools, etc.) are saved
under the session dir and referenced by id; the phone fetches them by URL.
- `Question { id, prompt, options }` — anything the session needs a human for:
Claude's AskUserQuestion, and **permission requests** (canUseTool) are the
same shape with approve/deny options. Answered via one endpoint.
- `Answered { id, answer }` — the manager's record of a question being
answered, so a rendered question card resolves on every connected device,
not just the one that answered (added 2026-08-24, same reasoning as
`UserMessage`).
- `Status { state }` — idle / running / awaiting-input / compacting / exited.
- `UsageDelta { tokens, context }` — what a turn cost, and how much the
model was holding when it ended, where the dialect reports them (both do).
`context` is prompt plus both cache figures, taken from the **last
assistant message** rather than the turn's `result`: measured 2026-08-30
against CLI 2.1.237, the result adds a turn's messages up, so its cache
read of 40,211 was the same conversation counted twice and no size the
model ever held. It is carried rather than summed by readers because it
goes *down* — a compaction replaces it with what the compaction reports,
and a clear leaves it unmeasured. `driver::context_after` is that rule,
and the phone folds with the same one (2026-08-30: this replaced a running
spend total, which could only climb and so kept reporting a context a
compaction or a clear had already taken away).
A session the server has no measurement of asks the CLI's own file
instead of waiting for a turn — `import::context_of`, the same three
fields the import list reads, in the background at load so a start never
waits on an ssh. A clear needs no special case: it gives the CLI a new
session id, so the lookup lands on a file with no usage in it and
answers "unknown", which is true.
- `Error { message }`.
Every event is appended to the session's transcript file with a sequence
number, then fanned out to any connected SSE subscribers. The phone renders
purely from this stream: reconnecting means "give me events after seq N" —
no separate "load history" path to drift from the live one.
Inbound, the driver trait is small:
```rust
trait Driver {
fn send_user_message(&self, text: String, images: Vec<ImageRef>);
fn answer_question(&self, id: QuestionId, answer: Answer);
fn interrupt(&self); // stop mid-run, session survives
fn set_model(&self, model: &str);
fn compact(&self); // pi: native; claude: /compact
fn shutdown(&self); // graceful process exit
}
```
`send_user_message` during a run is the point of the whole app: both dialects
queue it for injection at the next tool boundary rather than the end of the
turn. Claude's dialect: a `user` message on stdin mid-stream; pi's: `steer`.
### Claude driver specifics
- Spawn: `claude -p --verbose --input-format stream-json --output-format
stream-json --permission-mode <mode>` in the chosen working directory, plus
`--model` at spawn. Permission mode (default/plan/acceptEdits/
bypassPermissions) is chosen on the spawn screen.
- Interactive permissions: run with the stream-json control protocol's
permission request flow (the same mechanism the Agent SDK's `canUseTool`
uses) so tool approvals arrive as control requests, become `Question`
events, and our answer goes back as the control response. **Verify the
exact control-request wire format against the current CLI early in
implementation** — it's the least-documented part of this plan.
- Interrupt: control-protocol interrupt request.
- Model change mid-session: try the control protocol's set-model; if the
installed CLI doesn't support it, fall back to `shutdown` + respawn with
`--resume <session_id> --model <new>` — cheap, since Claude persists
sessions in `~/.claude/projects` anyway. That resume path is the recovery
story for a process that has genuinely died; a backend restart never takes
it — it adopts the process that is still there, and starts nothing for the
session that has none (see below).
**Resuming is only ever safe when nothing else has that session open.**
- Images in: base64 image content blocks in the stream-json user message.
- Working directory, host, and model are spawn-screen fields.
### Moving a session to another directory (decided 2026-08-31)
`POST /sessions/{id}/cwd {cwd}`, behind a field in the session settings
dialog. The directory is settled when the process is spawned -- the CLI is
launched with it as its cwd and there is no control request that changes one
-- so this records the new one and **ends** the process that is in the old
one. It does not start a replacement: a session with no process starts on
the next thing said to it or on Start, which is this app's rule for that
everywhere else, and "usually restarts" would be a worse control than
"always stops" (starting one here would have to wait for the recorded status
to catch up with a process already gone).
The path is checked against the session's own machine and **refused** if it
is not there, rather than corrected. The spawn path corrects instead,
because it is resuming a directory the *machine* recorded and that can be
gone through nobody's fault; a path somebody has just typed is different,
and a mistyped one accepted here would surface much later as a session that
would not start, with nothing pointing at the typo.
**Nothing of Claude Code's own is moved**, and that is a measurement rather
than an omission. Checked against CLI 2.1.237 on 2026-08-31: `claude
--resume <id>` finds a session from any working directory — an id that does
not exist answers "No conversation found with session ID", and a real one
resumed from an unrelated directory did not. So the conversation continues
in the new place with nothing relocated, and the session file stays under
the project directory the CLI made for it, which is where the CLI itself
looks. Relocating it would mean reproducing a rule this app cannot see the
whole of: the CLI's project directory is the path with every non-alphanumeric
character replaced by `-`, truncated at 200 characters with a hash of its own
appended, and an override can replace the name entirely.
While fixing this: `SessionInfo.cwd` came from the snapshot a session
launched with, so a moved session reported its *old* directory for as long
as the process lived. It is read from the config where the row is built now,
the same way `setup_name` already was and for the same reason.
### A message from another agent, on a live session (measured 2026-08-31)
Peer messages were only ever produced by the *import* path, reading them out
of the CLI's own session file — so a message another agent sent a session
this server was running never appeared at all, and the session simply
started working on something nobody on the phone had asked for.
Measured rather than guessed, by sending a real cross-session message to a
real `--input-format stream-json` session on CLI 2.1.237: the CLI emits **no
`user` record** for it, and nothing in the partial-message stream mentions
it. The whole of it arrives as an `origin` object on the turn's `result`, in
the same shape the session file records — `kind: "peer"`, the sending
session's `name`, and the message as `body` — so `import::peer_message` reads
both, and there is one function for one wire format. Only peer-caused turns
carry it: four ordinary results on a real session's stdout had no `origin`
between them.
**The cost was the position, and it is paid on the wire rather than on
screen** (2026-09-01). The event cannot be *recorded* in place: at no earlier
point in the turn does the CLI say why the turn started, and the transcript
is append-only, so by the time anyone knows, everything the message caused
has already been written above it. Reading it out of the CLI's own session
file instead — a second reader tailing the one record stdout does not carry —
was rejected then and stays rejected: two sources of truth for one
conversation and a poll per live session.
So the event carries **where it belongs** instead. `PeerMessage` has a
`turnStart`: the seq of the `Status` that opened the turn it started, stamped
by the pump, which is the only thing that knows a seq and the only thing that
sees every driver's turns. The phone gives the note that seq, so it sorts
into the transcript above the turn rather than being drawn out of order at
the end. That seq belongs to a status change, and a status draws no row, so
there is nothing for the note to collide with and the list stays sorted —
which is what the scroll anchor and paging depend on.
`turnStart` is absent where there is nothing to correct: a message replayed
out of a session file by `import` is already in the right place, and one that
opened no turn has no turn to sit above. Both are drawn where they arrive.
The echo driver models both shapes — `/peer` for the in-place one, and
`/peer-turn` for the live one, which reveals the note only after a reply and
a run of tool calls.
### Taking a queued message back (decided 2026-08-31)
A message sent into a running turn is drawn as a bubble waiting below the
transcript, and tapping it asks the server to drop it before the session
reads it — `POST /sessions/{id}/unqueue {messageId}`, answered by
`Driver::unqueue` and recorded as `Event::MessageDropped` so that every
device watching loses the bubble and a reconnect does not replay it back.
The answer has **three** states rather than a yes/no, and that is the whole
of the design: `Dropped`, `AlreadySent`, and `Unknown`. The reason is that
the Claude driver can only ever give the middle one. It writes a steer into
the CLI's stdin the instant it arrives — that is what makes a steer reach
the model at the next tool boundary instead of at the end of the turn, and
it was measured (see `Queue`'s doc comment) — so the line is gone before the
phone could ask for it back. What waits in `awaiting` is the *announcement*,
not the message.
Holding the write until a boundary was considered and rejected on 2026-08-31:
it would make the drop real everywhere, but it costs a steer one model call,
which is the latency the immediate write was introduced to remove. So the
refusal is the honest answer and it is reported where the reader pressed —
on the bubble itself, not in the screen's error row, which is under the
header a screen away. What a tap buys on a Claude session is therefore
knowing that the session has already been told; on a driver that really does
hold a queue (echo today) the message goes.
`Unknown` is not "we could not find out": a driver that is gone reported
everything it was holding when it closed, so there is nothing waiting.
### Session processes outlive the backend (decided 2026-08-29)
A session's process is **left running when the backend stops, and adopted
again when it starts.** Restarting the server — a rebuild, a service
restart, a crash — must not end a turn somebody is waiting on, and a turn
can easily be minutes long.
What this replaces: `shutdown_all` asked every driver to stop, then the
process exited immediately. The SIGKILL escape hatch was a timer inside the
runtime that died with it, so the stop was unreliable; whatever survived was
orphaned with nothing written down to find it by. Processes leaked either
way. The change is that they are now left on purpose and can be picked back
up.
How it works, all inside the session directory beside the transcript:
- `process.json` — the pid, the kernel's **start time** for that pid, and
how much of the output log has been read. The start time is what makes
the pid an identity: pids are reused, and adopting a stranger's would mean
never resuming the real conversation and signalling something unrelated.
- `stdin.fifo` — opened **read-write** and inherited by the process, so it
is its own last writer and never reads EOF when the server goes away.
Closing stdin therefore stops being the graceful-exit signal; ending a
process is a signal now, and only `Driver::stop` does it.
- `stdout.log` / `stderr.log` — plain appended files, read from a byte
offset. A fifo would fill its 64 KB buffer and block the process while
nothing was draining it, which would stall the very turn the leak exists
to protect. Measured: the CLI writes to a file unbuffered, so streaming is
unaffected.
Two consequences worth stating:
- **`--resume` is reachable only when nothing is running.** This is the same
rule as the import refusal below, and for the same reason: two CLIs on one
session file duplicate the conversation into it and bill the second for
re-reading all of it.
- **Remote sessions are adopted too, and the recorded pid is the `ssh`
client's.** This was written down as "local only" and that was wrong about
the code: `start` records a pid whatever the transport, and for a remote
session the process the backend owns *is* the ssh client. Adopting it is
coherent — the fifo still feeds it, its logs still capture the far end's
output, and `ssh` lives exactly as long as the remote command does, so its
liveness is the session's liveness.
The consequence worth knowing: **the remote `claude` always has an sshd
pipe on stdin, under old code and new alike**, because the fifo is on the
backend's side of the connection. So the far process's stdin says nothing
about which version of this server started it.
`Driver` therefore has two ways out rather than one: `detach` (the server is
going away and means to come back) and `stop` (the session is being deleted,
so the process must not survive). Every driver owes exactly one of them.
### Stopping and starting a session's process (decided 2026-08-30)
If a session outlives the backend, the person holding the phone needs the
other direction too: **end the process without ending the session, and start
it again on the same conversation.** `POST /sessions/:id/stop` and
`/start`.
Three decisions worth not undoing:
- **Stop signals the recorded process and says nothing else.** It does not
go through the driver and it does not announce `Exited`. The record is the
session's rather than any dialect's, so signalling it here works for a
session whose driver is in no state to be asked and adds no trait method a
new driver could implement wrongly — and the driver's own reader already
reports the death correctly, draining the last output and recording the
status. Announcing it from here would be a guess arriving ahead of the
measurement, and wrong for the grace period a process that ignores SIGTERM
keeps running.
- **Start replaces the driver and nothing else.** The transcript, the event
pump and the SSE stream every open phone is reading stay where they were,
so starting a session again is not a reconnect for anybody watching, and
there is still exactly one writer of the transcript — which relaunching
the whole `LiveSession` would not be, since the old pump outlives its
session and an imported session's sync task would go on feeding it.
`LiveSession` and `Commands` therefore share one `Mutex<Arc<dyn Driver>>`
rather than each holding a copy.
- **Start is refused unless the session is *known* to have exited.**
`Unknown` means nobody could find out whether the process is alive, and
starting one on that is exactly the two-CLIs-on-one-conversation fault
`session::process` exists to prevent.
That last rule found a real bug in the launch path, which is where the phone
would have hit it: a relaunched session took its status from the transcript,
so one whose process had died before a backend restart reported `Exited`
while the launch it had just gone through was starting a new process —
`Exited` there is not merely stale, it is the word that refuses every command
and invites somebody to start a second process against a live conversation.
**Who says so matters as much as what is said.** The first fix wrote `Idle`
straight into the manager's view, and that produced a second bug on the
phone: the session list reads the manager's status and the session screen
replays the transcript, so a status written in one and not the other is two
screens disagreeing about one session — visible as a stop button that turned
into a play button a moment after the screen opened. So the rule is that
**a driver announces the state it starts in, through the event sink**, which
is what `EchoDriver::new` and `LlamaDriver::attached` already did;
`ClaudeDriver` was the one that started a process silently. It says `Idle`
only when it *started* one — adopting says nothing, because a process that
was already running may be mid-turn and the transcript's last word is the
better answer until its output says otherwise. Coming from the driver also
orders it against the exit `follow` reports, which a status written from the
manager could not be.
**`Exited` is a claim about a process, and the record is what settles it.**
Adopting saying nothing left one word standing that a live process
contradicts. A session whose process was reported gone and then found again
at the next backend start kept `Exited` from the transcript — and `Exited` is
the word that draws a Start button. Start was then accepted every time it was
pressed, and since starting replaces the driver, each press attached *another*
reader to the one process: every line the CLI wrote was translated once per
reader, so three presses put three interleaved copies of one reply on screen.
Two rules come out of it, and neither is optional:
- **`Exited` is checked against `session::process` before it is believed** —
`corrected`, called in `launch` and again in `start_session`. A record that
is not known to be dead makes it false, and what replaces it is `Unknown`:
there is a process, and nothing here has heard from it, which is the answer
`status_of_unlaunched` already gave to the same question. Every other status
is left exactly as it was — those are the pump's, written from what the
process itself said, and none of them authorises starting anything. The
correction goes out through the sink for the reason above: written into the
manager's view alone it would be the list and the screen disagreeing again.
- **A driver that is replaced is detached.** Swapping the `Arc` does not end
the tasks the old one is running. `Driver::detach` is what does — it already
existed for the backend going away — and it is the whole of what a driver
whose process has exited is owed.
**A message or a command starts the process if there isn't one** (decided
2026-08-30). Refusing was work handed back: read the status word, find the
other button, press it, type the thing again. Both plainly mean "do this
now", and `--resume` puts the new process on the same conversation, so
nothing about what was typed changes — only whether there was anything there
to read it. A rename is included, and for a sharper reason than the rest:
Claude Code keeps its own copy of the name, that copy is what its session
picker shows and what other agents read when they list sessions, and a
session is only ever *given* a name at birth, since every later start is a
`--resume`. So a rename that reached no process would leave the two lists
disagreeing permanently, with this app's the only one that had moved — and
the cost of a resume buys the one thing renaming is for. It stays
`rename_session` rather than becoming a command like the others, because the
name is persisted and listed as well as forwarded and that is one operation;
the save happens first, so a failure to start reports that the telling
failed, not the rename. The manager's `send_message` and `run_command` and
the Start button ask one function
(`start_if_exited`) and want opposite answers from it: "there is already a
process" is a refusal worth showing to somebody who pressed Start, and
nothing at all to a message. Deciding it in one place under one write lock is
also what stops two requests arriving together from starting two CLIs. Only
`Exited` starts anything, for the reason above — `Unknown` has a process that
may well be reading its fifo, and what was typed goes to the driver as it
always did.
A command needs one thing a message does not. `Commands::submit` refuses on
`Exited`, and a driver that has just started a process announces `Idle`
through the sink rather than writing it — so a command judged against the
session's own status would be refused by the word the start had just
replaced. `start_if_exited` returning `Exited` is what says a process was
started, so `run_command` judges against `Idle` from there rather than
re-reading a status the pump may not have caught up with.
The phone's half is that the process button is disabled while its own request
is in flight, so a second press cannot be decided against a status the first
one has not changed yet. That is a courtesy rather than the fix: the server
refuses the second request either way, because a phone that has lost the
stream cannot be relied on to know.
On the phone this is one button in the composer, left of Send, whose mark and
colour say what pressing it would do now: an orange pause while a turn is
running (interrupt — the process stays), a red stop when it is not (end the
process), and a green play when it has exited (start it again). One button
rather than three that come and go, so its presence is never the signal.
### A backend start adopts, and starts nothing (decided 2026-08-30)
Starting the server is not something a session should be able to tell
happened. `SessionManager::new` takes charge of the processes that are still
running and **leaves every other session exactly as it found it** — listed,
with its transcript, its event pump and the SSE stream a phone reads, and no
driver at all until somebody asks for one.
What it did before was launch a driver for every session in the config, and
`ClaudeDriver::launch` starts a process when there is none to adopt. So a
session somebody had deliberately stopped came back at the next rebuild,
which is the decision Stop exists to make being undone by an unrelated
event — and since a driver announces `Idle` for a process it started, the
session was also stamped as active at the moment of the restart. On the
phone that read as *every* session idle and "just now" after every restart,
with the list — sorted by that time — in an order that meant nothing.
- **`Launching` is the parameter that says which it is**, and the seed an
import carries rides on the asked-for variant, because a restart re-seeding
a transcript would write the imported conversation into it twice.
- **A session with no process has no driver.** `DriverCell` is an option
rather than a driver whose requests go nowhere, so "nothing is running
this" is a state the code can be asked about instead of one it discovers by
sending into a dead fifo. `LiveSession::ask` is the one place that answers
it, with an `Event::Error` naming what could not happen — a request nobody
can carry out is reported, never swallowed.
- **`--resume` on a crashed session is now a press rather than a restart.**
That is the whole of what is given up, and it is small: a session whose CLI
died reports `Exited` and draws the Start button, and *sending it anything
at all* starts it (above). What is bought is that the two are told apart by
who asked, rather than a restart guessing that everything it found should be
running.
- **What a launch settles the status to is written into the transcript, at
the time of the last thing the session actually did.** Adopting, the
transcript's word stands except for the `Exited` a live process disproves.
Taking charge of nothing, every word but `Exited` is disproved at once — a
backend killed mid-turn leaves a transcript saying `Running`, and that
draws a stop button for a turn that ended hours ago. The correction goes in
the transcript because the list reads the manager's status and the session
screen replays the file; it is stamped with the transcript's own last time
because it is not something the session did — this server noticed, at a
moment of its own choosing, and `now` there is the same lie in the same
field that `Transcript::last_activity` exists to prevent.
- **A session that has never done anything reports when it was created.** Its
transcript is empty — a driver announcing the state it starts in is not
news, so nothing is written — which makes it the one session with no line to
read a time off. The clock was the fallback, so a session nobody had sent
anything to climbed to the top of the list at every restart. Not the
transcript file's mtime, which is the same instant for an empty file and a
worse answer for a shared checkout that can be copied or touched;
`SessionConfig::created` is recorded rather than inferred.
### Sessions spawned while testing clean themselves up (decided 2026-08-30)
`--throwaway-sessions`, **on by default in a debug build**. Every session
spawned by such a server is marked `throwaway` in the config, and a marked
session's process is stopped when the server exits or is signalled, instead
of being left for the next start to adopt.
Leaving processes running is the design and it is right for the sessions
somebody is using. It is exactly wrong for the ones a test made: those leave
a `claude` behind that every later server adopts, they cost tokens if
anything ever speaks to them, and nothing ever says they are there — twelve
accumulated on this machine in a day. An agent testing this app should not
have to remember a cleanup step, and "remember to" is not a mechanism.
- **The flag marks; the mark decides.** What a server was told at startup
governs only the sessions it spawns, and the mark is written into the
session, so it outlives that server. A session spawned deliberately keeps
running whichever server happens to be up when one exits, and a throwaway
one is cleaned away even by a server started without the flag. The
alternative — the exiting server stopping whatever it happens to have
marked in memory — makes cleanup depend on which process is up, which is
the thing that fails at exactly the wrong moment.
- **Stopping is not asking.** `process::stop` sends SIGTERM and leaves its
SIGKILL on a tokio timer, and a runtime that is shutting down never runs
it. That is precisely how the original `shutdown_all` leaked the processes
it reported stopping, so the exit path waits for them with
`process::wait_gone` — one deadline for all of them, since they were
signalled together — and kills whatever is left. `Driver::stop` is the
per-driver half, the same one a delete uses; only the waiting differs.
- **A zombie is dead.** Found by the test for the above: `/proc/<pid>/stat`
keeps the entry, with the same pid and the same start time, until the exit
status is collected — so a process that had plainly finished answered
"still there" for as long as nothing reaped it, and `Liveness::Alive` is
the word that makes `Exited` unsayable. The state field is read alongside
the start time now. This was reachable outside the test: anything that
blocks the runtime delays tokio's own reaping.
### Importing refuses a session that is already open (decided 2026-08-29)
Claude Code keeps a descriptor per live session at
`~/.claude/sessions/<pid>.json` carrying the `sessionId` and a `procStart`
— the same pid-plus-start-time identity used above. So "is this session
open right now" is a **measurement**, not a heuristic, and the import list
reports it as `no` / `yes` / `unknown`. Three answers because a machine that
keeps no such record cannot answer, and "could not check" is not "nobody is
using it".
`yes` is refused. This is not hypothetical: on 2026-08-29 an agent imported
the session it was itself running in. Two `claude --resume` processes then
edited one checkout and appended to one transcript, the whole 65 MB
conversation — 154 embedded screenshots — was duplicated into the file under
a new prompt id, and the adopted copy re-read all of it. It ended at the
account's session limit.
### pi driver specifics
- Spawn: `pi --mode rpc --provider openai-generic --model <name>` (endpoint =
the llama-server the LlamaServerManager provides), `--session-dir` under
our session storage so transcripts and pi's own session files live together.
- Auto-compaction on by default (`set_auto_compaction`), threshold
configurable per session; manual `compact` exposed as a button.
- `steer` for mid-run messages, `abort` for stop, `set_model` when the target
endpoint changes.
- pi's session JSONL gives resume-after-restart, same as Claude's.
### Models (built 2026-08-28)
The owner asked for listing and downloading models from HuggingFace and running
them with different parameters, which makes model management part of the
feature rather than something done by hand beforehand.
- **A download belongs to the model, not to the request.** Keyed by
`owner/repo/file.gguf` and owned by the server, so a second device can
watch one it did not start, and so an hour-long fetch survives a phone
locking its screen. Every run has an id and its outcome outlives it,
because "not downloading" otherwise means finished, never started, or
someone else's run ended while you were away.
- **Progress is measured.** `total` is Content-Length, or Content-Range's
last field on a resumed request, and absent when the server says
nothing — never an estimate.
- **Resume is guarded by identity, not by hope.** A partial carries the
ETag it was written against; a mismatch discards it. `If-Range` would be
the tidy mechanism but HuggingFace's CDN ignores it (probed
2026-08-28). The published sha256 is checked before the file is renamed.
- Parameters reach a driver as an untyped `params` map on the session, so
the shared schema does not grow llama.cpp's vocabulary.
### llama-server management
`config.ron` lists **models** (name → GGUF path or llama-server args, per
host) and **hosts**. The manager runs at most one llama-server per
`(host, model)`, spawned on demand when a session needs it:
- Spawn (local or `ssh host llama-server …`) on an allocated port, wait on
`/health`, hand the endpoint to the pi driver.
- Refcounted by sessions. The path out, written in the same change as the
spawn: the last session using an instance releasing it starts an idle
timer (configurable, e.g. 10 min), after which it's killed. Delete of the
last session kills it immediately.
- "Change model" on a llama session = acquire the new model's server,
`set_model` on pi, release the old one. Context carries over (it's
prompt-replayed by pi against the new endpoint).
- Remote llama-server output is only reachable from the backend host, and
binds localhost on the remote side with an SSH local port forward
(`ssh -L`) held by the manager — no LAN-exposed inference ports. Built
2026-09-04, held by the session's own ssh client rather than by a
manager: there is one server per session (not per `(host, model)`), so
the process that runs it is the process that owns the tunnel, and the
two die together.
- **The model file lives on the machine that serves it** (2026-09-04).
Each setup names its own models directory (`SshConfig::models_dir`,
default `~/.local/share/ai-app/models` expanded on that machine), and a
spawn resolves the key there — one round trip that answers "at
/abs/path" or "missing", so a model that is not there is refused at the
spawn instead of becoming a server that never becomes ready. The spawn
screen offers `GET /setups/{id}/models`, which is that machine's list,
rather than `GET /models`, which is the backend's downloads. Downloading
*to* another machine is deliberately not built: it would be a
multi-gigabyte transfer with no progress anywhere, and the file gets
there however anything else on that machine got there.
### SSH
- Host entries in `config.ron`: name, `user@host`, optional ssh options,
which capabilities it has (claude / pi / llama-server, with paths if not on
PATH). Key-based auth only, using the system `ssh` client via
`tokio::process` — no Rust SSH library; this inherits `~/.ssh/config`,
agents, and jump hosts for free. (Rule 23: openssh is already here and
battle-tested; a library buys nothing but a second config surface.)
- A remote session is exactly a local one with the command wrapped in
`ssh -T host …`. Process death ≙ connection death; the session shows as
`exited` and both dialects resume (`--resume` / pi session file) on respawn,
so a dropped SSH connection is an annoyance, not data loss.
- **The transport wraps the driver, not the other way round** (decided
2026-08-28). A driver says what to run — program, arguments, working
directory — and something above it turns that into a process, locally or
through ssh. Today `ClaudeDriver::spawn` calls `ssh::command` itself,
which puts transport knowledge inside a translator whose job is a wire
format, and means every future driver has to remember to do the same.
Inverting it also removes the "Run on" lie for free: a driver that emits
no command, like the echo one, has nothing for a transport to wrap, and
the picker can say so.
- The interface that inversion needs is **not just "run a command"**, and
llama.cpp is the case that shows it: a managed `llama-server` is started
as a process but then spoken to over HTTP, so a remote one needs a
forwarded port (`ssh -L`) as well as a spawned process. A transport is
therefore "run this" plus "reach this port", and the second operation is
a no-op locally. **Built 2026-09-04**: `Transport::reserve_port` returns
a `Forward { there, here }` — the port the program binds on its own
machine and the port that reaches it from the backend, the same number
when that machine is this one — and `Launch::reaching` carries it, so
the connection that runs the command also carries the tunnel. The far
end is a guess from a range below the ephemeral one, because no
portable way to ask a machine for a free port avoids racing with the
bind anyway; a collision is not silent, since the program fails to bind
and the readiness poll reports what its log said.
- **A forwarded launch gets a pty and every other one does not** (measured
2026-09-04). Killing the ssh client ends a CLI because it closes the
stdin that CLI is reading; `llama-server` never reads its stdin, so the
same kill left it running on the far machine with the model loaded —
one orphan per stopped session. With `-tt` the far side takes SIGHUP
when the connection goes. Its log then arrives through a line
discipline, which nothing parses. `-T` stays everywhere else, where a
pty would rewrite the JSONL.
- Images need no file transfer, contrary to what this section said
before: `attachment_block` base64s an uploaded image into the
stream-json message itself, and produced images come back the same way
for the translator to write out locally. Nothing has to exist on the
remote filesystem, so there is no `scp` step to get wrong.
- **Any other file is told to the session by path** (2026-09-03: a trace,
a log, a zip -- things a model cannot be shown and the CLI can read).
The upload is streamed to disk under the session's `attachments/` on
this machine -- the transcript references it there and the phone can
fetch it -- and the message ends with `Attached file: /abs/path`. For a
session on another machine the upload also copies the file there, in
the same request, over one `ssh` invocation (`cat` from stdin, then
`pwd -P` so the answer is the absolute path the CLI is told). It lands
in the setup's `attachmentsDir` if set, else the session's cwd, else
the login home; the resolved remote path is recorded beside the file
(`<name>.remote`) and is what the driver names. A copy that fails fails
the upload, so no message ever names a file that is not there. Images
are unaffected: they ride the message as base64.
### Usage limits (Claude)
Poll `https://api.anthropic.com/api/oauth/usage` — the same endpoint behind
Claude Code's `/usage` — with the OAuth access token from Claude Code's local
credential store (`~/.claude/.credentials.json`), headers
`anthropic-beta: oauth-2025-04-20` and `User-Agent: claude-code/<version>`
(without the User-Agent it lands in an aggressively rate-limited bucket).
Poll at ≥180 s, only while any Claude session exists or the usage screen is
open, cache the last answer. Surface: 5-hour and weekly window utilization %
and reset times. It's undocumented, so `usage.rs` treats every field as
optional and degrades rather than erroring. Structure it as one
`UsageProvider` per paid service so a second service later is a new impl,
not a parallel screen (rule 9).
**Per provider, not per machine (decided 2026-09-04).** A machine is not
what is metered; the provider a session runs is. One machine offers echo,
the Claude CLI and a local model side by side, and only the second of them
spends anything — so pairing a session with a snapshot by machine alone
drew the CLI's five-hour window under every echo session on it, reporting
a quota that session cannot spend and could never run down. A session now
names its meter (`usageProvider`, from `DriverKind::usage_provider`, which
`usage::providers_for` also reads so the two lists cannot disagree), and
`GET /usage` is matched on machine *and* provider. `None` is a session
that meters nothing, and the phone draws nothing at all for it — not a
zero, and not "unknown".
`DriverKind::Echo` names a meter of its own, and it exists only when a
test has asked for one: `/usage` in an echo session sets an invented
answer (`usage::Fixture`), and with none set there is no snapshot and no
bar. That is what makes the states of those screens reachable — a number
near the top, a window between blocks with no reset time, a machine
nobody logged into, one that could not be reached — without spending real
quota to arrange them, which is why none of them had ever been looked at.
**Per machine, not per backend (decided 2026-08-29).** The credential store
that matters is the one on the machine the session runs on, because that is
the account being billed. Reading this machine's was right only while the
backend and the CLI were the same box — and in the layout this is aiming
at they are not: `ai-server` belongs on the host, the host has no `claude`
CLI, and the CLI machine is a remote. So credentials are read through the
session `Transport` (`ssh host sh -c 'cat $HOME/…'`, `$HOME` expanded by
the far shell because a path built locally is the wrong home), one snapshot
per setup that offers Claude, cached per machine. The HTTP call stays on
the backend rather than running remotely, so the far end needs nothing but
a shell.
The snapshot says which of four things happened rather than carrying a flag
and a message: `ok`, `notLoggedIn`, `unreachable`, `failed`. The one that
matters is `notLoggedIn` — a machine nobody put an account on is working as
configured, and collapsing it into an error string made a healthy setup
read as broken. A machine with no Claude provider is not asked and gets no
row at all.
### HTTP surface (phone ⇄ backend)
REST for actions, one SSE stream per open session screen for events, all over
the pinned TLS listener. SSE over WebSocket because resume-by-cursor
(`Last-Event-ID` = transcript seq) is native to it and the inbound direction
is plain POSTs anyway.
```
GET /providers what can be spawned (name, kind, models)
GET /hosts machines a session can be run on
GET /sessions list (id, provider, host, title, model, status, last activity)
POST /sessions spawn {provider, host, model, cwd, permission_mode, title}
GET /sessions/:id/events?after=N SSE: transcript replay from N, then live
POST /sessions/:id/message {text, attachment_ids}
POST /sessions/:id/unqueue {message_id} take back one not read yet
POST /sessions/:id/answer {question_id, answer} (questions and permissions)
POST /sessions/:id/interrupt stop the running turn; the process stays
POST /sessions/:id/stop end the process; the session and transcript stay
POST /sessions/:id/start run the process again, continuing the conversation
POST /sessions/:id/model {model}
POST /sessions/:id/compact (llama sessions)
POST /sessions/:id/attachments multipart upload → id (referenced by /message)
GET /sessions/:id/files/:ref images the session produced or was sent
DELETE /sessions/:id kill process, release llama-server, delete transcript+files
GET /usage cached usage windows, per machine and provider
GET /setups/:id/models GGUFs on that machine, for a llama session there
GET /setups/:id/dir?path=P entries of directory P, and P resolved
GET /setups/:id/file?path=P content of file P, or why not
PUT /setups/:id/file {path, content, ifSha256}; 409 if it moved on
POST /setups/:id/file {path} create empty; refused if it exists
POST /setups/:id/dir {path} create; refused if it exists
GET/PUT /hosts, /models config editing from the phone
```
Sessions live in `config.ron` (`$XDG_CONFIG_HOME/ai-app/`) + a per-session
directory under `$XDG_DATA_HOME/ai-app/sessions/` (transcript.jsonl,
attachments, produced images), owner-only. Deleting a session is the
complete path out of everything spawning one created.
### The file explorer (decided 2026-09-03)
**`EXPLORER.md` holds this design**, decision by decision with what was
rejected, the same way this file does — it is long enough to be its own
document and it is where a change to it belongs. The one-line version: a
machine's filesystem, seen from the phone through the backend, keyed on
the **setup** rather than on a session (a session only says where to
start), with every operation one fixed shell script run through
`Transport` so the local and the ssh case are one implementation. The
security consequence is in the token paragraph below.
### Security
- TLS with a self-signed CA, pinned in the app — same
idempotent-CA/reissued-leaf scheme as dev-updater, same one-way-door
caveat about regenerating the CA, but generated **in process on first
start** (`certs.rs`) rather than by a shell script calling openssl
(2026-08-25). One place then decides the extensions, the file modes, and
which addresses the leaf covers — every local IPv4 plus loopback and the
emulator's host alias, so nobody maintains a hardcoded IP — and there is
no setup step to forget.
- Unlike dev-updater, the pinned CA is **not a constant in the source**:
the build reads `$XDG_CONFIG_HOME/ai-app/certs/ca.pem` from the machine
doing the build and generates the constant (`generatePinnedCert` in
`app/androidApp/build.gradle.kts`; `AI_APP_CA` overrides). Decided
2026-08-25, and it does three things at once — the trust anchor follows
the build machine, so an APK built on the backend host pins that host
and one built in the dev VM pins the VM's throwaway CA and is only good
for its emulator; there is no second anchor to add for development and
forget to remove; and regenerating a CA needs a rebuild rather than a
paste, so a stale constant can't quietly disagree with the server.
- **The dev VM is untrusted** (decided 2026-08-25): a machine that isn't
malicious but could become so. It matters because the repo is a
read-write virtiofs mount shared between the VM and the backend host, so
under this model everything in it — source, `server/target/` binaries,
and the shell scripts the host runs, some with sudo — is
attacker-writable. Two consequences:
- **Nothing secret lives in the repo.** Certificates are generated on
the machine that serves them and written to
`$XDG_CONFIG_HOME/ai-app/certs` (0700, keys 0600); `config.ron` and
session transcripts go to the XDG config and data directories, per
machine. A CA private key the VM could read would let it mint a leaf
the pinned app accepts, which is precisely the attack pinning exists
to stop — pinning against a CA the attacker holds is no pinning at
all. Transcripts move for a plainer reason: they are whole
conversations. As a bonus this ends the host and VM sharing one
config, which had already produced a test token live on the backend,
and takes state out of reach of `git clean -xdf`.
- **The host should not execute what the VM can write** — build and run
the backend from a host-only checkout rather than the shared mount.
Moving the keys closes the smaller door; this is the larger one.
- Development in the VM generates its own throwaway CA. Whatever is
installed on the real phone must pin only the host's.
- The CA key is not needed by the server at all (only `leaf.pem` and
`leaf-key.pem` are read), so it can move offline once the setup is
stable; reissuing a leaf is the only time it is wanted.
- Not addressed, and accepted: a compromised VM can return anything it
likes from the sessions it runs, since running an agent there is the
point. The blast radius is that session's content, not the backend.
- This server is strictly more dangerous than the updater: its API *is*
remote code execution (spawn a bypass-permissions Claude on any SSH host).
Pinning authenticates the server to the phone but not the phone to the
server, so a bearer token adds the other direction. Threat model: the token
gates LAN-reachable RCE; it does not (and cannot) defend a compromised
backend host or phone — those are inside the trust boundary, and a
compromised phone is handled by rotation.
- **The explorer's routes take a path, and that is deliberate**
(2026-09-03; see EXPLORER.md's decision 3). Elsewhere the rule is that
the phone picks an **id** and the server resolves which file it names —
the import listing is written that way so an enrolled token cannot
become "read me an arbitrary file". `/setups/{id}/dir` and
`/setups/{id}/file` take the path, because the path is the whole
feature. It grants nothing new: the same token already spawns a
bypass-permissions agent in any directory on any machine a setup names,
and that agent already reads and writes every file its user can, so
this is a shorter path to authority the token holds either way. The
import rule stands where it is, because there a path was unnecessary
and refusing one cost nothing. What is unchanged is the harder line:
**no route accepts a command.** Listing, reading and writing are fixed
scripts in `files.rs`; the phone chooses only the path and the bytes.
- **Generation**: 256 bits from the OS CSPRNG on first run, base64url. A
machine credential, never typed twice, so unguessable costs nothing; at
this entropy no key stretching is needed.
- **Enrollment**: printed once as a terminal QR code (`qrcode` crate,
ANSI), encoding `aiapp://enroll?host=…&port=…&token=…`. The CA stays
embedded in the APK (`PinnedCert.kt` pattern), so the QR carries no
trust material — photographing the terminal leaks only the token
(rotatable), never a way to weaken pinning. The app registers an intent
filter for the `aiapp://enroll` scheme as a fallback, for a camera app
that redirects a scanned URI straight to `MainActivity` (2026-08-24).
That was meant to be the only path — "the app side needs no QR library
at all" — but reversed the same day: not every phone's stock camera
redirects a scanned URI to an app reliably, so the Settings screen also
scans in-app via `zxing-android-embedded`'s `ScanContract` (a ready-made
scanner Activity reached through the AndroidX Activity Result API,
fully offline, no Play Services/ML Kit model download) and feeds the
decoded URI to the same `parseEnrollmentUri` (2026-08-25).
- **Storage**: server keeps only the SHA-256 in `config.ron` (plain hash
is enough for high-entropy random input; buys that a leaked config
doesn't leak the credential). No "show token again" — lost means rotate.
Phone side: sealed with an Android Keystore AES-GCM key (a small
hand-rolled helper in `ServerConfig.kt` — Jetpack's
EncryptedSharedPreferences is deprecated with no drop-in successor, and
Google's guidance is now "use Keystore directly"; 2026-08-24).
- **Transport**: `Authorization: Bearer` header on every request including
the SSE GET. Never a query parameter (URLs leak into logs). The tracing
layer must not log the header — covered by a test so a logging change
can't silently start leaking it.
- **Verification**: one middleware wrapping the entire router in `main.rs`,
never per-route, so a new route can't forget auth. Zero unauthenticated
endpoints, `/health` included. Hash-then-constant-time-compare
(`subtle`); failures logged with peer address plus a small fixed delay —
not against brute force (infeasible at 256 bits) but so scanners show up
in the log.
- **Rotation (the path out)**: `--rotate-token` regenerates, invalidates
the old hash immediately, reprints the QR. That's the whole lost-phone
story. Config stores a *list* of `{name, hash}` (of one, today) so
per-device tokens with individual revocation are a config entry later,
not a schema migration.
- **Why not mTLS**: stronger in theory (key never leaves the Keystore, no
bearer secret to exfiltrate), but given pinning the delta is only
"someone reads the token off a device already inside the trust
boundary", and it costs Android client-cert provisioning ceremony and a
worse new-phone story than a QR scan. Revisit if this outgrows
single-user-on-LAN.
- **Off-network access: plain WireGuard** (decided 2026-08-24; no third
party). The backend binds to the WireGuard interface (`wg0`) only; the
phone runs the official WireGuard app (always-on VPN, per-app tunneling),
enrolled by scanning its config as a terminal QR
(`qrencode -t ansiutf8 < phone.conf` — same gesture as token enrollment).
The only internet-visible thing is one forwarded UDP port that is silent
to unauthenticated packets — scanners see it as closed — so the app's
pre-auth surface (rustls handshake, hyper parsing, auth middleware) is
reachable only from enrolled peers, and the token becomes defense in depth
rather than the sole gate. Addressing stays single-path: the phone reaches
the backend at its WireGuard address (e.g. `10.66.0.1`) from everywhere —
one address in the app, one SAN in the leaf cert (`SERVER_IP=`/SAN
override in the cert script), no home/away distinction. Another machine
later is one keypair + one `[Peer]` block.
- Operational needs, accepted: a public endpoint hostname. The home IP is
mostly static but not guaranteed, so the phone's endpoint is a DDNS name
(free, e.g. DuckDNS, or the router's built-in client; a curl cron on the
backend host works too) that tracks changes automatically. One WireGuard
nuance: the phone app resolves the endpoint hostname when the tunnel
comes up and does not re-resolve on its own, so on the rare IP change
the fix is toggling the tunnel off/on once DDNS has caught up (minutes).
The symptom is obvious (app can't reach the backend) and lossless — the
SSE cursor design means reconnects replay whatever was missed. Also:
at-home traffic rides NAT hairpinning on the router (verify early; most
support it, and the fallback is toggling the tunnel off at home).
- Rejected: **Tailscale** — same WireGuard underneath with easier setup
(no port forward, LAN peer discovery), but it adds a third-party
coordination service and account this setup doesn't need at two or
three devices; **Headscale** — self-hosting that coordination server is
strictly more moving parts than one wg config per peer at this scale;
**forwarding the HTTPS port directly** — puts every internet scanner
one pre-auth bug away from RCE on a machine holding SSH keys.
- The server still refuses to start without TLS — no plaintext listener
exists even inside the tunnel, so the token can't travel unencrypted by
misconfiguration, and interface binding failing closed (refuse to start
if `wg0` is absent, rather than falling back to 0.0.0.0) is part of the
same guarantee. Development gets `--bind <ip>` as an *explicit, logged*
override (loopback for curl, a LAN address for a pre-WireGuard phone) —
a deliberate flag, never a fallback, so the fail-closed default is
untouched (2026-08-24).
- The bootstrap-over-HTTP trick from the updater is unnecessary here — the
app installs via Dev Updater.
## App (`app/`)
Kotlin + Compose Multiplatform, single `:androidApp` module, same versions as
dev-updater (Kotlin 2.4.x, CMP 1.11.x, JDK 21). Screens:
1. **Session list** — cards: kind icon, title, host, model, status
(running / awaiting answer / idle / exited), last activity. Spawn FAB;
swipe/long-press to delete (confirm). Sessions awaiting an answer sort to
the top — that's the "your turn" inbox.
2. **Spawn** — kind, host (from config), model (Claude list is static+editable;
llama list from config), working directory, permission mode (Claude),
title.
3. **Session screen** — the core:
- Transcript rendered from the event stream: markdown text, inline images,
collapsed-by-default tool cards (name + input summary, expandable to
output; a spinner while `ToolStart` has no matching `ToolEnd`).
- Question cards inline: option buttons for AskUserQuestion, allow/deny for
permissions, free-text where allowed.
- Expanding a row keeps still **the end nearest the tap**: touch a row's
upper half and its top edge holds, so it opens downwards; touch its
lower half and the bottom edge holds, as the list does by default.
Which half, rather than which control, so that everything that opens
behaves alike whether or not it has a control at each end — a group's
heading and foot bar simply fall in the halves they already occupy. The transcript is laid out from the
bottom, so a row's bottom edge is anchored for free and the top one
has to be arranged. The correction lives in the *layout* phase
(`Modifier.holdTopEdge`): the measurement that discovers the row's new
height asks the list to shift by that much, via
`requestScrollToItem`, before anything is drawn. Doing it from an
effect instead means the wrong position is drawn once first, which
reads as a flick and gets worse the faster the screen refreshes
(2026-08-30, asked for after groups opened upwards and sent their own
heading off the top of the screen).
- Input bar: text, attach (camera/gallery/file), send — **always enabled**;
mid-run sends become steering messages.
- Top bar: model chip (tap to change), stop button while running, token
count, compact button (llama), overflow → delete. (The count settled as
context held rather than tokens spent, and sits on the status row under
the transcript — see `UsageDelta` above.)
4. **Usage** — window bars for the 5-hour and weekly limits with reset times.
5. **Settings** — server address + token, hosts editor, llama model list
editor.
Networking mirrors dev-updater's app layer (`AppsApi.kt` style thin client +
pinned transport), plus an SSE client with `after=` resume driven by
connectivity/lifecycle. The app keeps no persistent transcript store — the
backend's transcript is the source of truth; the app caches only for the
screen it's showing.
### Notifications: two places, never both (decided 2026-08-30)
The backend's `GET /notifications` is one SSE stream of attention-wanting
moments, and the app decides where each one is said. Three outcomes, in one
place (`NotificationService.show`):
- **Nothing at all** if the session is the one on screen. The transcript in
front of the reader is already saying it.
- **A banner over the app** if the app is up — `SessionAlerts`, queued, one
per session replacing that session's own, dismissable by a push in either
direction, and otherwise retiring itself when the bar across its foot runs
out. Tapping one opens the session, through the same path a tapped
notification uses.
- **A row in Android's drawer** otherwise, which is what the foreground
service exists for.
Never two of them for one moment. A notification that has already been shown
in the app is not something to also find in the shade afterwards, and a
drawer that fills up behind an app that showed you each one is a drawer
nobody reads.
Which of the three applies is answered without a flag anybody has to keep
level: the session on screen is registered by the one composable that draws
one, and "the app is up" *is* the banner queue being collected, since it
collects only while it is on screen.
**What counts as finished** is decided in `notification_for`, and since
2026-08-31 it takes the number of messages the session has been given and
not started reading. With one waiting, a turn ending is not the work
ending: a message written into the tail of a turn is read the moment that
turn's `result` lands, so the session goes idle and immediately runs again
-- and the phone that sent it was told its work had finished, seconds
before any of it was done. The count is kept in `pump` from the recorded
events (`MessageQueued` up, the `UserMessage` that resolves it or a
`MessageDropped` down), because that is the one place that sees every event
in transcript order. It does not suppress *awaiting input*: a question is
worth saying whatever is queued behind it, and the queue is exactly what
will not move until it is answered.
The alternative considered and rejected was giving the app its own
connection to `/notifications` while it is in front. That is a second stream
per device saying the same thing, and it puts the "which of these two shows
it" decision in two processes' worth of code instead of one function.
### Deferred polish
Noticed and deliberately not fixed yet, so they are not re-found from
scratch. None is a defect; each is a decision waiting for the app to have
been used enough to say which way.
- **The session screen's header is lopsided.** The row is `padding(
horizontal = 8.dp)`, so the status on the right sits exactly 8dp from the
edge while "Back" on the left is a `TextButton` whose touch target is
wider than its text — the same 8dp reads as more. It is the "align the
mark, not the box" case: either align the button's content or size the
button to what it draws, rather than nudging with a hardcoded offset.
## Compaction: options explored
Context: raw llama-server has no conversation memory management; the context
window just fills.
1. **pi's auto-compaction** — *chosen*. When the prompt nears the model's
context limit, pi summarizes older history with the model itself and
replaces it with a structured summary; threshold configurable; manual
`compact` also exposed. Battle-tested, zero work for us.
2. **Manual compaction in a custom Rust loop** — *the later third driver*.
The design when we build it: every llama-server response reports prompt +
completion token counts; track them against `n_ctx` (from `/props`); at a
threshold (~75%), pause, run a summarization request over all but the last
few turns ("state of the task, decisions made, open items, relevant
file/tool state"), replace those turns with the summary as a system-adjacent
message, continue. Keep the full pre-compaction transcript on disk — the
phone view never loses history, only the model's view shrinks. Worth doing
eventually for control over the summarization prompt and for tool-loop
experiments pi doesn't allow.
3. **llama-server `--context-shift`** — *rejected* as the strategy. It
truncates old KV cache entries: silent forgetting, no summary, and it
corrupts the harness's view of what the model knows. Fine as a server-side
safety net; not memory management.
## Phases
1. **Skeleton** — *done 2026-08-24.* Repo layout, cert script, TLS + token
auth, wg0-bound listener (fail closed if the interface is missing),
config.ron, session registry with a fake `EchoDriver`, session list +
session screen in the app end-to-end over SSE. Proves the whole pipe
before any AI is involved. Verified: 10 server tests + clippy clean;
curl end-to-end over pinned TLS (auth rejection, spawn, SSE
replay/resume by cursor, question round trip, restart continuing seq
numbers, delete); the app on another checkout's emulator against the real
server (QR-style enrollment via deep link, spawn, streamed echo turn,
question answer, tool card).
2. **Claude local** — *done 2026-08-24.* ClaudeDriver: spawn, stream
text/tools, mid-run send, interrupt, permission questions,
AskUserQuestion, images both ways, delete.
*Milestone: daily-drivable Claude replacement on localhost.*
Wire-format notes live in `session/claude.rs`'s module doc (pinned
against CLI 2.1.237): permissions need the hidden
`--permission-prompt-tool stdio` flag; AskUserQuestion answers ride
`updatedInput.answers` keyed by question text; `set_model`/`interrupt`
are control requests; 2.x permission modes are acceptEdits / auto /
bypassPermissions / manual / dontAsk / plan (no more "default").
Attachments/files were re-homed under `/sessions/:id/…` (table above)
so their lifecycle is the session directory's — delete stays the
complete path out.
3. **Usage screen** — *done 2026-08-24.* The undocumented endpoint's
`limits[]` array parsed defensively into labeled window bars; cached
behind the ≥180 s minimum with no background polling.
4. **llama.cpp** — LlamaServerManager (local), PiDriver, model change,
compaction controls. *Deferred (2026-08-24): pi/llama-server aren't set
up in this VM, so this phase isn't testable here — Claude first; the
driver seam is ready when it is.*
5. **SSH** — host config, remote spawn for both kinds, remote llama-server
with port forward, attachment shipping. *Host config and remote spawn
done 2026-08-25* (any session of any provider can name a host; the
command is the identical one wrapped in `ssh -T`, with every argument
shell-quoted). Attachment shipping turned out to be unnecessary for
images — they ride the stdio JSONL as base64 in both directions, so
nothing needs `scp` — and was built on 2026-09-03 for files, which
are attached by path (see "Transport" above). Remote llama-server with
its port forward landed 2026-09-04 — see "Transport" above for the
forward and the pty, and "llama-server management" for where the model
file has to be.
Two things learned doing it: a remote session inherits ssh's non-login
PATH, which is narrower than an interactive shell's (point `command` at
an absolute path if a CLI isn't found), and the remote command is run
with `exec` so dropping the connection takes the CLI down rather than
orphaning it.
6. **Polish** — reconnect edges, notification when a session awaits an answer
(the "your turn" push), transcript search, whatever daily use surfaces.
Each phase ends runnable and verified against the real thing (rule 22); the
backend gets tests where logic is pure (event normalization, transcript
cursors, config persistence, refcounting) — the app is UI over the API and is
verified by running it, matching dev-updater's posture.
## Open questions / risks
- **Claude stream-json control protocol details** (permission requests,
set-model, interrupt wire format) are the least-documented dependency and
version-coupled to the installed CLI. Phase 2 starts by probing the
installed version and pinning what works; the `--resume` respawn fallback
covers whatever the control channel can't do.
- The **usage endpoint is undocumented** and has changed rate-limit behavior
before; treat as best-effort.
- **pi RPC schema drift** — pin a pi version; the translator is one file.
- Whether **notifications** need FCM or a foreground-service polling
connection — decide in phase 6; the SSE cursor design already supports
either.
- Claude sessions over SSH need the remote host **logged in to Claude**; usage
reporting reads only the backend host's credentials. Acceptable for now
(same account everywhere); revisit if not.
## References
Research behind the decisions above (verified 2026-08-24; re-check against
installed versions when each phase starts):
- pi RPC protocol: https://github.com/badlogic/pi-mono/blob/main/packages/coding-agent/docs/rpc.md
— commands (`prompt`, `steer`, `follow_up`, `abort`, `set_model`,
`compact`, `set_auto_compaction`, session ops) and the event stream.
- pi + llama-server in practice: https://medium.com/@tolgaeren/running-pi-with-local-llms-c596aa14b062
- llama-server API (`/health`, `/props`, OpenAI-compatible endpoints,
`--context-shift`): https://github.com/ggml-org/llama.cpp/blob/master/tools/server/README.md
and the offline-agentic-coding walkthrough:
https://github.com/ggml-org/llama.cpp/discussions/14758
- Usage endpoint (`GET https://api.anthropic.com/api/oauth/usage`, bearer
token from `~/.claude/.credentials.json`, headers
`anthropic-beta: oauth-2025-04-20` + `User-Agent: claude-code/<version>`,
≥180 s polling; wrong User-Agent → aggressive 429 bucket):
https://github.com/anthropics/claude-code/issues/31637 and
https://github.com/Maciek-roboblog/Claude-Code-Usage-Monitor/issues/202
- Sibling project this repo's conventions mirror: `../dev-updater`
(README.md + AGENTS.md — server/registry/routes layout, cert scheme,
testing posture, Android env notes).