Commit Graph
100 Commits
Author SHA1 Message Date
iris 6b4911c67b Give the session's own state a line, instead of the transcript's corner
The token total floated over the bottom-right of the transcript, where a
long message ran underneath it, and the working indicator was an item
inside the list -- so it scrolled away exactly when somebody reading back
wanted to know whether anything was still happening.

Both are facts about the session rather than turns in it, so they get one
row directly above the box you type into: the thing they report on is the
next thing you touch. `exited` moves with them, since it is the same kind
of fact and nothing else on the screen would have said it once the
indicator left the list.

The row is drawn whether or not it has anything to say. An empty one
costs a line; a row that came and went would move the text box under a
reader's thumb every time a turn started, and would make its own presence
the signal for a state it never names. For the same reason the compaction
case had to fit the same single line: its bar now takes the row's free
width between the label and the total rather than a row of its own, which
keeps it far wider than a spinner -- the reason it is a bar at all, since
nothing arrives in the transcript while a compaction runs and a small
moving thing there reads as a session that has hung.

Still no fraction to fill, re-measured today rather than assumed: a real
80,346-to-2,088-token compaction took 23 seconds and the CLI emitted not
one line between saying it had started and saying it had finished.
Elapsed seconds remain the only honest number.

Looked at on the emulator in all three states -- idle, working, and six
seconds into a real compaction -- and at 320dp, the narrowest width a
phone actually has, where the row still holds one line.
2026-08-29 21:35:30 -04:00
iris 549e49bc10 Record a steer where the model read it, not where it was typed
A message sent while an answer was streaming was recorded in the middle
of that answer and above the tool call it ended with. The model had
committed to that call in the same message it was already writing, so it
had read none of it -- and on screen the tool result underneath read as
something the steer had asked for. The answer also split into two
bubbles around a message that was not part of it.

The driver announced a steer at "the next assistant text or tool call",
on the reasoning that anything the CLI says next is proof it has been
round the model again. With --include-partial-messages that is not true:
the deltas and the tool_use block of a message already in flight keep
arriving afterwards, and none of them saw the steer.

`message_start` is what actually proves it. The CLI sends the previous
call's tool results back before it opens the next assistant message, so
that line is the first moment anything written since can have been read
-- and it carries no events of its own, which is what makes it a place
to put one. Verified against 2.1.237: message_start, the blocks, the
tool_result, then the next message_start.

The end of the turn stays as the other half, and is the case that must
not be lost: a message typed after the final model call has no later
message_start, and one that is only recorded when announced would
otherwise vanish while a phone drew it as still waiting.

Checked live on haiku, before and after. Before: the steer landed at
seq 37 among the essay's deltas, with the tool call at 45 and its result
at 46. After: essay whole, tool call 42, result 43, steer 44. Also
checked the case this had no reason to touch -- a steer sent during a
30-second Bash call, which was already correct -- and it still records
after the result. The two tests fail on the old rule; the failure prints
the old order, which is the bug.
2026-08-29 21:16:21 -04:00
iris a50d72960c Say how long is left, not that the window is five hours
The bar read "31% of 5h", which is the one thing about the window a
reader already knows. What decides whether to start something now is how
long what is left has to last: 80% with twenty minutes to go and 80%
with four hours to go are opposite answers, and the second number was a
screen away on the usage screen.

It now reads "31% - 2h 36m left", and the countdown is driven by a clock
the refresh loop advances rather than computed at draw time. A
percentage that comes back unchanged is an equal value, so Compose skips
the recomposition -- a "left" recomputed only when the quota happens to
move would have sat at a stale figure for hours while looking live.

A window can arrive with no reset time, so that keeps its own wording:
"reset time unknown" rather than "refresh soon", which would be a
recommendation nothing measured. Under a minute, including past the end,
is "refresh soon" -- "0m left" reads as a measurement.

The span arithmetic was already on the usage screen, so it moves into
`ResetCountdown.kt` and both callers supply their own sentence. That
screen still reads "resets in 2h 37m" and "resets in 5d 21h", checked on
the emulator alongside the bar it was not part of changing.

The fill is blue rather than the scheme's primary: the bar sits under
every session header, on a screen somebody opened to do something else,
and it reports a quantity rather than a verdict. The usage screen is
still where the same number turns yellow and then red, for a reader who
went there to be told where the limits are.

Also declares this project's resources for Dev Updater, whose
declaration schema changed in d27b5a3: `resources.ron` says ai-app keeps
its state as `ai-app`, so the Uninstall dialog offers the real
directories instead of saying it cannot tell where they are. Only the
name, because both XDG places are the conventional ones. What that
dialog's config toggle would delete includes the CA under `certs`, which
strands every phone running an APK pinned to it -- noted where somebody
would be standing when it matters.
2026-08-29 20:54:55 -04:00
irisandClaude Opus 5 894180de77 Take the CLI's word for a clear instead of inferring it
Bryan reported no divider when clearing. It was not this code -- the
backend serving him started at 17:06, three hours before `Event::Cleared`
existed, so it has no such event to send and `/clear` reaches it as an
unrecognised passthrough. Verified against a current build: the event is
recorded.

Probing the CLI to establish that turned up something better than what
was here. `/clear` in stream-json mode emits a dedicated
`conversation_reset` line and *then* a fresh `init` with the new session
id -- so watching the id be replaced, which is what this did, was reading
the event through one of its side effects. The announcement says it
directly, and it arrives first, so the divider now lands above the new
conversation rather than after its opening line.

That also removes the reasoning the previous commit needed about which id
changes count. There is one signal now instead of an inference with two
exceptions, and the test that used to pin those exceptions became
`an_init_alone_is_never_a_clear`, which covers all three ways an init
arrives: a session's first, the one a compaction re-announces with the
same id, and the one following a resume.

The resume token still follows the id, unchanged -- one CLI event with
two observable effects, and each half now reads the half it needs.

Verified end to end against a real claude-cli session: message, /clear,
message, and the transcript reads userMessage / assistantText / cleared /
userMessage, in that order. 74 tests, clippy and rustfmt clean.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VETa8afmpWaYezLCqJhDB8
2026-08-29 20:39:21 -04:00
irisandClaude Opus 5 71067275e2 Stop claiming a terminal, and stop calling a recoverable delete final
Two bugs Bryan hit, with one shape between them: a claim stronger than
the thing that was measured.

**The delete warning branched on `imported`.** It told him deleting
`ai-app` could be undone and deleting `manager` could not, when both are
claude-cli sessions whose conversations survive equally. `imported`
records how a session got into the app; what decides recoverability is
whether the *driver* keeps its own record -- the Claude Code CLI does,
under ~/.claude/projects, however the session started; echo and llama.cpp
do not, and for those the app's transcript is the only copy. So the fact
now sits on DriverKind and rides on SessionInfo, decided by the server
from the provider's kind rather than by the phone from its name, which a
person can change.

The comment above the branch asserted "a session started here has no copy
anywhere". That sentence was the bug written down and reasoned from, and
it is gone.

Neither branch promises a restore, which it should not: nothing here
checks the file is still on disk, and re-importing was never a restore
anyway -- this app's transcript holds images, peer messages and command
events the CLI's record never had. So the recoverable text says what is
known and names what goes either way. "Can't be undone" is now said only
where it is true, which is the point of saying it at all.

**"open in a terminal -- close it there first" named a place that need
not exist.** The detection is right and worth keeping: something live
holds that session, and importing it would reproduce the double-resume
incident. But which something was never measured. The live descriptors
here include two of this backend's own adopted sessions and a peer
agent's; none is a terminal, so the instruction sent the reader looking
for a window that was not there.

**And this app did not recognise its own spawned sessions.** The import
list filters out what the app is already driving, but it matched only the
import cursor -- which exists solely for imported sessions. Every session
the app spawned therefore stayed in the list, marked in use, telling the
reader to go and close it somewhere: here. Matching the resume token too,
which both kinds have, is the fix; `session_importing` is now
`session_driving`, because that is what it was always being asked.

Verified on the emulator against a scratch backend: a spawned claude-cli
session reports keepsOwnTranscript true with imported false -- Bryan's
`manager` case exactly -- and draws the recoverable warning; the echo
session draws "can't be undone"; and once the CLI named itself, the
spawned session's id was absent from the import list, where the old match
would have listed it.

75 tests, clippy and rustfmt clean; ktfmt, compileDebugKotlin and
lintDebug clean.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VETa8afmpWaYezLCqJhDB8
2026-08-29 20:34:03 -04:00
irisandClaude Opus 5 5f6fd34aba Record the clear that happened, not the one that was asked for
`ClaudeDriver::clear` emitted `Event::Cleared` beside the `/clear` it
sent, so the divider recorded a request. A reader scrolling back takes
that mark as a fact about the conversation -- the session no longer has
what is above this -- and a request is a different claim from a result.
Compaction already gets this right by taking its mark from the CLI's own
`compact_boundary` rather than from somebody pressing Compact; this is
the same rule, and it was the one place left applying it to the request.

Caught in review by the session this was measured against, which also
established that the CLI's `/clear` is declared `supportsNonInteractive`
and returns empty text with no result line -- so a fresh `init` bearing
a different `session_id` is the only trace it leaves. The reader already
watches for exactly that in order to persist the resume token, so the
mark now goes out there.

It has to be a *replacement* rather than any change, and the tests pin
both ways of getting that wrong. The first `init` sets the id from
nothing, which would otherwise open every session with a divider
announcing a clear that never happened. And a compaction re-announces
`init` carrying the *same* id, which would otherwise draw a clear on top
of the compaction's own mark -- that one was found by writing the test
rather than by reasoning about the change.

73 tests, clippy clean, rustfmt clean.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VETa8afmpWaYezLCqJhDB8
2026-08-29 20:20:49 -04:00
irisandClaude Opus 5 cf3ed68511 Tell "nothing to meter" apart from "couldn't find out"
Found by looking at the bar on an echo session rather than by reading
the diff: it said "5-hour usage unknown -- this machine reports no
usage", which is the failure the rest of this file was written to avoid,
one level up.

A machine with no metered provider is never asked by the backend, so it
returns no snapshot for it. The bar read that silence as a failed
lookup, because Unavailable was the nearest word it had -- and a session
on `echo`, or on a local llama.cpp, has no paid quota at all. That is a
fact about how somebody set the machine up, not a question that went
unanswered, and reporting it as unknown nags about a deliberate choice
on every screen forever.

So the state exists now: NotMetered, drawn as nothing, because there is
nothing. Unavailable keeps its words and its reason and still covers the
three ways an answer can fail -- nobody logged in, machine unreachable,
snapshot without the window.

Verified on the emulator against the real endpoint: a setup carrying
claude-cli draws the bar at 22% of 5h, selected by kind "session"; the
no-snapshot path was the one on screen before this change, so it is
reached, and this only changes what it draws.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VETa8afmpWaYezLCqJhDB8
2026-08-29 20:16:31 -04:00
irisandClaude Opus 5 6680fbb987 Clear from the phone, and put the numbers where they are read
Four changes to the session screen, three of them Bryan's and one that
fell out of them.

**/clear is offered like any other session command.** It joins
SESSION_COMMANDS, so it suggests itself while being typed and goes out
through the command endpoint that /compact already uses -- no new path,
and the boundary pump holds it mid-turn exactly as it holds a compaction.

**A clear draws a divider, not a deletion.** `Event::Cleared` becomes a
ClearedNote row saying that everything above stays here and is no longer
sent. That sentence is the row's whole job: the reader can see the
conversation is still on screen, so without it the divider reads as
something having been thrown away, which is the one thing it is not. It
carries no counts, because nothing was measured -- a compaction's
numbers are real and there is no equivalent here to report.

Compaction and clear now share `TranscriptDivider`. They are the same
kind of mark to somebody scrolling back -- "the session no longer has
what is above this" -- and the difference belongs in the words rather
than in how they are drawn, so the styling is written once and cannot
drift.

**The five-hour usage bar sits under the session header.** It reports
the paid service's own metering for the machine this session runs on,
fetched from that machine, refreshed every minute off the backend's
cache. It is never derived from the transcript's token counts: those are
a different quantity measured differently, and a quota-shaped bar built
out of them would be a guess wearing a measurement's clothes. Not
knowing has its own appearance and its own words -- "unknown" and why --
because a bar resting at zero because a machine is unreachable reads as
plenty of headroom, which is the opposite of the truth. The window is
selected by the API's own `kind` ("session"), added to UsageWindow in
this change, rather than by matching the label a person reads.

**The token total moved from the header to the bottom right of the
transcript**, pinned above the input rather than scrolling with it. In
the header it was one item in a run of dot-separated facts about the
session and read as another of them, rather than as the running total it
is.

SessionSummary now carries the setup id, which it deliberately did not.
The stated reason was that nothing here addressed a setup and holding
both id and name invited showing the wrong one; the usage bar addresses
one, so the reason lapsed rather than being overruled, and the comment
now carries the rule that replaces it: never display it. The server has
always sent the field, so nothing changed on the wire.

ktfmt, compileDebugKotlin and lintDebug all clean.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VETa8afmpWaYezLCqJhDB8
2026-08-29 20:11:26 -04:00
irisandClaude Opus 5 1e3eec8cf0 Let a session drop its context without ending
Adds `Event::Cleared`, `SessionCommand::Clear`, and `Driver::clear`, so
`POST /sessions/{id}/command {"text": "/clear"}` does for a session what
the CLI's own `/clear` does for a terminal.

The marker is a divider, not a truncation: everything above it stays in
the transcript, because that is the only copy of the conversation the
phone has and a person scrolling back is a different question from what
the model is given. It also makes clearing mean one thing across
drivers -- `claude` sends `/clear` and the CLI answers with a fresh
`init` whose new session_id the reader already persists as the resume
token, so the next launch resumes the cleared conversation with nothing
to keep in step; `llama` needs no state at all, since `conversation()`
already folds the transcript and now folds from the last marker; `echo`
emits the marker alone, so the phone's divider and scroll behaviour can
be exercised without spending a real session's context.

That fold is why `Cleared` is documented as load-bearing rather than
decorative. For any driver that rebuilds its conversation from the
transcript, this marker decides what the model sees, and treating it as
something only the phone draws would silently put the cleared
conversation back in front of the model at full price.

Clear rides the existing boundary pump like any other SessionCommand, so
one arriving mid-turn waits exactly as a compaction does, and nothing
grows a second way to wait.

Removes `--autocompact` in the same change, because clearing is the
cheaper answer to the problem it was added for and Bryan would rather
manage context that way. Keeping the measurement here, since it was the
reason for the constant and is worth more than the constant was:
context returned to 70-85k within ten calls of a compaction; a
compaction took 104,346 to 147,671 ms; compaction cost that session
2,655,508 tokens across six boundaries, of which the single automatic
one at the 1M ceiling was 1,696,870. Clearing costs nothing, because
nothing is sent.

70 tests, clippy clean, rustfmt clean.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VETa8afmpWaYezLCqJhDB8
2026-08-29 20:06:04 -04:00
irisandClaude Opus 5 797513bb86 Raise the compaction window to 200k
100k is the cheapest window on tokens and the wrong one to sit in front
of. Measured on the session this was written against: context returns to
70-85k within ten calls of a compaction, so a 100k window compacts about
every thirteen calls, and a compaction takes roughly two minutes
(durationMs 104,346 to 147,671 across the six recorded). A 130-call
request would have spent some twenty minutes compacting -- optimising
the number that was asked about while making the thing somebody actually
waits for on a phone considerably worse.

200k keeps most of the saving against the 1M ceiling and halves the
stalls.

The comment now also says what the window does not do, because measuring
this turned up the opposite of what the byte counts suggested. Images are
93% of the bytes that tool calls put into that transcript but only 8% of
the context growth -- the adb wrapper's downscaling holds a screenshot to
a median of 476 tokens, while text-only calls add a median of 740 and a
mean of 1,139. So the file is large because of screenshots and the
context is large because of ordinary tool output, and only the second one
is what this constant governs.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VETa8afmpWaYezLCqJhDB8
2026-08-29 19:43:43 -04:00
irisandClaude Opus 5 4f586d6551 Attribute the drift to the CLI's ceiling, not to a manual compaction
The reasoning on AUTOCOMPACT_WINDOW cited 491,562 tokens as where "the
CLI compacted it". That was a manual /compact somebody ran, not the
CLI's own trigger, so the comment credited a person's intervention to
the automatic behaviour it was arguing about. Caught in review by the
session whose transcript it was measured from.

Corrected from that transcript's compaction boundaries: the window left
to `auto` was 1M, and the one automatic compaction fired at preTokens
1,000,184 with the API context peaking at 999,668. So the drift ceiling
is twice what the comment said, and near it a single tool call bills
about 100k tokens rather than 49k.

The correction strengthens the case, but it also changes what the
example is evidence *of*, which is why it was worth fixing rather than
just raising the number: what held that session together was the person
in it running /compact by hand four times, and the 4.2-million-token
request happened at the merely-large contexts left between those. The
constant is for the sessions where nobody is doing that.

No behaviour change. 68 tests, clippy clean, rustfmt clean.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VETa8afmpWaYezLCqJhDB8
2026-08-29 19:11:51 -04:00
irisandClaude Opus 5 36984a9e1b Put the compaction window's reasoning on the compaction window
4a40578 inserted AUTOCOMPACT_WINDOW between an existing doc comment and
the constants it described, so rustdoc attached "the session directory's
copies of the process's standard streams" to the compaction window and
left STDIN_FIFO, STDOUT_LOG and STDERR_LOG undocumented. Moving the new
constant below them restores both.

Verified: 68 tests, clippy clean, rustfmt clean.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-29 19:10:39 -04:00
irisandClaude Opus 5 4a4057886d Compact at 100k rather than letting a phone session drift
Left at `auto` the CLI picks a very large window, which suits a terminal
session somebody closes at the end of the day and does not suit this app
at all: these run for hours, nobody closes them, and the transcript
carries screenshots. One session here reached 491,562 tokens of context
before the CLI compacted it.

That matters because every API call re-reads the whole context, and one
request is not one call. At half a million tokens a single tool call
bills about 49k before it does anything, so "can you make it so you can
rename a session?" cost 4.2 million tokens across the 130 calls it took.
Measured over that session's life: 2,498 calls, 1.08 billion cache-read
tokens.

100k is the smallest window the CLI accepts and roughly the cheapest.
Per-call cost falls with the cap, while the compaction it forces costs
about the same in total either way -- a smaller window compacts more
often, but each pass is proportionally smaller. What it trades is how
much detail survives a compaction, which is a real cost to the work and
the reason this is one named constant with the reasoning written down
rather than a computed value.

Passed before the resume/name branch, so it applies to adopted and
imported sessions too -- which are the large ones, and the ones this is
for.

Verified: the exact argument list the app now spawns starts, accepts an
empty stream-json stdin and exits 0, so the flag combination is good
without spending a token. 68 tests, clippy clean, rustfmt clean.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VETa8afmpWaYezLCqJhDB8
2026-08-29 19:09:27 -04:00
irisandClaude Opus 5 620a7d0a83 Trim what every session pays to load
Measured: an app-spawned session starts at ~33,200 tokens of context, of
which ~20,000 is written fresh on every spawn -- the always-loaded rule
files and this file -- and only ~13,200 comes from a shared cache. That
20,000 is billed at 1.25x on every single session start.

This file drops to 19,882 bytes from 21,293. What went is narrative that
PLAN.md already carries in more detail (the phase history, the submodule
drift story) and the parts of "Where things run" that MACHINE.md states
once for every project. What stayed is every operational fact: the
commands, the llama.cpp and ssh test recipes, the import rules, and
everything under "Things that have bitten".

The global chain was trimmed in the same pass, 43,039 -> 34,069 bytes,
mostly by moving the Gentoo host build profile out of the @import chain
into ~/.claude/HOST_BUILD.md, which MACHINE.md now points at. Nothing was
deleted there either; it is referenced rather than loaded, the same
arrangement this file has with PLAN.md.

Worth being honest about the size of the win: ~10,400 bytes is roughly
2,200 tokens off each session start. It is real and permanent, but it is
not what makes a long session expensive.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VETa8afmpWaYezLCqJhDB8
2026-08-29 19:00:59 -04:00
iris 749b2db287 Make a command a thing the app knows, and hold it until it can run
Typing "/" now suggests what this app understands -- `/compact` and
`/rename <name>` -- with a line each about what they do, and anything
else beginning with a slash is passed to whatever runs the session,
because a dialect's own vocabulary grows without this list.

None of them are messages, and that is the substance of the change. A
line written into a running turn is read by the *model*, so a command
sent mid-turn either does nothing or arrives as text somebody has to
puzzle over. They now wait for the turn to end. The waiting is done once
for every provider, in the pump that already watches every event for the
boundary, rather than in each driver where a new provider could get it
wrong by leaving it out.

Waiting is a state, so it is on screen: the command sits at the reader's
end of the conversation in blue, with a spinner and "waiting for this
turn to end", and becomes an ordinary blue row when it goes. Blue
because these are about the session rather than about the task -- the
same blue a compaction already used, which is now one colour with one
name rather than two.

Renaming from the settings screen sends exactly this, so it waits and
draws the same way. The name itself is not held: it is this server's own
datum, so the list and the header change at once and only telling the
session waits.

Echo grew the same split, which is where the bug in it showed: its
commands are its messages, so running one announced a `MessageTaken` as
well, and the same line drew twice -- once blue, once purple. A command
owes no announcement; the manager has already recorded that it was sent.

Watched rather than reasoned about: `/compact` during a 25 second turn
held with its bubble up, went out when the turn ended, and the
compaction that followed reported what it recovered.
2026-08-29 17:10:21 -04:00
iris bebaae7a94 Carry a question in the event model, not in one provider's JSON
A question is now fully described by the event that reports it: the tag
it was asked under, each option's label, what it means, and the sample of
what picking it would produce, plus whether several may be picked at
once. The app renders from that alone.

It had been reading Claude Code's tool input to find the parts the event
dropped -- that dialect's schema, written out a second time in Kotlin,
where no other provider could reach it and where it would drift the
first time the schema moved. Echo could not describe an option at all,
and llama never will.

Answers travel as a list for the same reason. A question that takes one
answer sends a list of one rather than being a different shape, and the
one place that flattens it is where the CLI is spoken to: its answers
map holds a string, so several choices are joined there. That join was
in the phone.

Also here because it is the same rule: the permission ask reuses the
question body rather than owning a second one, so Allow/Deny renders and
resolves through exactly the code an AskUserQuestion does.

Verified against both, since a refactor that only satisfies the case it
was written for has been tried on the half that cannot fail: a two
question `/ask` answered from the phone, one option and then two, and a
real sonnet session's `rm -f` permission asked, allowed, and run.
2026-08-29 16:46:43 -04:00
iris fea8e7e92b Show every option a question offers, on the call that asked
Reported by Iris through the dev-updater session: a two-question
AskUserQuestion arrived with only one option visible per question, so
the answer she sent was the only one she had been offered.

The cause was a `Row`. It hands out intrinsic widths in order and clips
whatever runs past the edge, so the first option or two drew and the
rest went off the side of the screen -- which does not read as a bug, it
reads as those having been the only choices. The same Row was in the
permission ask beside it; both wrap now. That pairing is the reason to
look: a rule stated on one member of a set is usually missing from the
others.

The rest of what she asked for, and what each was:

- It drew twice, as the tool call and again as loose question cards,
  because the backend marked these questions as belonging to no call.
  They belong to the call that asked, and now say so.
- So it renders like any other tool: one card, its own heading, opened
  because a decision cannot be made from a closed row.
- Each option shows its description and its `preview` block, which is
  the part a reader is deciding on and none of which was reaching them.
- "Other" is a field on every question. The harness always offers it, so
  leaving it out narrowed a question that was never that narrow.
- A multi-select sends the labels it collected as one string, which is
  the tool's own schema rather than a guess -- its answers map is
  string-valued.
- No spinner while it waits. A spinner says the machine is working; here
  the machine is idle and the turn is stopped on the person, so the card
  says "your turn" in the colour this app already uses for that.

Verified against a real session as well as the echo fixture: haiku asked
two questions with three described options each, both were answered from
the phone, and the model carried on with the answers. Echo grew `/ask`
so the shape can be looked at without paying a model to produce one, and
its option cards are outlined rather than tinted -- as one surface step
up they were three paragraphs where three things to press should be.
2026-08-29 16:22:15 -04:00
iris cae04c2559 Delete one session without putting the rest through loading
Pressing Delete refetched the whole list on success, so every other row
went back through its loading state and the reader got a blank screen
for the length of a round trip -- to report on something that was never
in doubt. Now the row being deleted fades, says so where its status
goes, and stops responding to taps; when the server answers, that one
row is removed and nothing else moves.

A refusal keeps the row, because it is still there: the server answered
and said no, so the session it said no about is exactly as it was, and
the error goes on its own card as it already did.

Faded rather than removed on the way out, deliberately. Taking the row
away when Delete is pressed is a promise about a request that has not
been answered, and putting it back when the server refuses is worse than
never having taken it away.

Looked at rather than reasoned about: the in-flight state lasts
milliseconds against a local server, so I slowed the delete route to
four seconds, watched the faded row and its spinner, watched it removed
on success, then killed the server and watched a refusal leave the row
in place with the reason on it.
2026-08-29 16:11:18 -04:00
iris d3fff3d229 Let a session be renamed, under the same name everywhere
A gear at the end of the session's own bar opens what can be changed
about that session; the name is the first thing there. Compact is gone
from that bar -- `/compact` typed into the message box is the CLI's own
way to ask and it already worked, so the button was a second way to say
one thing. Echo takes the typed word too now, since it is the rig the
compaction display is checked against and losing the button would have
taken that with it.

The name is this server's, not a driver's: it is what the list shows, it
exists before any process does, and every provider has one. So it is
settled in the config and the driver is *told* -- which is the opposite
of the model and the permission mode, and the difference is written down
at `Driver::set_title`. A driver whose process has no notion of a name
does nothing and says nothing, because there is no failure to report.

Claude Code has one, so the name reaches it: `--name` for a session we
create, and `/rename` afterwards, which is a local command rather than a
control request -- `set_session_name` is not a subtype it knows, which I
established by asking it. A resumed session is deliberately not renamed
at launch: an import already has a name, quite possibly one the person
typing in it chose, and taking that would be helping itself to something
the app was only shown.

Verified end to end rather than argued: renaming from the phone put
"Session renamed to: paging and scroll" in the CLI's own session file,
and the session now lists under that name to other agents.

The gear is drawn rather than set in a font, for the reason Chevron
gives. It was a sun on the first attempt -- thin teeth standing clear of
a thin hub -- which no amount of reading the diff would have shown.
2026-08-29 16:06:15 -04:00
iris 1629e0911e Say what the line splitter would do with a bare carriage return
`complete_lines` splits on `\n` only, which is right -- this stream is
JSONL, and a record terminated by a bare `\r` would not be a record --
but the doc comment said why the remainder is held without saying what
decides where a line ends.

Worth the sentence because of what the failure would look like if the
CLI ever wrote such a line: the session goes quiet, the process is
healthy, nothing errors, and the cause is a line splitter. The
dev-updater session hit exactly this shape today reading cargo's
progress line, which is `\r`-terminated for redrawing in place, and
lost a whole build's worth of output to it.
2026-08-29 15:42:32 -04:00
irisandClaude Opus 5 f842d0e512 Call the server component "server", to match dev-updater
READ BEFORE PULLING. A Managed component's service unit is named
<config key>-<component name>, so this rename moves the unit from
ai-app-backend to ai-app-server and nothing points at the old one
afterwards. Uninstall the backend component from its card *first*, while
it is still called "backend"; then pull, accept the new declaration --
.dev-updater.ron is a request, so the card shows it as pending -- and
build. The unit installs under the new name.

Two things reset rather than break, both keyed by component name: the
per-component built_from sha, and the build and runtime logs. One build
makes the sha current again.

Also drops the claim that the components list is walked in order. They
have built in parallel since 2026-08-28, so the reasoning the comment
gave -- backend first, so a failing APK leaves the phone what it had --
no longer describes what happens.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VETa8afmpWaYezLCqJhDB8
2026-08-29 15:40:31 -04:00
iris 3eccf7e443 Show the model and mode the session has, not the ones it was asked for
Picking either from the phone wrote the choice straight into the
session's state and then sent the request. Asking and having are
different things, and the difference is not rare: `auto` is a permission
mode the CLI accepts on the command line, silently resolves to
`default`, and refuses outright over the control channel -- "auto mode
unavailable for this model" -- so a session spawned in auto was in
default and one switched to auto stayed where it was, with the phone
reporting auto in both cases.

So the drivers report what they are set to and the manager follows that.
Measured, because the confirmations are not uniform: a model change
answers success with no value, so what was asked is remembered until the
answer arrives; a mode change echoes the mode it became, and that answer
wins over the request; and `init` names both -- resolving `haiku` to
claude-haiku-4-5-20251001 -- which also covers a session adopted from a
terminal that set them outside this app. A driver that cannot change
either already says so with an error, and now that error is the whole
story rather than a note beside a display that changed anyway.

The config keeps the requested value, deliberately: that answers a
different question, which is what to launch this session with next time.

Two things fall out. Control request ids are random rather than the
clock, because two in the same second shared an id and something now
looks them up. And the phone shortens a resolved name for the button --
`haiku-4-5` -- since the full one is what the CLI reports and roughly
twice the room that row has once Stop is in it.
2026-08-29 15:36:53 -04:00
iris 404066fa7d Say when a session is working, and what it was told
Three things a phone could not see, all of them the same shape: the
session was doing something and nothing on screen said so.

A turn nobody here started never reported itself. `Running` was sent
where a message was *sent*, so a session picked up mid-turn, one
compacting on its own, or one another agent wrote to sat there reading
as idle until it finished. The driver now says it from what it observes
-- output that could only come from a turn in flight -- which is the
same set of events that already announced a steer, with the ends
swapped.

An imported session had it worse: nothing but replayed lines ever
reaches it, and a status was not among them, so it was permanently
whatever it was when it was adopted. Its file does not record a turn
ending, but it does record why each assistant message stopped, and
`tool_use` versus anything else answers it. A record that says nothing
leaves the status alone rather than voting for idle.

Messages from other agents were dropped outright: the CLI marks them
meta, and this replayed everything except meta. They are now a row of
their own, closed by default like a tool call, named for the session
that sent it -- not the reader's own bubble, because they did not say
it, and a session working on something this phone never asked for is
exactly what one of these explains.

Measured against a real session file rather than guessed: the peer
record carries the sender's name and the message body in `origin`,
beside a copy wrapped for the model to read.
2026-08-29 15:27:46 -04:00
iris f18639e4b1 Count the far end of the list in rows, not events
Scrolling back stopped dead at the top of what was loaded, and no older
page ever arrived. Bryan spotted the cause from the outside: it had to do
with tool calls being collapsed.

The trigger compared an index into the list being drawn against
`items.size`, the number of transcript events. Those were the same number
when it was written. They stopped being the same when adjacent tool calls
started folding into one row, and the queued bubble and the working
indicator are two more rows with no event behind them. In this session
645 rows stood in for 720 events, so the last visible index could reach
646 and the threshold it needed was 717. It was not close; it was
unreachable, and the further a session went the worse it got.

Both numbers now come from the list itself, which is the only place they
are commensurable, and `totalItemsCount` counts whatever gets added to it
next.

Checked on the emulator against the case it was breaking on rather than a
clean one: five collapsed "Called 8 tools" groups in front of a 720-event
transcript, scrolled from the bottom to seq 1, which is the beginning of
the session. It stops there because that is the top, and holds position
while each page arrives.
2026-08-29 14:55:19 -04:00
irisandClaude Opus 5 fc71cb4403 Say "unknown" for a session we are not driving but cannot bury
A session in the config with no live entry reported `Exited`, whatever the
reason. That covers three different situations -- one that genuinely
ended, one that failed to relaunch, and one whose process could not be
checked -- and the wrong one is the expensive one.

`Exited` reads as "this conversation is over", and what a reader does about
it is start a fresh session. If the process is in fact still running, that
is a second CLI against a conversation that already has one: the exact
fault `session::process` exists to prevent, arriving through the status
field instead of through a spawn.

So it is said only when the process is known to be gone. A record that
cannot be checked reports `Unknown`, and so does one that is still alive --
this server is not driving it, so it genuinely does not know what that
process is doing, and the honest word is the one meaning "wait" rather than
the one meaning "act". A session with no record at all is still `Exited`:
an echo session, or one already stopped and cleaned up, and known to be.

The distinction was available all along -- `process::recorded` returns the
liveness -- which makes this the same mistake as the other five today:
reporting what was convenient to compute rather than what was measured.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VETa8afmpWaYezLCqJhDB8
2026-08-29 14:50:20 -04:00
iris 279c76e8a1 Hold the reader's place when a message arrives
Scrolled back through a conversation, every new message dragged the view
with it -- which reads as the screen scrolling down on its own, at the
exact moment somebody is trying to read something else.

The list had no keys, so its rows were identified by position. The
transcript is drawn newest-first, so a new message is an insertion at
index 0: every existing row shifts up one index, the viewport stays on
the index it was on, and the content slides through it. The working
indicator appearing and disappearing did the same thing at the same end.

So rows now carry the identity they always had in the data. Every
transcript event has a seq, which is what the transcript is ordered by
and never changes, and every row keeps the seq of the first event behind
it -- a streaming message keeps the seq of its first delta, so it holds
still for the whole answer rather than becoming a new row on every
frame, and a tool call keeps its start's.

Paging older history is the same insertion from the other end, and it is
the thing this could plausibly have broken. Checked on the emulator:
scrolled back mid-turn, the view sat still through twenty seconds of
streamed deltas, and scrolling to the far end still fetched earlier
pages and stayed where it was while they arrived.
2026-08-29 14:47:31 -04:00
iris 5396da76c7 Show a compaction happening, and what it recovered
The Compacting status had been declared, rendered in four places, and
never once emitted: no driver produced it, and the app had no control to
ask for a compaction in the first place. Pressing nothing for two
minutes and then quietly having less context was the whole experience.

The CLI turns out to announce all of it, which was worth measuring
rather than guessing at. Driven through /compact against 2.1.237 it
emits a `status: "compacting"` line at the start, a `status: null`
carrying `compact_result` at the end -- `"failed"` with a sentence
saying why, when it does -- and then a `compact_boundary` with the token
counts. The same records appear in the CLI's own transcript file with
camelCase keys, which is the obvious place to read the shape off and
gets every field name wrong.

So none of it is inferred here. The driver writes the line and says
nothing; the translator reports what the CLI reports. A failed
compaction surfaces the CLI's own sentence, which is specific enough to
act on.

The counts are the part worth keeping afterwards, so they land in the
transcript rather than only in a status that vanishes: a session that
went from 128,402 tokens to 9,617 has just been given its context back.
They are optional throughout, because a compaction whose size nobody
reported has to be able to say so -- a zero would read as "recovered
nothing".

Also here, all found on the way:

- `rename_all` renames variants; fields need `rename_all_fields`. Every
  field in Event was a single word until `pre_tokens`, which went out as
  snake_case, was not found by the app, and rendered as the "no counts
  reported" case -- a state it is allowed to be in, so nothing looked
  wrong. There is now a test on the wire names.
- The unparseable-line warning sliced bytes, not chars, on output that
  is full of em dashes. A panic there kills the task reading the
  session's stdout, and the session goes deaf with nothing on screen.
  The other three truncations in the tree already did this correctly.
- Echo compacts too, with invented numbers and a real shape, so this
  screen can be looked at without spending two minutes of somebody's
  account to reach the state.
2026-08-29 14:42:08 -04:00
irisandClaude Opus 5 42131c75d6 Send a steer into the running turn, and put an image under its call
**The queue was holding messages the CLI would have taken.** Two claims
in this file contradicted each other: the module header said a mid-turn
message is injected at the next tool boundary -- "the behavior this app
exists for" -- and `Queue`'s own doc said a line written mid-turn simply
becomes the next turn. The code followed the second, parking every
message until `Status::Idle`, which is the end of the whole turn.

Measured rather than argued, twice. Writing a line straight into a live
session's stdin fifo mid-turn produced one `result` for the whole thing,
so it was consumed inside that turn, not as a new one. The header was
right and the queue was built on the wrong claim.

The cost was exactly what Bryan reported: he steered after the second
tool call and it sat unread until every remaining call had finished.
Measured before and after on the same three-step turn -- steer sent at
+13s, recorded at +24.7s before this change and at +14.1s after, which
is the next tool boundary.

So the line goes out immediately. What stays behind is the
*announcement*: the CLI says nothing on stdout about having read a
message, so `MessageTaken` now waits for the next assistant text or tool
call, which is proof another model call happened and the steer was in
it. That keeps a held message drawn below the working indicator until
the session has actually taken it -- the thing that mattered when this
was last changed -- without delaying the message to get it. Idle counts
too, and is the case that must not be missed: a message written after a
turn's last model call has no later output to prove anything.

`closed` is untouched, and `Queue::close` still reports held messages by
name rather than dropping them.

**An image now names the call that produced it.** `Event::Image` gains
`about`, the `tool_use_id` from the tool result it came out of, so a
screenshot is drawn inside that call's card instead of floating beside
it -- pairing them by position is what a page boundary breaks. `None`
for a person's own attachment, which belongs to no call. The import path
threads it through as well, so replayed history reads the same as live.
Images show whether the card is open or closed: a call whose result *is*
a picture says less closed than the one line it replaced.

Verified on the emulator against a real haiku turn: the checkerboard sits
inside `Read /tmp/tiny.png`, and the steer sits between that call and the
next, where it was taken. 53 tests, clippy, rustfmt, lint and ktfmt clean.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-29 14:17:22 -04:00
irisandClaude Opus 5 184b6fc6a6 Correct the claim that reattach is local only
Written down as "an ssh session's child dies with its connection, so it
takes the ordinary --resume path". The code never had that branch: `start`
records a pid whatever the transport, and for a remote session the process
the backend owns is the ssh client. Adopting it is right -- the fifo feeds
it, its logs capture the far end, and ssh lives exactly as long as the
remote command, so its liveness is the session's.

The docs claimed less than the code does, which is the safe direction to be
wrong in but still wrong, and it was about to mislead someone: a remote
`claude` has an sshd pipe on stdin under every version of this server,
because the fifo is on the backend's side of the connection. Reading a
remote session's stdin therefore says nothing about which backend started
it, and we were an inch from concluding otherwise.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VETa8afmpWaYezLCqJhDB8
2026-08-29 13:57:12 -04:00
irisandClaude Opus 5 4cbd567c35 Take ktfmt's formatting
Committed unformatted again: I piped the check through grep, so the
task's failure never reached the shell's exit status and the chained
commit ran anyway. Check exit codes, not output.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-29 13:46:34 -04:00
irisandClaude Opus 5 73923dbbc3 Make jump-to-latest the same chevron, pointing down
It was a labelled button beside a tool group that collapses with a
drawn chevron -- two controls doing the same kind of thing in two
visual languages. Now one `Chevron` composable serves both directions,
parameterised rather than copied, since a pair that differs by a minus
sign drifts and the drift is a bug in exactly one direction.

The comment it replaces argued against an arrow here, on the grounds
that the list is laid out upside down. That reasoning was about the
code: nobody reading the screen knows the list is reversed, and on
screen the newest message is at the bottom, which is where this goes.

It draws no text, so the name lives in its content description -- the
whole of what a screen reader has, and the answer to "what was that
arrow for" later.

Verified on the emulator: scrolled up, the chevron appears bottom
centre matching the group's; tapped, it returns to the newest message
and takes itself away.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-29 13:46:13 -04:00
irisandClaude Opus 5 011ed0d0e1 Hold an image's place, open it full screen, and fold a run of calls
Five changes to how a transcript reads.

**Images no longer move the page.** The row was as tall as whatever had
loaded, so it grew when the bytes arrived and pushed everything below it
-- and in a bottom-anchored list, an image loading above the viewport
moved the text under the reader's eyes. The height is now decided before
the fetch and never changes: four lines of the body style, measured from
the type so it stays four lines when the reader has scaled their fonts.
Nothing to see when loading finishes, which is the point.

**A small image is enlarged with nearest neighbour**, a large one shrunk
smoothly -- decided per image from its actual size rather than set once,
since blowing a 16px sprite up with interpolation turns it into a blur
of exactly the thing being looked at.

**Tapping one opens it full screen**, fitted so the whole image is
visible first, with two-finger zoom to 8x and pan once zoomed. A dialog
rather than a screen, so back returns to the transcript.

**A tool call is one line closed**: the tool's name and what the call is
for. The command is not on it, because a wrapped command turns one row
into four. Open, it shows the command, the rest of the input and the
output, with the timeout at the top right -- a limit on the call rather
than part of what it does, worth seeing beside the command it constrains.
A call waiting on permission is shown open regardless, since the command
is the thing being decided.

**Adjacent calls fold into "Called n tools"**, closed by default, and it
closes again from either end -- a long group's heading scrolls away while
its last call is still on screen, and the reader who wants it shut is
looking at the bottom. The calls keep their full width; what says they
belong together is the surface behind them, one cue rather than two
half-cues. Grouping happens at display time, not in the fold: the
transcript's own order is what paging and the stream depend on.

Echo gains `/tools [n]` so a run of calls can be produced without paying
for one.

Verified on the emulator: four calls folded and expanded, one opened
inside the group showing `timeout 5000` top right, a 16px checkerboard
enlarged with hard pixel edges beside a shrunk screenshot at the same
height, the screen byte-identical between one second and six after
opening, full screen fitted, and back returning to the same scroll
position. Pinch itself is the one thing not verified here -- `adb input`
cannot inject a two-finger gesture.

53 tests, clippy, rustfmt, Android lint and ktfmt clean.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-29 13:43:46 -04:00
irisandClaude Opus 5 8fe13634cb Take ktfmt's formatting on the files just added
Four files went in unformatted: I ran the formatter mid-change and then
kept editing. ktfmtCheck is part of finishing, not part of starting.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-29 06:28:24 -04:00
irisandClaude Opus 5 d580461bd5 Show what was remembered as a note, not as markup
Claude Code marks a sentence taken from its stored memory by wrapping it
in `<cc-memory filenames="...">`. Markdown has nothing to say about that,
so it arrived as literal angle brackets mid-sentence and read as the
model having emitted broken HTML. It is the opposite: a claim about
where something came from, and "I was told this before" and "I worked
this out just now" are different things the reader could not otherwise
tell apart.

Each one becomes a card naming the files it came from, with the prose
either side of it left as prose. Named rather than merely tinted, since
a colour can say "this one is different" but not what kind of different.

A tag still arriving is left alone: streaming means the closing half may
be seconds away, and a half-written marker is not a marker yet.

Verified on the emulator with two notes in one message, one of them
citing two files, and prose before, between and after.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-29 06:27:00 -04:00
irisandClaude Opus 5 c5dbd1535d Ask for permission on the call it is about, and read the input
A bash permission request arrived as a second card repeating the tool
call's input verbatim, so the same command appeared twice and the reader
had to work out it was one event. `Event::Question` now carries `about`:
the `tool_use_id` the CLI's `can_use_tool` request already names. That
makes the pairing a measured fact rather than a match on input text --
and it stays `Option`, because AskUserQuestion is not permission for
anything and an echo session's question is about no tool at all. Those
still draw as their own card, which is what every question did before.

The card also reads the input instead of dumping it. Every tool's input
is JSON, and showing it raw makes the reader parse `{"command":"…",
"timeout":5000}` to find the line they care about. A small table says
which field is the subject of which tool -- Bash's `command`, Read's
`file_path` -- and the rest is still listed, since dropping a field
would claim the tool had no other input when it might. The subject is
syntax-highlighted with dev.snipme:highlights, for the reason the
markdown renderer is a library: lexical rules are somebody else's
specification. Its theme is Catppuccin, mapped in Theme.kt beside the
rest of the palette rather than taken from the library's defaults.

The input shows whether or not the card is expanded. A row that says
only "Bash" says nothing anyone can act on, least of all when it is
asking to run something.

Verified on the emulator against a real haiku session: one card, the
description, `grep -rn "needle" /tmp | head -3` highlighted, `timeout:
5000` pulled out, and "Allow Bash?" with its buttons inside the card --
then Allow, which resolved in place and ran.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-29 06:24:05 -04:00
irisandClaude Opus 5 206319045d Show usage per machine, and say why a machine has none
The server now reports limits per machine, so the screen has to as well:
one card per machine that offers a paid service, named by the machine
first, because these are one account's numbers and which account is decided
by which box ran the session.

It also has to say which of four things happened, and the reason for
splitting them shows up here rather than in the data. A machine nobody has
logged in on is working exactly as somebody set it up, so it reads as a
plain statement in ordinary text -- marking it would be the interface
nagging about a decision already made, and would dilute the marks that do
mean something. Only "couldn't reach it" and "the endpoint refused" are
coloured as faults, and they say different things because they need
different things done. The old screen drew all three in the error colour.

No machine offering a paid service is not an error either: it says so
instead of drawing nothing.

The app also stopped parsing: `available` no longer exists and
`getBoolean` on a missing key throws, so this had to land with the server
change rather than after it. An older backend sending no `state` is read as
"failed" rather than "ok", since an empty card drawn as healthy is the
worse failure.

Looked at running, against five machines: local reporting notLoggedIn with
the backend's HOME emptied, loopback-over-ssh returning real windows beside
it, an unreachable host showing ssh's own message in red, and a machine
with no Claude provider correctly absent.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VETa8afmpWaYezLCqJhDB8
2026-08-29 06:18:46 -04:00
irisandClaude Opus 5 6166b1f626 State the deny-unknown-fields rule where it governs all the bodies
It had landed appended to `SshRequest`'s doc comment, so a rule about
every request body in the module read as something about how to describe
a machine. Moved to the module doc beside the route table, where the set
it governs is what a reader is already looking at, and worded so a new
request body knows it is expected to carry the attribute too.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-29 06:14:32 -04:00
irisandClaude Opus 5 446fb92ba3 Refuse a request field the server does not know
A misspelled field was accepted and dropped. Sending `permission_mode`
instead of `permissionMode` produced a 200 and a session running in the
default permission mode -- so the caller's setting was gone, and nothing
anywhere said so. That is the expensive shape: indistinguishable from
success at the place you are looking. It cost an hour here, chasing a
"startup race" that was a key serde had silently discarded; with the name
spelled the way the API asks, a bypassPermissions session runs a `sleep`
loop with no prompt at all.

So every request body refuses unknown fields, not just the one that bit.
Axum's message names the offending field and lists what was expected,
which is the whole of what the caller needs.

Query strings are deliberately left permissive: a stale link carrying an
extra parameter is not a mistake worth failing a request over.

53 tests, clippy and rustfmt clean; verified against the running server
that the misspelling is now a 422 naming the field and the correct
spelling still spawns.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-29 06:13:18 -04:00
irisandClaude Opus 5 d8570f4d5a Ask each machine about its own limits, not this one about all of them
`ClaudeUsage` read `~/.claude/.credentials.json` on the machine running the
backend and asked the API about that account, once, globally. But a session
runs wherever its setup says, so the numbers on the usage screen belonged
to the backend's account rather than to the account that spent the tokens.

That is not a rounding error in the layout this project is aiming at.
`ai-server` belongs on the host; the host has no `claude` CLI at all and
the VM is a remote. So the screen would have reported "is Claude Code
logged in on this machine?" while every session ran fine on a machine whose
limits nobody could see. It looked correct only because backend and CLI
happen to be the same box today.

Usage is now per machine, asked through the same `Transport` the sessions
use -- `ssh host sh -c 'cat $HOME/...'` for a remote, unchanged for the
local one. `$HOME` is left for the far shell to expand, since a path built
here is this machine's home directory and over ssh that is somebody else's.
Machines with no Claude provider are not asked and get no row: they have no
Claude limits, and a row about them would be a fact about nothing.

The snapshot gains the states it could not say. `available` plus an `error`
string made three different situations look identical, and the one that
suffered was the harmless one: a machine nobody has logged in on is a
decision somebody made, with nothing to fix, and it read as broken.
`notLoggedIn`, `unreachable` and `failed` are now distinct, and which one a
failed read is gets decided in `why_no_credentials` rather than at the call
site.

Supporting changes: `ssh::command` builds a `std::process::Command` that
tokio converts from, so a blocking caller can use the one place that knows
what a correct ssh invocation is; `Transport::capture_blocking` is that
caller's door. The cache is keyed by machine and service rather than by
position, since the set is no longer fixed at startup -- a positional cache
would hand one machine's numbers to another the moment a setup was added.
A cached snapshot still picks up a rename immediately, because the name has
nothing to do with the poll interval.

Verified against a real ssh setup (loopback, per AGENTS.md) with five
machines, all four states seen: local `ok`, loopback-over-ssh `ok`,
unreachable host `unreachable` carrying ssh's own message, a claude-less
machine correctly absent, and -- with the backend's HOME emptied -- local
`notLoggedIn` while the ssh machine still reported real windows, which is
the production shape.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VETa8afmpWaYezLCqJhDB8
2026-08-29 06:12:17 -04:00
irisandClaude Opus 5 9f403cddab Render markdown, and stop wiping messages still waiting to be read
Two things, both about the transcript telling the truth about itself.

Markdown is rendered rather than shown as its source. The parsing is
mikepenz/multiplatform-markdown-renderer, not something written here:
markdown is somebody else's specification, and a hand-written subset of
one disagrees with it at the edges, which is where the bug reports come
from. `Markdown.kt` is only the mapping onto this app's palette, so code,
links and rules take the Catppuccin values the rest of the app uses
rather than the renderer's defaults.

The queued-message list was cleared wholesale whenever a turn ended. But
the backend holds a queue of its own and takes one message per turn, so a
turn ending is precisely the moment the *rest* are still waiting -- the
bubbles vanished while the messages were on their way, which reads as
everything after the first having been dropped. Now a held message
leaves the list exactly two ways: the session reads it, which arrives as
a UserMessage, or its send failed and there is nothing to wait for.

Measured first, because the report was that the backend dropped them:
three messages sent behind one long turn were all delivered in order
(ONE, TWO, THREE) against current main, so the loss was in the display.

Verified on the emulator: headings, emphasis, inline code, nested lists,
a quote bar, a fenced block, a rule and a link all render, and the three
queued messages sit through their turn.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-29 06:10:22 -04:00
iris 90a57ca7e9 Revert "Default a new session to bypassPermissions"
This reverts commit d81c9a7d65.
2026-08-29 05:58:17 -04:00
irisandClaude Opus 5 d81c9a7d65 Default a new session to bypassPermissions
Auto already allows most of what a session does, so the prompts it did
raise were mostly the interruption without the choice -- and on a phone
each one is a round trip to a question card. Measured rather than
assumed: a haiku session spawned this way ran a bare `echo` and a
`sleep` loop with no prompt at all.

The stricter modes stay one tap away in the same picker, which is where
a session that warrants one picks it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-29 05:57:59 -04:00
irisandClaude Opus 5 19e3525131 Say what continuing a session will cost, not just how big it is
The import list reported a file size, which predicts the wrong thing. Most
of a large transcript is history from before a compaction, and the model is
not given it again: of the 133 MB session behind the 2026-08-29 incident,
99% of the bytes sat before its last compaction summary.

So each row now carries the tokens the model was actually holding at the
last turn -- the input side of the most recent assistant message's usage,
prompt plus both cache figures, which the CLI records itself rather than
anything inferred from the file. The two disagree in exactly the way that
makes the size misleading. Measured on this machine: `ai-app` is an 80 MB
file with 150k of context, while `ai-app-backup` is 3 MB with 481k. The
smaller file is the more expensive one to continue.

Absent rather than zero when no turn has recorded usage, since a session
with no turns has no figure rather than a figure of none.

The row is three lines instead of one run of separators:

  path      cut at the head, keeping the tail, and the only thing here that
            is cut -- one long value with no natural break, where the lines
            below it are short enough to wrap
  stats     named, context, lines, size
  warning   only when there is one, in the warning colour

The warning gets its own line and its own colour because it differs in kind
from the stats rather than in degree: those describe the session, it says
whether taking it is safe at all. Colour makes it findable, the words make
it actionable -- "open somewhere else" and "we could not check" are not
distinguishable by shade.

Titles no longer ellipse either; they wrap.

Looked at on the emulator rather than reasoned about, including the states
that are not the default: a long path truncating, a row with no warning,
and a row with no context figure.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VETa8afmpWaYezLCqJhDB8
2026-08-29 05:54:05 -04:00
irisandClaude Opus 5 c1a468432d Read a session's transcript once at launch, not twice
Reading the last status back from the transcript -- added so a restart
stops claiming an exited session is idle -- walked the whole file a second
time, after `Transcript::open` had just walked it for the sequence number.
Both answers are wanted at the same moment by the same caller, so the cost
was paid per session at exactly the point a restart is trying to be quick.

`Transcript::open` now finds both in its one pass and reports the status it
saw. The free function goes; a transcript knowing what it last recorded is
where that belongs anyway.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VETa8afmpWaYezLCqJhDB8
2026-08-29 05:39:54 -04:00
irisandClaude Opus 5 3a74bd9c35 Make the process record survive a crash mid-write
Reviewing the reattach code found the fault it exists to prevent, sitting
in its own save point.

`process::write` used `fs::write`, which truncates before it fills. A crash
inside that window leaves no readable record -- and a missing record reads
as "nothing is running", which is the single answer that makes the next
launch start a *second* CLI against a conversation that already has one.
The window is not rare: the record is rewritten on every read that makes
progress, so many times a second while a turn is producing output.

Written to a neighbouring file and renamed over the real name now. The
rename is atomic, so a reader sees the whole old record or the whole new
one. That also makes the fixed-size padding pointless -- a rename replaces
the file rather than overwriting part of it -- so it goes.

Two more from the same pass:

- A failed read of the stdout log was logged and nothing else. The session
  then went deaf with nothing on screen: no more output, no error, a status
  that stayed wherever it was. It now says so, closes the queue rather than
  stranding messages in it, and reports `Unknown` -- not `Exited`, because
  the process may well still be running; what failed is this server's
  ability to hear it.
- Sizing the stderr log by reading it. `read_from` with a large offset
  answers "how long is it" by allocating the whole file first, which on a
  chatty process is a large pointless read on every reattach. `size_of`
  asks the filesystem.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VETa8afmpWaYezLCqJhDB8
2026-08-29 05:37:42 -04:00
irisandClaude Opus 5 aff2cb90e2 Stop reporting the app being switched away from as a failure
Backgrounding the app left "Lost the event stream
(SocketTimeoutException: null)" waiting at the top on return. Android
stops the activity, the socket dies with it, and the reconnect loop --
which kept running on a phone nobody was looking at -- recorded the
failure. Switching apps is a choice somebody made, not a fault to report.

Worse, it could not clear. `streamError` was reset when an event arrived,
so a session that reconnected and then sat idle displayed a connection
error it had already recovered from, indefinitely. That is the expensive
half: a stale failure is indistinguishable from a live one.

So the stream now runs only while the screen is at least STARTED, which
makes the drop a deliberate close rather than an error (EventStream
already distinguishes them), and resuming reconnects from the same
cursor. What takes a failure off the screen is `onOpen` -- the measured
moment the server accepted the connection -- rather than the first event
to follow it.

The message that does get shown leads with what will happen next rather
than with the exception's class name, which named nothing the reader
could act on.

lifecycle-runtime-compose is declared rather than inherited from
activity-compose, for the reason core-ktx already is: this code calls
repeatOnLifecycle and LocalLifecycleOwner directly now, and a transitive
could change under it. 2.11.0, the current stable.

Verified on the emulator against an idle session, which is the case the
old code could never clear: backgrounded 35s, returned, no banner -- and
a message sent afterwards arrived live, so the reconnect genuinely
reattached rather than merely staying quiet. Build, lint and ktfmt clean.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-29 05:36:04 -04:00
irisandClaude Opus 5 d7c692a4ec Show a replayed session's images instead of dropping them
An imported session showed no screenshots. `text_of` kept only `text`
blocks, so every image in the replayed tail was silently discarded -- while
the *live* translator has always saved them into the session's `files/` and
referenced them. Two readings of the same records, and the one used for
history was the lesser.

`save_image` moves out of `Translator` to a free function both paths call,
since the naming scheme for that directory should exist once. `events_from`
now takes the session directory to write into, which means the conversion
has to happen where that directory exists -- so `Seed` carries the raw
JSONL and `launch` turns it into events, rather than `routes` doing it
before the session is created.

Costs nothing in tokens, which is the point worth recording: this writes
into ai-app's own session directory and the phone fetches a reference only
when it draws one. Nothing here is ever written to the CLI's stdin -- it
reads its own session file, and the only things this app sends it are typed
messages, control requests and `/compact`.

Verified against the 133 MB session behind the 2026-08-29 incident: 45
images in the replayed tail, written as real PNGs and served over the files
route, with the transcript itself staying at 756 KB of references.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VETa8afmpWaYezLCqJhDB8
2026-08-29 05:01:48 -04:00
irisandClaude Opus 5 d96bc041a7 Say how big a session is before it is imported
The import list reported a line count, which is the wrong axis: these
transcripts embed screenshots as base64, so one line can be a megabyte.
On this machine a 69 MB session has 3,427 lines while a 44 MB one has
6,792 — the number on the row said nothing about what continuing the
session would cost, and size is the only thing there that predicts it.
The session behind the 2026-08-29 incident was 65 MB across 13,000 lines,
a line count that looks unremarkable.

Shown beside the line count rather than instead of it, since a short file
of long lines is exactly the expensive case. Not warned about and not
marked: importing a large session is a choice somebody is entitled to
make, and flagging it would be the interface nagging about a decision
already taken.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VETa8afmpWaYezLCqJhDB8
2026-08-29 04:54:11 -04:00
irisandClaude Opus 5 362d436d4f Let sessions outlive the backend, and never resume one twice
Three `claude` processes ended up running against this checkout on
2026-08-29, and the account hit its session limit. One cause, several
ways in.

An agent imported the Claude Code session it was *itself* running in.
That is an ordinary import, and importing runs `--resume` -- so a second
CLI attached to a file the first was still writing. The whole 65 MB
conversation, 154 embedded screenshots included, was re-appended to the
transcript under a new prompt id; both copies then read each other's
writes as work done elsewhere, and the adopted one was billed for
re-reading all of it. Meanwhile `shutdown_all` asked each session to stop
and the process exited immediately, so the SIGKILL timer died with the
runtime, the stop was unreliable, and whatever survived was orphaned with
nothing written down to find it by.

The processes leaked either way. So leak them on purpose, and be able to
pick them back up.

A session's process now outlives the backend and is adopted again on the
way up, which is worth having for its own sake: restarting the server no
longer ends a turn somebody is waiting on. Its stdio lives in the session
directory -- a fifo opened read-write so the process is its own last
writer and never reads EOF, plus stdout/stderr logs read from a byte
offset. `session::process` records the pid *and* the kernel's start time
for it, because a pid alone is reused and adopting a stranger's would mean
never resuming the real conversation.

That makes the fix structural rather than a check: everything goes through
`ClaudeDriver::launch`, which adopts if it can and starts if it cannot,
and `--resume` is reachable only on the second path. `Driver` gains two
ways out where it had one -- `detach` (coming back) and `stop` (the
session is being deleted, so the process must not survive).

Importing a session that is open is now refused outright. Claude Code
keeps `~/.claude/sessions/<pid>.json` for every live session, so this is a
measurement rather than a guess; it reports no/yes/unknown, because a
machine that keeps no such record cannot answer and "could not check" is
not "nobody is using it". `SessionStatus` gains `Unknown` for the same
reason.

Also here, found on the way:

- A reconnecting phone was sent the entire backlog. Opening a session was
  bounded to a page but reconnecting was not, so a long disconnect
  delivered thousands of events one frame at a time. Past `CATCH_UP_LIMIT`
  the stream sends a `reset` frame and the newest window, and the client
  rebuilds from it as it does on open -- without the reset the window is
  spliced onto rows no longer adjacent to it.
- A session's status was assumed idle at launch. Read from the transcript
  instead, so a restart stops claiming an exited session is waiting for
  you.
- `llama-server`'s stdout was piped and never drained, so a chatty one
  blocked on a full pipe buffer mid-load. It goes to a log now.
- A turn that exited or errored never emitted `Idle`, so the queue stayed
  "running" for good: every later message was held forever and, since a
  message is only recorded when taken, vanished with nothing on screen.
- Two doc comments had drifted onto the wrong functions.

Verified by killing the server mid-turn: the process survived, finished
its turn unattended (12.8 KB of output nothing was reading), and the
restarted server adopted it -- one process, all 700 lines in the
transcript, no hole, and it still took a new message afterwards. Deleting
a session stops its process; a 266-event backlog resets while a 16-event
one streams. 46 tests, clippy and rustfmt clean, app compiles and lints.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VETa8afmpWaYezLCqJhDB8
2026-08-29 04:47:43 -04:00
iris 9791afcfd6 Hold a queued message below the indicator until the turn takes it
The backend records a message the moment it is sent, so one sent into a
running turn landed in the transcript at the time *we* spoke -- above the
working indicator, in among things the session had already read. It had
not read it. Showing it there says otherwise.

Messages sent while a turn is in flight are now held below the indicator,
drawn quieter, and take their place in the conversation when the turn
ends. That is an approximation and worth naming: the CLI injects a queued
message at a tool boundary, and tells nobody when it does, so the end of
the turn is the first moment anything here can honestly say the message
was taken. It errs toward "not yet read", which is the direction that
cannot mislead.

Also in this change, from the same pass over the screen: the app draws
above the gesture strip rather than under it, and the three status colours
that were literals in two other files -- an amber, a green and a red off
Material's defaults -- are now Catppuccin members in Theme.kt beside the
rest of the scheme.

The ordering is verified by construction rather than photographed: the
list is bottom-anchored, so the first item emitted is the lowest on
screen, and the queued block is emitted before the indicator. Staging a
real long-running turn to photograph cost four model turns and never
produced one, because the model kept declining to sleep -- which is its
own finding, and the reason the next change is a test command in the echo
driver.
2026-08-28 23:49:55 -04:00
iris f402a1f6c8 Hold the bottom while typing, and put the status where the answer goes
**Typing moved the newest message out of sight.** The list re-pinned on a
new item and nothing else, but the two things that shrink it while
somebody writes a reply are the field growing from one line to four and
the keyboard opening under it -- neither of which is a new item. So the
message being replied to drifted upward, and the view only came back when
the reply was finally sent, which is the one moment it did not matter.

It now watches the viewport as well as the item count, and only acts while
already pinned.

**The status has moved out of the corner.** As a label up there it said
something about the session; at the end of the transcript it says
something about a place -- this is where the next thing appears -- and
that is where the reader is already looking, because it is where the last
message is.

`exited` is still said, in the same place. Removing the corner label
without it would have left a session whose process is gone looking exactly
like one waiting for input, on a screen whose whole purpose is typing at
it.

Verified on the phone-sized case rather than reasoned about: three lines
of text with the keyboard open keeps the newest message directly above the
field, and sending shows "working" immediately under the sent message.
2026-08-28 23:32:35 -04:00
iris e03b757dec Stop losing tool calls at the seam between transcript pages
A tool's start and end are two events folded into one row, and the fold
only ever *updated* an existing row -- so an end whose start was not in
the same fold changed nothing and vanished. Not a broken row: no row at
all, which is indistinguishable from a tool that never ran, and is what
Iris saw as gaps where she remembered work happening.

Paging made it routine. Each page was folded on its own and prepended, so
every seam split whatever spanned it: 30 tool ends in the first page of
this conversation, one of them already orphaned before a single scroll.

Two changes, because the two halves fail differently.

Pages are re-folded from the events they came from rather than folded
apart and stitched together. That needs the events kept beside the rows,
since folding is one-way. One pass over everything loaded, paid only when
somebody scrolls back, which is the moment they asked for it.

And an end with no start now creates a row instead of disappearing. Its
name is unknown from an end alone, so it says "tool" until the earlier
page arrives and replaces it -- a row that admits what it does not know
beats silence, because silence is a claim that nothing happened.

Verified across a real seam: scrolling back through the 80-event boundary
of an 864-event import is continuous, with no gaps where tool calls were.
2026-08-28 23:25:13 -04:00
iris 2fe34176c0 Give the transcript a way back, and stop losing an import's name
**Every imported session was called "claude-cli session".** The app has
nothing to say about the title -- the server names it after the session it
is continuing -- so it sent `""`. That is `Some`, which satisfied the
`or_else` meant to catch "no title given", so the imported name was
computed and then thrown away in favour of the `<provider> session`
fallback.

Normalised at the boundary instead: blank means absent, because that is
what it means to the person who left it blank. Both the client's value and
the imported one go through the same trim, so neither can be a string of
spaces standing in for a name.

**And a jump-to-latest button**, shown only while the newest message is
off-screen. Reading back through a conversation is a place to be rather
than a state to be rescued from, so it waits to be wanted and leaves once
there is nowhere to jump to.

It says where it goes instead of drawing an arrow. The list is laid out
from the bottom, so "down" in the data is up on the screen, and an arrow
would be asking the reader to hold that in their head to press a button.

Verified on screen: an import now arrives titled "ai-app" rather than
"claude-cli session", and the button appears on scrolling back, returns to
the newest message, and disappears on arrival.
2026-08-28 23:19:07 -04:00
iris dcb158ee44 Fetch the transcript instead of replaying it one event at a time
**The five seconds of loading top-down.** Opening a session subscribed to
the event stream from sequence zero, so the backlog arrived as one SSE
frame per event -- 864 of them for an imported conversation, rendered as
they landed. That is not a slow list; it is a conversation being replayed
at network speed, and it looks like loading from the top because it is.

The newest page now comes as one request, and the stream starts from where
that page ended, carrying live events only -- which is what a stream is
good at. Scrolling back fetches the page before it, so history costs
something only when somebody actually reads it. 80 events instead of 864,
and the first frame is already the end of the conversation.

I had called this fixed after anchoring the list at the bottom, on the
strength of an emulator on the same machine as the server. That test could
not have shown the problem: the whole backlog arrived in one frame's worth
of time over loopback. Iris's phone, over a tunnel, took five seconds.

**Send disappearing while running.** It was never conditional -- the row
simply ran out of width. A Row hands out intrinsic widths in order and
clips the overflow, so when Stop appeared the pickers I had added pushed
Send off the screen: the app's central control, gone at exactly the moment
the app is most in use. The settings now share what is left after the
actions have taken what they need.

While a turn is in flight the button says **Queue**, because that is what
sending then does -- the message is injected at the next tool boundary
rather than starting a turn of its own. The backend has always done this;
the button was describing something else.

**And the model picker no longer dismisses the keyboard**, which it did by
taking focus. Changing the model mid-sentence is an aside, not a departure
from what you were typing.

Verified on the 864-event import: at the newest message within a second,
history paging back continuously past the first page, and Stop beside
Queue while running.
2026-08-28 23:12:34 -04:00
iris ba25a5cacf Open the transcript at the bottom instead of travelling there
The list was built oldest-first and then scrolled to the end, so opening a
session started at the top and raced downward through everything in it. On
an imported conversation that is nine hundred items measured before a word
is readable, and it was visible every single time.

Laying the list out from the bottom removes the journey rather than hiding
it. The newest message is index 0, so the first frame is already the right
one, and older items are composed only when somebody scrolls back to them
-- which is also what makes a long history cheap rather than something to
load up front.

Following the tail gets simpler as a result. There is no longer a moment
where new content pushes the anchor away, so "am I pinned" is read
straight from the scroll position instead of being remembered across
scrolls, and a new message is one step back to index 0 rather than a jump
across the transcript.

Checked on the 863-event import of this very conversation: a screenshot
one second after opening is already at the newest message, and scrolling
back reaches older ones in the order they happened.
2026-08-28 23:02:56 -04:00
iris 3f8805a610 Say which kind of delete this is, and stop offering what is already open
**Deleting was one word for two different acts.** An imported session's
real transcript belongs to Claude Code and outlives anything this app
does, so removing it here undoes a view. A session started here has no
copy anywhere, and removing it ends the conversation. The dialog warned
"this can't be undone" of both, which makes the warning worthless on the
one where it is true -- and frightening on the one where it is not, since
what it actually deletes is a cache of a conversation still sitting on
the machine.

Sessions now report whether they were imported, and the dialog says which
act this is. No new mechanism: the soft delete already existed, it was
just indistinguishable from the hard one.

**And a session already open here is no longer offered for import.**
Importing one twice would leave two `--resume` processes appending to the
same transcript, each seeing the other's writes as work done elsewhere and
replaying them -- both sessions then showing a conversation neither is
having. The route refuses it as well, so the rule holds for anything not
going through the app.

Left out of the list rather than shown and disabled. The usual argument
says absence is ambiguous, and it is wrong here: an imported session has
not disappeared, it has moved to the session list, which is where it now
belongs. Absence means "already somewhere you can reach it". Deleting the
app's copy puts it straight back -- verified: 68 offered, 67 after
importing one, 68 again after the soft delete, which is also the clearest
demonstration that a soft delete keeps the conversation.
2026-08-28 22:53:42 -04:00
iris d2915c12fa Record how the import sync tells its own writes apart
The reasoning belongs where somebody would look before changing the poll:
status is the obvious discriminator and is wrong, and the failure it
produces reads as the model repeating itself rather than as a bug.
2026-08-28 22:45:01 -04:00
iris a9ea84c96c Stop choosing a model, and keep an imported session up to date
**Why the model became fable.** `spawn_session` fell back to the
provider's first listed model when none was given. That list is a shortcut
for the spawn screen, written in whatever order somebody typed it, and its
first entry is `fable` -- so every session spawned without a model, which
is every import, silently became a fable session. It looked like a default
and was an artefact of list order. Absent now means absent: no `--model`
flag, and the CLI uses whatever the person configured for themselves.

**Model and permission mode are now visible and changeable** from the
session, as buttons that read as their current value rather than labels
beside one. The mode was spawn-only; the CLI turns out to accept
`control_request{subtype:set_permission_mode}` and echo the mode back,
probed against 2.1.237 the same way the rest of the protocol record was.
Both default to `auto` -- on a phone every ask is a round trip to a
question card, which is how "allow Bash?" became the most-answered
question in the app.

The mode is reported by the API so the picker shows what the session is
actually set to, and it is kept in the live session beside the model for
the reason the model already was: `meta` is the shape a session was
*launched* with, so reporting from it shows the value a change replaced.

**And an imported session keeps itself level with its source file**, so
work done at a terminal arrives without a button. `--resume` appends to
the same transcript rather than forking -- measured, not assumed -- so the
only hard question is which new lines came from here.

Answered by counting the events this session has recorded. Status is the
obvious signal and is wrong, which cost a round trip to find: a turn that
starts and finishes between two polls reads as idle at both, so its output
is replayed on top of itself. It showed up on screen as `donedone`, and
only because the reply was one word -- with a longer answer it would have
looked like the model repeating itself.

Verified against both halves: text appended to the source file the way a
terminal writes it appears within one interval, and a message sent through
the app appears exactly once, before and after a turn.
2026-08-28 22:44:41 -04:00
iris c3e7f07a5d Pin the transcript to its tail, and stop asking about every command
**The scroll.** The transcript scrolled on new items and nothing else,
which missed the two cases that matter most. An imported session's history
arrived and left the view wherever it landed; the keyboard opening shrank
the viewport and slid the newest messages under the IME, so typing meant
typing into a view showing the middle of something.

The view is now pinned to the tail, and it is the reader's scroll that
decides: settling anywhere above the bottom releases the pin, settling
back at the bottom re-arms it. The pin is written only when a scroll
*ends*, so it survives the moment when new content has just pushed the
bottom away but the reader never moved -- deriving it continuously from
"is the bottom visible" would release it on every append, which is the
race that makes naive follow-the-tail implementations let go at random.
New items and viewport resizes both re-scroll; the jump is instant rather
than animated, because an imported session appends hundreds of items at
once and animating through them is a light show.

**The input field gets a row of its own**, above the buttons. Sharing one
row put the full width behind three controls, so the thing being typed
into was the narrowest thing on the row.

**Permissions default to auto, and importing can choose.** The spawn
screen defaulted to "manual" and imports passed no mode at all, so the
CLI asked about everything -- and on a phone every ask is a round trip to
a question card, which is how "allow Bash?" became the most-answered
question in the app. Both paths now default to auto, with the other modes
one tap away for a session that warrants caution.

Looked at running, all three: an imported session opens at its bottom,
the tail stays visible while typing with the keyboard open, and the mode
picker shows auto selected.
2026-08-28 22:24:37 -04:00
iris 233689ced6 Name a session in the import list, and let one be deleted
Three things about finding a session in a list of seventy, and one about
getting rid of it.

**A name beats anything inferred.** `/rename` appends a `custom-title`
record, so if somebody has said what a session is, that is the row. Eleven
of the seventy here turned out to be named already and none of it showed.

**Otherwise the last thing said, not the first.** The question this list
answers is "which one was I just in", and a session's opening line is the
least distinctive thing about it -- several of these begin with the same
slash command.

Finding that last message took three tries, and the two wrong ones are
worth recording because they failed in opposite directions. Grepping the
user record type caught tool results, which are *also* user records -- so
a session that ended mid-tool showed a tail of empty records and a row
saying nothing was said, when plenty had been. Narrowing to a string
`content` fixed those two and broke twenty others, because a message
carrying an attachment stores its text in a list. Excluding `tool_use_id`
keeps both shapes of a real message and drops the one that is not: seven
rows still have nothing to show, and those are sessions that really are
empty.

**Sorted by when it was last used**, and the time is on the row. Naming
was tried as the first sort key and is a worse list -- it buries what
somebody was just doing under everything they ever named. A name is for
recognising a row, not for ordering it, so it stays as the title and as a
word beside it.

**And a session can be deleted**, which is asked before it is done. The
transcript *is* the session, so this ends any chance of resuming that
conversation, and the dialog says exactly that rather than "are you
sure?". Deletion resolves the id against what the machine reported, like
importing, so no path crosses the wire in either direction.

Looked at on the emulator, including the dialog -- which is where the
delete button turned out to be missing entirely after a patch that
compiled fine, and where the row layout got its second look.
2026-08-28 22:08:43 -04:00
iris 6bbc829a3e Import a Claude Code session the machine already has
Claude Code keeps every session as JSONL under `~/.claude/projects/`, and
the CLI continues one with `--resume <id>`. `claude.rs` already resumes
whenever it finds a resume token in the session directory, for crash
recovery -- so importing is that same path with the token written before
the driver starts, and there is deliberately no second way to begin a
session. The seed goes through `launch` with the ordinary spawn, so the
driver never learns which kind it got.

Two things the machine answers and the phone does not.

**Which sessions exist.** One command per setup rather than one per file,
for the reason discovery already gives: over ssh each would be its own
connection. Titles come from the first few user records rather than the
first, because a session opens with records the CLI injected -- slash
commands, caveats around local command output -- which are stored as
ordinary user records without the meta flag, so titling by "first user
record" produced a list where most rows read `<command-name>/clear`.

**Which file an id names.** The phone sends an id and never a path; the
server looks it up again among the sessions it enumerated. An enrolled
token must not be able to turn a spawn into "read me this file", which is
the same rule that keeps a provider's command out of `POST /setups`.

Only the tail is replayed. The imported conversation is for reading --
continuing it is the CLI's job, and it reads the whole file itself -- so
this is a display budget, and it has to be one: the session this was
written in is 39 MB, and all of it would otherwise cross a tunnel to a
phone.

A recorded working directory can outlive itself, which this found
immediately: every session from before the checkouts moved to `~/repos`
still records `~/host/repos/...`. Resuming into one fails at `cd` before
the CLI starts -- a confusing way to meet a feature whose promise is
"carry on where you left off" -- so the directory is checked, and a
missing one is dropped with a log line naming it rather than being passed
on to fail.

Verified against this very session: 905 events replayed from the tail
(351 tool calls, 350 results, 185 assistant messages, 19 mine), the resume
token pointing at its id, and the stale directory reported and dropped.
The list was read on the emulator, where the top row is that session under
its opening sentence.
2026-08-28 21:45:14 -04:00
iris 2a1bc84c1e Expand a leading ~, and say what the failing command said
Three things, two of which are the same failure seen from opposite ends.

**A working directory of `~/repos/ai-app` never worked.** Everything
crossing to the remote side is single-quoted, which is right for paths,
model names and prompts alike -- unquoted they would be shell syntax
rather than data. It is wrong for exactly one character: `~` means "expand
me", and quoting is what stops expansion. So the remote shell was handed
the literal four-character directory `~` and correctly said it did not
exist, which reads as the path being wrong rather than the quoting.

Paths now go through `quote_path`, which emits `"$HOME"` for a leading
`~/` and single-quotes the rest. The variable expands, the expansion is
not re-split or globbed because it is double-quoted, and nothing after it
gains a meaning -- there is a test that pushes a quote-and-semicolon
injection through the tilde branch and gets back one absurd path rather
than three commands. `$HOME` is set by every shell this can land in, so
this does not depend on the remote side being POSIX; verified by running
the generated script under both sh and fish, which is what the dev VM
actually uses.

**The phone could not have told you any of that.** The exit report kept
the last line of stderr, and a shell's error message ends with a blank
line -- so the last line was empty, the report was a bare exit status, and
the seven lines of fish complaining sat in the server's log where nobody
holding a phone is looking. It now keeps the last 50 lines in a ring and
reports them with blank lines trimmed from both ends. The tests use the
real fish `cd` failure as their fixture.

**The status bar was unreadable.** `isAppearanceLightStatusBars` was
hardcoded to `true` -- dark icons -- which was right against the default
light surface and wrong the moment the app wore Mocha. It now asks the
scheme's own background for its luminance, so changing the palette cannot
reintroduce it.

**And the address field takes `user@host:port`.** One field rather than
two, because that is how an address is written everywhere else and a port
that is nearly always 22 does not deserve its own box on a phone keyboard.
Absent means absent rather than 22: the backend already decides that
default, and writing it here would be a second answer in a second place.
A colon only means "port" when it can -- brackets for IPv6 as ssh writes
them, otherwise exactly one colon followed by digits.

Looked at on the emulator: the status bar, and the form, whose label I
then shortened because it wrapped onto a second line and made that field
taller than the two beside it.
2026-08-28 21:05:17 -04:00
iris 31135e3f22 Wear Catppuccin Mocha, and move Usage to where the provider is
Two changes to the app, plus the one they turned up.

**The theme.** Catppuccin Mocha, copied from dev-updater rather than
shared: wg-app-link is the *link* -- the tunnel, the pinned CA, enrollment
-- and a palette is not that. The two apps looking alike is a preference
rather than a contract, and the moment one wants a different accent a
shared version becomes a thing to fight. dev-updater's ActionTone and its
ANSI table did not come across; nothing here draws a log or a destructive
button yet, and copying a vocabulary with no speakers is how a file starts
lying about what the app does.

**Usage is no longer a global button.** It belongs to the provider, and
the session view is the only place a provider is currently named, so that
is where the control sits -- beside the line that names it, rather than
collected with the app-wide controls where its scope had to be guessed. It
carries the session back with it, so Back returns to that session rather
than dumping the reader on the list.

Its real home is that provider's own settings, which do not exist yet.

**And the bit that only running it could find.** I first subtitled the
usage screen with the session's provider, which on an echo session put
"echo" directly above a card reading "claude" -- Claude's account-wide
numbers labelled as echo's, a claim about echo that nothing measured. The
subtitle is gone; each card names the service that answered, which is the
true scope, and the reason is written where the subtitle was so it does
not get re-added.

Looked at on the emulator rather than read: the palette, the session
header, the usage screen, and Back landing on the session it came from.
ktfmt, compile and Lint clean.
2026-08-28 20:26:12 -04:00
iris 4370c467ca Discover this machine's providers instead of asserting them
A fresh install wrote a `claude-cli` provider into the local setup
unconditionally. Nothing looked for `claude`; the list was hardcoded in
`Config::seed`, so on any machine without it -- which is every machine but
the dev VM -- the phone was offered a provider that cannot spawn, stated
with exactly the confidence of one that had been checked.

Discovery already existed and was already right: `setups::discover` probes
with `command -v` over the transport, includes echo for the local one
because it runs in-process, and records the resolved path rather than the
bare name. Only the local setup skipped it, which is the one place the
answer felt obvious enough not to ask.

So `seed` now takes the providers it is given, and seeding asks this
machine the same question it asks any other. It moved out of
`SessionManager::new` into an awaited step in main, because asking is I/O
and a constructor that quietly spawns a subprocess surprises every caller.
A discovery that fails seeds `echo` alone and says so, since echo is true
wherever this server runs -- falling back to the hardcoded list would be
the same bug with an extra step.

The test that covered this agreed with the bug, because both were written
from the same assumption: it asserted the seed contains `claude-cli`. It
now asserts the opposite -- that the seed invents nothing -- and the
session tests seed echo explicitly rather than relying on a constructor
that would make them pass or fail on whether `claude` happens to be
installed on whoever runs them.

Verified by running a server on a PATH holding only `sh`: it seeds `echo`
alone. With claude and llama-server present it finds both. 34 tests.
2026-08-28 20:15:44 -04:00
iris 3cf6925d90 Let Dev Updater supervise the backend, now that the switch can be sequenced
Re-applies the change reverted in d5a0f67. It was right then and only
mistimed: something was reading this working tree live, so deleting the
script took the backend card to "couldn't check" immediately rather than
on a pull. Iris has now uninstalled the service while the script was still
the declared one, which is the step that stops the old unit being orphaned
under a name nothing points at any more.

`Managed` runs the command with dev-updater's built-in service script,
resolved against the component's cwd. ai-app's own script was the generic
case exactly -- ExecStart=$BINARY with no arguments and no environment --
so it was 233 lines of init-system detection kept in step with an
identical copy next door.

Deleting it loses nothing, and that was checked rather than assumed: all
three findings the test guest produced today are in the built-in --
reading OpenRC's status exit code instead of grepping text it prints to
stderr, treating an uninitialised user softlevel as "couldn't find out"
rather than as stopped, and assigning through `|| code=$?` so `set -e`
cannot kill the script before it reads the code.

The one thing the script said that the built-in cannot is already in
AGENTS.md: Stop on this card takes down the server a phone reaches through
the tunnel, while Dev Updater itself is unaffected because it uses its own
port.

Install from the backend card after taking this.
2026-08-28 19:52:42 -04:00
iris a8b6c13b01 Report a crashed service as failed, which OpenRC was never telling us
The OpenRC branch of this script had never run anywhere. A guest built to
reproduce the host says it was wrong in the way the `failed` state exists
to prevent: a service that fell over reported `stopped`, which reads as a
decision somebody made.

Two causes, both measured rather than reasoned about.

`rc-service status` prints `* status: crashed` to **stderr**. The check
was `status 2>/dev/null | grep -qw crashed`, which discards precisely the
word it is searching for, finds nothing, and falls through to `stopped`.
The old comment argued for reading the word rather than the exit code, and
that argument was sound except that the code turns out to be specific
rather than merely non-zero.

So it now reads the code, which says more than the text did: 0 started,
3 stopped, 32 crashed, and 1 for every way the question cannot be
answered -- an unknown service, XDG_RUNTIME_DIR unset, or a user softlevel
that was never initialised. That last group is a real state and not one of
the other three, so it exits non-zero and says so instead of guessing.

`set -e` is the second cause, found by running the first fix: every answer
except "running" is a non-zero exit, so a bare invocation killed the
script before the code could be looked at. It prints nothing and exits 3,
which is indistinguishable from a crash of this script itself.

Verified in the guest, all four states: not-installed, stopped, failed
after a real crash, and a non-zero exit with nothing on stdout when the
softlevel is removed. Before the fix the crashed case printed `stopped`.
2026-08-28 18:55:05 -04:00
iris 297e85c68d Delete the pre-setups migration, which has done its job
The rule is that migration code goes once the update carrying it has been
received, because there is one backend and one phone: once they are past
a shape, nothing anywhere is still on it, and a second parsing path that
nothing exercises only constrains later changes to the schema. The
module's own comment said to delete it "once the host has started on a
build containing it", and that has happened -- it is in the pushed commit
the host reports itself up to date with, and the server has been starting
on it.

Out: the `legacy` module, `migrate_from_pre_setups`, the branch in
`Config::load` that reached it, and the test. `Config::load` is now one
expression.

PLAN.md keeps the history rather than reverting to what it said before,
because the interesting part is not the migration but the decision it
replaced: refusing to start on an old config was the wrong trade and
proved it on Iris's host, as a crash loop that could not explain itself
because the crashing process is how the phone reaches the machine at all.

Verified by running it, since the point of this change is what happens at
startup: a server with no existing state starts, generates its CA, prints
its enrollment QR and writes a config that reads back. 34 tests, clippy
silent, rustfmt clean.
2026-08-28 18:40:24 -04:00
iris af41d86186 Pin the crate at the local-network note
Nothing in this repo changes; the submodule moves to the commit carrying
the measured finding that ACCESS_LOCAL_NETWORK is still required through
the tunnel, contrary to Android's own documentation.
2026-08-28 18:19:41 -04:00
iris d5a0f67a3a Put the service script back until the switch can be sequenced
Reverts the switch to `service: Managed(...)`. The switch is still right
and the reasoning in that commit still holds; what was wrong was doing it
now, unilaterally, to a checkout something is reading live.

A dev-updater is running against this working tree, so deleting
`server/service` did not wait for a pull to take effect -- the backend
card went to "couldn't check -- failed to run the service script: No such
file or directory" immediately, and the pushed declaration still names the
script, so the tree and the declaration disagreed in the one direction
that breaks things. My own commit message had said this change was not
safe to pull blind; it turned out not to need a pull at all.

The switch needs three steps in order, and only the middle one is mine:
Uninstall from the backend card while the script is still declared, then
take the change, then Install. Re-apply when Iris is ready to do that,
which is also when dev-updater's conversion path can be deleted.
2026-08-28 18:18:38 -04:00
iris 295602adfe Save the config through the shared crate as well
`Config::save` was the same nine lines as dev-updater's, so it is now
`format::write(path, self)`. The reasoning that made those nine lines
correct -- the leftover temp file that keeps its old mode and is then
renamed over the token hashes -- lives with the code and its test rather
than in two places that could stop agreeing.

Verified: 35 tests, clippy silent, rustfmt clean.
2026-08-28 18:12:30 -04:00
iris b6b33dc9c5 Take the app half from wg-app-link as well
The four Kotlin files that were the link rather than this product now come
from the submodule: the pinned TrustManager, the enrollment store and its
Keystore sealing, the QR capture activity, and the local-network permission
check. `:link` is a subproject resolved by path, so the app half is
version-locked to the same commit the Rust half already was.

What stays here is the two facts that are actually about this app, and
both are load-bearing in a way that would fail quietly if got wrong: the
`aiapp` URI scheme, and the Keystore alias `aiapp-token-key` that every
enrolled phone's token is already sealed under. A wrong alias would leave
those phones reading as not enrolled with nothing on screen to explain it,
so the value is carried over exactly and the reason is written beside it.

Call sites are unchanged. `ServerSettings` stays available unqualified as
a typealias and `applyPinnedTls()` stays an extension, so the diff is the
three files that bind the product-specific values plus two imports --
rather than every screen that happens to use a setting.

Also clears a warning the build had been printing: `setup?.id.orEmpty()`
where the compiler already knows `setup` is non-null, because `chosen`
came from that setup's own provider list.

Verified by running the build, not only by reading it: ktfmt, Kotlin
compile and Android Lint are all clean with no warnings, and the APK still
builds -- which exercises the pinned-CA generator, since that is the step
that reads the CA off this machine.

Still unpushed, per the hold until the rebuild bug is proven fixed. Note
the submodule: a checkout of this commit needs `git submodule update
--init` before `app/` or `server/` will build.
2026-08-28 17:57:45 -04:00
iris c2dfaab349 Let Dev Updater supervise the backend instead of shipping a script
dev-updater now carries a built-in service implementation, generated from
a template and driven through the identical interface a project-supplied
script uses, so a project whose service is unremarkable no longer writes
one. ai-app's was unremarkable: `ExecStart=$BINARY` and `command="$BINARY"`
with no arguments and no environment. 233 lines of it, and the half that
matters most -- the OpenRC branch, which neither project can exercise from
a systemd machine -- existed twice, so a fix found by testing would have
had two places to land and no way to notice the second.

The field keeps its name; `Managed` takes the command, resolved against
the component's `cwd`.

The one thing the script said that the built-in cannot is kept, in
AGENTS.md rather than lost: Stop on this card takes down the server a
phone reaches through the tunnel, while Dev Updater itself is unaffected
because it uses its own port -- which is exactly what makes that button
easy to press and easy to regret.

NOT SAFE TO PULL BLIND. A managed service is named after the component, so
this one becomes `app-backend` while the installed one still has the name
the script gave it. Uninstall from the backend card *before* taking this
change, then Install after; pulling first orphans a service that stays
enabled and starts at boot with nothing pointing at it.
2026-08-28 17:35:27 -04:00
iris 2c925a679f Take XDG resolution from the shared crate too
Sixth and last of the modules that were the link rather than this
product. main.rs loses config_home, data_home and xdg_dir, and its test
module with them -- it held one test, which moved to the crate that now
holds the code.

The helpers gained a `product` parameter, matching certs::ensure and
netif::wg_address, which is what keeps two products' state apart while
resolving it identically.

Verified by running it: with only XDG_CONFIG_HOME and XDG_DATA_HOME set
and no --config or --data-dir, the server puts its certificates in
$XDG_CONFIG_HOME/ai-app/certs and its sessions in
$XDG_DATA_HOME/ai-app/sessions, and still prints an aiapp:// enrollment
URI. 35 tests here and 19 in the crate, clippy silent, rustfmt clean.
2026-08-28 17:33:47 -04:00
iris aa05ff9336 Take the link from wg-app-link instead of keeping a second copy
The five modules underneath this backend that were never about AI
sessions -- the pinned CA and leaf, QR enrollment and the bearer token,
wg0 binding and the certificate's SANs, owner-only files, and the RON
house rules -- were written twice, once here and once in dev-updater,
and had drifted. They now come from the submodule, as a path dependency
so both projects stay locked to one commit.

What stayed is what makes this project itself: the routes, the drivers,
the config schema, and the auth middleware, which is generic over this
server's state. Sharing a transport is worth doing; sharing an API would
mean inventing a vocabulary neither project wants.

Four dependencies go with the code -- rcgen, qrcode, subtle and if-addrs
are no longer named here at all -- and the three that remain are now
described by what still uses them rather than by what used to.

Verified by running it, not only by building: a fresh server generates
its CA, prints an `aiapp://enroll` QR with the scheme now passed as a
parameter, covers 127.0.0.1, 10.0.2.2 and wg0's 10.66.0.1 in the leaf,
answers an enrolled token and returns 401 without one, and writes
config.ron in the house rules with every file owner-only. 36 tests pass,
clippy is silent, rustfmt is clean.
2026-08-28 17:14:33 -04:00
irisandClaude Opus 5 a83dbcff6a Say when the server fell over, and where to read why
Iris found the backend crash-looping by checking rc-service by hand,
because the card could only say `stopped` -- which reads as a state
somebody chose. dev-updater's contract now has a fourth word, `failed`,
and this script implements it.

The OpenRC detail is the one worth not rederiving: it prints `crashed`
*and* exits non-zero, so the word is read rather than the exit code.
Leaning on the code would report "couldn't check", which is a different
and less useful claim. The `running` check stays on its exit code, which
already worked and does not depend on wording.

The script also arranges the logging rather than only reporting it,
because neither unit wrote a file: systemd went to the journal and
OpenRC's `command_background=true` discarded output entirely, which is why
a crash left nothing to read. Output now goes to
$XDG_DATA_HOME/ai-server/ai-server.log -- generated data, outliving any one
build, and not in a repository shared with a machine that should not read
it. `start` rotates one generation aside, so what is kept is exactly this
run and the one before: the pair worth having after a crash and a restart.
`logs` prints the paths, newest first, and nothing else.

Verified on systemd by causing the failure rather than reasoning about it:
installed, started, confirmed `running` on the wg0 bind, wrote an
unparseable config, restarted, and watched status settle on **failed**
rather than stopped -- with the reason, line and column, in the file `logs`
points at, and the crash preserved in .1 after recovery. Then restored,
confirmed `running` again, and uninstalled.

**The OpenRC branch is written from the documentation and is untested**,
here and in dev-updater, since neither machine that can run it is one
either of us can test on. It is also the branch that actually matters,
since the backend runs under OpenRC on the host. `output_log`/`error_log`
in the openrc-run script are the parts to distrust first.

One thing that bit while writing it: the systemd heredoc is unquoted so
$LOG expands, which makes a backtick in a comment inside it run as command
substitution. A comment saying "`start` rotates" executed `start`, and the
unit was written without ever being valid. There is now a note in the
heredoc saying why it contains no backticks.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017xn8nHw1tw1R6PtiY1eEtw
2026-08-28 15:57:05 -04:00
irisandClaude Opus 5 0c55b809b1 Drop the reader for the kebab-case driver kind
Iris has already moved past that spelling, so nothing will ever present it
again -- there is one backend and one phone, and both are past it. The
alias and the enum that carried it are gone; the legacy provider is just a
`ProviderConfig` now.

The rest of the migration stays until it has actually run on the host,
because deleting it before then would strand the install it was written
for. Its doc now says that outright, along with what to delete and when:
this module and the branch in `Config::load` that reaches it, once the host
has started on a build containing it.

That is the general rule Iris gave, not a judgement about this migration:
a reader for a superseded format has a defined end, because his population
is one machine he controls, and leaving it keeps a second parsing path
alive that nothing exercises and that constrains every later change to the
schema.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017xn8nHw1tw1R6PtiY1eEtw
2026-08-28 15:35:20 -04:00
irisandClaude Opus 5 5ccadffeaa Migrate an old config instead of refusing to start
The AI Sessions backend was crash-looping on the host, and I caused it. A
config written before setups existed makes `Config::load` bail, the process
exits 1 immediately, and under OpenRC's `command_background=true` that
presents as a service that will not stay up.

The refusal was deliberate and it was the wrong trade. I chose it to avoid
silently emptying a config and re-seeding over it -- a real hazard -- but
weighed it against the wrong cost. This process is how a phone reaches that
machine at all, so refusing to run strands the person who would have to fix
it, at a terminal, on the machine they were trying to avoid needing. And
what it was protecting is the cheap half: providers and hosts are
rediscoverable now, while the half that genuinely cannot be recovered --
the enrolled token hashes -- survives a migration untouched.

So it migrates. Each old host becomes a setup keeping its name, since that
is what sessions referenced; the top-level providers belong to the machine
this server runs on; and every session's host becomes its setup, so
conversations keep working. The original is copied to
`config.ron.pre-setups` first, because this is a one-way conversion of the
only record of what was configured and one file makes it reversible by
hand.

**Migrated hosts arrive with no providers, deliberately.** The old file
never recorded which machine had which program -- that was the flaw the
setups model exists to fix -- so inventing an answer would recreate exactly
the impossible pairings it was meant to end. Rediscover asks the machine.

Both driver-kind spellings are read. The kebab rename and the RON move
landed on the same day, so a file written that morning says
`r#claude-cli` and one from the afternoon says `claude_cli`; reading only
one would have turned this fix into a different crash.

Verified against a host-shaped config: the server starts, the token and
both sessions survive, the remote session points at the migrated setup and
the local one at `local`, the original is kept, and a second start is an
ordinary load that neither migrates again nor overwrites the backup.

Found by Iris, who had to check `rc-service` by hand because the card
reported it as merely stopped -- dev-updater's session is adding a `failed`
state for that separately.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017xn8nHw1tw1R6PtiY1eEtw
2026-08-28 15:33:25 -04:00
irisandClaude Opus 5 4d162331e3 Say that llama.cpp reaches the phone now
The status section still said "server side" and "no app screen yet", both
of which stopped being true today. Also records the discovery trade that
would otherwise be rediscovered: command -v follows a non-interactive ssh
PATH, so llama.cpp unpacked into ~/.local/opt is invisible until it is
symlinked onto PATH.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017xn8nHw1tw1R6PtiY1eEtw
2026-08-28 13:38:54 -04:00
irisandClaude Opus 5 6187958de3 Let a llama session actually be started from the phone
The driver worked and the models could be downloaded, but the spawn screen
had no idea llama.cpp existed: the model field and every extra setting were
gated behind `isClaude`, so a llama provider offered nothing, `model`
arrived null, and the driver refused with "a llama.cpp session needs a
model". The feature was reachable only by curl, which is not what was asked
for.

A llama provider now gets the models this backend has downloaded, as a
picker rather than free text -- there is nothing sensible to type, and a
name that is not on disk is a session that cannot start. Context size and
temperature are there too, blank meaning llama.cpp's own default rather
than a zero. Spawn stays disabled until a model is chosen, because without
one the button could only fail.

**Two bugs that only appeared by pressing the button**, both mine, both
from changing the server without re-driving the app:

- The app sent the setup's *label* where the server had started resolving
  by *id*. The failure was almost self-diagnosing -- `no setup named "this
  machine" -- configured: this machine` -- and that message now says "no
  setup with id" and lists ids, since listing labels was what made it read
  as a contradiction.
- The session header showed `on local`, the id, because the app read
  `setup` where the server had begun sending both `setup` (id) and
  `setupName` (label). The app now carries only the label: nothing in it
  addresses a setup, and holding both is what let it show the wrong one.

Verified by doing it: rediscovered the local machine from the phone so
`local-llama` appeared, spawned a session on Qwen3-0.6B-Q8_0 with a 4096
context, sent "Reply with exactly one word: ready", and it replied "ready"
with 125 tokens counted.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017xn8nHw1tw1R6PtiY1eEtw
2026-08-28 13:38:00 -04:00
irisandClaude Opus 5 d0b6b66a44 Narrow a file this server rewrites, not only one it creates
`OpenOptions::mode` applies to a file the call creates and to nothing else,
so rewriting a file that already existed kept whatever permissions it had.
Three functions above, `create_dir` has carried a comment about exactly
this hazard since it was written -- the file path never got the same
treatment.

This is not hypothetical here. `certs.rs` reissues the TLS leaf and rewrites
its **private key on every start**, so a key that ever existed
world-readable would have stayed that way for the rest of its life, with
every subsequent start looking like it was setting the mode. The config's
temp file is the other one: normally fresh, but a leftover from a crashed
save would be reused with its old mode and then renamed over the real
config, which holds the enrolled token hashes.

Set through the open handle rather than the path, deliberately:
`set_permissions` on a path re-resolves it, so between the open and the
chmod something could put a different file -- or a symlink to one -- where
this was, and the mode would land there instead. A handle cannot be
redirected.

Three tests, and the first was checked against the bug rather than only
against the fix: with the new line commented out it fails with "rewriting
left it at 644".

Found by dev-updater's session, which had taken this module for a shared
crate and read it as a unit. I had spotted the same line being wrong in
their new `append_file` and missed that `create_file` -- the one I wrote --
had it too.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017xn8nHw1tw1R6PtiY1eEtw
2026-08-28 13:24:53 -04:00
irisandClaude Opus 5 9f54a80ca4 Tell the difference between a blocked app and a missing server
Two findings from dev-updater's session reading this codebase, both checked
against the code here before acting on them, and both real.

**A denied local-network permission was invisible.** The manifest requests
ACCESS_LOCAL_NETWORK and MainActivity asks for it, but nothing ever checked
whether it was granted -- and on Android 17 a denial is indistinguishable
from an unreachable server at the socket, because the OS simply drops the
traffic. So every screen would have shown "is ai-server running, and is
this device able to reach that address (WireGuard up)?", blaming two things
that were both fine.

Stated once at the root as a standing condition rather than appended to
each failure it might have caused: it is not a property of any one request,
and repeating it per error is how a message ends up saying the same thing
twice, which this app has already done once today.

**The Keystore read path was creating keys.** `unseal` called the
get-or-create key function, so a sealed token whose key had been lost -- a
device reset, or the app's data restored onto a device the key cannot
travel to -- generated a fresh key, then failed to decrypt with it, leaving
a key nothing had ever sealed with. The behaviour was already right by
accident (it fails soft to "not enrolled"), but the read side now asks for
the key without making one, which is what it meant all along.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017xn8nHw1tw1R6PtiY1eEtw
2026-08-28 13:19:58 -04:00
irisandClaude Opus 5 3c144f8070 A screen for the machines, and failures a phone can act on
The other half of making setups editable: add, rename, rediscover and
remove, with a Test that tries a machine before anything is saved.

The screen cannot name a program, which is the point rather than an
omission -- providers are what the server found when it asked, so this app
has no way to introduce something to run. The dialog says so, because
"what it can run is discovered, not typed" is the answer to the question a
person will otherwise ask when they look for a command field.

Two things running it changed. The card showed "this machine / this
machine", because the seeded setup is *called* that and my fallback line
for a local setup said the same -- the line now says something the name
cannot also be. And the header row absorbed a fifth action without
complaint, which is the earlier title-and-actions split paying off exactly
as its comment predicted.

**Host key verification is the failure that would have made this look
broken.** Every machine fails it the first time, because its key is not in
known_hosts yet, and ssh's own words -- "Host key verification failed." --
are written for somebody at a terminal on the backend, which is exactly who
is not reading a phone. It now says what to do: ssh to it once from the
backend and try again. Permission denied gets the same treatment.

Deliberately *not* fixed by relaxing StrictHostKeyChecking. Accepting a new
key is a decision somebody should make with the key in front of them, not
something this app does quietly on their behalf while adding a machine.

Verified on the emulator against a running server: the seeded setup renders
with what was discovered on it, the add dialog explains itself, and Test
against an untrusted machine produces the full explanation rather than
ssh's four words.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017xn8nHw1tw1R6PtiY1eEtw
2026-08-28 13:17:08 -04:00
irisandClaude Opus 5 19e3531c5d Add, rename and remove machines from the phone -- without letting it name commands
The gap PLAN.md recorded: setups were readable but only hand-editable, so
adding a machine meant a shell on the backend.

**The design decision, made with Bryan, is that the phone never composes a
command.** A setup carries providers, and a provider carries something to
run -- so a route that accepted a command from the request body would make
the enrolled token arbitrary code execution on every machine a setup names,
and the transport already reaches those over ssh. Instead the phone sends
connection details, and the server asks the machine itself what it has:
one `command -v` round trip per setup, matched against a table of the
drivers this server knows. The phone's authority is "add this machine",
never "run this".

Worth recording that this was a narrower change than it first appeared: the
token could already run anything on the backend, because the spawn screen
offers `bypassPermissions` with a free-text working directory. Discovery
does not close that door. What it does is keep the *list of what can run*
out of the phone's reach, and make adding a machine a thing you cannot get
wrong by typing.

It is also simply better to use. Nobody wants to type an absolute path on a
phone keyboard, and a machine whose binaries have moved answers correctly
on the next probe. The cost is that a program somewhere unusual is
invisible -- `command -v` follows PATH under a non-interactive ssh session,
which is not the PATH a person sees when they log in. That is the trade,
and the escape hatch is editing config.ron on the backend, which is exactly
the authority the phone is not being given.

Setups now have an **id separate from their label**, so renaming a machine
does not orphan the sessions that name it; a session stores the id, and
every row resolves the current label when it is built. `POST /setups/probe`
tries a machine without saving anything, so a wrong address or an
unauthorised key is caught while the form that caused it is still on
screen. Deleting is refused while sessions still run there, and says which
ones rather than cascading.

Every mutation goes through one `update`: clone, apply, save, then commit,
so a failed write leaves the previous state intact and reports why.

Verified against a running server, including a real ssh machine (this VM,
via a throwaway loopback key since removed): probing here found echo and
claude-cli; probing over ssh found claude-cli and correctly no echo, which
runs in-process and exists only where this server does; an unreachable
machine came back with ssh's own words ("connect to host ... Connection
timed out"); adding derived the id `loopback-vm` from "loopback vm";
renaming kept the id; deleting was refused while a session used it, naming
it, and succeeded once nothing did.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017xn8nHw1tw1R6PtiY1eEtw
2026-08-28 13:08:44 -04:00
irisandClaude Opus 5 ef019a7aea Mark the setups model as built, and say what is left
The three questions the plan left open are answered by having built it:
where echo lives (seeded into the local setup, not implicit), what happens
to an old config (refused with instructions, because the silent version
loses everything), and that editing setups from the phone is still
missing.

That last one is the honest gap: GET /setups exists, writing them does
not, so adding a machine is still a hand edit on the backend.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017xn8nHw1tw1R6PtiY1eEtw
2026-08-28 12:56:35 -04:00
irisandClaude Opus 5 ecac404fd4 A setup is a machine, and it carries what that machine can run
Providers and hosts were two independent lists, and a session named one of
each. They were never independent: a provider only exists on a machine
where that program is installed, so the spawn screen offered the whole
cross-product, including "the Claude CLI on the box that hasn't got it".
The picker could not know, because nothing in the model said.

Now a setup is a machine -- optional ssh, plus the providers it has -- and
spawning is two choices in order: pick a setup, then one of its providers.
The impossible pairs stop being expressible rather than being validated
against. Provider names are unique within a setup and only within one, so
two machines can each have a `claude-cli`, which was previously either a
name collision or two entries called things like "claude" and "claude on
the vm".

It also settles the "Run on" problem properly. That control was offered for
every provider but honoured only by the Claude driver -- an echo session
sent to a host ran locally and said otherwise. There is no such control
now: the machine is chosen first, and echo is a provider of the setup with
no ssh, where it belongs, since it runs in-process and has no transport to
cross.

The built-in echo provider is gone as a concept. It used to be conjured at
read time and never written to the file, which meant a provider nobody
could see or edit; it is now seeded into the config on first run alongside
claude-cli. What the file says is what there is, and deleting it is a
choice rather than a state to be repaired.

A config in the old shape is refused with instructions rather than loaded.
`Config` defaults unknown fields away, so `providers:` and `hosts:` would
otherwise have vanished into an empty config that was then seeded over --
a migration nobody would notice until their setups were gone.

Verified against a running server and on the emulator: a fresh install
seeds "this machine" with echo and claude-cli and the file reads cleanly;
a two-setup config lists both with their own providers; spawning on a
setup works and the session row names it; asking for a provider a setup
lacks says which it offers, and an unknown setup says which exist. On the
phone, selecting "dev vm" narrows the provider chips to that machine's one
and shows its address.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017xn8nHw1tw1R6PtiY1eEtw
2026-08-28 12:56:14 -04:00
irisandClaude Opus 5 7cf36005ae Browse, download and delete models from the phone
The point of the llama.cpp work was that models are managed from the app,
not by editing the backend's filesystem, so this is the screen for it:
search HuggingFace, expand a repository to see its GGUFs with sizes,
download one and watch it, cancel it, delete what is no longer wanted.

Everything shown is the server's state rather than the screen's. A
download started here keeps going when the screen closes, is visible from
any enrolled device, and its outcome outlives it -- demonstrated by
accident while testing, when a 538 MB download finished during an app
rebuild and was still there, complete, after reinstalling.

Polled rather than streamed, at 1.5s. A download belongs to the machine
rather than to any session, so it has no event stream of its own; this is
the one screen in the app that asks repeatedly instead of being told.

Three things the screenshots decided rather than the diff:

- **The list header no longer squeezes its title.** Adding a fourth action
  to the row wrapped "AI Sessions" onto three lines. Title and actions now
  have a row each, so a fifth costs nothing and the title is never what
  gives.
- **A repository's files render inside its own card**, not as a section
  after the list -- drawn after every card they read as belonging to
  whichever was last.
- **A file already downloading says so** and is disabled, rather than
  offering a Download button whose effect nobody can see.

The progress bar is determinate only when the server reported a size, and
says "total size unknown" otherwise rather than inventing a position.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017xn8nHw1tw1R6PtiY1eEtw
2026-08-28 05:32:38 -04:00
irisandClaude Opus 5 3be25c2f64 Write down what the llama.cpp work actually does
Phase 4 is no longer deferred and phase 5 is exercised, so the status
section says so. The two decisions worth not undoing by accident get
named: the conversation lives in the transcript rather than the driver,
and a llama session is refused on an ssh host rather than half-working.

Also the local testing recipe, including the trap that cost me twenty
minutes -- a 2-bit quant produces fluent nonsense that reads exactly like
a broken driver, and llama-cli on the same file is how to tell the two
apart.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017xn8nHw1tw1R6PtiY1eEtw
2026-08-28 05:23:47 -04:00
irisandClaude Opus 5 3deeffd1e7 Run GGUF models through llama-server, and stop orphaning them
The second half of the llama.cpp work: a session can now name a downloaded
model and talk to it. `llama-server` is spawned through the same transport
as any other driver, polled until the model is loaded, then driven over its
OpenAI-compatible streaming endpoint and translated into the same events
the Claude driver emits -- so the transcript, the SSE stream and the phone
need to know nothing new.

**The conversation is rebuilt from the transcript, not held in the driver.**
llama-server is stateless between requests, so the whole history goes with
every one, and the obvious place to keep it is a Vec in the driver. That
fails the requirement: memory in a driver is invisible to a second device
and gone on restart, and this app is meant to work across devices. Reading
it back also means the model is prompted with exactly what the phone was
shown -- including a reply that was interrupted half way, which is in the
transcript because the deltas were already emitted.

That leaves the Claude driver as the odd one out rather than this one: the
CLI's memory of a conversation is a cache in front of the same transcript,
not a second truth. Said so at the top of llama.rs, because it is the sort
of inconsistency that gets "fixed" in the wrong direction.

Session settings arrive as a driver-interpreted `params` map rather than
new typed fields, so the shared schema does not grow one dialect's
vocabulary. Context size, gpu layers and threads become server flags;
temperature and the rest ride on each request, so changing them need not
reload a model.

**Also fixes an orphan this feature would have created.** Drivers set
kill_on_drop, which covers a session being deleted -- but nothing drops on
the way out of a SIGTERM, so signalling the server left its children
running. For the Claude CLI that is untidy; for a llama-server holding a
model it is gigabytes belonging to nobody. The server now stops its
sessions on SIGTERM and SIGINT. Found by killing a test server and noticing
two 600 MB processes still resident.

Remote llama sessions are refused rather than half-working: the model is
reached over HTTP, and forwarding that port to an ssh host is the "reach
this port" operation the transport does not have yet.

Verified end to end against a real model: downloaded Qwen3-0.6B Q8_0
through the app's own download route, spawned a session on it, and held a
two-turn conversation -- "my favourite colour is teal" then "what is my
favourite colour?", answered "teal", which is the transcript replay doing
its job. Token counts arrive. An earlier attempt with the IQ2_XXS quant
produced fluent nonsense, which turned out to be the quantisation rather
than the pipeline: llama-cli produces the same from that file directly.
Four unit tests cover the fold and the path guard.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017xn8nHw1tw1R6PtiY1eEtw
2026-08-28 05:23:12 -04:00
irisandClaude Opus 5 e50d1a2bbf Take the resumed total from Content-Range, not Content-Length
On a 206 those two headers answer different questions: Content-Length is
the length of the range, so a resume at 162 MB reports 72 MB and a bar
drawn from it fills at a third of the model. The arithmetic that was here
(`have + length`) happened to be right, but only because the range always
starts exactly at what is on disk -- it was correct by coincidence of two
things agreeing rather than by asking for the number wanted. Content-Range
carries the whole size as its last field and does not care where the range
began.

Confirmed against HuggingFace: `content-range: bytes
162000000-234074815/234074816` beside `content-length: 72074816`, and a
resumed download now reports 234.1 MB rather than 72.

Two other things checked rather than assumed, both fine as they stood.
Downloads are already single-flight per file -- the check and the insert
happen under one lock, keyed by the model, so a second client asking for
the same file joins the running download instead of starting a second
writer onto the same partial. And HuggingFace's ETag is stable across
requests, with no weak prefix or per-edge variation, so the identity check
will not discard good partials and refetch gigabytes for nothing.

One hypothesis worth recording as false: HF's ETag is not the content
sha256 for these files (`db6593d0…` against a published `55e0d0b8…`), so
the published-hash check cannot collapse into the identity check. Both
earn their place.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017xn8nHw1tw1R6PtiY1eEtw
2026-08-28 05:10:05 -04:00
irisandClaude Opus 5 6f149398d0 Don't resume onto a partial from a different revision
A resume splices: it appends bytes from wherever the server is now onto
whatever is already on disk. If the file changed upstream in between, the
result is exactly the failure that survives every cheap check -- the right
number of bytes, the wrong contents, and no error anywhere. HuggingFace
files do get updated, so this is a real path rather than a theoretical one.

`If-Range` is the header for this and would have been the tidy answer, but
HuggingFace's CDN ignores it: probed today, a deliberately stale validator
still answers 206 with the ranged bytes rather than 200 with the whole
body. So the check is done here instead. A partial now has an identity file
beside it holding the ETag it was written against, written before the body
so an interrupted download still knows what it is a piece of. On resume,
the response's ETag is compared against it, and a mismatch throws the
partial away and asks again from zero. A partial with no identity at all is
not resumed either -- it could be a fragment of anything.

The sha256 HuggingFace publishes is now also checked before the file gets
its real name, so a bad one is never offered to be run. That is
belt-and-braces after the above rather than the primary defence, which is
the right order: detecting corruption after downloading gigabytes is worth
far less than not creating it.

Verified by planting one: a 60 MB partial of random bytes with an
identity file naming a revision that does not exist. The server logged
"changed upstream since the partial was written -- starting again",
restarted from zero rather than appending, and the finished file's sha256
matches the published one. Repeated the honest resume too -- cancel at
145 MB, restart, resume at 162 MB, correct hash.

Thanks to dev-updater's session for the If-Range idea and for saying to
confirm the CDN honours it rather than assume, which is exactly what it
turned out not to do.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017xn8nHw1tw1R6PtiY1eEtw
2026-08-28 05:06:45 -04:00
irisandClaude Opus 5 9d29776f02 Download GGUF models from HuggingFace, watchably
The first half of the llama.cpp work Bryan asked for: browse HuggingFace,
fetch a model, and see how far it has got from any device.

The design is dev-updater's build-progress shape with the four changes its
author recommended after living with it, since a model download is an hour
where a build is two minutes:

- **A run has an id.** Without one "not downloading" means three different
  things -- finished, never started, or someone else's run ended while you
  were away -- and over an hour that ambiguity is certain rather than
  theoretical. A device compares the run it was watching to the run
  reported now.
- **Outcomes outlive their run**, so a phone that was asleep at the moment
  of completion can still find out what happened.
- **Cancel exists.** Retrofitting cancellation into a blocking loop is
  miserable, and several gigabytes over someone's data plan is not
  something to have no answer for.
- **Progress is bytes, not a parsed marker.** We own the loop, so it counts
  directly; `total` is whatever Content-Length said and nothing else, and
  stays absent when the server sends none rather than becoming a bar drawn
  from a guess.

The download owns its own thread rather than the blocking pool, which
exists for short work. It resumes through HTTP Range, and trusts the 206
rather than the request -- a server that ignores Range answers 200 with the
whole file, and appending to that would corrupt it. `truncate(false)` on
the open is load-bearing for the same reason and says so.

Searching is proxied through the server rather than done from the phone,
because the app trusts exactly one certificate -- this one -- and the
machine that must do the downloading is also the one whose view of what
exists matters.

Verified against the real HuggingFace, not a mock: searched, listed a
repository's GGUFs, downloaded 234 MB with live byte progress, cancelled
mid-flight, confirmed the partial survived, restarted and watched it resume
at 162 MB rather than 0, and let it finish. The result's sha256 matches the
one HuggingFace publishes for that file, so the resume is byte-correct and
not merely the right length. llama.cpp then loaded it and ran inference.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017xn8nHw1tw1R6PtiY1eEtw
2026-08-28 05:01:53 -04:00
irisandClaude Opus 5 4004183acf Answer the "how do we test SSH here" question by testing it
The status section had carried an open question since phase 3: SSH could
not be exercised in this VM because there is no second machine and no key
in ~/.ssh. There is a second machine, though -- this one. Ssh it to itself
with a throwaway key and a host of bob@127.0.0.1, point the provider's
command at /bin/echo rather than claude, and the whole path runs:
connection, remote exec, and the process's death arriving as
`status: exited` in the transcript. It costs no tokens and touches nothing
real, and the key comes back out afterwards.

Done that way just now against the new transport, so the technique is
written down as something that worked rather than something that should.

Also recorded: the login shell in this VM is fish. `cd '…' && exec '…'`
is valid there and the POSIX single-quote escaping happens to mean the
same thing, but both are luck, and a non-POSIX remote shell is the first
thing to suspect if an argument is ever mangled on the way over.

Phase 4 is no longer deferred -- Bryan asked for llama.cpp today -- so the
status paragraph stops saying it is.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017xn8nHw1tw1R6PtiY1eEtw
2026-08-28 04:51:57 -04:00
irisandClaude Opus 5 4cdcbd204a Put the transport above the drivers instead of inside one
ClaudeDriver::spawn called ssh::command itself, so a translator whose job
is a wire format also knew how sessions reach other machines, and every
future driver would have had to remember the same. It now emits a `Launch`
-- program, arguments, working directory -- and hands it to a `Transport`
the manager chose from the session's host.

This is the inversion Bryan asked for, and it pays for itself immediately
in a place I had reported as a UI bug: "Run on" is offered for every
provider but only the Claude driver honoured it, so choosing a host for an
echo session silently ran it locally. With the transport above the driver
that cannot be written -- EchoDriver builds no Launch, so there is nothing
to wrap and nothing to misreport. The picker still needs to stop offering
it, but the code no longer lies underneath.

crate::ssh keeps the quoting, the forced options and the remote script,
with its tests; transport.rs only decides which of the two it is. The two
failure messages move with it, since they are transport-specific -- a
missing ssh client here is a different thing to check than a program
missing from a remote PATH.

Noted in transport.rs rather than built, because nothing needs it yet: a
remote llama-server is spawned as a process but spoken to over HTTP, so a
transport eventually needs "reach this port" as well as "run this".

Verified: cargo test (35), clippy, fmt. Nothing outside transport.rs
mentions ssh now.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017xn8nHw1tw1R6PtiY1eEtw
2026-08-28 04:49:55 -04:00
irisandClaude Opus 5 ba6c770be3 Record the setups model, and put the transport above the drivers
Two decisions from Bryan today, written down before they are built, since
this file is where the reasoning is supposed to live rather than in a
conversation.

A setup is a machine carrying the providers that machine has, and spawning
picks a setup then one of its providers. The independent providers × hosts
model it replaces is left in place below it, because its reasoning is worth
keeping and the code still implements it. What that model got wrong is that
the axes are not independent: it offers combinations that cannot work, and
"Run on" is already a control that does nothing for the echo driver, which
takes no host at all.

The transport wraps the driver rather than the driver reaching for the
transport. ClaudeDriver::spawn calls ssh::command itself today, which puts
transport knowledge inside a translator whose job is a wire format, and
obliges every future driver to remember the same. Inverted, a driver that
emits no command has nothing to wrap, which is the same "Run on" problem
solved structurally instead of by a special case.

Also noted: llama.cpp is the case where "wrap a command" is not enough on
its own, since a managed llama-server is spawned but then spoken to over
HTTP, so a transport is "run this" plus "reach this port". And this
section's claim that remote attachments need scp was never true of the
code -- images are base64 inside the stream-json message in both
directions, so nothing has to exist on the remote filesystem.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017xn8nHw1tw1R6PtiY1eEtw
2026-08-28 04:30:50 -04:00
irisandClaude Opus 5 8bf99f4b76 Write down the UI polish noticed while cleaning up
One item so far: the session header's status sits tight against the right
edge while Back looks roomier, because the row's 8dp padding is measured
against a TextButton whose touch target is wider than its text. Not a
defect and not urgent, but the kind of thing that gets re-noticed and
re-diagnosed every few months unless it is written down once.

PLAN.md rather than a new file, since that is where this project's
decisions live, and this is a decision to defer rather than a bug to
track.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017xn8nHw1tw1R6PtiY1eEtw
2026-08-28 04:15:05 -04:00
irisandClaude Opus 5 2cefede450 Compose Multiplatform 1.12.0
The last four lint warnings were all this one thing, and they were right:
1.11.1 with 1.12.0 out. Nothing else here is behind -- material3 1.9.0 and
activity-compose 1.13.0 are both current, checked against Maven Central and
Google Maven today, which is also why the header comment's date moves.

Android Lint now reports no issues at all, from 11 errors and 8 warnings
this morning.

Verified by running it rather than by the build succeeding, since a Compose
minor can change how things draw: the session list, the usage screen and a
session transcript with a message sent through it all render as before.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017xn8nHw1tw1R6PtiY1eEtw
2026-08-28 04:09:44 -04:00
irisandClaude Opus 5 87a0ef1a2a Put the formatter and the linter in the routine that people actually run
AGENTS.md's "Checking your work" listed a typecheck for the app and tests
plus clippy for the server. Neither half mentioned a formatter, and nothing
mentioned Android Lint at all -- which is how a linter that was in the
build the whole time went unrun long enough to accumulate a crash.

Both lines now say the whole thing, and the app's reads as the counterpart
of the server's rather than a shorter version of it. The note about lint is
there because it is the one step a build does not do for you: nothing fails
if you skip it, which is exactly why it needs writing down.

The command in the app bullet is the one I ran to verify this commit.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017xn8nHw1tw1R6PtiY1eEtw
2026-08-28 04:03:48 -04:00
irisandClaude Opus 5 0c13090d70 Take ktfmt's defaults for the Kotlin half
The server half went to rustfmt's defaults earlier today; this is the same
move for the app, and the same reasoning. Kotlin ships no formatter with
the Gradle build, so the question was which to adopt: ktfmt is Kotlin-org
owned now (it moved from facebook/ktfmt), is a formatter rather than a
configurable linter, and has essentially nothing to tune -- which is what
rule 27 is asking for. ktlint's .editorconfig surface is the thing that
rule warns against, and detekt is static analysis, whose job Android Lint
already does here.

One setting, and it is a choice between the tool's own two styles rather
than a tuning: kotlinLangStyle() is the 4-space one, which is what this
code already was. The 2-space default would have reindented every file to
say nothing.

  ./gradlew :androidApp:ktfmtFormat   to apply
  ./gradlew :androidApp:ktfmtCheck    to verify

Formatting only. The one thing worth checking by hand was the generated
PEM constant, since a leading newline there costs Android's
CertificateFactory its preamble sniff and fails at runtime nowhere near
the cause: ktfmt moved `.trimMargin()` onto its own line and left the
template alone, and the regenerated constant still starts at the opening
quotes.

Verified after: ktfmtCheck, compileDebugKotlin and lintDebug all pass, and
the APK builds.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017xn8nHw1tw1R6PtiY1eEtw
2026-08-28 04:03:14 -04:00
irisandClaude Opus 5 dde5042b12 Run the linter the app build already had, and fix what it found
`./gradlew :androidApp:lintDebug` had apparently never been run. It
reported 11 errors, 8 warnings and 3 hints, and one of the errors was a
crash: UsageScreen formats its reset countdown with java.time --
OffsetDateTime and Duration, both API 26 -- while minSdk is 24 and core
library desugaring was off. On 24 and 25 that is a NoClassDefFoundError,
and the `catch (_: Exception)` around the code does not stop an Error, so
the usage screen would have died rather than degraded.

Core library desugaring is now on, with desugar_jdk_libs 2.1.5. Verified
by looking in the built APK rather than trusting the flag: it now carries
Lj$/time/Duration and Lj$/time/OffsetDateTime, the backported classes the
call sites are rewritten against.

The rest: the two KTX suggestions taken (SharedPreferences.edit's block
form, which cannot forget its apply(), and String.toUri), the three
autoboxing hints taken (mutableIntStateOf/mutableLongStateOf), and
androidx.core:core-ktx declared at 1.19.0 rather than inherited through
activity-compose, since this code now calls its extensions directly.

Two are suppressed with their reasons, both scoped to the one element.
DiscouragedApi on the scanner's screenOrientation, which is not a pin but
the removal of the library's landscape lock. MissingApplicationIcon,
because there is no icon yet and that is a decision to make later, not an
oversight -- an app with no icon is obvious to anyone who opens a
launcher, so the warning tells nobody here anything.

0 errors now. The 4 warnings left are one thing: Compose Multiplatform
1.12.0 is out and this is on 1.11.1. That is an upgrade to decide on, not
a defect, so it is left for its own change.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017xn8nHw1tw1R6PtiY1eEtw
2026-08-28 04:00:57 -04:00
irisandClaude Opus 5 78c054918f Don't offer an empty spawn form when the options never arrived
SpawnScreen kept its own `loading` flag and `error` string, the third
mechanism in this app for a state two screens already share. That was not
just untidy: on a failed fetch it set the error, left `providers` and
`hosts` at the empty lists they started as, and rendered the form anyway --
so "couldn't reach the server" arrived as a Provider row with no providers
in it, which is what a server offering nothing would also look like. The
error sat below both empty pickers.

It now holds `LoadState<SpawnOptions>` like the others: loading shows the
spinner, a failure reports and stops, and the form exists only where there
is something to fill it with.

The spawn action keeps its own error, renamed `spawnError` so the two can't
be confused again. They are different in kind and the distinction is the
one the session list just learned: a fetch that never answered leaves no
form worth showing, while a spawn the server refused leaves a filled-in
form the user still wants, so that one stays beside the button that
produced it.

Verified on the emulator: the form loading with both providers, the Claude
fields appearing when claude-cli is selected (which also exercises the
snake_case kind the phone now compares against), and the failure state with
the server stopped -- reporting alone, with Cancel still working.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017xn8nHw1tw1R6PtiY1eEtw
2026-08-28 03:45:42 -04:00