Commit Graph
150 Commits
Author SHA1 Message Date
iris a49120b0c8 Merge remote-tracking branch 'origin/main' 2026-08-29 22:45:22 -04:00
iris 69ef6f068a Colour the composer by what its buttons do, and let the read-out breathe less
Queue is the paper plane with a clock on it (`md-send_clock`) rather than
the plain plane plus the word: the pair is now told apart by the mark, which
is what an icon is for, and the word survives as the button's accessible
name where a screen reader still needs it.

The three composer buttons take their colour from what pressing one does --
green sends now, blue sends later, red takes the running turn away -- and
Stop becomes a filled button like the other two. Outlined said it was a
qualifier on the primary action; it is a second thing you can do about the
turn, and what separates them is the colour and the mark. The fills are
named in Theme.kt with their content colour stated beside them, because a
semantic colour has to carry its own contrast: these do not change with the
surface, so nothing will rescue a foreground that stops being readable.
Worth knowing when reading that file: the action greens and reds sit next to
a `runningColor` green and a `failedColor` red, which are *states*. Nothing
in one set is pressable and nothing in the other is a state, so a reader
never has to tell them apart.

The usage read-out loses its per-machine cards. A card is a step up the
surface ladder and inside a dialog -- already a raised surface -- the step
barely rendered while costing 16dp on every side. The machine and the
service it answered for are one small quiet line instead of a heading over a
subtitle, since the numbers underneath are what somebody opened this to see.

The gaps between the bars now go *between* them rather than after each,
which is what put a band of empty dialog above Close. The rest of that band
was AlertDialog's own spacing, fixed at sizes meant for a sentence of prose
and a decision, so this is a plain Dialog with the same container colour and
corner and spacing chosen for a dense read-out.

Looked at on the emulator: green send, then blue queue beside red stop
during a `/slow 20` echo turn, and the dialog over the live session. The
account had risen to 78% by then, which showed the five-hour bar and the
header glyph going yellow on real numbers rather than forced ones.
2026-08-29 22:45:18 -04:00
iris 1635fe97c8 Take the formatter's line wrapping in NerdIcons
Left over from running ktfmt across the merge: the comment reflows two
lines. No change to what it says.
2026-08-29 22:33:46 -04:00
iris bb191eec21 Merge remote-tracking branch 'origin/main' 2026-08-29 22:29:18 -04:00
iris e37e90a579 Let the server say what is waiting, instead of the phone remembering
A message sent into a running turn was drawn as a pending bubble from
screen state, so leaving the session or restarting the app showed nothing
waiting while the queue was full. Nothing waiting is what "there is
nothing" looks like -- the reader had no way to tell it from a queue that
had already drained, and Bryan hit exactly that: a message he sent
arrived, and his phone stopped showing it after a restart.

The server now records the waiting. `MessageQueued { id, text }` goes into
the transcript when a driver takes a message it cannot deliver yet, and
is resolved by the `UserMessage` carrying the same id -- the same shape
`CommandQueued` and `CommandSent` already had, so this is one more
instance of a mechanism rather than a second one beside it.

The message itself still lands where the session read it, which is what
the last change was about; only the *waiting* is recorded early. The two
are different facts and now have different events.

Paired by id rather than by text. The old code removed the bubble whose
text matched, so sending the same thing twice cleared the wrong one and
left a message on screen that had already been read.

Both drivers that can queue do it: the echo driver too, because the phone
now draws pending bubbles from the stream and a rig that skipped the
event would exercise a state the real app never sees.

Checked on the emulator: two messages sent into a `/slow` turn, then the
app force-stopped and relaunched -- both still drawn as waiting, in the
pending style, and both resolved into ordinary bubbles when the turn
ended and the session read them.

Still outstanding, and worth knowing: an entry outlives a *server*
restart in the transcript but not in the driver's memory, so a backend
restarted mid-queue would leave the bubble drawn with nothing coming to
resolve it. Before this change that message vanished from the transcript
entirely, so the failure is now visible rather than silent -- but it is
not yet right.
2026-08-29 22:29:12 -04:00
iris ff39ef5cf9 Say it in icons, and put the whole backend behind four tabs
Six things Bryan asked for, which turned out to be one change: the app had
no icon set, so every one of them was blocked on having somewhere for icons
to come from.

That somewhere is dev-updater's arrangement, ported: a Nerd Fonts subset
committed as an asset, drawn as text. `Gear.kt`'s hand-drawn canvas gear
argued against icon fonts because a system font may not have the glyph and
whoever gets the empty box is never the person who wrote it. The objection
is right about *relying* on a system font and the answer is to ship the
glyph, so the file is gone and its reasoning is restated in `NerdIcons.kt`
rather than deleted -- otherwise the next reader re-derives it. `md-cog` and
`md-refresh` are dev-updater's own codepoints, because a cog means the same
thing in both apps.

The root screen's four words under the title are now four tabs, and the two
that act on the whole screen -- settings and refresh -- moved up onto the
title row as glyphs. That row's old comment recorded that a fifth word would
have had nowhere to go; tabs also say something the words did not, which is
that sessions, import, models and setups are four views of one backend
rather than four errands. Refresh feeds whichever tab is showing. Import,
models and setups lose their headings and their Back buttons, since the tab
row is now both.

Usage is a dialog. It is checked *against* what you were reading -- "can I
start this" is asked with the transcript still on screen -- and it had no
navigation of its own, so the only thing its Back could mean was "put this
away". The button that opens it is a chart glyph coloured by the worst of
the machine's windows, so the row says whether the limits are worth opening
before anybody opens them.

One `quotaColor` now colours every bar that measures a quota: blue, yellow
at 75%, red at 95%. The session bar escalates where it used to sit blue at
every level, and the dialog's thresholds moved out of it. A download keeps
plain blue at every value -- it has no limit to approach, and colouring it
like one would say the opposite of what is happening. States that are not
measurements take the ordinary control colour, since blue is the low end of
this scale and would read as "checked, and fine" about a machine nobody
could reach.

Send and stop are the filled paper plane and the filled square. Send keeps
the word "Queue" while a turn is in flight, because that is what pressing it
then does, and an icon that does two things while looking identical would
promise something immediate and do something that waits.

Looked at on the emulator: all six glyphs render, the tabs and the system
back gesture between them, the dialog over a live session, and the bar
bands at 82% and 97% forced through a scratch build, since this account is
at 72/31/5 and would only ever have shown blue.
2026-08-29 22:27:54 -04:00
iris ba71c798f5 Keep what was typed, and ask before a switch that re-reads everything
Two things about the box at the bottom of a session.

**A half-typed message survived nothing.** It lived in `remember`, so
leaving the screen threw it away, and so did the system reclaiming the
app. `Drafts.kt` keeps it per session id and the box is seeded from it.
On the device rather than the backend, which is where this app otherwise
puts state so every device sees it: this is the contents of a text box on
the phone somebody is holding, written on every keystroke, and half a
sentence surfacing on another device would be a surprise rather than a
convenience. What has been *sent* is the server's, and that is the part
which has to outlive this phone.

**Switching model quietly re-reads the whole conversation.** The picker
did it on the tap, and the cost only showed up as the next turn being
expensive. It now asks first, in words, with no number: what it will cost
depends on how long this conversation is, and the screen does not know
that -- the running total beside it counts what has been spent, which is
a different quantity, and a figure derived from it would be a guess
wearing a measurement's clothes.

The picker beside it deliberately gets no dialog, and that is measured
rather than assumed. Driving one session through both changes and reading
the CLI's own usage: a warm turn read 30,771 tokens from cache and
created 87; after a *permission mode* change it read 30,858 and created
75 -- still a hit; after a *model* change it read nothing at all and
created 41,509. So the model picker is the whole of the set, and warning
on both would teach that these dialogs can be clicked through, which is
what stops the one that matters from working.

Nothing is asked when there is nothing to lose either: choosing the model
already set, or switching before the session has said anything, applies
straight through.

Checked on the emulator. A draft survived leaving the session and a
force-stop; the dialog names both models and both buttons; declining left
the model where it was; and the permission picker still applies on the
tap with no dialog in the way.
2026-08-29 22:16:02 -04:00
iris 76ba24993c Make the permission mode the picker offers actually selectable
The session screen's mode picker listed `bypassPermissions`, and on any
session not born in it, choosing it failed:

    Cannot set permission mode to bypassPermissions because the session
    was not launched with --dangerously-skip-permissions

The CLI is asymmetric about that mode and it is not obvious. It will
launch straight into it on `--permission-mode` alone -- so spawning into
it from the phone has always worked -- but it refuses to switch into it
afterwards unless the process was started with the flag. So the picker
offered a state the session could not reach, and the failure arrived
after the fact as an error line in the transcript.

Sessions now launch with `--allow-dangerously-skip-permissions`, which
makes that mode reachable without selecting it: the session still starts
in whatever mode it was asked for and only moves when somebody moves it.
Deliberately the `--allow-` form; `--dangerously-skip-permissions` is the
one that turns bypassing on for everything, which would take the choice
away from whoever is holding the phone. Since the mode was already
reachable at spawn, this withholds nothing new -- it makes the two routes
to it agree.

Measured both ways round against 2.1.237, driving the control request
directly: without the flag the response is `subtype: error` with the
message above, with it `subtype: success, mode: bypassPermissions`. Then
through the app's own route on a session spawned `manual`, whose argv
reads `--permission-mode manual … --allow-dangerously-skip-permissions`
and which reported `permissionMode: bypassPermissions` when asked to
change.
2026-08-29 22:05:26 -04:00
iris 2bf90daada Let the markdown renderer draw tables, and stop headings shouting
Two things a reply could not render, both from the same cause: the
renderer was pinned nineteen releases back.

Tables arrived in the library at 0.30.0. On 0.26.0 a GFM table was not a
table at all -- the rows fell through as text and ran together, pipes and
all. They now draw as a table, and scroll sideways when they are wider
than the phone rather than losing the last column.

Headings took the renderer's defaults, which are the Material *display*
styles: `#` came out at 57sp and `##` at 45sp, both larger than this
app's own screen titles, so any reply with a heading in it read as
shouting. They now descend from headlineSmall to labelSmall -- six steps,
every one a different size, so two levels of nesting never draw the same.

The pin was not carelessness, which is the part worth recording: the
version comment says Maven Central was checked on 2026-08-29 and 0.26.0
was the newest stable. It still answers that, because
`search.maven.org/solrsearch` is stale for this artifact -- it knows
nothing past 0.27.0-rc02. `maven-metadata.xml` in the repository itself
lists up to 0.45.0, updated 2026-08-28. The comment now says to read the
metadata rather than the search API, since the same check will otherwise
be made the same way next time.

The colour mapping moved with the API: `markdownColor` no longer carries
`codeText`, `inlineCodeText` or `linkText`, which now ride on the
typography as the style's own colour and a `TextLinkStyles`. Same
Catppuccin values as before. `tableBackground` is set to the tint code
blocks use rather than the library's 2%-alpha default, which on this
surface was invisible.

Checked on the emulator against a reply carrying all six heading levels,
inline code, a link, and a three-column table -- including scrolling the
table to confirm the clipped last column is reachable rather than lost.
2026-08-29 21:54:16 -04:00
iris 6b4911c67b Give the session's own state a line, instead of the transcript's corner
The token total floated over the bottom-right of the transcript, where a
long message ran underneath it, and the working indicator was an item
inside the list -- so it scrolled away exactly when somebody reading back
wanted to know whether anything was still happening.

Both are facts about the session rather than turns in it, so they get one
row directly above the box you type into: the thing they report on is the
next thing you touch. `exited` moves with them, since it is the same kind
of fact and nothing else on the screen would have said it once the
indicator left the list.

The row is drawn whether or not it has anything to say. An empty one
costs a line; a row that came and went would move the text box under a
reader's thumb every time a turn started, and would make its own presence
the signal for a state it never names. For the same reason the compaction
case had to fit the same single line: its bar now takes the row's free
width between the label and the total rather than a row of its own, which
keeps it far wider than a spinner -- the reason it is a bar at all, since
nothing arrives in the transcript while a compaction runs and a small
moving thing there reads as a session that has hung.

Still no fraction to fill, re-measured today rather than assumed: a real
80,346-to-2,088-token compaction took 23 seconds and the CLI emitted not
one line between saying it had started and saying it had finished.
Elapsed seconds remain the only honest number.

Looked at on the emulator in all three states -- idle, working, and six
seconds into a real compaction -- and at 320dp, the narrowest width a
phone actually has, where the row still holds one line.
2026-08-29 21:35:30 -04:00
iris 549e49bc10 Record a steer where the model read it, not where it was typed
A message sent while an answer was streaming was recorded in the middle
of that answer and above the tool call it ended with. The model had
committed to that call in the same message it was already writing, so it
had read none of it -- and on screen the tool result underneath read as
something the steer had asked for. The answer also split into two
bubbles around a message that was not part of it.

The driver announced a steer at "the next assistant text or tool call",
on the reasoning that anything the CLI says next is proof it has been
round the model again. With --include-partial-messages that is not true:
the deltas and the tool_use block of a message already in flight keep
arriving afterwards, and none of them saw the steer.

`message_start` is what actually proves it. The CLI sends the previous
call's tool results back before it opens the next assistant message, so
that line is the first moment anything written since can have been read
-- and it carries no events of its own, which is what makes it a place
to put one. Verified against 2.1.237: message_start, the blocks, the
tool_result, then the next message_start.

The end of the turn stays as the other half, and is the case that must
not be lost: a message typed after the final model call has no later
message_start, and one that is only recorded when announced would
otherwise vanish while a phone drew it as still waiting.

Checked live on haiku, before and after. Before: the steer landed at
seq 37 among the essay's deltas, with the tool call at 45 and its result
at 46. After: essay whole, tool call 42, result 43, steer 44. Also
checked the case this had no reason to touch -- a steer sent during a
30-second Bash call, which was already correct -- and it still records
after the result. The two tests fail on the old rule; the failure prints
the old order, which is the bug.
2026-08-29 21:16:21 -04:00
iris a50d72960c Say how long is left, not that the window is five hours
The bar read "31% of 5h", which is the one thing about the window a
reader already knows. What decides whether to start something now is how
long what is left has to last: 80% with twenty minutes to go and 80%
with four hours to go are opposite answers, and the second number was a
screen away on the usage screen.

It now reads "31% - 2h 36m left", and the countdown is driven by a clock
the refresh loop advances rather than computed at draw time. A
percentage that comes back unchanged is an equal value, so Compose skips
the recomposition -- a "left" recomputed only when the quota happens to
move would have sat at a stale figure for hours while looking live.

A window can arrive with no reset time, so that keeps its own wording:
"reset time unknown" rather than "refresh soon", which would be a
recommendation nothing measured. Under a minute, including past the end,
is "refresh soon" -- "0m left" reads as a measurement.

The span arithmetic was already on the usage screen, so it moves into
`ResetCountdown.kt` and both callers supply their own sentence. That
screen still reads "resets in 2h 37m" and "resets in 5d 21h", checked on
the emulator alongside the bar it was not part of changing.

The fill is blue rather than the scheme's primary: the bar sits under
every session header, on a screen somebody opened to do something else,
and it reports a quantity rather than a verdict. The usage screen is
still where the same number turns yellow and then red, for a reader who
went there to be told where the limits are.

Also declares this project's resources for Dev Updater, whose
declaration schema changed in d27b5a3: `resources.ron` says ai-app keeps
its state as `ai-app`, so the Uninstall dialog offers the real
directories instead of saying it cannot tell where they are. Only the
name, because both XDG places are the conventional ones. What that
dialog's config toggle would delete includes the CA under `certs`, which
strands every phone running an APK pinned to it -- noted where somebody
would be standing when it matters.
2026-08-29 20:54:55 -04:00
irisandClaude Opus 5 894180de77 Take the CLI's word for a clear instead of inferring it
Bryan reported no divider when clearing. It was not this code -- the
backend serving him started at 17:06, three hours before `Event::Cleared`
existed, so it has no such event to send and `/clear` reaches it as an
unrecognised passthrough. Verified against a current build: the event is
recorded.

Probing the CLI to establish that turned up something better than what
was here. `/clear` in stream-json mode emits a dedicated
`conversation_reset` line and *then* a fresh `init` with the new session
id -- so watching the id be replaced, which is what this did, was reading
the event through one of its side effects. The announcement says it
directly, and it arrives first, so the divider now lands above the new
conversation rather than after its opening line.

That also removes the reasoning the previous commit needed about which id
changes count. There is one signal now instead of an inference with two
exceptions, and the test that used to pin those exceptions became
`an_init_alone_is_never_a_clear`, which covers all three ways an init
arrives: a session's first, the one a compaction re-announces with the
same id, and the one following a resume.

The resume token still follows the id, unchanged -- one CLI event with
two observable effects, and each half now reads the half it needs.

Verified end to end against a real claude-cli session: message, /clear,
message, and the transcript reads userMessage / assistantText / cleared /
userMessage, in that order. 74 tests, clippy and rustfmt clean.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VETa8afmpWaYezLCqJhDB8
2026-08-29 20:39:21 -04:00
irisandClaude Opus 5 71067275e2 Stop claiming a terminal, and stop calling a recoverable delete final
Two bugs Bryan hit, with one shape between them: a claim stronger than
the thing that was measured.

**The delete warning branched on `imported`.** It told him deleting
`ai-app` could be undone and deleting `manager` could not, when both are
claude-cli sessions whose conversations survive equally. `imported`
records how a session got into the app; what decides recoverability is
whether the *driver* keeps its own record -- the Claude Code CLI does,
under ~/.claude/projects, however the session started; echo and llama.cpp
do not, and for those the app's transcript is the only copy. So the fact
now sits on DriverKind and rides on SessionInfo, decided by the server
from the provider's kind rather than by the phone from its name, which a
person can change.

The comment above the branch asserted "a session started here has no copy
anywhere". That sentence was the bug written down and reasoned from, and
it is gone.

Neither branch promises a restore, which it should not: nothing here
checks the file is still on disk, and re-importing was never a restore
anyway -- this app's transcript holds images, peer messages and command
events the CLI's record never had. So the recoverable text says what is
known and names what goes either way. "Can't be undone" is now said only
where it is true, which is the point of saying it at all.

**"open in a terminal -- close it there first" named a place that need
not exist.** The detection is right and worth keeping: something live
holds that session, and importing it would reproduce the double-resume
incident. But which something was never measured. The live descriptors
here include two of this backend's own adopted sessions and a peer
agent's; none is a terminal, so the instruction sent the reader looking
for a window that was not there.

**And this app did not recognise its own spawned sessions.** The import
list filters out what the app is already driving, but it matched only the
import cursor -- which exists solely for imported sessions. Every session
the app spawned therefore stayed in the list, marked in use, telling the
reader to go and close it somewhere: here. Matching the resume token too,
which both kinds have, is the fix; `session_importing` is now
`session_driving`, because that is what it was always being asked.

Verified on the emulator against a scratch backend: a spawned claude-cli
session reports keepsOwnTranscript true with imported false -- Bryan's
`manager` case exactly -- and draws the recoverable warning; the echo
session draws "can't be undone"; and once the CLI named itself, the
spawned session's id was absent from the import list, where the old match
would have listed it.

75 tests, clippy and rustfmt clean; ktfmt, compileDebugKotlin and
lintDebug clean.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VETa8afmpWaYezLCqJhDB8
2026-08-29 20:34:03 -04:00
irisandClaude Opus 5 5f6fd34aba Record the clear that happened, not the one that was asked for
`ClaudeDriver::clear` emitted `Event::Cleared` beside the `/clear` it
sent, so the divider recorded a request. A reader scrolling back takes
that mark as a fact about the conversation -- the session no longer has
what is above this -- and a request is a different claim from a result.
Compaction already gets this right by taking its mark from the CLI's own
`compact_boundary` rather than from somebody pressing Compact; this is
the same rule, and it was the one place left applying it to the request.

Caught in review by the session this was measured against, which also
established that the CLI's `/clear` is declared `supportsNonInteractive`
and returns empty text with no result line -- so a fresh `init` bearing
a different `session_id` is the only trace it leaves. The reader already
watches for exactly that in order to persist the resume token, so the
mark now goes out there.

It has to be a *replacement* rather than any change, and the tests pin
both ways of getting that wrong. The first `init` sets the id from
nothing, which would otherwise open every session with a divider
announcing a clear that never happened. And a compaction re-announces
`init` carrying the *same* id, which would otherwise draw a clear on top
of the compaction's own mark -- that one was found by writing the test
rather than by reasoning about the change.

73 tests, clippy clean, rustfmt clean.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VETa8afmpWaYezLCqJhDB8
2026-08-29 20:20:49 -04:00
irisandClaude Opus 5 cf3ed68511 Tell "nothing to meter" apart from "couldn't find out"
Found by looking at the bar on an echo session rather than by reading
the diff: it said "5-hour usage unknown -- this machine reports no
usage", which is the failure the rest of this file was written to avoid,
one level up.

A machine with no metered provider is never asked by the backend, so it
returns no snapshot for it. The bar read that silence as a failed
lookup, because Unavailable was the nearest word it had -- and a session
on `echo`, or on a local llama.cpp, has no paid quota at all. That is a
fact about how somebody set the machine up, not a question that went
unanswered, and reporting it as unknown nags about a deliberate choice
on every screen forever.

So the state exists now: NotMetered, drawn as nothing, because there is
nothing. Unavailable keeps its words and its reason and still covers the
three ways an answer can fail -- nobody logged in, machine unreachable,
snapshot without the window.

Verified on the emulator against the real endpoint: a setup carrying
claude-cli draws the bar at 22% of 5h, selected by kind "session"; the
no-snapshot path was the one on screen before this change, so it is
reached, and this only changes what it draws.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VETa8afmpWaYezLCqJhDB8
2026-08-29 20:16:31 -04:00
irisandClaude Opus 5 6680fbb987 Clear from the phone, and put the numbers where they are read
Four changes to the session screen, three of them Bryan's and one that
fell out of them.

**/clear is offered like any other session command.** It joins
SESSION_COMMANDS, so it suggests itself while being typed and goes out
through the command endpoint that /compact already uses -- no new path,
and the boundary pump holds it mid-turn exactly as it holds a compaction.

**A clear draws a divider, not a deletion.** `Event::Cleared` becomes a
ClearedNote row saying that everything above stays here and is no longer
sent. That sentence is the row's whole job: the reader can see the
conversation is still on screen, so without it the divider reads as
something having been thrown away, which is the one thing it is not. It
carries no counts, because nothing was measured -- a compaction's
numbers are real and there is no equivalent here to report.

Compaction and clear now share `TranscriptDivider`. They are the same
kind of mark to somebody scrolling back -- "the session no longer has
what is above this" -- and the difference belongs in the words rather
than in how they are drawn, so the styling is written once and cannot
drift.

**The five-hour usage bar sits under the session header.** It reports
the paid service's own metering for the machine this session runs on,
fetched from that machine, refreshed every minute off the backend's
cache. It is never derived from the transcript's token counts: those are
a different quantity measured differently, and a quota-shaped bar built
out of them would be a guess wearing a measurement's clothes. Not
knowing has its own appearance and its own words -- "unknown" and why --
because a bar resting at zero because a machine is unreachable reads as
plenty of headroom, which is the opposite of the truth. The window is
selected by the API's own `kind` ("session"), added to UsageWindow in
this change, rather than by matching the label a person reads.

**The token total moved from the header to the bottom right of the
transcript**, pinned above the input rather than scrolling with it. In
the header it was one item in a run of dot-separated facts about the
session and read as another of them, rather than as the running total it
is.

SessionSummary now carries the setup id, which it deliberately did not.
The stated reason was that nothing here addressed a setup and holding
both id and name invited showing the wrong one; the usage bar addresses
one, so the reason lapsed rather than being overruled, and the comment
now carries the rule that replaces it: never display it. The server has
always sent the field, so nothing changed on the wire.

ktfmt, compileDebugKotlin and lintDebug all clean.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VETa8afmpWaYezLCqJhDB8
2026-08-29 20:11:26 -04:00
irisandClaude Opus 5 1e3eec8cf0 Let a session drop its context without ending
Adds `Event::Cleared`, `SessionCommand::Clear`, and `Driver::clear`, so
`POST /sessions/{id}/command {"text": "/clear"}` does for a session what
the CLI's own `/clear` does for a terminal.

The marker is a divider, not a truncation: everything above it stays in
the transcript, because that is the only copy of the conversation the
phone has and a person scrolling back is a different question from what
the model is given. It also makes clearing mean one thing across
drivers -- `claude` sends `/clear` and the CLI answers with a fresh
`init` whose new session_id the reader already persists as the resume
token, so the next launch resumes the cleared conversation with nothing
to keep in step; `llama` needs no state at all, since `conversation()`
already folds the transcript and now folds from the last marker; `echo`
emits the marker alone, so the phone's divider and scroll behaviour can
be exercised without spending a real session's context.

That fold is why `Cleared` is documented as load-bearing rather than
decorative. For any driver that rebuilds its conversation from the
transcript, this marker decides what the model sees, and treating it as
something only the phone draws would silently put the cleared
conversation back in front of the model at full price.

Clear rides the existing boundary pump like any other SessionCommand, so
one arriving mid-turn waits exactly as a compaction does, and nothing
grows a second way to wait.

Removes `--autocompact` in the same change, because clearing is the
cheaper answer to the problem it was added for and Bryan would rather
manage context that way. Keeping the measurement here, since it was the
reason for the constant and is worth more than the constant was:
context returned to 70-85k within ten calls of a compaction; a
compaction took 104,346 to 147,671 ms; compaction cost that session
2,655,508 tokens across six boundaries, of which the single automatic
one at the 1M ceiling was 1,696,870. Clearing costs nothing, because
nothing is sent.

70 tests, clippy clean, rustfmt clean.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VETa8afmpWaYezLCqJhDB8
2026-08-29 20:06:04 -04:00
irisandClaude Opus 5 797513bb86 Raise the compaction window to 200k
100k is the cheapest window on tokens and the wrong one to sit in front
of. Measured on the session this was written against: context returns to
70-85k within ten calls of a compaction, so a 100k window compacts about
every thirteen calls, and a compaction takes roughly two minutes
(durationMs 104,346 to 147,671 across the six recorded). A 130-call
request would have spent some twenty minutes compacting -- optimising
the number that was asked about while making the thing somebody actually
waits for on a phone considerably worse.

200k keeps most of the saving against the 1M ceiling and halves the
stalls.

The comment now also says what the window does not do, because measuring
this turned up the opposite of what the byte counts suggested. Images are
93% of the bytes that tool calls put into that transcript but only 8% of
the context growth -- the adb wrapper's downscaling holds a screenshot to
a median of 476 tokens, while text-only calls add a median of 740 and a
mean of 1,139. So the file is large because of screenshots and the
context is large because of ordinary tool output, and only the second one
is what this constant governs.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VETa8afmpWaYezLCqJhDB8
2026-08-29 19:43:43 -04:00
irisandClaude Opus 5 4f586d6551 Attribute the drift to the CLI's ceiling, not to a manual compaction
The reasoning on AUTOCOMPACT_WINDOW cited 491,562 tokens as where "the
CLI compacted it". That was a manual /compact somebody ran, not the
CLI's own trigger, so the comment credited a person's intervention to
the automatic behaviour it was arguing about. Caught in review by the
session whose transcript it was measured from.

Corrected from that transcript's compaction boundaries: the window left
to `auto` was 1M, and the one automatic compaction fired at preTokens
1,000,184 with the API context peaking at 999,668. So the drift ceiling
is twice what the comment said, and near it a single tool call bills
about 100k tokens rather than 49k.

The correction strengthens the case, but it also changes what the
example is evidence *of*, which is why it was worth fixing rather than
just raising the number: what held that session together was the person
in it running /compact by hand four times, and the 4.2-million-token
request happened at the merely-large contexts left between those. The
constant is for the sessions where nobody is doing that.

No behaviour change. 68 tests, clippy clean, rustfmt clean.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VETa8afmpWaYezLCqJhDB8
2026-08-29 19:11:51 -04:00
irisandClaude Opus 5 36984a9e1b Put the compaction window's reasoning on the compaction window
4a40578 inserted AUTOCOMPACT_WINDOW between an existing doc comment and
the constants it described, so rustdoc attached "the session directory's
copies of the process's standard streams" to the compaction window and
left STDIN_FIFO, STDOUT_LOG and STDERR_LOG undocumented. Moving the new
constant below them restores both.

Verified: 68 tests, clippy clean, rustfmt clean.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-29 19:10:39 -04:00
irisandClaude Opus 5 4a4057886d Compact at 100k rather than letting a phone session drift
Left at `auto` the CLI picks a very large window, which suits a terminal
session somebody closes at the end of the day and does not suit this app
at all: these run for hours, nobody closes them, and the transcript
carries screenshots. One session here reached 491,562 tokens of context
before the CLI compacted it.

That matters because every API call re-reads the whole context, and one
request is not one call. At half a million tokens a single tool call
bills about 49k before it does anything, so "can you make it so you can
rename a session?" cost 4.2 million tokens across the 130 calls it took.
Measured over that session's life: 2,498 calls, 1.08 billion cache-read
tokens.

100k is the smallest window the CLI accepts and roughly the cheapest.
Per-call cost falls with the cap, while the compaction it forces costs
about the same in total either way -- a smaller window compacts more
often, but each pass is proportionally smaller. What it trades is how
much detail survives a compaction, which is a real cost to the work and
the reason this is one named constant with the reasoning written down
rather than a computed value.

Passed before the resume/name branch, so it applies to adopted and
imported sessions too -- which are the large ones, and the ones this is
for.

Verified: the exact argument list the app now spawns starts, accepts an
empty stream-json stdin and exits 0, so the flag combination is good
without spending a token. 68 tests, clippy clean, rustfmt clean.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VETa8afmpWaYezLCqJhDB8
2026-08-29 19:09:27 -04:00
irisandClaude Opus 5 620a7d0a83 Trim what every session pays to load
Measured: an app-spawned session starts at ~33,200 tokens of context, of
which ~20,000 is written fresh on every spawn -- the always-loaded rule
files and this file -- and only ~13,200 comes from a shared cache. That
20,000 is billed at 1.25x on every single session start.

This file drops to 19,882 bytes from 21,293. What went is narrative that
PLAN.md already carries in more detail (the phase history, the submodule
drift story) and the parts of "Where things run" that MACHINE.md states
once for every project. What stayed is every operational fact: the
commands, the llama.cpp and ssh test recipes, the import rules, and
everything under "Things that have bitten".

The global chain was trimmed in the same pass, 43,039 -> 34,069 bytes,
mostly by moving the Gentoo host build profile out of the @import chain
into ~/.claude/HOST_BUILD.md, which MACHINE.md now points at. Nothing was
deleted there either; it is referenced rather than loaded, the same
arrangement this file has with PLAN.md.

Worth being honest about the size of the win: ~10,400 bytes is roughly
2,200 tokens off each session start. It is real and permanent, but it is
not what makes a long session expensive.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VETa8afmpWaYezLCqJhDB8
2026-08-29 19:00:59 -04:00
iris 749b2db287 Make a command a thing the app knows, and hold it until it can run
Typing "/" now suggests what this app understands -- `/compact` and
`/rename <name>` -- with a line each about what they do, and anything
else beginning with a slash is passed to whatever runs the session,
because a dialect's own vocabulary grows without this list.

None of them are messages, and that is the substance of the change. A
line written into a running turn is read by the *model*, so a command
sent mid-turn either does nothing or arrives as text somebody has to
puzzle over. They now wait for the turn to end. The waiting is done once
for every provider, in the pump that already watches every event for the
boundary, rather than in each driver where a new provider could get it
wrong by leaving it out.

Waiting is a state, so it is on screen: the command sits at the reader's
end of the conversation in blue, with a spinner and "waiting for this
turn to end", and becomes an ordinary blue row when it goes. Blue
because these are about the session rather than about the task -- the
same blue a compaction already used, which is now one colour with one
name rather than two.

Renaming from the settings screen sends exactly this, so it waits and
draws the same way. The name itself is not held: it is this server's own
datum, so the list and the header change at once and only telling the
session waits.

Echo grew the same split, which is where the bug in it showed: its
commands are its messages, so running one announced a `MessageTaken` as
well, and the same line drew twice -- once blue, once purple. A command
owes no announcement; the manager has already recorded that it was sent.

Watched rather than reasoned about: `/compact` during a 25 second turn
held with its bubble up, went out when the turn ended, and the
compaction that followed reported what it recovered.
2026-08-29 17:10:21 -04:00
iris bebaae7a94 Carry a question in the event model, not in one provider's JSON
A question is now fully described by the event that reports it: the tag
it was asked under, each option's label, what it means, and the sample of
what picking it would produce, plus whether several may be picked at
once. The app renders from that alone.

It had been reading Claude Code's tool input to find the parts the event
dropped -- that dialect's schema, written out a second time in Kotlin,
where no other provider could reach it and where it would drift the
first time the schema moved. Echo could not describe an option at all,
and llama never will.

Answers travel as a list for the same reason. A question that takes one
answer sends a list of one rather than being a different shape, and the
one place that flattens it is where the CLI is spoken to: its answers
map holds a string, so several choices are joined there. That join was
in the phone.

Also here because it is the same rule: the permission ask reuses the
question body rather than owning a second one, so Allow/Deny renders and
resolves through exactly the code an AskUserQuestion does.

Verified against both, since a refactor that only satisfies the case it
was written for has been tried on the half that cannot fail: a two
question `/ask` answered from the phone, one option and then two, and a
real sonnet session's `rm -f` permission asked, allowed, and run.
2026-08-29 16:46:43 -04:00
iris fea8e7e92b Show every option a question offers, on the call that asked
Reported by Iris through the dev-updater session: a two-question
AskUserQuestion arrived with only one option visible per question, so
the answer she sent was the only one she had been offered.

The cause was a `Row`. It hands out intrinsic widths in order and clips
whatever runs past the edge, so the first option or two drew and the
rest went off the side of the screen -- which does not read as a bug, it
reads as those having been the only choices. The same Row was in the
permission ask beside it; both wrap now. That pairing is the reason to
look: a rule stated on one member of a set is usually missing from the
others.

The rest of what she asked for, and what each was:

- It drew twice, as the tool call and again as loose question cards,
  because the backend marked these questions as belonging to no call.
  They belong to the call that asked, and now say so.
- So it renders like any other tool: one card, its own heading, opened
  because a decision cannot be made from a closed row.
- Each option shows its description and its `preview` block, which is
  the part a reader is deciding on and none of which was reaching them.
- "Other" is a field on every question. The harness always offers it, so
  leaving it out narrowed a question that was never that narrow.
- A multi-select sends the labels it collected as one string, which is
  the tool's own schema rather than a guess -- its answers map is
  string-valued.
- No spinner while it waits. A spinner says the machine is working; here
  the machine is idle and the turn is stopped on the person, so the card
  says "your turn" in the colour this app already uses for that.

Verified against a real session as well as the echo fixture: haiku asked
two questions with three described options each, both were answered from
the phone, and the model carried on with the answers. Echo grew `/ask`
so the shape can be looked at without paying a model to produce one, and
its option cards are outlined rather than tinted -- as one surface step
up they were three paragraphs where three things to press should be.
2026-08-29 16:22:15 -04:00
iris cae04c2559 Delete one session without putting the rest through loading
Pressing Delete refetched the whole list on success, so every other row
went back through its loading state and the reader got a blank screen
for the length of a round trip -- to report on something that was never
in doubt. Now the row being deleted fades, says so where its status
goes, and stops responding to taps; when the server answers, that one
row is removed and nothing else moves.

A refusal keeps the row, because it is still there: the server answered
and said no, so the session it said no about is exactly as it was, and
the error goes on its own card as it already did.

Faded rather than removed on the way out, deliberately. Taking the row
away when Delete is pressed is a promise about a request that has not
been answered, and putting it back when the server refuses is worse than
never having taken it away.

Looked at rather than reasoned about: the in-flight state lasts
milliseconds against a local server, so I slowed the delete route to
four seconds, watched the faded row and its spinner, watched it removed
on success, then killed the server and watched a refusal leave the row
in place with the reason on it.
2026-08-29 16:11:18 -04:00
iris d3fff3d229 Let a session be renamed, under the same name everywhere
A gear at the end of the session's own bar opens what can be changed
about that session; the name is the first thing there. Compact is gone
from that bar -- `/compact` typed into the message box is the CLI's own
way to ask and it already worked, so the button was a second way to say
one thing. Echo takes the typed word too now, since it is the rig the
compaction display is checked against and losing the button would have
taken that with it.

The name is this server's, not a driver's: it is what the list shows, it
exists before any process does, and every provider has one. So it is
settled in the config and the driver is *told* -- which is the opposite
of the model and the permission mode, and the difference is written down
at `Driver::set_title`. A driver whose process has no notion of a name
does nothing and says nothing, because there is no failure to report.

Claude Code has one, so the name reaches it: `--name` for a session we
create, and `/rename` afterwards, which is a local command rather than a
control request -- `set_session_name` is not a subtype it knows, which I
established by asking it. A resumed session is deliberately not renamed
at launch: an import already has a name, quite possibly one the person
typing in it chose, and taking that would be helping itself to something
the app was only shown.

Verified end to end rather than argued: renaming from the phone put
"Session renamed to: paging and scroll" in the CLI's own session file,
and the session now lists under that name to other agents.

The gear is drawn rather than set in a font, for the reason Chevron
gives. It was a sun on the first attempt -- thin teeth standing clear of
a thin hub -- which no amount of reading the diff would have shown.
2026-08-29 16:06:15 -04:00
iris 1629e0911e Say what the line splitter would do with a bare carriage return
`complete_lines` splits on `\n` only, which is right -- this stream is
JSONL, and a record terminated by a bare `\r` would not be a record --
but the doc comment said why the remainder is held without saying what
decides where a line ends.

Worth the sentence because of what the failure would look like if the
CLI ever wrote such a line: the session goes quiet, the process is
healthy, nothing errors, and the cause is a line splitter. The
dev-updater session hit exactly this shape today reading cargo's
progress line, which is `\r`-terminated for redrawing in place, and
lost a whole build's worth of output to it.
2026-08-29 15:42:32 -04:00
irisandClaude Opus 5 f842d0e512 Call the server component "server", to match dev-updater
READ BEFORE PULLING. A Managed component's service unit is named
<config key>-<component name>, so this rename moves the unit from
ai-app-backend to ai-app-server and nothing points at the old one
afterwards. Uninstall the backend component from its card *first*, while
it is still called "backend"; then pull, accept the new declaration --
.dev-updater.ron is a request, so the card shows it as pending -- and
build. The unit installs under the new name.

Two things reset rather than break, both keyed by component name: the
per-component built_from sha, and the build and runtime logs. One build
makes the sha current again.

Also drops the claim that the components list is walked in order. They
have built in parallel since 2026-08-28, so the reasoning the comment
gave -- backend first, so a failing APK leaves the phone what it had --
no longer describes what happens.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VETa8afmpWaYezLCqJhDB8
2026-08-29 15:40:31 -04:00
iris 3eccf7e443 Show the model and mode the session has, not the ones it was asked for
Picking either from the phone wrote the choice straight into the
session's state and then sent the request. Asking and having are
different things, and the difference is not rare: `auto` is a permission
mode the CLI accepts on the command line, silently resolves to
`default`, and refuses outright over the control channel -- "auto mode
unavailable for this model" -- so a session spawned in auto was in
default and one switched to auto stayed where it was, with the phone
reporting auto in both cases.

So the drivers report what they are set to and the manager follows that.
Measured, because the confirmations are not uniform: a model change
answers success with no value, so what was asked is remembered until the
answer arrives; a mode change echoes the mode it became, and that answer
wins over the request; and `init` names both -- resolving `haiku` to
claude-haiku-4-5-20251001 -- which also covers a session adopted from a
terminal that set them outside this app. A driver that cannot change
either already says so with an error, and now that error is the whole
story rather than a note beside a display that changed anyway.

The config keeps the requested value, deliberately: that answers a
different question, which is what to launch this session with next time.

Two things fall out. Control request ids are random rather than the
clock, because two in the same second shared an id and something now
looks them up. And the phone shortens a resolved name for the button --
`haiku-4-5` -- since the full one is what the CLI reports and roughly
twice the room that row has once Stop is in it.
2026-08-29 15:36:53 -04:00
iris 404066fa7d Say when a session is working, and what it was told
Three things a phone could not see, all of them the same shape: the
session was doing something and nothing on screen said so.

A turn nobody here started never reported itself. `Running` was sent
where a message was *sent*, so a session picked up mid-turn, one
compacting on its own, or one another agent wrote to sat there reading
as idle until it finished. The driver now says it from what it observes
-- output that could only come from a turn in flight -- which is the
same set of events that already announced a steer, with the ends
swapped.

An imported session had it worse: nothing but replayed lines ever
reaches it, and a status was not among them, so it was permanently
whatever it was when it was adopted. Its file does not record a turn
ending, but it does record why each assistant message stopped, and
`tool_use` versus anything else answers it. A record that says nothing
leaves the status alone rather than voting for idle.

Messages from other agents were dropped outright: the CLI marks them
meta, and this replayed everything except meta. They are now a row of
their own, closed by default like a tool call, named for the session
that sent it -- not the reader's own bubble, because they did not say
it, and a session working on something this phone never asked for is
exactly what one of these explains.

Measured against a real session file rather than guessed: the peer
record carries the sender's name and the message body in `origin`,
beside a copy wrapped for the model to read.
2026-08-29 15:27:46 -04:00
iris f18639e4b1 Count the far end of the list in rows, not events
Scrolling back stopped dead at the top of what was loaded, and no older
page ever arrived. Bryan spotted the cause from the outside: it had to do
with tool calls being collapsed.

The trigger compared an index into the list being drawn against
`items.size`, the number of transcript events. Those were the same number
when it was written. They stopped being the same when adjacent tool calls
started folding into one row, and the queued bubble and the working
indicator are two more rows with no event behind them. In this session
645 rows stood in for 720 events, so the last visible index could reach
646 and the threshold it needed was 717. It was not close; it was
unreachable, and the further a session went the worse it got.

Both numbers now come from the list itself, which is the only place they
are commensurable, and `totalItemsCount` counts whatever gets added to it
next.

Checked on the emulator against the case it was breaking on rather than a
clean one: five collapsed "Called 8 tools" groups in front of a 720-event
transcript, scrolled from the bottom to seq 1, which is the beginning of
the session. It stops there because that is the top, and holds position
while each page arrives.
2026-08-29 14:55:19 -04:00
irisandClaude Opus 5 fc71cb4403 Say "unknown" for a session we are not driving but cannot bury
A session in the config with no live entry reported `Exited`, whatever the
reason. That covers three different situations -- one that genuinely
ended, one that failed to relaunch, and one whose process could not be
checked -- and the wrong one is the expensive one.

`Exited` reads as "this conversation is over", and what a reader does about
it is start a fresh session. If the process is in fact still running, that
is a second CLI against a conversation that already has one: the exact
fault `session::process` exists to prevent, arriving through the status
field instead of through a spawn.

So it is said only when the process is known to be gone. A record that
cannot be checked reports `Unknown`, and so does one that is still alive --
this server is not driving it, so it genuinely does not know what that
process is doing, and the honest word is the one meaning "wait" rather than
the one meaning "act". A session with no record at all is still `Exited`:
an echo session, or one already stopped and cleaned up, and known to be.

The distinction was available all along -- `process::recorded` returns the
liveness -- which makes this the same mistake as the other five today:
reporting what was convenient to compute rather than what was measured.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VETa8afmpWaYezLCqJhDB8
2026-08-29 14:50:20 -04:00
iris 279c76e8a1 Hold the reader's place when a message arrives
Scrolled back through a conversation, every new message dragged the view
with it -- which reads as the screen scrolling down on its own, at the
exact moment somebody is trying to read something else.

The list had no keys, so its rows were identified by position. The
transcript is drawn newest-first, so a new message is an insertion at
index 0: every existing row shifts up one index, the viewport stays on
the index it was on, and the content slides through it. The working
indicator appearing and disappearing did the same thing at the same end.

So rows now carry the identity they always had in the data. Every
transcript event has a seq, which is what the transcript is ordered by
and never changes, and every row keeps the seq of the first event behind
it -- a streaming message keeps the seq of its first delta, so it holds
still for the whole answer rather than becoming a new row on every
frame, and a tool call keeps its start's.

Paging older history is the same insertion from the other end, and it is
the thing this could plausibly have broken. Checked on the emulator:
scrolled back mid-turn, the view sat still through twenty seconds of
streamed deltas, and scrolling to the far end still fetched earlier
pages and stayed where it was while they arrived.
2026-08-29 14:47:31 -04:00
iris 5396da76c7 Show a compaction happening, and what it recovered
The Compacting status had been declared, rendered in four places, and
never once emitted: no driver produced it, and the app had no control to
ask for a compaction in the first place. Pressing nothing for two
minutes and then quietly having less context was the whole experience.

The CLI turns out to announce all of it, which was worth measuring
rather than guessing at. Driven through /compact against 2.1.237 it
emits a `status: "compacting"` line at the start, a `status: null`
carrying `compact_result` at the end -- `"failed"` with a sentence
saying why, when it does -- and then a `compact_boundary` with the token
counts. The same records appear in the CLI's own transcript file with
camelCase keys, which is the obvious place to read the shape off and
gets every field name wrong.

So none of it is inferred here. The driver writes the line and says
nothing; the translator reports what the CLI reports. A failed
compaction surfaces the CLI's own sentence, which is specific enough to
act on.

The counts are the part worth keeping afterwards, so they land in the
transcript rather than only in a status that vanishes: a session that
went from 128,402 tokens to 9,617 has just been given its context back.
They are optional throughout, because a compaction whose size nobody
reported has to be able to say so -- a zero would read as "recovered
nothing".

Also here, all found on the way:

- `rename_all` renames variants; fields need `rename_all_fields`. Every
  field in Event was a single word until `pre_tokens`, which went out as
  snake_case, was not found by the app, and rendered as the "no counts
  reported" case -- a state it is allowed to be in, so nothing looked
  wrong. There is now a test on the wire names.
- The unparseable-line warning sliced bytes, not chars, on output that
  is full of em dashes. A panic there kills the task reading the
  session's stdout, and the session goes deaf with nothing on screen.
  The other three truncations in the tree already did this correctly.
- Echo compacts too, with invented numbers and a real shape, so this
  screen can be looked at without spending two minutes of somebody's
  account to reach the state.
2026-08-29 14:42:08 -04:00
irisandClaude Opus 5 42131c75d6 Send a steer into the running turn, and put an image under its call
**The queue was holding messages the CLI would have taken.** Two claims
in this file contradicted each other: the module header said a mid-turn
message is injected at the next tool boundary -- "the behavior this app
exists for" -- and `Queue`'s own doc said a line written mid-turn simply
becomes the next turn. The code followed the second, parking every
message until `Status::Idle`, which is the end of the whole turn.

Measured rather than argued, twice. Writing a line straight into a live
session's stdin fifo mid-turn produced one `result` for the whole thing,
so it was consumed inside that turn, not as a new one. The header was
right and the queue was built on the wrong claim.

The cost was exactly what Bryan reported: he steered after the second
tool call and it sat unread until every remaining call had finished.
Measured before and after on the same three-step turn -- steer sent at
+13s, recorded at +24.7s before this change and at +14.1s after, which
is the next tool boundary.

So the line goes out immediately. What stays behind is the
*announcement*: the CLI says nothing on stdout about having read a
message, so `MessageTaken` now waits for the next assistant text or tool
call, which is proof another model call happened and the steer was in
it. That keeps a held message drawn below the working indicator until
the session has actually taken it -- the thing that mattered when this
was last changed -- without delaying the message to get it. Idle counts
too, and is the case that must not be missed: a message written after a
turn's last model call has no later output to prove anything.

`closed` is untouched, and `Queue::close` still reports held messages by
name rather than dropping them.

**An image now names the call that produced it.** `Event::Image` gains
`about`, the `tool_use_id` from the tool result it came out of, so a
screenshot is drawn inside that call's card instead of floating beside
it -- pairing them by position is what a page boundary breaks. `None`
for a person's own attachment, which belongs to no call. The import path
threads it through as well, so replayed history reads the same as live.
Images show whether the card is open or closed: a call whose result *is*
a picture says less closed than the one line it replaced.

Verified on the emulator against a real haiku turn: the checkerboard sits
inside `Read /tmp/tiny.png`, and the steer sits between that call and the
next, where it was taken. 53 tests, clippy, rustfmt, lint and ktfmt clean.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-29 14:17:22 -04:00
irisandClaude Opus 5 184b6fc6a6 Correct the claim that reattach is local only
Written down as "an ssh session's child dies with its connection, so it
takes the ordinary --resume path". The code never had that branch: `start`
records a pid whatever the transport, and for a remote session the process
the backend owns is the ssh client. Adopting it is right -- the fifo feeds
it, its logs capture the far end, and ssh lives exactly as long as the
remote command, so its liveness is the session's.

The docs claimed less than the code does, which is the safe direction to be
wrong in but still wrong, and it was about to mislead someone: a remote
`claude` has an sshd pipe on stdin under every version of this server,
because the fifo is on the backend's side of the connection. Reading a
remote session's stdin therefore says nothing about which backend started
it, and we were an inch from concluding otherwise.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VETa8afmpWaYezLCqJhDB8
2026-08-29 13:57:12 -04:00
irisandClaude Opus 5 4cbd567c35 Take ktfmt's formatting
Committed unformatted again: I piped the check through grep, so the
task's failure never reached the shell's exit status and the chained
commit ran anyway. Check exit codes, not output.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-29 13:46:34 -04:00
irisandClaude Opus 5 73923dbbc3 Make jump-to-latest the same chevron, pointing down
It was a labelled button beside a tool group that collapses with a
drawn chevron -- two controls doing the same kind of thing in two
visual languages. Now one `Chevron` composable serves both directions,
parameterised rather than copied, since a pair that differs by a minus
sign drifts and the drift is a bug in exactly one direction.

The comment it replaces argued against an arrow here, on the grounds
that the list is laid out upside down. That reasoning was about the
code: nobody reading the screen knows the list is reversed, and on
screen the newest message is at the bottom, which is where this goes.

It draws no text, so the name lives in its content description -- the
whole of what a screen reader has, and the answer to "what was that
arrow for" later.

Verified on the emulator: scrolled up, the chevron appears bottom
centre matching the group's; tapped, it returns to the newest message
and takes itself away.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-29 13:46:13 -04:00
irisandClaude Opus 5 011ed0d0e1 Hold an image's place, open it full screen, and fold a run of calls
Five changes to how a transcript reads.

**Images no longer move the page.** The row was as tall as whatever had
loaded, so it grew when the bytes arrived and pushed everything below it
-- and in a bottom-anchored list, an image loading above the viewport
moved the text under the reader's eyes. The height is now decided before
the fetch and never changes: four lines of the body style, measured from
the type so it stays four lines when the reader has scaled their fonts.
Nothing to see when loading finishes, which is the point.

**A small image is enlarged with nearest neighbour**, a large one shrunk
smoothly -- decided per image from its actual size rather than set once,
since blowing a 16px sprite up with interpolation turns it into a blur
of exactly the thing being looked at.

**Tapping one opens it full screen**, fitted so the whole image is
visible first, with two-finger zoom to 8x and pan once zoomed. A dialog
rather than a screen, so back returns to the transcript.

**A tool call is one line closed**: the tool's name and what the call is
for. The command is not on it, because a wrapped command turns one row
into four. Open, it shows the command, the rest of the input and the
output, with the timeout at the top right -- a limit on the call rather
than part of what it does, worth seeing beside the command it constrains.
A call waiting on permission is shown open regardless, since the command
is the thing being decided.

**Adjacent calls fold into "Called n tools"**, closed by default, and it
closes again from either end -- a long group's heading scrolls away while
its last call is still on screen, and the reader who wants it shut is
looking at the bottom. The calls keep their full width; what says they
belong together is the surface behind them, one cue rather than two
half-cues. Grouping happens at display time, not in the fold: the
transcript's own order is what paging and the stream depend on.

Echo gains `/tools [n]` so a run of calls can be produced without paying
for one.

Verified on the emulator: four calls folded and expanded, one opened
inside the group showing `timeout 5000` top right, a 16px checkerboard
enlarged with hard pixel edges beside a shrunk screenshot at the same
height, the screen byte-identical between one second and six after
opening, full screen fitted, and back returning to the same scroll
position. Pinch itself is the one thing not verified here -- `adb input`
cannot inject a two-finger gesture.

53 tests, clippy, rustfmt, Android lint and ktfmt clean.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-29 13:43:46 -04:00
irisandClaude Opus 5 8fe13634cb Take ktfmt's formatting on the files just added
Four files went in unformatted: I ran the formatter mid-change and then
kept editing. ktfmtCheck is part of finishing, not part of starting.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-29 06:28:24 -04:00
irisandClaude Opus 5 d580461bd5 Show what was remembered as a note, not as markup
Claude Code marks a sentence taken from its stored memory by wrapping it
in `<cc-memory filenames="...">`. Markdown has nothing to say about that,
so it arrived as literal angle brackets mid-sentence and read as the
model having emitted broken HTML. It is the opposite: a claim about
where something came from, and "I was told this before" and "I worked
this out just now" are different things the reader could not otherwise
tell apart.

Each one becomes a card naming the files it came from, with the prose
either side of it left as prose. Named rather than merely tinted, since
a colour can say "this one is different" but not what kind of different.

A tag still arriving is left alone: streaming means the closing half may
be seconds away, and a half-written marker is not a marker yet.

Verified on the emulator with two notes in one message, one of them
citing two files, and prose before, between and after.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-29 06:27:00 -04:00
irisandClaude Opus 5 c5dbd1535d Ask for permission on the call it is about, and read the input
A bash permission request arrived as a second card repeating the tool
call's input verbatim, so the same command appeared twice and the reader
had to work out it was one event. `Event::Question` now carries `about`:
the `tool_use_id` the CLI's `can_use_tool` request already names. That
makes the pairing a measured fact rather than a match on input text --
and it stays `Option`, because AskUserQuestion is not permission for
anything and an echo session's question is about no tool at all. Those
still draw as their own card, which is what every question did before.

The card also reads the input instead of dumping it. Every tool's input
is JSON, and showing it raw makes the reader parse `{"command":"…",
"timeout":5000}` to find the line they care about. A small table says
which field is the subject of which tool -- Bash's `command`, Read's
`file_path` -- and the rest is still listed, since dropping a field
would claim the tool had no other input when it might. The subject is
syntax-highlighted with dev.snipme:highlights, for the reason the
markdown renderer is a library: lexical rules are somebody else's
specification. Its theme is Catppuccin, mapped in Theme.kt beside the
rest of the palette rather than taken from the library's defaults.

The input shows whether or not the card is expanded. A row that says
only "Bash" says nothing anyone can act on, least of all when it is
asking to run something.

Verified on the emulator against a real haiku session: one card, the
description, `grep -rn "needle" /tmp | head -3` highlighted, `timeout:
5000` pulled out, and "Allow Bash?" with its buttons inside the card --
then Allow, which resolved in place and ran.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-29 06:24:05 -04:00
irisandClaude Opus 5 206319045d Show usage per machine, and say why a machine has none
The server now reports limits per machine, so the screen has to as well:
one card per machine that offers a paid service, named by the machine
first, because these are one account's numbers and which account is decided
by which box ran the session.

It also has to say which of four things happened, and the reason for
splitting them shows up here rather than in the data. A machine nobody has
logged in on is working exactly as somebody set it up, so it reads as a
plain statement in ordinary text -- marking it would be the interface
nagging about a decision already made, and would dilute the marks that do
mean something. Only "couldn't reach it" and "the endpoint refused" are
coloured as faults, and they say different things because they need
different things done. The old screen drew all three in the error colour.

No machine offering a paid service is not an error either: it says so
instead of drawing nothing.

The app also stopped parsing: `available` no longer exists and
`getBoolean` on a missing key throws, so this had to land with the server
change rather than after it. An older backend sending no `state` is read as
"failed" rather than "ok", since an empty card drawn as healthy is the
worse failure.

Looked at running, against five machines: local reporting notLoggedIn with
the backend's HOME emptied, loopback-over-ssh returning real windows beside
it, an unreachable host showing ssh's own message in red, and a machine
with no Claude provider correctly absent.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VETa8afmpWaYezLCqJhDB8
2026-08-29 06:18:46 -04:00
irisandClaude Opus 5 6166b1f626 State the deny-unknown-fields rule where it governs all the bodies
It had landed appended to `SshRequest`'s doc comment, so a rule about
every request body in the module read as something about how to describe
a machine. Moved to the module doc beside the route table, where the set
it governs is what a reader is already looking at, and worded so a new
request body knows it is expected to carry the attribute too.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-29 06:14:32 -04:00
irisandClaude Opus 5 446fb92ba3 Refuse a request field the server does not know
A misspelled field was accepted and dropped. Sending `permission_mode`
instead of `permissionMode` produced a 200 and a session running in the
default permission mode -- so the caller's setting was gone, and nothing
anywhere said so. That is the expensive shape: indistinguishable from
success at the place you are looking. It cost an hour here, chasing a
"startup race" that was a key serde had silently discarded; with the name
spelled the way the API asks, a bypassPermissions session runs a `sleep`
loop with no prompt at all.

So every request body refuses unknown fields, not just the one that bit.
Axum's message names the offending field and lists what was expected,
which is the whole of what the caller needs.

Query strings are deliberately left permissive: a stale link carrying an
extra parameter is not a mistake worth failing a request over.

53 tests, clippy and rustfmt clean; verified against the running server
that the misspelling is now a 422 naming the field and the correct
spelling still spawns.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-29 06:13:18 -04:00
irisandClaude Opus 5 d8570f4d5a Ask each machine about its own limits, not this one about all of them
`ClaudeUsage` read `~/.claude/.credentials.json` on the machine running the
backend and asked the API about that account, once, globally. But a session
runs wherever its setup says, so the numbers on the usage screen belonged
to the backend's account rather than to the account that spent the tokens.

That is not a rounding error in the layout this project is aiming at.
`ai-server` belongs on the host; the host has no `claude` CLI at all and
the VM is a remote. So the screen would have reported "is Claude Code
logged in on this machine?" while every session ran fine on a machine whose
limits nobody could see. It looked correct only because backend and CLI
happen to be the same box today.

Usage is now per machine, asked through the same `Transport` the sessions
use -- `ssh host sh -c 'cat $HOME/...'` for a remote, unchanged for the
local one. `$HOME` is left for the far shell to expand, since a path built
here is this machine's home directory and over ssh that is somebody else's.
Machines with no Claude provider are not asked and get no row: they have no
Claude limits, and a row about them would be a fact about nothing.

The snapshot gains the states it could not say. `available` plus an `error`
string made three different situations look identical, and the one that
suffered was the harmless one: a machine nobody has logged in on is a
decision somebody made, with nothing to fix, and it read as broken.
`notLoggedIn`, `unreachable` and `failed` are now distinct, and which one a
failed read is gets decided in `why_no_credentials` rather than at the call
site.

Supporting changes: `ssh::command` builds a `std::process::Command` that
tokio converts from, so a blocking caller can use the one place that knows
what a correct ssh invocation is; `Transport::capture_blocking` is that
caller's door. The cache is keyed by machine and service rather than by
position, since the set is no longer fixed at startup -- a positional cache
would hand one machine's numbers to another the moment a setup was added.
A cached snapshot still picks up a rename immediately, because the name has
nothing to do with the poll interval.

Verified against a real ssh setup (loopback, per AGENTS.md) with five
machines, all four states seen: local `ok`, loopback-over-ssh `ok`,
unreachable host `unreachable` carrying ssh's own message, a claude-less
machine correctly absent, and -- with the backend's HOME emptied -- local
`notLoggedIn` while the ssh machine still reported real windows, which is
the production shape.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VETa8afmpWaYezLCqJhDB8
2026-08-29 06:12:17 -04:00
irisandClaude Opus 5 9f403cddab Render markdown, and stop wiping messages still waiting to be read
Two things, both about the transcript telling the truth about itself.

Markdown is rendered rather than shown as its source. The parsing is
mikepenz/multiplatform-markdown-renderer, not something written here:
markdown is somebody else's specification, and a hand-written subset of
one disagrees with it at the edges, which is where the bug reports come
from. `Markdown.kt` is only the mapping onto this app's palette, so code,
links and rules take the Catppuccin values the rest of the app uses
rather than the renderer's defaults.

The queued-message list was cleared wholesale whenever a turn ended. But
the backend holds a queue of its own and takes one message per turn, so a
turn ending is precisely the moment the *rest* are still waiting -- the
bubbles vanished while the messages were on their way, which reads as
everything after the first having been dropped. Now a held message
leaves the list exactly two ways: the session reads it, which arrives as
a UserMessage, or its send failed and there is nothing to wait for.

Measured first, because the report was that the backend dropped them:
three messages sent behind one long turn were all delivered in order
(ONE, TWO, THREE) against current main, so the loss was in the display.

Verified on the emulator: headings, emphasis, inline code, nested lists,
a quote bar, a fenced block, a rule and a link all render, and the three
queued messages sit through their turn.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-29 06:10:22 -04:00
iris 90a57ca7e9 Revert "Default a new session to bypassPermissions"
This reverts commit d81c9a7d65.
2026-08-29 05:58:17 -04:00