Commit Graph
21 Commits
Author SHA1 Message Date
iris 68704fce7c Ask the driver whether a command can go, not the status it reported
A `/clear` that did nothing, traced to the end. There was no race to lose:
the driver sees every line it writes and every line that comes back, so it
always knew. What it knew was being asked of the wrong thing.

Two views of "is a turn running" had grown apart. The driver's moves the
instant it writes a line; `SessionStatus` moves when output is *recorded*.
Messages ask the driver -- which is why they behave -- and commands asked
the status, which for a command is stale for its whole round trip: a
command's reply carries no assistant text, so nothing proved a turn had
started and the recorded status stayed idle from the moment it went out
until the moment it came back. A second command in that window went
straight out too, landing inside the turn the first one had started, where
the CLI reads it as text instead of running it. Nothing anywhere says so:
a command read as a message looks like a message.

So `Commands` asks `Driver::between_turns()` now, and asks again when it
releases a held one -- the recorded idle that woke it is a moment in the
past by then. `local_command` says `Running` when it writes, which is both
true and what makes the next idle a change worth recording; without it the
idle at the end of a command was equal to the idle before it, and nothing
behind it was ever released.

The other half was a turn nobody here started. The CLI picks the
conversation back up on its own -- measured: a backgrounded `sleep`
finished nine seconds after the turn's result and it began again unprompted
-- and it announces that with a `system/init` about a second and a half
before its first assistant text. We had been ignoring that line and
learning about the turn from the text, so for that second and a half the
session read as idle. It is a turn now, told apart from the `init` at
startup by the translator already having a session id, and from our own
`/clear` by `running` already being true.

Measured against the real CLI, not argued: two `/clear`s sent back to back
on one connection now record `commandSent`, `running`, `commandQueued`,
`cleared`, `idle`, `commandSent`, `cleared` -- held, then run, in order,
both of them. Before this the second was swallowed. The self-started turn
shows as `running` eleven seconds after the previous turn's idle, which is
the window a command used to disappear into.

Also measured on the way, and worth writing down: a message written into a
running turn is *folded into it* -- one `result`, `num_turns: 2`, both
things answered -- so an idle after one is honest and there was nothing to
fix there. A command written when the CLI is genuinely between turns is
executed even ten milliseconds after the result, so the boundary itself was
never the problem.
2026-08-30 01:42:47 -04:00
iris da65c1571f Merge remote-tracking branch 'origin/main'
# Conflicts:
#	app/androidApp/src/main/kotlin/com/example/aiapp/SessionScreen.kt
2026-08-30 00:49:34 -04:00
iris b0629f77ca Shrink a photo to what the provider takes, and put it in its own bubble
Sending an image was broken in the way that is hardest to see from the
phone: a camera photo is twelve megapixels and several megabytes, the Claude
API resizes anything past 1568px on its long edge before looking at it and
refuses far larger outright, so the picture was uploaded whole over the
tunnel to be thrown away or rejected at the other end.

Shrunk on the phone, to a limit the server states. Which number it is comes
from the provider's *kind* -- `DriverKind::max_image_edge`, reported on the
session row -- because that is where a provider's requirements are known,
and a phone carrying its own copy of them would be a second place to update
when one changes. `None` where nothing cares, rather than a large number:
"no limit" and "a limit that happens to be big" are different answers and
only one of them stays true. Doing it before the upload rather than after is
the point -- the expensive part on a phone is the tunnel, not the decode --
and an image already inside the limit is uploaded byte for byte rather than
being round-tripped through JPEG for nothing.

EXIF orientation is applied while scaling. The camera writes which way up
the picture is into a tag rather than into the pixels, and re-encoding drops
it, so a portrait photo would have arrived at the model on its side with
nothing anywhere saying so.

**What is attached is now visible before it is sent**, in a row directly
above the box it will be sent from: the count on the "+" button said how
many and never which, so the only way to find out what you had picked was to
send it. It scrolls sideways rather than shrinking, and tapping one takes it
back off -- an image picked by mistake could otherwise only be dealt with by
sending it. The tile is outlined as well as filled, because most of what
gets attached here is a screenshot of a dark app and a cropped one is
near-black: without an edge the only thing on screen saying an image was
attached was the cross drawn on top of nothing.

**And the picture is inside the bubble that sent it.** Attachments used to
be their own `Image` events emitted just before the message, which drew
somebody's screenshot as a row floating above the bubble and left the phone
deciding from adjacency alone which message an image belonged to -- a thing
the sender knew and could simply say. `UserMessage`, `MessageQueued` and
`MessageTaken` carry the refs now, so a waiting message keeps its picture
for as long as the turn runs, and a replay puts it back in the same place.

Verified on a real claude-cli session rather than an echo one, since the
limit only exists for that kind: a 3000x4000 image arrived as 1176x1568
JPEG -- long edge exactly the limit, aspect ratio intact -- and haiku
answered "AI Sessions displays idle Photo", which is what the picture was.
No error, and the transcript records the message with `images` on it.
2026-08-30 00:47:16 -04:00
iris de049bcec6 Hold the transcript still while somebody is reading further back
Two separate defects, both of which moved the list under the reader.

The first: a row that grows drags the view toward the newest end. The
list is laid out from the bottom, so it anchors on the first visible
item's *bottom* edge -- and a reply streaming in extends that row
upwards, pushing everything already on screen with it. Measured against
a reply streamed in four hundred pieces: scrolling back one screen and
waiting six seconds ended at the very bottom, forty lines further on
than where it was left. So the transcript now only changes while the
reader is at the newest end; anything arriving before then waits in
order and lands when they return. Status, tokens and the model still
update live, because none of those are drawn in the list and freezing
them would trade a jumping transcript for a status row that lies.

The second: every markdown row was measured at nothing before it was
measured at its real height. The renderer's `content: String` overload
parses in a coroutine and draws an empty loading slot until it finishes,
so a row composes with no height and springs open a frame later. Seen
with five replies on screen at once, all blank, the whole conversation
shrunk to a single screen. Parsing in the composition costs a few
milliseconds on the main thread and is worth it: no scroll anchoring can
survive a row that lies about its height first.

`/stream N` in the echo driver is what made the first one reproducible
-- `/slow` emits a line a second, and the growth has to be continuous
for the anchor row to drag.

Verified on the emulator: scrolled back through a whole 400-piece
stream, the transcript region is pixel-identical across ten seconds
while the status row goes from working to idle; returning to the bottom
brings the backlog in one go. Checked the tool-call rig too, which this
change had no reason to touch -- paging back still works and every group
still reads "Called 8 tools".
2026-08-30 00:42:28 -04:00
iris 47d6b84265 Report what a session last did and what it cost, not what this page holds
Three readings that were each a part presented as the whole.

**"just now", everywhere, after a restart.** A relaunched session took its
last-activity from the clock, so every session the backend brought back
claimed to have been active that instant. On the phone that is every row
reading "just now" and the list -- which sorts by it -- coming back in an
order that means nothing, with the conversation somebody was in the middle
of buried among sessions untouched for days. It comes from the transcript
now, in the pass `Transcript::open` already makes, which is the same
correction `last_status` got and for the same reason: a server that has just
started has been told nothing, and the file is the only thing it knows. The
test backdates a transcript by a day, so it cannot pass by the test being
fast; it fails on the old code with the clock's answer in the message.

**The token total was the newest page's.** The phone added up the
`UsageDelta`s it had received, and it opens a session on the newest page of
the transcript -- so a long conversation reported its last few turns as the
total, and a page with no turn in it reported nothing at all, since zero is
drawn as blank. That is the reading Bryan saw: no tokens, on sessions that
had certainly spent some.

The count belongs to the server, which is the only side that sees every
turn. `UsageDelta` now carries the running total beside the delta, filled in
by the pump rather than by each driver -- a driver knows what its own turn
cost and nothing else does, so a new one cannot get this wrong by leaving it
out -- and the session row reports it for a screen that has not opened the
stream yet. The phone takes the largest total it has seen instead of
accumulating, which also means paging older history cannot move it, and
leaves the seeded figure alone for transcripts recorded before the field
existed. Seeded by summing deltas at startup for exactly that reason.

**The header said the model twice and the machine backwards.** A session's
subtitle now reads `machine · provider`, in that order and with no "on"
joining them, matching the list and the usage dialog -- the "on" made it a
phrase, which works in one order and stops working the moment the same pair
is shown somewhere else. The model is gone from it: the footer's picker
already shows what the session is set to, and two places showing it meant
two things to keep in step, which disagreed for a moment on every switch
since one follows the request and the other the session's own answer.

Checked on the emulator against a twelve-turn session whose visible page
held the last six: the header reads "this machine · echo", the status row
reads "idle", and the total reads 42 tok, which is what `GET
/sessions/{id}` says rather than what the page adds up to.
2026-08-29 23:52:58 -04:00
iris 2dc61c5780 Stop drawing one tool call twice where a page of history begins
A page boundary lands wherever it lands, and about half the time that is
between a tool call and its result. The newer page then holds a `ToolEnd`
whose start it never saw, which the fold draws as a row of its own --
correctly, since a call rendering as nothing is indistinguishable from
one that never happened. But when the older page arrived it brought the
real `ToolStart`, and the two lists were concatenated, so the call was
left on screen twice: once as a proper card and once as a nameless
placeholder.

`joinPages` merges the two halves by the call's own id instead, which is
the one thing a page boundary cannot destroy. The older half wins on what
a start knows -- the tool's name, its input -- and the newer on what an
end knows, its output and whether it finished.

The miscount was the visible part; the moving was the point. The extra
row sits exactly at the join, which is where the reader is looking when
the page loads, so everything below it stepped down by a row at the
moment they scrolled into it.

Demonstrated both ways round on a rig of twelve `/tools 8` runs, whose
groups are eight calls each and whose page boundary falls inside the
second one: without this the transcript reads "Called 9 tools" there and
eight everywhere else, with it every group reads eight.

That rig is `/mixed N` in the echo driver, added here: N beats of
paragraphs at three lengths, single tool calls, runs of adjacent ones,
images and peer messages -- every row shape the app draws, in one
session, from a command that costs nothing and produces the same
transcript every time. The paragraphs are deliberately ragged, because a
wall of identical lines looks the same at every offset and makes a scroll
of one line indistinguishable from a scroll of ten, by eye or by
comparing frames.
2026-08-29 23:34:11 -04:00
iris e37e90a579 Let the server say what is waiting, instead of the phone remembering
A message sent into a running turn was drawn as a pending bubble from
screen state, so leaving the session or restarting the app showed nothing
waiting while the queue was full. Nothing waiting is what "there is
nothing" looks like -- the reader had no way to tell it from a queue that
had already drained, and Bryan hit exactly that: a message he sent
arrived, and his phone stopped showing it after a restart.

The server now records the waiting. `MessageQueued { id, text }` goes into
the transcript when a driver takes a message it cannot deliver yet, and
is resolved by the `UserMessage` carrying the same id -- the same shape
`CommandQueued` and `CommandSent` already had, so this is one more
instance of a mechanism rather than a second one beside it.

The message itself still lands where the session read it, which is what
the last change was about; only the *waiting* is recorded early. The two
are different facts and now have different events.

Paired by id rather than by text. The old code removed the bubble whose
text matched, so sending the same thing twice cleared the wrong one and
left a message on screen that had already been read.

Both drivers that can queue do it: the echo driver too, because the phone
now draws pending bubbles from the stream and a rig that skipped the
event would exercise a state the real app never sees.

Checked on the emulator: two messages sent into a `/slow` turn, then the
app force-stopped and relaunched -- both still drawn as waiting, in the
pending style, and both resolved into ordinary bubbles when the turn
ended and the session read them.

Still outstanding, and worth knowing: an entry outlives a *server*
restart in the transcript but not in the driver's memory, so a backend
restarted mid-queue would leave the bubble drawn with nothing coming to
resolve it. Before this change that message vanished from the transcript
entirely, so the failure is now visible rather than silent -- but it is
not yet right.
2026-08-29 22:29:12 -04:00
irisandClaude Opus 5 1e3eec8cf0 Let a session drop its context without ending
Adds `Event::Cleared`, `SessionCommand::Clear`, and `Driver::clear`, so
`POST /sessions/{id}/command {"text": "/clear"}` does for a session what
the CLI's own `/clear` does for a terminal.

The marker is a divider, not a truncation: everything above it stays in
the transcript, because that is the only copy of the conversation the
phone has and a person scrolling back is a different question from what
the model is given. It also makes clearing mean one thing across
drivers -- `claude` sends `/clear` and the CLI answers with a fresh
`init` whose new session_id the reader already persists as the resume
token, so the next launch resumes the cleared conversation with nothing
to keep in step; `llama` needs no state at all, since `conversation()`
already folds the transcript and now folds from the last marker; `echo`
emits the marker alone, so the phone's divider and scroll behaviour can
be exercised without spending a real session's context.

That fold is why `Cleared` is documented as load-bearing rather than
decorative. For any driver that rebuilds its conversation from the
transcript, this marker decides what the model sees, and treating it as
something only the phone draws would silently put the cleared
conversation back in front of the model at full price.

Clear rides the existing boundary pump like any other SessionCommand, so
one arriving mid-turn waits exactly as a compaction does, and nothing
grows a second way to wait.

Removes `--autocompact` in the same change, because clearing is the
cheaper answer to the problem it was added for and Bryan would rather
manage context that way. Keeping the measurement here, since it was the
reason for the constant and is worth more than the constant was:
context returned to 70-85k within ten calls of a compaction; a
compaction took 104,346 to 147,671 ms; compaction cost that session
2,655,508 tokens across six boundaries, of which the single automatic
one at the 1M ceiling was 1,696,870. Clearing costs nothing, because
nothing is sent.

70 tests, clippy clean, rustfmt clean.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VETa8afmpWaYezLCqJhDB8
2026-08-29 20:06:04 -04:00
iris 749b2db287 Make a command a thing the app knows, and hold it until it can run
Typing "/" now suggests what this app understands -- `/compact` and
`/rename <name>` -- with a line each about what they do, and anything
else beginning with a slash is passed to whatever runs the session,
because a dialect's own vocabulary grows without this list.

None of them are messages, and that is the substance of the change. A
line written into a running turn is read by the *model*, so a command
sent mid-turn either does nothing or arrives as text somebody has to
puzzle over. They now wait for the turn to end. The waiting is done once
for every provider, in the pump that already watches every event for the
boundary, rather than in each driver where a new provider could get it
wrong by leaving it out.

Waiting is a state, so it is on screen: the command sits at the reader's
end of the conversation in blue, with a spinner and "waiting for this
turn to end", and becomes an ordinary blue row when it goes. Blue
because these are about the session rather than about the task -- the
same blue a compaction already used, which is now one colour with one
name rather than two.

Renaming from the settings screen sends exactly this, so it waits and
draws the same way. The name itself is not held: it is this server's own
datum, so the list and the header change at once and only telling the
session waits.

Echo grew the same split, which is where the bug in it showed: its
commands are its messages, so running one announced a `MessageTaken` as
well, and the same line drew twice -- once blue, once purple. A command
owes no announcement; the manager has already recorded that it was sent.

Watched rather than reasoned about: `/compact` during a 25 second turn
held with its bubble up, went out when the turn ended, and the
compaction that followed reported what it recovered.
2026-08-29 17:10:21 -04:00
iris bebaae7a94 Carry a question in the event model, not in one provider's JSON
A question is now fully described by the event that reports it: the tag
it was asked under, each option's label, what it means, and the sample of
what picking it would produce, plus whether several may be picked at
once. The app renders from that alone.

It had been reading Claude Code's tool input to find the parts the event
dropped -- that dialect's schema, written out a second time in Kotlin,
where no other provider could reach it and where it would drift the
first time the schema moved. Echo could not describe an option at all,
and llama never will.

Answers travel as a list for the same reason. A question that takes one
answer sends a list of one rather than being a different shape, and the
one place that flattens it is where the CLI is spoken to: its answers
map holds a string, so several choices are joined there. That join was
in the phone.

Also here because it is the same rule: the permission ask reuses the
question body rather than owning a second one, so Allow/Deny renders and
resolves through exactly the code an AskUserQuestion does.

Verified against both, since a refactor that only satisfies the case it
was written for has been tried on the half that cannot fail: a two
question `/ask` answered from the phone, one option and then two, and a
real sonnet session's `rm -f` permission asked, allowed, and run.
2026-08-29 16:46:43 -04:00
iris fea8e7e92b Show every option a question offers, on the call that asked
Reported by Iris through the dev-updater session: a two-question
AskUserQuestion arrived with only one option visible per question, so
the answer she sent was the only one she had been offered.

The cause was a `Row`. It hands out intrinsic widths in order and clips
whatever runs past the edge, so the first option or two drew and the
rest went off the side of the screen -- which does not read as a bug, it
reads as those having been the only choices. The same Row was in the
permission ask beside it; both wrap now. That pairing is the reason to
look: a rule stated on one member of a set is usually missing from the
others.

The rest of what she asked for, and what each was:

- It drew twice, as the tool call and again as loose question cards,
  because the backend marked these questions as belonging to no call.
  They belong to the call that asked, and now say so.
- So it renders like any other tool: one card, its own heading, opened
  because a decision cannot be made from a closed row.
- Each option shows its description and its `preview` block, which is
  the part a reader is deciding on and none of which was reaching them.
- "Other" is a field on every question. The harness always offers it, so
  leaving it out narrowed a question that was never that narrow.
- A multi-select sends the labels it collected as one string, which is
  the tool's own schema rather than a guess -- its answers map is
  string-valued.
- No spinner while it waits. A spinner says the machine is working; here
  the machine is idle and the turn is stopped on the person, so the card
  says "your turn" in the colour this app already uses for that.

Verified against a real session as well as the echo fixture: haiku asked
two questions with three described options each, both were answered from
the phone, and the model carried on with the answers. Echo grew `/ask`
so the shape can be looked at without paying a model to produce one, and
its option cards are outlined rather than tinted -- as one surface step
up they were three paragraphs where three things to press should be.
2026-08-29 16:22:15 -04:00
iris d3fff3d229 Let a session be renamed, under the same name everywhere
A gear at the end of the session's own bar opens what can be changed
about that session; the name is the first thing there. Compact is gone
from that bar -- `/compact` typed into the message box is the CLI's own
way to ask and it already worked, so the button was a second way to say
one thing. Echo takes the typed word too now, since it is the rig the
compaction display is checked against and losing the button would have
taken that with it.

The name is this server's, not a driver's: it is what the list shows, it
exists before any process does, and every provider has one. So it is
settled in the config and the driver is *told* -- which is the opposite
of the model and the permission mode, and the difference is written down
at `Driver::set_title`. A driver whose process has no notion of a name
does nothing and says nothing, because there is no failure to report.

Claude Code has one, so the name reaches it: `--name` for a session we
create, and `/rename` afterwards, which is a local command rather than a
control request -- `set_session_name` is not a subtype it knows, which I
established by asking it. A resumed session is deliberately not renamed
at launch: an import already has a name, quite possibly one the person
typing in it chose, and taking that would be helping itself to something
the app was only shown.

Verified end to end rather than argued: renaming from the phone put
"Session renamed to: paging and scroll" in the CLI's own session file,
and the session now lists under that name to other agents.

The gear is drawn rather than set in a font, for the reason Chevron
gives. It was a sun on the first attempt -- thin teeth standing clear of
a thin hub -- which no amount of reading the diff would have shown.
2026-08-29 16:06:15 -04:00
iris 404066fa7d Say when a session is working, and what it was told
Three things a phone could not see, all of them the same shape: the
session was doing something and nothing on screen said so.

A turn nobody here started never reported itself. `Running` was sent
where a message was *sent*, so a session picked up mid-turn, one
compacting on its own, or one another agent wrote to sat there reading
as idle until it finished. The driver now says it from what it observes
-- output that could only come from a turn in flight -- which is the
same set of events that already announced a steer, with the ends
swapped.

An imported session had it worse: nothing but replayed lines ever
reaches it, and a status was not among them, so it was permanently
whatever it was when it was adopted. Its file does not record a turn
ending, but it does record why each assistant message stopped, and
`tool_use` versus anything else answers it. A record that says nothing
leaves the status alone rather than voting for idle.

Messages from other agents were dropped outright: the CLI marks them
meta, and this replayed everything except meta. They are now a row of
their own, closed by default like a tool call, named for the session
that sent it -- not the reader's own bubble, because they did not say
it, and a session working on something this phone never asked for is
exactly what one of these explains.

Measured against a real session file rather than guessed: the peer
record carries the sender's name and the message body in `origin`,
beside a copy wrapped for the model to read.
2026-08-29 15:27:46 -04:00
iris 5396da76c7 Show a compaction happening, and what it recovered
The Compacting status had been declared, rendered in four places, and
never once emitted: no driver produced it, and the app had no control to
ask for a compaction in the first place. Pressing nothing for two
minutes and then quietly having less context was the whole experience.

The CLI turns out to announce all of it, which was worth measuring
rather than guessing at. Driven through /compact against 2.1.237 it
emits a `status: "compacting"` line at the start, a `status: null`
carrying `compact_result` at the end -- `"failed"` with a sentence
saying why, when it does -- and then a `compact_boundary` with the token
counts. The same records appear in the CLI's own transcript file with
camelCase keys, which is the obvious place to read the shape off and
gets every field name wrong.

So none of it is inferred here. The driver writes the line and says
nothing; the translator reports what the CLI reports. A failed
compaction surfaces the CLI's own sentence, which is specific enough to
act on.

The counts are the part worth keeping afterwards, so they land in the
transcript rather than only in a status that vanishes: a session that
went from 128,402 tokens to 9,617 has just been given its context back.
They are optional throughout, because a compaction whose size nobody
reported has to be able to say so -- a zero would read as "recovered
nothing".

Also here, all found on the way:

- `rename_all` renames variants; fields need `rename_all_fields`. Every
  field in Event was a single word until `pre_tokens`, which went out as
  snake_case, was not found by the app, and rendered as the "no counts
  reported" case -- a state it is allowed to be in, so nothing looked
  wrong. There is now a test on the wire names.
- The unparseable-line warning sliced bytes, not chars, on output that
  is full of em dashes. A panic there kills the task reading the
  session's stdout, and the session goes deaf with nothing on screen.
  The other three truncations in the tree already did this correctly.
- Echo compacts too, with invented numbers and a real shape, so this
  screen can be looked at without spending two minutes of somebody's
  account to reach the state.
2026-08-29 14:42:08 -04:00
irisandClaude Opus 5 011ed0d0e1 Hold an image's place, open it full screen, and fold a run of calls
Five changes to how a transcript reads.

**Images no longer move the page.** The row was as tall as whatever had
loaded, so it grew when the bytes arrived and pushed everything below it
-- and in a bottom-anchored list, an image loading above the viewport
moved the text under the reader's eyes. The height is now decided before
the fetch and never changes: four lines of the body style, measured from
the type so it stays four lines when the reader has scaled their fonts.
Nothing to see when loading finishes, which is the point.

**A small image is enlarged with nearest neighbour**, a large one shrunk
smoothly -- decided per image from its actual size rather than set once,
since blowing a 16px sprite up with interpolation turns it into a blur
of exactly the thing being looked at.

**Tapping one opens it full screen**, fitted so the whole image is
visible first, with two-finger zoom to 8x and pan once zoomed. A dialog
rather than a screen, so back returns to the transcript.

**A tool call is one line closed**: the tool's name and what the call is
for. The command is not on it, because a wrapped command turns one row
into four. Open, it shows the command, the rest of the input and the
output, with the timeout at the top right -- a limit on the call rather
than part of what it does, worth seeing beside the command it constrains.
A call waiting on permission is shown open regardless, since the command
is the thing being decided.

**Adjacent calls fold into "Called n tools"**, closed by default, and it
closes again from either end -- a long group's heading scrolls away while
its last call is still on screen, and the reader who wants it shut is
looking at the bottom. The calls keep their full width; what says they
belong together is the surface behind them, one cue rather than two
half-cues. Grouping happens at display time, not in the fold: the
transcript's own order is what paging and the stream depend on.

Echo gains `/tools [n]` so a run of calls can be produced without paying
for one.

Verified on the emulator: four calls folded and expanded, one opened
inside the group showing `timeout 5000` top right, a 16px checkerboard
enlarged with hard pixel edges beside a shrunk screenshot at the same
height, the screen byte-identical between one second and six after
opening, full screen fitted, and back returning to the same scroll
position. Pinch itself is the one thing not verified here -- `adb input`
cannot inject a two-finger gesture.

53 tests, clippy, rustfmt, Android lint and ktfmt clean.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-29 13:43:46 -04:00
irisandClaude Opus 5 c5dbd1535d Ask for permission on the call it is about, and read the input
A bash permission request arrived as a second card repeating the tool
call's input verbatim, so the same command appeared twice and the reader
had to work out it was one event. `Event::Question` now carries `about`:
the `tool_use_id` the CLI's `can_use_tool` request already names. That
makes the pairing a measured fact rather than a match on input text --
and it stays `Option`, because AskUserQuestion is not permission for
anything and an echo session's question is about no tool at all. Those
still draw as their own card, which is what every question did before.

The card also reads the input instead of dumping it. Every tool's input
is JSON, and showing it raw makes the reader parse `{"command":"…",
"timeout":5000}` to find the line they care about. A small table says
which field is the subject of which tool -- Bash's `command`, Read's
`file_path` -- and the rest is still listed, since dropping a field
would claim the tool had no other input when it might. The subject is
syntax-highlighted with dev.snipme:highlights, for the reason the
markdown renderer is a library: lexical rules are somebody else's
specification. Its theme is Catppuccin, mapped in Theme.kt beside the
rest of the palette rather than taken from the library's defaults.

The input shows whether or not the card is expanded. A row that says
only "Bash" says nothing anyone can act on, least of all when it is
asking to run something.

Verified on the emulator against a real haiku session: one card, the
description, `grep -rn "needle" /tmp | head -3` highlighted, `timeout:
5000` pulled out, and "Allow Bash?" with its buttons inside the card --
then Allow, which resolved in place and ran.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-29 06:24:05 -04:00
irisandClaude Opus 5 362d436d4f Let sessions outlive the backend, and never resume one twice
Three `claude` processes ended up running against this checkout on
2026-08-29, and the account hit its session limit. One cause, several
ways in.

An agent imported the Claude Code session it was *itself* running in.
That is an ordinary import, and importing runs `--resume` -- so a second
CLI attached to a file the first was still writing. The whole 65 MB
conversation, 154 embedded screenshots included, was re-appended to the
transcript under a new prompt id; both copies then read each other's
writes as work done elsewhere, and the adopted one was billed for
re-reading all of it. Meanwhile `shutdown_all` asked each session to stop
and the process exited immediately, so the SIGKILL timer died with the
runtime, the stop was unreliable, and whatever survived was orphaned with
nothing written down to find it by.

The processes leaked either way. So leak them on purpose, and be able to
pick them back up.

A session's process now outlives the backend and is adopted again on the
way up, which is worth having for its own sake: restarting the server no
longer ends a turn somebody is waiting on. Its stdio lives in the session
directory -- a fifo opened read-write so the process is its own last
writer and never reads EOF, plus stdout/stderr logs read from a byte
offset. `session::process` records the pid *and* the kernel's start time
for it, because a pid alone is reused and adopting a stranger's would mean
never resuming the real conversation.

That makes the fix structural rather than a check: everything goes through
`ClaudeDriver::launch`, which adopts if it can and starts if it cannot,
and `--resume` is reachable only on the second path. `Driver` gains two
ways out where it had one -- `detach` (coming back) and `stop` (the
session is being deleted, so the process must not survive).

Importing a session that is open is now refused outright. Claude Code
keeps `~/.claude/sessions/<pid>.json` for every live session, so this is a
measurement rather than a guess; it reports no/yes/unknown, because a
machine that keeps no such record cannot answer and "could not check" is
not "nobody is using it". `SessionStatus` gains `Unknown` for the same
reason.

Also here, found on the way:

- A reconnecting phone was sent the entire backlog. Opening a session was
  bounded to a page but reconnecting was not, so a long disconnect
  delivered thousands of events one frame at a time. Past `CATCH_UP_LIMIT`
  the stream sends a `reset` frame and the newest window, and the client
  rebuilds from it as it does on open -- without the reset the window is
  spliced onto rows no longer adjacent to it.
- A session's status was assumed idle at launch. Read from the transcript
  instead, so a restart stops claiming an exited session is waiting for
  you.
- `llama-server`'s stdout was piped and never drained, so a chatty one
  blocked on a full pipe buffer mid-load. It goes to a log now.
- A turn that exited or errored never emitted `Idle`, so the queue stayed
  "running" for good: every later message was held forever and, since a
  message is only recorded when taken, vanished with nothing on screen.
- Two doc comments had drifted onto the wrong functions.

Verified by killing the server mid-turn: the process survived, finished
its turn unattended (12.8 KB of output nothing was reading), and the
restarted server adopted it -- one process, all 700 lines in the
transcript, no hole, and it still took a new message afterwards. Deleting
a session stops its process; a 266-event backlog resets while a 16-event
one streams. 46 tests, clippy and rustfmt clean, app compiles and lints.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VETa8afmpWaYezLCqJhDB8
2026-08-29 04:47:43 -04:00
iris a9ea84c96c Stop choosing a model, and keep an imported session up to date
**Why the model became fable.** `spawn_session` fell back to the
provider's first listed model when none was given. That list is a shortcut
for the spawn screen, written in whatever order somebody typed it, and its
first entry is `fable` -- so every session spawned without a model, which
is every import, silently became a fable session. It looked like a default
and was an artefact of list order. Absent now means absent: no `--model`
flag, and the CLI uses whatever the person configured for themselves.

**Model and permission mode are now visible and changeable** from the
session, as buttons that read as their current value rather than labels
beside one. The mode was spawn-only; the CLI turns out to accept
`control_request{subtype:set_permission_mode}` and echo the mode back,
probed against 2.1.237 the same way the rest of the protocol record was.
Both default to `auto` -- on a phone every ask is a round trip to a
question card, which is how "allow Bash?" became the most-answered
question in the app.

The mode is reported by the API so the picker shows what the session is
actually set to, and it is kept in the live session beside the model for
the reason the model already was: `meta` is the shape a session was
*launched* with, so reporting from it shows the value a change replaced.

**And an imported session keeps itself level with its source file**, so
work done at a terminal arrives without a button. `--resume` appends to
the same transcript rather than forking -- measured, not assumed -- so the
only hard question is which new lines came from here.

Answered by counting the events this session has recorded. Status is the
obvious signal and is wrong, which cost a round trip to find: a turn that
starts and finishes between two polls reads as idle at both, so its output
is replayed on top of itself. It showed up on screen as `donedone`, and
only because the reply was one word -- with a longer answer it would have
looked like the model repeating itself.

Verified against both halves: text appended to the source file the way a
terminal writes it appears within one interval, and a message sent through
the app appears exactly once, before and after a turn.
2026-08-28 22:44:41 -04:00
irisandClaude Opus 5 c12ab7f098 Take rustfmt's defaults
The code was hand-formatted -- close to rustfmt's output but not it, mostly
in keeping chains and call arguments on one line where the formatter would
break them. That is a per-line decision every future change has to make
again, and reproducing it would mean a config whose only job is to preserve
how the code already looks.

So this is `cargo fmt` at its defaults, with no rustfmt.toml, which is
where the sibling dev-updater checkout already sits: it is clean at the
defaults today, so the two repos now agree on layout without either of them
configuring it.

Formatting only -- no behaviour, no renames, nothing reordered. Verified
after: cargo test (35 pass), cargo clippy --all-targets clean, cargo fmt
--check clean.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017xn8nHw1tw1R6PtiY1eEtw
2026-08-28 03:13:36 -04:00
irisandClaude Opus 5 99bcc341c1 Cleanup pass: one home for duplicated logic, stale comments out
Nothing behavioral except two status codes; mostly removing places where
the same rule was written down more than once and could drift.

- server/src/private.rs: the owner-only create/write helpers, which
  config.rs, certs.rs, and the session dirs each had their own copy of
  (certs.rs even duplicated the explanatory comment). One module owns the
  modes now, so the "nothing this server writes is readable by anyone
  else" property is checkable in one place.
- server/src/media.rs: the image media-type/extension table, which the
  four places that have to agree on it each spelled out separately --
  storing an upload, serving it back, building a content block, saving a
  produced image. The differing *defaults* stay at the call sites with
  the reasoning, since they genuinely differ by direction.
- routes.rs: a missing file was a 400 and an unreadable one a 400 with a
  hand-rolled log line; they are now 404 and Internal respectively.
  UnknownSession became NotFound, since it was the only 404-with-message.
- main.rs: xdg_dir takes the variable's value instead of reading the
  environment, which drops the unsafe set_var from its test and lets the
  test actually assert the relative-path rule.
- echo.rs had its own 4-byte hex generator beside session::random_hex.
- claude.rs: the two impl Translator blocks were one type's methods.
- Stale comments: phase-2 markers on shipped work, a permission-mode list
  that had drifted from the CLI's, "dev-updater" as the leaf certificate's
  fallback common name, a half-written sentence in build-apk.sh.
- App: the JSONArray walk written out in four fetchers, the four
  near-identical BackHandlers in AppRoot, and SessionScreen's inline
  fully-qualified names where the file otherwise imports.
- server/wg-test.log was committed by accident; *.log is ignored now, and
  the gitignore comments describe where state actually lives.
- PLAN.md's backend layout gains the new modules and drops hosts.rs for
  the ssh.rs that was built instead.

Verified: 35 server tests, clippy clean, app compiles warning-free, and a
scratch server driven over curl -- attachment upload/serve round-trip with
both a known and an unknown content type, the new 404s, transcript and
session-dir deletion, plus a real claude-cli session answering a prompt.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017xn8nHw1tw1R6PtiY1eEtw
2026-08-25 16:10:33 -04:00
irisandClaude Fable 5 967fc814ab Phase 1 server: TLS + token auth, session registry, EchoDriver, SSE with cursors
The whole pipe behind one Driver trait and a common event model:
spawn/list/delete sessions, message + question answering, append-only
JSONL transcripts whose sequence numbers are the phone's resume cursor
(surviving backend restarts), bearer-token middleware wrapping every
route including the fallback, wg0-only binding that fails closed, and
first-run token enrollment via a terminal QR.

Verified: cargo test (10), clippy clean, and curl end-to-end over pinned
TLS -- auth rejection, spawn, streamed SSE replay/resume, /question
round trip, restart continuing seq numbers, delete removing everything.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017xn8nHw1tw1R6PtiY1eEtw
2026-08-24 20:51:34 -04:00