e0eaa4b3f8c9281554eb2cf73eb1df5e6417ae01
100
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
e0eaa4b3f8 | Merge branch 'main' of git.arirex.me:iris/ai-app | ||
|
|
9d5a7cf602 |
Leave the session you are reading alone, and put its menus on their buttons
Three things about the session screen. A notification is no longer posted about the session in front of you: the transcript is already saying it, and one that was posted before you opened it is cancelled, since a row in the drawer for the conversation on screen is the same duplication. Bound to RESUMED rather than STARTED, so a session left on this screen behind another app still reports. The model and permission menus opened 142px clear of the buttons that opened them -- the status bar's height, exactly. Compose measures the anchor in window coordinates, which for an edge-to-edge activity is the whole display, but asks whether the menu fits inside the visible frame, which is that less the system bars; sitting just above a control near the bottom then reads as an overflow and Material3 parks the menu near the bottom of the visible frame instead. Turning clipping off makes both questions about the same window. The model picker now offers "default". The button has always been able to say it -- that is what a session with no model of its own reads as -- but the list could not, so choosing any model was a one-way trip. It is the Claude CLI's own word for "whatever is configured", which its set_model accepts, so it is a request rather than a name invented here. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
0d623b7073 |
Recover a session's context from the CLI's own file
A restarted server has been told nothing, so a Claude session that has not taken a turn since reported its context as unknown -- which was true, and useless, since the CLI had written the figure down at the time and it was sitting in the session's file the whole while. It now reads it from there at load, over the session's transport, in the background: the same three input fields the import list already reads, so it is a measurement rather than a guess. Only when nothing else has answered, and only for a provider that keeps such a file. A clear needs no special case even though it makes the last usage in a file stale, because clearing gives the CLI a new session id -- so the lookup lands on a file with no usage yet and answers unknown, which is what it is. |
||
|
|
bc0a48799c |
Put a question to the reader on a row of its own
An AskUserQuestion arrived in the middle of a run of tool calls and was folded into the collapsed card with them, so the one row where somebody was asked something -- and the answer they gave -- was hidden behind "Called 6 tools" like any other grep. It now starts a run of its own and ends the one before it, which needs no change to the grouping: a run of one is drawn as itself. The calls around it become a group before and a group after, so where the work stopped to ask is legible from the shape of the transcript without opening anything. Echo's `/ask` now runs three ordinary calls on each side of the question, because that is the shape this has to be looked at in and there was no way to produce it. |
||
|
|
09f7f8d203 | Merge branch 'main' of git.arirex.me:iris/ai-app | ||
|
|
5e11b9da80 |
Report the context a session holds, not what it has spent
The number on the status row was a running total of tokens spent, so it could only ever climb: a session compacted from 128k down to 10k, or cleared outright, went on reporting the larger figure, and disagreed with the divider directly above it saying what the compaction had recovered. It now reports what the model is holding -- prompt plus both cache figures -- folded through `driver::context_after`, which is the one rule the pump, the transcript and the phone all use: a turn sets it, a compaction replaces it with what the compaction measured, and a clear leaves it unmeasured. Unmeasured says so in words, because an empty context and one nobody has counted used to look identical. Taken from the turn's last assistant message rather than its `result`: measured against CLI 2.1.237, a two-message turn reported a cache read of 40,211, being 14,259 and 25,952 -- the same conversation counted twice, and no size the model ever held. |
||
|
|
68704fce7c |
Ask the driver whether a command can go, not the status it reported
A `/clear` that did nothing, traced to the end. There was no race to lose: the driver sees every line it writes and every line that comes back, so it always knew. What it knew was being asked of the wrong thing. Two views of "is a turn running" had grown apart. The driver's moves the instant it writes a line; `SessionStatus` moves when output is *recorded*. Messages ask the driver -- which is why they behave -- and commands asked the status, which for a command is stale for its whole round trip: a command's reply carries no assistant text, so nothing proved a turn had started and the recorded status stayed idle from the moment it went out until the moment it came back. A second command in that window went straight out too, landing inside the turn the first one had started, where the CLI reads it as text instead of running it. Nothing anywhere says so: a command read as a message looks like a message. So `Commands` asks `Driver::between_turns()` now, and asks again when it releases a held one -- the recorded idle that woke it is a moment in the past by then. `local_command` says `Running` when it writes, which is both true and what makes the next idle a change worth recording; without it the idle at the end of a command was equal to the idle before it, and nothing behind it was ever released. The other half was a turn nobody here started. The CLI picks the conversation back up on its own -- measured: a backgrounded `sleep` finished nine seconds after the turn's result and it began again unprompted -- and it announces that with a `system/init` about a second and a half before its first assistant text. We had been ignoring that line and learning about the turn from the text, so for that second and a half the session read as idle. It is a turn now, told apart from the `init` at startup by the translator already having a session id, and from our own `/clear` by `running` already being true. Measured against the real CLI, not argued: two `/clear`s sent back to back on one connection now record `commandSent`, `running`, `commandQueued`, `cleared`, `idle`, `commandSent`, `cleared` -- held, then run, in order, both of them. Before this the second was swallowed. The self-started turn shows as `running` eleven seconds after the previous turn's idle, which is the window a command used to disappear into. Also measured on the way, and worth writing down: a message written into a running turn is *folded into it* -- one `result`, `num_turns: 2`, both things answered -- so an idle after one is honest and there was nothing to fix there. A command written when the CLI is genuinely between turns is executed even ten milliseconds after the result, so the boundary itself was never the problem. |
||
|
|
81c8a57181 |
Colour the divider rules to match their words
A compaction line and a clear line each read as one mark now, rather than a coloured phrase sitting in a grey rule that looked unrelated to it. |
||
|
|
451afb50d5 | Merge branch 'main' of git.arirex.me:iris/ai-app | ||
|
|
ea2da0896d |
Open the session a notification is about, and say the two dividers plainly
Tapping a notification landed on whatever the app was last showing. It now opens the session it named. The id rides in the intent's data rather than an extra, because PendingIntent identity is Intent.filterEquals -- with an extra every session's notification would share one PendingIntent and every tap would open whichever session was notified last. MainActivity sorts the aiapp:// URI by host, so enrollment and this are one entry point rather than two. The notification carries only an id, so the session is fetched before there is a screen; a fetch that fails says so and offers to try again, since somebody deliberately tapped and an app that opens to the list explains nothing. That made session-to-session navigation reachable for the first time, and it crashed: SessionScreen remembers a transcript and an event stream, and without a key Compose kept both across the change and merged two conversations into duplicate list keys. Keyed on the session id. The two transcript dividers now say only what they are, centred between two rules: "Compacted <bullet> 128,402 -> 9,617 tok" in blue, and "Context cleared" in red. The rules stay the ordinary divider colour -- they are framing, and the words are what carries the meaning. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
a5659393e9 | Merge remote-tracking branch 'origin/main' | ||
|
|
f1a185a9cd |
Answer a command a session can never run, and put attachments under the text
Investigating a `/clear` that did nothing. What I could measure says the basic path is sound: the CLI honours `/clear` in stream-json mode -- it emits `conversation_reset`, opens a fresh session id, and the model then answers "NO CONTEXT" to a question about something it was told a moment before -- and a `/clear` sent into a running turn is queued here and applied at the boundary, with the model losing context, in two reproductions. What the investigation did find is a command that can wait forever. Held commands drain at the next idle, and a session whose process is gone has no next idle, so `/clear` sent to one sat in the queue with a waiting bubble on the phone that nothing could resolve and nothing anywhere saying why. The *message* path has always answered this case -- a message to the same session reports the exit at once -- which is what made the silence visible: one session answered one and swallowed the other. A command owes the same answer, since what makes it unanswerable is the same fact. `Unknown` still waits. It means nobody could find out whether the process is there and it resolves itself, so refusing on it would turn "we don't know" into "it's gone". `local_command` gets the `closed` check `send_user_message` has had all along, for the window between the status being read and the line being written -- a line into a fifo nothing is reading goes nowhere and looks exactly like one that arrived. Attachments now draw under the message text rather than above it: what somebody wrote is what the bubble is, and it keeps the first line of every bubble in the same place down the transcript whether or not there is an image in it. |
||
|
|
03a8d7d3d3 | Merge branch 'main' of git.arirex.me:iris/ai-app | ||
|
|
ad791a0d84 |
Rejoin a message the page boundary cut in two
joinPages healed a tool call split across a page boundary but not a message split across one, so a long reply came back as two rows with a paragraph break through the middle of a sentence -- visible on any session whose replies are longer than an eighty-event page. Same cause, same cure, and the rule was already written down one member of the set: `foldEvent` never leaves two assistant messages adjacent inside a page, since deltas accumulate into the message before them, so two meeting at a join are always halves of one reply. The newer half keeps its identity for the reason adoptRun gives -- it is the row already on screen. It grows by what the older half brings, which is safe at this join and nowhere else: the join is at the oldest end of what is loaded, so the growth extends off the top, away from the row the list anchors to. |
||
|
|
8a1621a207 |
Load history in one go, and go back to the newest instantly
Two things that made scrolling back feel like work. The jump-to-newest button animated. An animated scroll travels the whole transcript, so the further back somebody has read the longer the press takes -- the one control whose cost grows with how much there is to skip, which is backwards. It goes straight there now. History loaded a page per gesture, and a page is eighty *events*. Eighty events are routinely one row: a reply arrives as hundreds of text deltas that fold into a single message. So a page could land and leave the far end exactly where it was -- and since the far end moving is what asks for the next page, nothing did. The list then only loaded when somebody dragged it again, which is what "it only loads when you touch the top" was. It now keeps fetching until there are rows behind the reader again, and starts doing that a cushion before the end rather than at it. Measured on a session of five very long replies, about two thousand events: reaching the oldest message used to stall at every drag; it now takes flings alone, and the jump back to the newest end is one frame. |
||
|
|
da65c1571f |
Merge remote-tracking branch 'origin/main'
# Conflicts: # app/androidApp/src/main/kotlin/com/example/aiapp/SessionScreen.kt |
||
|
|
b0629f77ca |
Shrink a photo to what the provider takes, and put it in its own bubble
Sending an image was broken in the way that is hardest to see from the phone: a camera photo is twelve megapixels and several megabytes, the Claude API resizes anything past 1568px on its long edge before looking at it and refuses far larger outright, so the picture was uploaded whole over the tunnel to be thrown away or rejected at the other end. Shrunk on the phone, to a limit the server states. Which number it is comes from the provider's *kind* -- `DriverKind::max_image_edge`, reported on the session row -- because that is where a provider's requirements are known, and a phone carrying its own copy of them would be a second place to update when one changes. `None` where nothing cares, rather than a large number: "no limit" and "a limit that happens to be big" are different answers and only one of them stays true. Doing it before the upload rather than after is the point -- the expensive part on a phone is the tunnel, not the decode -- and an image already inside the limit is uploaded byte for byte rather than being round-tripped through JPEG for nothing. EXIF orientation is applied while scaling. The camera writes which way up the picture is into a tag rather than into the pixels, and re-encoding drops it, so a portrait photo would have arrived at the model on its side with nothing anywhere saying so. **What is attached is now visible before it is sent**, in a row directly above the box it will be sent from: the count on the "+" button said how many and never which, so the only way to find out what you had picked was to send it. It scrolls sideways rather than shrinking, and tapping one takes it back off -- an image picked by mistake could otherwise only be dealt with by sending it. The tile is outlined as well as filled, because most of what gets attached here is a screenshot of a dark app and a cropped one is near-black: without an edge the only thing on screen saying an image was attached was the cross drawn on top of nothing. **And the picture is inside the bubble that sent it.** Attachments used to be their own `Image` events emitted just before the message, which drew somebody's screenshot as a row floating above the bubble and left the phone deciding from adjacency alone which message an image belonged to -- a thing the sender knew and could simply say. `UserMessage`, `MessageQueued` and `MessageTaken` carry the refs now, so a waiting message keeps its picture for as long as the turn runs, and a replay puts it back in the same place. Verified on a real claude-cli session rather than an echo one, since the limit only exists for that kind: a 3000x4000 image arrived as 1176x1568 JPEG -- long edge exactly the limit, aspect ratio intact -- and haiku answered "AI Sessions displays idle Photo", which is what the picture was. No error, and the transcript records the message with `images` on it. |
||
|
|
c60b0cbac9 | Merge branch 'main' of git.arirex.me:iris/ai-app | ||
|
|
de049bcec6 |
Hold the transcript still while somebody is reading further back
Two separate defects, both of which moved the list under the reader. The first: a row that grows drags the view toward the newest end. The list is laid out from the bottom, so it anchors on the first visible item's *bottom* edge -- and a reply streaming in extends that row upwards, pushing everything already on screen with it. Measured against a reply streamed in four hundred pieces: scrolling back one screen and waiting six seconds ended at the very bottom, forty lines further on than where it was left. So the transcript now only changes while the reader is at the newest end; anything arriving before then waits in order and lands when they return. Status, tokens and the model still update live, because none of those are drawn in the list and freezing them would trade a jumping transcript for a status row that lies. The second: every markdown row was measured at nothing before it was measured at its real height. The renderer's `content: String` overload parses in a coroutine and draws an empty loading slot until it finishes, so a row composes with no height and springs open a frame later. Seen with five replies on screen at once, all blank, the whole conversation shrunk to a single screen. Parsing in the composition costs a few milliseconds on the main thread and is worth it: no scroll anchoring can survive a row that lies about its height first. `/stream N` in the echo driver is what made the first one reproducible -- `/slow` emits a line a second, and the growth has to be continuous for the anchor row to drag. Verified on the emulator: scrolled back through a whole 400-piece stream, the transcript region is pixel-identical across ten seconds while the status row goes from working to idle; returning to the bottom brings the backlog in one go. Checked the tool-call rig too, which this change had no reason to touch -- paging back still works and every group still reads "Called 8 tools". |
||
|
|
5d47a1ec89 | Merge remote-tracking branch 'origin/main' | ||
|
|
5616ed8aba |
Say which emulator, now that this checkout has its own
`run-android.sh` derives its AVD from the checkout, so two clones of this repo run two emulators, and with both attached a bare `adb` call stops working. The failures do not say so: `adb shell` and `adb get-state` fail with `more than one device/emulator`, and `adb shell pm list packages` comes back empty -- which reads as the app having been uninstalled, and sent dev-updater's session looking for a wipe that had not happened. Its enrolment script mis-diagnosed the same ambiguity as "no emulator is running", whose advice is to start a third. The enrolment line in this file was itself a bare `adb shell`, so it was the instruction that would have produced the confusion. |
||
|
|
49d3439ec4 | Merge remote-tracking branch 'origin/main' | ||
|
|
50670863f4 |
Name a run of tool calls once, instead of after whichever call is first
A group of adjacent tool calls was identified by its first call, and the list is keyed by that identity. But a run can gain members at *either* end -- a new call arriving beside it, or a page of history arriving in front of it -- so its first member is not a name, it is a description that changes. Every time it changed, the row was a different row as far as the list was concerned: the anchor went with it, and the transcript stepped under whoever was reading. Each call now carries the run it belongs to, decided once when it is folded in and never recomputed, and the row keys on that. A lone call that gains a neighbour becomes a group *without* changing identity, which the old key got wrong in the other direction too -- one row was replaced by another rather than updated. `joinPages` hands the arriving older calls the name of the run they are joining, rather than renaming that run after them. The obvious way round is the wrong one: the newer half is the part already on screen, so naming the joined run after the older half renames the row the reader is looking at, which is the whole failure this is meant to remove. Checked against the same twelve-`/tools 8` rig, whose page boundary falls inside the second group: every group still reads eight, so the grouping is unchanged -- what changed is that none of their identities move. Toward the standing rule for this screen, which is that it may only move when the reader is at the newest end and something new arrives. |
||
|
|
2f4dff1435 |
Stop painting code green, and give each checkout its own emulator
Code is not a literal. Green is what this palette colours a literal, so painting a whole fenced block green said the block *was* one -- and it disagreed with the syntax highlighter a tool call's input already gets, where green means a string and peach means a number. Code blocks and inline spans now take the ordinary text colour; the monospace face and the tinted background are what say "this is code", which is the part colour was not doing. `codeColor` goes with it, since nothing else wanted a colour for code. Where a literal really does appear inside code, the thing that should colour it is a highlighter reading the code, not a rule about the container. `run-android.sh` derives its AVD name from the checkout instead of defaulting to a machine-wide `tdep`. That default made the emulator the one thing here that cannot be worked on in parallel: two clones of this repo meant asking whoever had it, waiting, and handing it back, and installing onto a running one takes the foreground from whatever they were looking at. Derived rather than written down, so neither clone names the other's, and `AVD_NAME=` still overrides for sharing one deliberately. Also: two things reported as markdown defects yesterday were not defects, and are worth recording so nobody fixes them twice. The table is not clipped -- it scrolls horizontally, which the renderer does whenever the columns are wider than the screen; a screenshot of one looks exactly like a clipped table, and swiping it shows the rest. The paragraph that appeared to break around an inline code span was an artifact of how the test text was sent through the echo driver, not of the renderer: sent as one message it flows correctly. |
||
|
|
dde592ecb7 | Merge remote-tracking branch 'origin/main' | ||
|
|
47d6b84265 |
Report what a session last did and what it cost, not what this page holds
Three readings that were each a part presented as the whole.
**"just now", everywhere, after a restart.** A relaunched session took its
last-activity from the clock, so every session the backend brought back
claimed to have been active that instant. On the phone that is every row
reading "just now" and the list -- which sorts by it -- coming back in an
order that means nothing, with the conversation somebody was in the middle
of buried among sessions untouched for days. It comes from the transcript
now, in the pass `Transcript::open` already makes, which is the same
correction `last_status` got and for the same reason: a server that has just
started has been told nothing, and the file is the only thing it knows. The
test backdates a transcript by a day, so it cannot pass by the test being
fast; it fails on the old code with the clock's answer in the message.
**The token total was the newest page's.** The phone added up the
`UsageDelta`s it had received, and it opens a session on the newest page of
the transcript -- so a long conversation reported its last few turns as the
total, and a page with no turn in it reported nothing at all, since zero is
drawn as blank. That is the reading Bryan saw: no tokens, on sessions that
had certainly spent some.
The count belongs to the server, which is the only side that sees every
turn. `UsageDelta` now carries the running total beside the delta, filled in
by the pump rather than by each driver -- a driver knows what its own turn
cost and nothing else does, so a new one cannot get this wrong by leaving it
out -- and the session row reports it for a screen that has not opened the
stream yet. The phone takes the largest total it has seen instead of
accumulating, which also means paging older history cannot move it, and
leaves the seeded figure alone for transcripts recorded before the field
existed. Seeded by summing deltas at startup for exactly that reason.
**The header said the model twice and the machine backwards.** A session's
subtitle now reads `machine · provider`, in that order and with no "on"
joining them, matching the list and the usage dialog -- the "on" made it a
phrase, which works in one order and stops working the moment the same pair
is shown somewhere else. The model is gone from it: the footer's picker
already shows what the session is set to, and two places showing it meant
two things to keep in step, which disagreed for a moment on every switch
since one follows the request and the other the session's own answer.
Checked on the emulator against a twelve-turn session whose visible page
held the last six: the header reads "this machine · echo", the status row
reads "idle", and the total reads 42 tok, which is what `GET
/sessions/{id}` says rather than what the page adds up to.
|
||
|
|
2264723ee3 |
Ask again when the app comes back, so a failure cannot outstay it
Leaving the app and returning left "Couldn't reach the server" sitting at the top of a list the server would by then answer perfectly well, and nothing took it off until somebody pressed Refresh. The four tabs draw a snapshot of a backend they are not connected to, so what they show is only as fresh as the last answer. A stale *list* is a small thing. A stale failure is not: it is a claim about right now, and it is wrong in the direction that makes somebody go looking for a problem that has already gone. Returning to the foreground now bumps the same token the Refresh button uses. One instruction the tabs already understand rather than a second path into each of them -- which is also what makes this cover Import, Models and Setups rather than only the list the report came from. Not on first entry, since the tab composing already asks and bumping there would make every cold start fetch twice. Reproduced and fixed against the same sequence: server stopped, app opened so the load fails, server started, app backgrounded and resumed from the launcher. Before, the error is still there; after, the list is drawn and current. |
||
|
|
694535badc | Merge remote-tracking branch 'origin/main' | ||
|
|
135950c8ed |
Say when a session wants you, and stop calling a stop an error
Three things Bryan asked for, and one the second of them exposed.
**Notifications.** A session that asks a question or finishes a turn now
says so on the phone, per session, switchable from its settings screen.
The switch is stored on the backend rather than the phone, because it is a
fact about the session: one that runs unattended overnight should be quiet
on every device, and answering that question again on each device is how two
of them come to disagree. It is on by default -- a notification nobody
wanted is turned off in one tap, where one that never arrived is not
diagnosable at all.
Which moments count is `notification_for`, and the asymmetry in it is the
point. *Waiting on a person* is worth saying however it was reached.
*Finished* is only worth saying when this server watched the work happen:
sessions settle into idle for several reasons that are not "your work
ended", including every one of them being adopted at startup, and announcing
those would put "finished" on the phone for the whole config on every
backend restart. That is the failure that makes somebody switch the feature
off, so it has a test naming every transition rather than the two that
work.
The stream is `GET /notifications`, live only and with no cursor -- the one
place this server does not offer to catch a client up. A notification is a
claim about now; replaying "your turn" from an hour ago sends somebody to a
session that may have been answered from another device since, and a
notification that is wrong costs the trip *and* the credibility of the next
one. What was missed is still on the session list, which says what is
waiting without claiming to be news.
On the phone it is a foreground service, because Android has had no
long-lived background service since 8.0 -- it is what Syncthing does, and
Discord is not a counter-example since it takes a push from Google, which
would mean this backend talking to Google about somebody's sessions. The
ongoing notification Android charges for it sits on an `IMPORTANCE_MIN`
channel: no sound, no status-bar icon, bottom of the shade. `specialUse`
rather than `dataSync`, which is what it looks like: Android 15 caps
dataSync at six hours a day, and a connection that stops listening after six
hours misses the overnight run it exists for.
**A stop is not an error.** The CLI reports an interrupted turn exactly as
it reports a broken one -- `is_error` on a `result` -- so pressing Stop
showed "the turn ended with an error" for doing what the button says. The
line cannot distinguish them; what does is that this side asked, so the
driver says so before the request goes out and the translator spends that on
the next result. The test's second half is the one that matters: the naive
fix passes the first half and silences every genuine failure after it.
**Every status says which one it is.** The session screen's status row named
only `exited` and left the rest blank, so idle and "nobody could read it"
looked identical -- and a just-stopped turn showed nothing, which reads as
the app having lost the session rather than as the stop having worked. The
words are the session list's own, so a state is not called two things
depending which screen you are on. Red on a quota bar now starts at 90%.
**`GET /sessions/{id}`**, which the notification switch found missing. A
screen opened from a list row carries the row the list last fetched: fine
for a title, wrong for a switch, which is *set to* something. Caught on the
emulator, where the switch read on against a backend that said off, with
nothing on screen to say which was true. The screen now reads the session
when it opens, and until that answers the switch is disabled and says so --
a two-position control cannot say "I do not know", so it does not pretend
to.
Verified on the emulator with the app backgrounded: the service holds the
stream (`isForeground=true types=0x40000000`), a finished turn posts
"Finished" and a question replaces it with "Waiting for you" on the same
tag, turning the switch off silences it with no restart, and turning it back
on from the phone reaches config.ron. The interrupt is a translator test
rather than a live turn, which is where that logic is anyway.
|
||
|
|
2dc61c5780 |
Stop drawing one tool call twice where a page of history begins
A page boundary lands wherever it lands, and about half the time that is between a tool call and its result. The newer page then holds a `ToolEnd` whose start it never saw, which the fold draws as a row of its own -- correctly, since a call rendering as nothing is indistinguishable from one that never happened. But when the older page arrived it brought the real `ToolStart`, and the two lists were concatenated, so the call was left on screen twice: once as a proper card and once as a nameless placeholder. `joinPages` merges the two halves by the call's own id instead, which is the one thing a page boundary cannot destroy. The older half wins on what a start knows -- the tool's name, its input -- and the newer on what an end knows, its output and whether it finished. The miscount was the visible part; the moving was the point. The extra row sits exactly at the join, which is where the reader is looking when the page loads, so everything below it stepped down by a row at the moment they scrolled into it. Demonstrated both ways round on a rig of twelve `/tools 8` runs, whose groups are eight calls each and whose page boundary falls inside the second one: without this the transcript reads "Called 9 tools" there and eight everywhere else, with it every group reads eight. That rig is `/mixed N` in the echo driver, added here: N beats of paragraphs at three lengths, single tool calls, runs of adjacent ones, images and peer messages -- every row shape the app draws, in one session, from a command that costs nothing and produces the same transcript every time. The paragraphs are deliberately ragged, because a wall of identical lines looks the same at every offset and makes a scroll of one line indistinguishable from a scroll of ten, by eye or by comparing frames. |
||
|
|
a49120b0c8 | Merge remote-tracking branch 'origin/main' | ||
|
|
69ef6f068a |
Colour the composer by what its buttons do, and let the read-out breathe less
Queue is the paper plane with a clock on it (`md-send_clock`) rather than the plain plane plus the word: the pair is now told apart by the mark, which is what an icon is for, and the word survives as the button's accessible name where a screen reader still needs it. The three composer buttons take their colour from what pressing one does -- green sends now, blue sends later, red takes the running turn away -- and Stop becomes a filled button like the other two. Outlined said it was a qualifier on the primary action; it is a second thing you can do about the turn, and what separates them is the colour and the mark. The fills are named in Theme.kt with their content colour stated beside them, because a semantic colour has to carry its own contrast: these do not change with the surface, so nothing will rescue a foreground that stops being readable. Worth knowing when reading that file: the action greens and reds sit next to a `runningColor` green and a `failedColor` red, which are *states*. Nothing in one set is pressable and nothing in the other is a state, so a reader never has to tell them apart. The usage read-out loses its per-machine cards. A card is a step up the surface ladder and inside a dialog -- already a raised surface -- the step barely rendered while costing 16dp on every side. The machine and the service it answered for are one small quiet line instead of a heading over a subtitle, since the numbers underneath are what somebody opened this to see. The gaps between the bars now go *between* them rather than after each, which is what put a band of empty dialog above Close. The rest of that band was AlertDialog's own spacing, fixed at sizes meant for a sentence of prose and a decision, so this is a plain Dialog with the same container colour and corner and spacing chosen for a dense read-out. Looked at on the emulator: green send, then blue queue beside red stop during a `/slow 20` echo turn, and the dialog over the live session. The account had risen to 78% by then, which showed the five-hour bar and the header glyph going yellow on real numbers rather than forced ones. |
||
|
|
1635fe97c8 |
Take the formatter's line wrapping in NerdIcons
Left over from running ktfmt across the merge: the comment reflows two lines. No change to what it says. |
||
|
|
bb191eec21 | Merge remote-tracking branch 'origin/main' | ||
|
|
e37e90a579 |
Let the server say what is waiting, instead of the phone remembering
A message sent into a running turn was drawn as a pending bubble from
screen state, so leaving the session or restarting the app showed nothing
waiting while the queue was full. Nothing waiting is what "there is
nothing" looks like -- the reader had no way to tell it from a queue that
had already drained, and Bryan hit exactly that: a message he sent
arrived, and his phone stopped showing it after a restart.
The server now records the waiting. `MessageQueued { id, text }` goes into
the transcript when a driver takes a message it cannot deliver yet, and
is resolved by the `UserMessage` carrying the same id -- the same shape
`CommandQueued` and `CommandSent` already had, so this is one more
instance of a mechanism rather than a second one beside it.
The message itself still lands where the session read it, which is what
the last change was about; only the *waiting* is recorded early. The two
are different facts and now have different events.
Paired by id rather than by text. The old code removed the bubble whose
text matched, so sending the same thing twice cleared the wrong one and
left a message on screen that had already been read.
Both drivers that can queue do it: the echo driver too, because the phone
now draws pending bubbles from the stream and a rig that skipped the
event would exercise a state the real app never sees.
Checked on the emulator: two messages sent into a `/slow` turn, then the
app force-stopped and relaunched -- both still drawn as waiting, in the
pending style, and both resolved into ordinary bubbles when the turn
ended and the session read them.
Still outstanding, and worth knowing: an entry outlives a *server*
restart in the transcript but not in the driver's memory, so a backend
restarted mid-queue would leave the bubble drawn with nothing coming to
resolve it. Before this change that message vanished from the transcript
entirely, so the failure is now visible rather than silent -- but it is
not yet right.
|
||
|
|
ff39ef5cf9 |
Say it in icons, and put the whole backend behind four tabs
Six things Bryan asked for, which turned out to be one change: the app had no icon set, so every one of them was blocked on having somewhere for icons to come from. That somewhere is dev-updater's arrangement, ported: a Nerd Fonts subset committed as an asset, drawn as text. `Gear.kt`'s hand-drawn canvas gear argued against icon fonts because a system font may not have the glyph and whoever gets the empty box is never the person who wrote it. The objection is right about *relying* on a system font and the answer is to ship the glyph, so the file is gone and its reasoning is restated in `NerdIcons.kt` rather than deleted -- otherwise the next reader re-derives it. `md-cog` and `md-refresh` are dev-updater's own codepoints, because a cog means the same thing in both apps. The root screen's four words under the title are now four tabs, and the two that act on the whole screen -- settings and refresh -- moved up onto the title row as glyphs. That row's old comment recorded that a fifth word would have had nowhere to go; tabs also say something the words did not, which is that sessions, import, models and setups are four views of one backend rather than four errands. Refresh feeds whichever tab is showing. Import, models and setups lose their headings and their Back buttons, since the tab row is now both. Usage is a dialog. It is checked *against* what you were reading -- "can I start this" is asked with the transcript still on screen -- and it had no navigation of its own, so the only thing its Back could mean was "put this away". The button that opens it is a chart glyph coloured by the worst of the machine's windows, so the row says whether the limits are worth opening before anybody opens them. One `quotaColor` now colours every bar that measures a quota: blue, yellow at 75%, red at 95%. The session bar escalates where it used to sit blue at every level, and the dialog's thresholds moved out of it. A download keeps plain blue at every value -- it has no limit to approach, and colouring it like one would say the opposite of what is happening. States that are not measurements take the ordinary control colour, since blue is the low end of this scale and would read as "checked, and fine" about a machine nobody could reach. Send and stop are the filled paper plane and the filled square. Send keeps the word "Queue" while a turn is in flight, because that is what pressing it then does, and an icon that does two things while looking identical would promise something immediate and do something that waits. Looked at on the emulator: all six glyphs render, the tabs and the system back gesture between them, the dialog over a live session, and the bar bands at 82% and 97% forced through a scratch build, since this account is at 72/31/5 and would only ever have shown blue. |
||
|
|
ba71c798f5 |
Keep what was typed, and ask before a switch that re-reads everything
Two things about the box at the bottom of a session. **A half-typed message survived nothing.** It lived in `remember`, so leaving the screen threw it away, and so did the system reclaiming the app. `Drafts.kt` keeps it per session id and the box is seeded from it. On the device rather than the backend, which is where this app otherwise puts state so every device sees it: this is the contents of a text box on the phone somebody is holding, written on every keystroke, and half a sentence surfacing on another device would be a surprise rather than a convenience. What has been *sent* is the server's, and that is the part which has to outlive this phone. **Switching model quietly re-reads the whole conversation.** The picker did it on the tap, and the cost only showed up as the next turn being expensive. It now asks first, in words, with no number: what it will cost depends on how long this conversation is, and the screen does not know that -- the running total beside it counts what has been spent, which is a different quantity, and a figure derived from it would be a guess wearing a measurement's clothes. The picker beside it deliberately gets no dialog, and that is measured rather than assumed. Driving one session through both changes and reading the CLI's own usage: a warm turn read 30,771 tokens from cache and created 87; after a *permission mode* change it read 30,858 and created 75 -- still a hit; after a *model* change it read nothing at all and created 41,509. So the model picker is the whole of the set, and warning on both would teach that these dialogs can be clicked through, which is what stops the one that matters from working. Nothing is asked when there is nothing to lose either: choosing the model already set, or switching before the session has said anything, applies straight through. Checked on the emulator. A draft survived leaving the session and a force-stop; the dialog names both models and both buttons; declining left the model where it was; and the permission picker still applies on the tap with no dialog in the way. |
||
|
|
76ba24993c |
Make the permission mode the picker offers actually selectable
The session screen's mode picker listed `bypassPermissions`, and on any
session not born in it, choosing it failed:
Cannot set permission mode to bypassPermissions because the session
was not launched with --dangerously-skip-permissions
The CLI is asymmetric about that mode and it is not obvious. It will
launch straight into it on `--permission-mode` alone -- so spawning into
it from the phone has always worked -- but it refuses to switch into it
afterwards unless the process was started with the flag. So the picker
offered a state the session could not reach, and the failure arrived
after the fact as an error line in the transcript.
Sessions now launch with `--allow-dangerously-skip-permissions`, which
makes that mode reachable without selecting it: the session still starts
in whatever mode it was asked for and only moves when somebody moves it.
Deliberately the `--allow-` form; `--dangerously-skip-permissions` is the
one that turns bypassing on for everything, which would take the choice
away from whoever is holding the phone. Since the mode was already
reachable at spawn, this withholds nothing new -- it makes the two routes
to it agree.
Measured both ways round against 2.1.237, driving the control request
directly: without the flag the response is `subtype: error` with the
message above, with it `subtype: success, mode: bypassPermissions`. Then
through the app's own route on a session spawned `manual`, whose argv
reads `--permission-mode manual … --allow-dangerously-skip-permissions`
and which reported `permissionMode: bypassPermissions` when asked to
change.
|
||
|
|
2bf90daada |
Let the markdown renderer draw tables, and stop headings shouting
Two things a reply could not render, both from the same cause: the renderer was pinned nineteen releases back. Tables arrived in the library at 0.30.0. On 0.26.0 a GFM table was not a table at all -- the rows fell through as text and ran together, pipes and all. They now draw as a table, and scroll sideways when they are wider than the phone rather than losing the last column. Headings took the renderer's defaults, which are the Material *display* styles: `#` came out at 57sp and `##` at 45sp, both larger than this app's own screen titles, so any reply with a heading in it read as shouting. They now descend from headlineSmall to labelSmall -- six steps, every one a different size, so two levels of nesting never draw the same. The pin was not carelessness, which is the part worth recording: the version comment says Maven Central was checked on 2026-08-29 and 0.26.0 was the newest stable. It still answers that, because `search.maven.org/solrsearch` is stale for this artifact -- it knows nothing past 0.27.0-rc02. `maven-metadata.xml` in the repository itself lists up to 0.45.0, updated 2026-08-28. The comment now says to read the metadata rather than the search API, since the same check will otherwise be made the same way next time. The colour mapping moved with the API: `markdownColor` no longer carries `codeText`, `inlineCodeText` or `linkText`, which now ride on the typography as the style's own colour and a `TextLinkStyles`. Same Catppuccin values as before. `tableBackground` is set to the tint code blocks use rather than the library's 2%-alpha default, which on this surface was invisible. Checked on the emulator against a reply carrying all six heading levels, inline code, a link, and a three-column table -- including scrolling the table to confirm the clipped last column is reachable rather than lost. |
||
|
|
6b4911c67b |
Give the session's own state a line, instead of the transcript's corner
The token total floated over the bottom-right of the transcript, where a long message ran underneath it, and the working indicator was an item inside the list -- so it scrolled away exactly when somebody reading back wanted to know whether anything was still happening. Both are facts about the session rather than turns in it, so they get one row directly above the box you type into: the thing they report on is the next thing you touch. `exited` moves with them, since it is the same kind of fact and nothing else on the screen would have said it once the indicator left the list. The row is drawn whether or not it has anything to say. An empty one costs a line; a row that came and went would move the text box under a reader's thumb every time a turn started, and would make its own presence the signal for a state it never names. For the same reason the compaction case had to fit the same single line: its bar now takes the row's free width between the label and the total rather than a row of its own, which keeps it far wider than a spinner -- the reason it is a bar at all, since nothing arrives in the transcript while a compaction runs and a small moving thing there reads as a session that has hung. Still no fraction to fill, re-measured today rather than assumed: a real 80,346-to-2,088-token compaction took 23 seconds and the CLI emitted not one line between saying it had started and saying it had finished. Elapsed seconds remain the only honest number. Looked at on the emulator in all three states -- idle, working, and six seconds into a real compaction -- and at 320dp, the narrowest width a phone actually has, where the row still holds one line. |
||
|
|
549e49bc10 |
Record a steer where the model read it, not where it was typed
A message sent while an answer was streaming was recorded in the middle of that answer and above the tool call it ended with. The model had committed to that call in the same message it was already writing, so it had read none of it -- and on screen the tool result underneath read as something the steer had asked for. The answer also split into two bubbles around a message that was not part of it. The driver announced a steer at "the next assistant text or tool call", on the reasoning that anything the CLI says next is proof it has been round the model again. With --include-partial-messages that is not true: the deltas and the tool_use block of a message already in flight keep arriving afterwards, and none of them saw the steer. `message_start` is what actually proves it. The CLI sends the previous call's tool results back before it opens the next assistant message, so that line is the first moment anything written since can have been read -- and it carries no events of its own, which is what makes it a place to put one. Verified against 2.1.237: message_start, the blocks, the tool_result, then the next message_start. The end of the turn stays as the other half, and is the case that must not be lost: a message typed after the final model call has no later message_start, and one that is only recorded when announced would otherwise vanish while a phone drew it as still waiting. Checked live on haiku, before and after. Before: the steer landed at seq 37 among the essay's deltas, with the tool call at 45 and its result at 46. After: essay whole, tool call 42, result 43, steer 44. Also checked the case this had no reason to touch -- a steer sent during a 30-second Bash call, which was already correct -- and it still records after the result. The two tests fail on the old rule; the failure prints the old order, which is the bug. |
||
|
|
a50d72960c |
Say how long is left, not that the window is five hours
The bar read "31% of 5h", which is the one thing about the window a reader already knows. What decides whether to start something now is how long what is left has to last: 80% with twenty minutes to go and 80% with four hours to go are opposite answers, and the second number was a screen away on the usage screen. It now reads "31% - 2h 36m left", and the countdown is driven by a clock the refresh loop advances rather than computed at draw time. A percentage that comes back unchanged is an equal value, so Compose skips the recomposition -- a "left" recomputed only when the quota happens to move would have sat at a stale figure for hours while looking live. A window can arrive with no reset time, so that keeps its own wording: "reset time unknown" rather than "refresh soon", which would be a recommendation nothing measured. Under a minute, including past the end, is "refresh soon" -- "0m left" reads as a measurement. The span arithmetic was already on the usage screen, so it moves into `ResetCountdown.kt` and both callers supply their own sentence. That screen still reads "resets in 2h 37m" and "resets in 5d 21h", checked on the emulator alongside the bar it was not part of changing. The fill is blue rather than the scheme's primary: the bar sits under every session header, on a screen somebody opened to do something else, and it reports a quantity rather than a verdict. The usage screen is still where the same number turns yellow and then red, for a reader who went there to be told where the limits are. Also declares this project's resources for Dev Updater, whose declaration schema changed in d27b5a3: `resources.ron` says ai-app keeps its state as `ai-app`, so the Uninstall dialog offers the real directories instead of saying it cannot tell where they are. Only the name, because both XDG places are the conventional ones. What that dialog's config toggle would delete includes the CA under `certs`, which strands every phone running an APK pinned to it -- noted where somebody would be standing when it matters. |
||
|
|
894180de77 |
Take the CLI's word for a clear instead of inferring it
Bryan reported no divider when clearing. It was not this code -- the backend serving him started at 17:06, three hours before `Event::Cleared` existed, so it has no such event to send and `/clear` reaches it as an unrecognised passthrough. Verified against a current build: the event is recorded. Probing the CLI to establish that turned up something better than what was here. `/clear` in stream-json mode emits a dedicated `conversation_reset` line and *then* a fresh `init` with the new session id -- so watching the id be replaced, which is what this did, was reading the event through one of its side effects. The announcement says it directly, and it arrives first, so the divider now lands above the new conversation rather than after its opening line. That also removes the reasoning the previous commit needed about which id changes count. There is one signal now instead of an inference with two exceptions, and the test that used to pin those exceptions became `an_init_alone_is_never_a_clear`, which covers all three ways an init arrives: a session's first, the one a compaction re-announces with the same id, and the one following a resume. The resume token still follows the id, unchanged -- one CLI event with two observable effects, and each half now reads the half it needs. Verified end to end against a real claude-cli session: message, /clear, message, and the transcript reads userMessage / assistantText / cleared / userMessage, in that order. 74 tests, clippy and rustfmt clean. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01VETa8afmpWaYezLCqJhDB8 |
||
|
|
71067275e2 |
Stop claiming a terminal, and stop calling a recoverable delete final
Two bugs Bryan hit, with one shape between them: a claim stronger than the thing that was measured. **The delete warning branched on `imported`.** It told him deleting `ai-app` could be undone and deleting `manager` could not, when both are claude-cli sessions whose conversations survive equally. `imported` records how a session got into the app; what decides recoverability is whether the *driver* keeps its own record -- the Claude Code CLI does, under ~/.claude/projects, however the session started; echo and llama.cpp do not, and for those the app's transcript is the only copy. So the fact now sits on DriverKind and rides on SessionInfo, decided by the server from the provider's kind rather than by the phone from its name, which a person can change. The comment above the branch asserted "a session started here has no copy anywhere". That sentence was the bug written down and reasoned from, and it is gone. Neither branch promises a restore, which it should not: nothing here checks the file is still on disk, and re-importing was never a restore anyway -- this app's transcript holds images, peer messages and command events the CLI's record never had. So the recoverable text says what is known and names what goes either way. "Can't be undone" is now said only where it is true, which is the point of saying it at all. **"open in a terminal -- close it there first" named a place that need not exist.** The detection is right and worth keeping: something live holds that session, and importing it would reproduce the double-resume incident. But which something was never measured. The live descriptors here include two of this backend's own adopted sessions and a peer agent's; none is a terminal, so the instruction sent the reader looking for a window that was not there. **And this app did not recognise its own spawned sessions.** The import list filters out what the app is already driving, but it matched only the import cursor -- which exists solely for imported sessions. Every session the app spawned therefore stayed in the list, marked in use, telling the reader to go and close it somewhere: here. Matching the resume token too, which both kinds have, is the fix; `session_importing` is now `session_driving`, because that is what it was always being asked. Verified on the emulator against a scratch backend: a spawned claude-cli session reports keepsOwnTranscript true with imported false -- Bryan's `manager` case exactly -- and draws the recoverable warning; the echo session draws "can't be undone"; and once the CLI named itself, the spawned session's id was absent from the import list, where the old match would have listed it. 75 tests, clippy and rustfmt clean; ktfmt, compileDebugKotlin and lintDebug clean. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01VETa8afmpWaYezLCqJhDB8 |
||
|
|
5f6fd34aba |
Record the clear that happened, not the one that was asked for
`ClaudeDriver::clear` emitted `Event::Cleared` beside the `/clear` it sent, so the divider recorded a request. A reader scrolling back takes that mark as a fact about the conversation -- the session no longer has what is above this -- and a request is a different claim from a result. Compaction already gets this right by taking its mark from the CLI's own `compact_boundary` rather than from somebody pressing Compact; this is the same rule, and it was the one place left applying it to the request. Caught in review by the session this was measured against, which also established that the CLI's `/clear` is declared `supportsNonInteractive` and returns empty text with no result line -- so a fresh `init` bearing a different `session_id` is the only trace it leaves. The reader already watches for exactly that in order to persist the resume token, so the mark now goes out there. It has to be a *replacement* rather than any change, and the tests pin both ways of getting that wrong. The first `init` sets the id from nothing, which would otherwise open every session with a divider announcing a clear that never happened. And a compaction re-announces `init` carrying the *same* id, which would otherwise draw a clear on top of the compaction's own mark -- that one was found by writing the test rather than by reasoning about the change. 73 tests, clippy clean, rustfmt clean. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01VETa8afmpWaYezLCqJhDB8 |
||
|
|
cf3ed68511 |
Tell "nothing to meter" apart from "couldn't find out"
Found by looking at the bar on an echo session rather than by reading the diff: it said "5-hour usage unknown -- this machine reports no usage", which is the failure the rest of this file was written to avoid, one level up. A machine with no metered provider is never asked by the backend, so it returns no snapshot for it. The bar read that silence as a failed lookup, because Unavailable was the nearest word it had -- and a session on `echo`, or on a local llama.cpp, has no paid quota at all. That is a fact about how somebody set the machine up, not a question that went unanswered, and reporting it as unknown nags about a deliberate choice on every screen forever. So the state exists now: NotMetered, drawn as nothing, because there is nothing. Unavailable keeps its words and its reason and still covers the three ways an answer can fail -- nobody logged in, machine unreachable, snapshot without the window. Verified on the emulator against the real endpoint: a setup carrying claude-cli draws the bar at 22% of 5h, selected by kind "session"; the no-snapshot path was the one on screen before this change, so it is reached, and this only changes what it draws. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01VETa8afmpWaYezLCqJhDB8 |
||
|
|
6680fbb987 |
Clear from the phone, and put the numbers where they are read
Four changes to the session screen, three of them Bryan's and one that
fell out of them.
**/clear is offered like any other session command.** It joins
SESSION_COMMANDS, so it suggests itself while being typed and goes out
through the command endpoint that /compact already uses -- no new path,
and the boundary pump holds it mid-turn exactly as it holds a compaction.
**A clear draws a divider, not a deletion.** `Event::Cleared` becomes a
ClearedNote row saying that everything above stays here and is no longer
sent. That sentence is the row's whole job: the reader can see the
conversation is still on screen, so without it the divider reads as
something having been thrown away, which is the one thing it is not. It
carries no counts, because nothing was measured -- a compaction's
numbers are real and there is no equivalent here to report.
Compaction and clear now share `TranscriptDivider`. They are the same
kind of mark to somebody scrolling back -- "the session no longer has
what is above this" -- and the difference belongs in the words rather
than in how they are drawn, so the styling is written once and cannot
drift.
**The five-hour usage bar sits under the session header.** It reports
the paid service's own metering for the machine this session runs on,
fetched from that machine, refreshed every minute off the backend's
cache. It is never derived from the transcript's token counts: those are
a different quantity measured differently, and a quota-shaped bar built
out of them would be a guess wearing a measurement's clothes. Not
knowing has its own appearance and its own words -- "unknown" and why --
because a bar resting at zero because a machine is unreachable reads as
plenty of headroom, which is the opposite of the truth. The window is
selected by the API's own `kind` ("session"), added to UsageWindow in
this change, rather than by matching the label a person reads.
**The token total moved from the header to the bottom right of the
transcript**, pinned above the input rather than scrolling with it. In
the header it was one item in a run of dot-separated facts about the
session and read as another of them, rather than as the running total it
is.
SessionSummary now carries the setup id, which it deliberately did not.
The stated reason was that nothing here addressed a setup and holding
both id and name invited showing the wrong one; the usage bar addresses
one, so the reason lapsed rather than being overruled, and the comment
now carries the rule that replaces it: never display it. The server has
always sent the field, so nothing changed on the wire.
ktfmt, compileDebugKotlin and lintDebug all clean.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VETa8afmpWaYezLCqJhDB8
|
||
|
|
1e3eec8cf0 |
Let a session drop its context without ending
Adds `Event::Cleared`, `SessionCommand::Clear`, and `Driver::clear`, so
`POST /sessions/{id}/command {"text": "/clear"}` does for a session what
the CLI's own `/clear` does for a terminal.
The marker is a divider, not a truncation: everything above it stays in
the transcript, because that is the only copy of the conversation the
phone has and a person scrolling back is a different question from what
the model is given. It also makes clearing mean one thing across
drivers -- `claude` sends `/clear` and the CLI answers with a fresh
`init` whose new session_id the reader already persists as the resume
token, so the next launch resumes the cleared conversation with nothing
to keep in step; `llama` needs no state at all, since `conversation()`
already folds the transcript and now folds from the last marker; `echo`
emits the marker alone, so the phone's divider and scroll behaviour can
be exercised without spending a real session's context.
That fold is why `Cleared` is documented as load-bearing rather than
decorative. For any driver that rebuilds its conversation from the
transcript, this marker decides what the model sees, and treating it as
something only the phone draws would silently put the cleared
conversation back in front of the model at full price.
Clear rides the existing boundary pump like any other SessionCommand, so
one arriving mid-turn waits exactly as a compaction does, and nothing
grows a second way to wait.
Removes `--autocompact` in the same change, because clearing is the
cheaper answer to the problem it was added for and Bryan would rather
manage context that way. Keeping the measurement here, since it was the
reason for the constant and is worth more than the constant was:
context returned to 70-85k within ten calls of a compaction; a
compaction took 104,346 to 147,671 ms; compaction cost that session
2,655,508 tokens across six boundaries, of which the single automatic
one at the 1M ceiling was 1,696,870. Clearing costs nothing, because
nothing is sent.
70 tests, clippy clean, rustfmt clean.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VETa8afmpWaYezLCqJhDB8
|
||
|
|
797513bb86 |
Raise the compaction window to 200k
100k is the cheapest window on tokens and the wrong one to sit in front of. Measured on the session this was written against: context returns to 70-85k within ten calls of a compaction, so a 100k window compacts about every thirteen calls, and a compaction takes roughly two minutes (durationMs 104,346 to 147,671 across the six recorded). A 130-call request would have spent some twenty minutes compacting -- optimising the number that was asked about while making the thing somebody actually waits for on a phone considerably worse. 200k keeps most of the saving against the 1M ceiling and halves the stalls. The comment now also says what the window does not do, because measuring this turned up the opposite of what the byte counts suggested. Images are 93% of the bytes that tool calls put into that transcript but only 8% of the context growth -- the adb wrapper's downscaling holds a screenshot to a median of 476 tokens, while text-only calls add a median of 740 and a mean of 1,139. So the file is large because of screenshots and the context is large because of ordinary tool output, and only the second one is what this constant governs. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01VETa8afmpWaYezLCqJhDB8 |
||
|
|
4f586d6551 |
Attribute the drift to the CLI's ceiling, not to a manual compaction
The reasoning on AUTOCOMPACT_WINDOW cited 491,562 tokens as where "the CLI compacted it". That was a manual /compact somebody ran, not the CLI's own trigger, so the comment credited a person's intervention to the automatic behaviour it was arguing about. Caught in review by the session whose transcript it was measured from. Corrected from that transcript's compaction boundaries: the window left to `auto` was 1M, and the one automatic compaction fired at preTokens 1,000,184 with the API context peaking at 999,668. So the drift ceiling is twice what the comment said, and near it a single tool call bills about 100k tokens rather than 49k. The correction strengthens the case, but it also changes what the example is evidence *of*, which is why it was worth fixing rather than just raising the number: what held that session together was the person in it running /compact by hand four times, and the 4.2-million-token request happened at the merely-large contexts left between those. The constant is for the sessions where nobody is doing that. No behaviour change. 68 tests, clippy clean, rustfmt clean. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01VETa8afmpWaYezLCqJhDB8 |
||
|
|
36984a9e1b |
Put the compaction window's reasoning on the compaction window
|
||
|
|
4a4057886d |
Compact at 100k rather than letting a phone session drift
Left at `auto` the CLI picks a very large window, which suits a terminal session somebody closes at the end of the day and does not suit this app at all: these run for hours, nobody closes them, and the transcript carries screenshots. One session here reached 491,562 tokens of context before the CLI compacted it. That matters because every API call re-reads the whole context, and one request is not one call. At half a million tokens a single tool call bills about 49k before it does anything, so "can you make it so you can rename a session?" cost 4.2 million tokens across the 130 calls it took. Measured over that session's life: 2,498 calls, 1.08 billion cache-read tokens. 100k is the smallest window the CLI accepts and roughly the cheapest. Per-call cost falls with the cap, while the compaction it forces costs about the same in total either way -- a smaller window compacts more often, but each pass is proportionally smaller. What it trades is how much detail survives a compaction, which is a real cost to the work and the reason this is one named constant with the reasoning written down rather than a computed value. Passed before the resume/name branch, so it applies to adopted and imported sessions too -- which are the large ones, and the ones this is for. Verified: the exact argument list the app now spawns starts, accepts an empty stream-json stdin and exits 0, so the flag combination is good without spending a token. 68 tests, clippy clean, rustfmt clean. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01VETa8afmpWaYezLCqJhDB8 |
||
|
|
620a7d0a83 |
Trim what every session pays to load
Measured: an app-spawned session starts at ~33,200 tokens of context, of which ~20,000 is written fresh on every spawn -- the always-loaded rule files and this file -- and only ~13,200 comes from a shared cache. That 20,000 is billed at 1.25x on every single session start. This file drops to 19,882 bytes from 21,293. What went is narrative that PLAN.md already carries in more detail (the phase history, the submodule drift story) and the parts of "Where things run" that MACHINE.md states once for every project. What stayed is every operational fact: the commands, the llama.cpp and ssh test recipes, the import rules, and everything under "Things that have bitten". The global chain was trimmed in the same pass, 43,039 -> 34,069 bytes, mostly by moving the Gentoo host build profile out of the @import chain into ~/.claude/HOST_BUILD.md, which MACHINE.md now points at. Nothing was deleted there either; it is referenced rather than loaded, the same arrangement this file has with PLAN.md. Worth being honest about the size of the win: ~10,400 bytes is roughly 2,200 tokens off each session start. It is real and permanent, but it is not what makes a long session expensive. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01VETa8afmpWaYezLCqJhDB8 |
||
|
|
749b2db287 |
Make a command a thing the app knows, and hold it until it can run
Typing "/" now suggests what this app understands -- `/compact` and `/rename <name>` -- with a line each about what they do, and anything else beginning with a slash is passed to whatever runs the session, because a dialect's own vocabulary grows without this list. None of them are messages, and that is the substance of the change. A line written into a running turn is read by the *model*, so a command sent mid-turn either does nothing or arrives as text somebody has to puzzle over. They now wait for the turn to end. The waiting is done once for every provider, in the pump that already watches every event for the boundary, rather than in each driver where a new provider could get it wrong by leaving it out. Waiting is a state, so it is on screen: the command sits at the reader's end of the conversation in blue, with a spinner and "waiting for this turn to end", and becomes an ordinary blue row when it goes. Blue because these are about the session rather than about the task -- the same blue a compaction already used, which is now one colour with one name rather than two. Renaming from the settings screen sends exactly this, so it waits and draws the same way. The name itself is not held: it is this server's own datum, so the list and the header change at once and only telling the session waits. Echo grew the same split, which is where the bug in it showed: its commands are its messages, so running one announced a `MessageTaken` as well, and the same line drew twice -- once blue, once purple. A command owes no announcement; the manager has already recorded that it was sent. Watched rather than reasoned about: `/compact` during a 25 second turn held with its bubble up, went out when the turn ended, and the compaction that followed reported what it recovered. |
||
|
|
bebaae7a94 |
Carry a question in the event model, not in one provider's JSON
A question is now fully described by the event that reports it: the tag it was asked under, each option's label, what it means, and the sample of what picking it would produce, plus whether several may be picked at once. The app renders from that alone. It had been reading Claude Code's tool input to find the parts the event dropped -- that dialect's schema, written out a second time in Kotlin, where no other provider could reach it and where it would drift the first time the schema moved. Echo could not describe an option at all, and llama never will. Answers travel as a list for the same reason. A question that takes one answer sends a list of one rather than being a different shape, and the one place that flattens it is where the CLI is spoken to: its answers map holds a string, so several choices are joined there. That join was in the phone. Also here because it is the same rule: the permission ask reuses the question body rather than owning a second one, so Allow/Deny renders and resolves through exactly the code an AskUserQuestion does. Verified against both, since a refactor that only satisfies the case it was written for has been tried on the half that cannot fail: a two question `/ask` answered from the phone, one option and then two, and a real sonnet session's `rm -f` permission asked, allowed, and run. |
||
|
|
fea8e7e92b |
Show every option a question offers, on the call that asked
Reported by Iris through the dev-updater session: a two-question AskUserQuestion arrived with only one option visible per question, so the answer she sent was the only one she had been offered. The cause was a `Row`. It hands out intrinsic widths in order and clips whatever runs past the edge, so the first option or two drew and the rest went off the side of the screen -- which does not read as a bug, it reads as those having been the only choices. The same Row was in the permission ask beside it; both wrap now. That pairing is the reason to look: a rule stated on one member of a set is usually missing from the others. The rest of what she asked for, and what each was: - It drew twice, as the tool call and again as loose question cards, because the backend marked these questions as belonging to no call. They belong to the call that asked, and now say so. - So it renders like any other tool: one card, its own heading, opened because a decision cannot be made from a closed row. - Each option shows its description and its `preview` block, which is the part a reader is deciding on and none of which was reaching them. - "Other" is a field on every question. The harness always offers it, so leaving it out narrowed a question that was never that narrow. - A multi-select sends the labels it collected as one string, which is the tool's own schema rather than a guess -- its answers map is string-valued. - No spinner while it waits. A spinner says the machine is working; here the machine is idle and the turn is stopped on the person, so the card says "your turn" in the colour this app already uses for that. Verified against a real session as well as the echo fixture: haiku asked two questions with three described options each, both were answered from the phone, and the model carried on with the answers. Echo grew `/ask` so the shape can be looked at without paying a model to produce one, and its option cards are outlined rather than tinted -- as one surface step up they were three paragraphs where three things to press should be. |
||
|
|
cae04c2559 |
Delete one session without putting the rest through loading
Pressing Delete refetched the whole list on success, so every other row went back through its loading state and the reader got a blank screen for the length of a round trip -- to report on something that was never in doubt. Now the row being deleted fades, says so where its status goes, and stops responding to taps; when the server answers, that one row is removed and nothing else moves. A refusal keeps the row, because it is still there: the server answered and said no, so the session it said no about is exactly as it was, and the error goes on its own card as it already did. Faded rather than removed on the way out, deliberately. Taking the row away when Delete is pressed is a promise about a request that has not been answered, and putting it back when the server refuses is worse than never having taken it away. Looked at rather than reasoned about: the in-flight state lasts milliseconds against a local server, so I slowed the delete route to four seconds, watched the faded row and its spinner, watched it removed on success, then killed the server and watched a refusal leave the row in place with the reason on it. |
||
|
|
d3fff3d229 |
Let a session be renamed, under the same name everywhere
A gear at the end of the session's own bar opens what can be changed about that session; the name is the first thing there. Compact is gone from that bar -- `/compact` typed into the message box is the CLI's own way to ask and it already worked, so the button was a second way to say one thing. Echo takes the typed word too now, since it is the rig the compaction display is checked against and losing the button would have taken that with it. The name is this server's, not a driver's: it is what the list shows, it exists before any process does, and every provider has one. So it is settled in the config and the driver is *told* -- which is the opposite of the model and the permission mode, and the difference is written down at `Driver::set_title`. A driver whose process has no notion of a name does nothing and says nothing, because there is no failure to report. Claude Code has one, so the name reaches it: `--name` for a session we create, and `/rename` afterwards, which is a local command rather than a control request -- `set_session_name` is not a subtype it knows, which I established by asking it. A resumed session is deliberately not renamed at launch: an import already has a name, quite possibly one the person typing in it chose, and taking that would be helping itself to something the app was only shown. Verified end to end rather than argued: renaming from the phone put "Session renamed to: paging and scroll" in the CLI's own session file, and the session now lists under that name to other agents. The gear is drawn rather than set in a font, for the reason Chevron gives. It was a sun on the first attempt -- thin teeth standing clear of a thin hub -- which no amount of reading the diff would have shown. |
||
|
|
1629e0911e |
Say what the line splitter would do with a bare carriage return
`complete_lines` splits on `\n` only, which is right -- this stream is JSONL, and a record terminated by a bare `\r` would not be a record -- but the doc comment said why the remainder is held without saying what decides where a line ends. Worth the sentence because of what the failure would look like if the CLI ever wrote such a line: the session goes quiet, the process is healthy, nothing errors, and the cause is a line splitter. The dev-updater session hit exactly this shape today reading cargo's progress line, which is `\r`-terminated for redrawing in place, and lost a whole build's worth of output to it. |
||
|
|
f842d0e512 |
Call the server component "server", to match dev-updater
READ BEFORE PULLING. A Managed component's service unit is named <config key>-<component name>, so this rename moves the unit from ai-app-backend to ai-app-server and nothing points at the old one afterwards. Uninstall the backend component from its card *first*, while it is still called "backend"; then pull, accept the new declaration -- .dev-updater.ron is a request, so the card shows it as pending -- and build. The unit installs under the new name. Two things reset rather than break, both keyed by component name: the per-component built_from sha, and the build and runtime logs. One build makes the sha current again. Also drops the claim that the components list is walked in order. They have built in parallel since 2026-08-28, so the reasoning the comment gave -- backend first, so a failing APK leaves the phone what it had -- no longer describes what happens. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01VETa8afmpWaYezLCqJhDB8 |
||
|
|
3eccf7e443 |
Show the model and mode the session has, not the ones it was asked for
Picking either from the phone wrote the choice straight into the session's state and then sent the request. Asking and having are different things, and the difference is not rare: `auto` is a permission mode the CLI accepts on the command line, silently resolves to `default`, and refuses outright over the control channel -- "auto mode unavailable for this model" -- so a session spawned in auto was in default and one switched to auto stayed where it was, with the phone reporting auto in both cases. So the drivers report what they are set to and the manager follows that. Measured, because the confirmations are not uniform: a model change answers success with no value, so what was asked is remembered until the answer arrives; a mode change echoes the mode it became, and that answer wins over the request; and `init` names both -- resolving `haiku` to claude-haiku-4-5-20251001 -- which also covers a session adopted from a terminal that set them outside this app. A driver that cannot change either already says so with an error, and now that error is the whole story rather than a note beside a display that changed anyway. The config keeps the requested value, deliberately: that answers a different question, which is what to launch this session with next time. Two things fall out. Control request ids are random rather than the clock, because two in the same second shared an id and something now looks them up. And the phone shortens a resolved name for the button -- `haiku-4-5` -- since the full one is what the CLI reports and roughly twice the room that row has once Stop is in it. |
||
|
|
404066fa7d |
Say when a session is working, and what it was told
Three things a phone could not see, all of them the same shape: the session was doing something and nothing on screen said so. A turn nobody here started never reported itself. `Running` was sent where a message was *sent*, so a session picked up mid-turn, one compacting on its own, or one another agent wrote to sat there reading as idle until it finished. The driver now says it from what it observes -- output that could only come from a turn in flight -- which is the same set of events that already announced a steer, with the ends swapped. An imported session had it worse: nothing but replayed lines ever reaches it, and a status was not among them, so it was permanently whatever it was when it was adopted. Its file does not record a turn ending, but it does record why each assistant message stopped, and `tool_use` versus anything else answers it. A record that says nothing leaves the status alone rather than voting for idle. Messages from other agents were dropped outright: the CLI marks them meta, and this replayed everything except meta. They are now a row of their own, closed by default like a tool call, named for the session that sent it -- not the reader's own bubble, because they did not say it, and a session working on something this phone never asked for is exactly what one of these explains. Measured against a real session file rather than guessed: the peer record carries the sender's name and the message body in `origin`, beside a copy wrapped for the model to read. |
||
|
|
f18639e4b1 |
Count the far end of the list in rows, not events
Scrolling back stopped dead at the top of what was loaded, and no older page ever arrived. Bryan spotted the cause from the outside: it had to do with tool calls being collapsed. The trigger compared an index into the list being drawn against `items.size`, the number of transcript events. Those were the same number when it was written. They stopped being the same when adjacent tool calls started folding into one row, and the queued bubble and the working indicator are two more rows with no event behind them. In this session 645 rows stood in for 720 events, so the last visible index could reach 646 and the threshold it needed was 717. It was not close; it was unreachable, and the further a session went the worse it got. Both numbers now come from the list itself, which is the only place they are commensurable, and `totalItemsCount` counts whatever gets added to it next. Checked on the emulator against the case it was breaking on rather than a clean one: five collapsed "Called 8 tools" groups in front of a 720-event transcript, scrolled from the bottom to seq 1, which is the beginning of the session. It stops there because that is the top, and holds position while each page arrives. |
||
|
|
fc71cb4403 |
Say "unknown" for a session we are not driving but cannot bury
A session in the config with no live entry reported `Exited`, whatever the reason. That covers three different situations -- one that genuinely ended, one that failed to relaunch, and one whose process could not be checked -- and the wrong one is the expensive one. `Exited` reads as "this conversation is over", and what a reader does about it is start a fresh session. If the process is in fact still running, that is a second CLI against a conversation that already has one: the exact fault `session::process` exists to prevent, arriving through the status field instead of through a spawn. So it is said only when the process is known to be gone. A record that cannot be checked reports `Unknown`, and so does one that is still alive -- this server is not driving it, so it genuinely does not know what that process is doing, and the honest word is the one meaning "wait" rather than the one meaning "act". A session with no record at all is still `Exited`: an echo session, or one already stopped and cleaned up, and known to be. The distinction was available all along -- `process::recorded` returns the liveness -- which makes this the same mistake as the other five today: reporting what was convenient to compute rather than what was measured. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01VETa8afmpWaYezLCqJhDB8 |
||
|
|
279c76e8a1 |
Hold the reader's place when a message arrives
Scrolled back through a conversation, every new message dragged the view with it -- which reads as the screen scrolling down on its own, at the exact moment somebody is trying to read something else. The list had no keys, so its rows were identified by position. The transcript is drawn newest-first, so a new message is an insertion at index 0: every existing row shifts up one index, the viewport stays on the index it was on, and the content slides through it. The working indicator appearing and disappearing did the same thing at the same end. So rows now carry the identity they always had in the data. Every transcript event has a seq, which is what the transcript is ordered by and never changes, and every row keeps the seq of the first event behind it -- a streaming message keeps the seq of its first delta, so it holds still for the whole answer rather than becoming a new row on every frame, and a tool call keeps its start's. Paging older history is the same insertion from the other end, and it is the thing this could plausibly have broken. Checked on the emulator: scrolled back mid-turn, the view sat still through twenty seconds of streamed deltas, and scrolling to the far end still fetched earlier pages and stayed where it was while they arrived. |
||
|
|
5396da76c7 |
Show a compaction happening, and what it recovered
The Compacting status had been declared, rendered in four places, and never once emitted: no driver produced it, and the app had no control to ask for a compaction in the first place. Pressing nothing for two minutes and then quietly having less context was the whole experience. The CLI turns out to announce all of it, which was worth measuring rather than guessing at. Driven through /compact against 2.1.237 it emits a `status: "compacting"` line at the start, a `status: null` carrying `compact_result` at the end -- `"failed"` with a sentence saying why, when it does -- and then a `compact_boundary` with the token counts. The same records appear in the CLI's own transcript file with camelCase keys, which is the obvious place to read the shape off and gets every field name wrong. So none of it is inferred here. The driver writes the line and says nothing; the translator reports what the CLI reports. A failed compaction surfaces the CLI's own sentence, which is specific enough to act on. The counts are the part worth keeping afterwards, so they land in the transcript rather than only in a status that vanishes: a session that went from 128,402 tokens to 9,617 has just been given its context back. They are optional throughout, because a compaction whose size nobody reported has to be able to say so -- a zero would read as "recovered nothing". Also here, all found on the way: - `rename_all` renames variants; fields need `rename_all_fields`. Every field in Event was a single word until `pre_tokens`, which went out as snake_case, was not found by the app, and rendered as the "no counts reported" case -- a state it is allowed to be in, so nothing looked wrong. There is now a test on the wire names. - The unparseable-line warning sliced bytes, not chars, on output that is full of em dashes. A panic there kills the task reading the session's stdout, and the session goes deaf with nothing on screen. The other three truncations in the tree already did this correctly. - Echo compacts too, with invented numbers and a real shape, so this screen can be looked at without spending two minutes of somebody's account to reach the state. |
||
|
|
42131c75d6 |
Send a steer into the running turn, and put an image under its call
**The queue was holding messages the CLI would have taken.** Two claims in this file contradicted each other: the module header said a mid-turn message is injected at the next tool boundary -- "the behavior this app exists for" -- and `Queue`'s own doc said a line written mid-turn simply becomes the next turn. The code followed the second, parking every message until `Status::Idle`, which is the end of the whole turn. Measured rather than argued, twice. Writing a line straight into a live session's stdin fifo mid-turn produced one `result` for the whole thing, so it was consumed inside that turn, not as a new one. The header was right and the queue was built on the wrong claim. The cost was exactly what Bryan reported: he steered after the second tool call and it sat unread until every remaining call had finished. Measured before and after on the same three-step turn -- steer sent at +13s, recorded at +24.7s before this change and at +14.1s after, which is the next tool boundary. So the line goes out immediately. What stays behind is the *announcement*: the CLI says nothing on stdout about having read a message, so `MessageTaken` now waits for the next assistant text or tool call, which is proof another model call happened and the steer was in it. That keeps a held message drawn below the working indicator until the session has actually taken it -- the thing that mattered when this was last changed -- without delaying the message to get it. Idle counts too, and is the case that must not be missed: a message written after a turn's last model call has no later output to prove anything. `closed` is untouched, and `Queue::close` still reports held messages by name rather than dropping them. **An image now names the call that produced it.** `Event::Image` gains `about`, the `tool_use_id` from the tool result it came out of, so a screenshot is drawn inside that call's card instead of floating beside it -- pairing them by position is what a page boundary breaks. `None` for a person's own attachment, which belongs to no call. The import path threads it through as well, so replayed history reads the same as live. Images show whether the card is open or closed: a call whose result *is* a picture says less closed than the one line it replaced. Verified on the emulator against a real haiku turn: the checkerboard sits inside `Read /tmp/tiny.png`, and the steer sits between that call and the next, where it was taken. 53 tests, clippy, rustfmt, lint and ktfmt clean. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
184b6fc6a6 |
Correct the claim that reattach is local only
Written down as "an ssh session's child dies with its connection, so it takes the ordinary --resume path". The code never had that branch: `start` records a pid whatever the transport, and for a remote session the process the backend owns is the ssh client. Adopting it is right -- the fifo feeds it, its logs capture the far end, and ssh lives exactly as long as the remote command, so its liveness is the session's. The docs claimed less than the code does, which is the safe direction to be wrong in but still wrong, and it was about to mislead someone: a remote `claude` has an sshd pipe on stdin under every version of this server, because the fifo is on the backend's side of the connection. Reading a remote session's stdin therefore says nothing about which backend started it, and we were an inch from concluding otherwise. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01VETa8afmpWaYezLCqJhDB8 |
||
|
|
4cbd567c35 |
Take ktfmt's formatting
Committed unformatted again: I piped the check through grep, so the task's failure never reached the shell's exit status and the chained commit ran anyway. Check exit codes, not output. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
73923dbbc3 |
Make jump-to-latest the same chevron, pointing down
It was a labelled button beside a tool group that collapses with a drawn chevron -- two controls doing the same kind of thing in two visual languages. Now one `Chevron` composable serves both directions, parameterised rather than copied, since a pair that differs by a minus sign drifts and the drift is a bug in exactly one direction. The comment it replaces argued against an arrow here, on the grounds that the list is laid out upside down. That reasoning was about the code: nobody reading the screen knows the list is reversed, and on screen the newest message is at the bottom, which is where this goes. It draws no text, so the name lives in its content description -- the whole of what a screen reader has, and the answer to "what was that arrow for" later. Verified on the emulator: scrolled up, the chevron appears bottom centre matching the group's; tapped, it returns to the newest message and takes itself away. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
011ed0d0e1 |
Hold an image's place, open it full screen, and fold a run of calls
Five changes to how a transcript reads. **Images no longer move the page.** The row was as tall as whatever had loaded, so it grew when the bytes arrived and pushed everything below it -- and in a bottom-anchored list, an image loading above the viewport moved the text under the reader's eyes. The height is now decided before the fetch and never changes: four lines of the body style, measured from the type so it stays four lines when the reader has scaled their fonts. Nothing to see when loading finishes, which is the point. **A small image is enlarged with nearest neighbour**, a large one shrunk smoothly -- decided per image from its actual size rather than set once, since blowing a 16px sprite up with interpolation turns it into a blur of exactly the thing being looked at. **Tapping one opens it full screen**, fitted so the whole image is visible first, with two-finger zoom to 8x and pan once zoomed. A dialog rather than a screen, so back returns to the transcript. **A tool call is one line closed**: the tool's name and what the call is for. The command is not on it, because a wrapped command turns one row into four. Open, it shows the command, the rest of the input and the output, with the timeout at the top right -- a limit on the call rather than part of what it does, worth seeing beside the command it constrains. A call waiting on permission is shown open regardless, since the command is the thing being decided. **Adjacent calls fold into "Called n tools"**, closed by default, and it closes again from either end -- a long group's heading scrolls away while its last call is still on screen, and the reader who wants it shut is looking at the bottom. The calls keep their full width; what says they belong together is the surface behind them, one cue rather than two half-cues. Grouping happens at display time, not in the fold: the transcript's own order is what paging and the stream depend on. Echo gains `/tools [n]` so a run of calls can be produced without paying for one. Verified on the emulator: four calls folded and expanded, one opened inside the group showing `timeout 5000` top right, a 16px checkerboard enlarged with hard pixel edges beside a shrunk screenshot at the same height, the screen byte-identical between one second and six after opening, full screen fitted, and back returning to the same scroll position. Pinch itself is the one thing not verified here -- `adb input` cannot inject a two-finger gesture. 53 tests, clippy, rustfmt, Android lint and ktfmt clean. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
8fe13634cb |
Take ktfmt's formatting on the files just added
Four files went in unformatted: I ran the formatter mid-change and then kept editing. ktfmtCheck is part of finishing, not part of starting. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
d580461bd5 |
Show what was remembered as a note, not as markup
Claude Code marks a sentence taken from its stored memory by wrapping it in `<cc-memory filenames="...">`. Markdown has nothing to say about that, so it arrived as literal angle brackets mid-sentence and read as the model having emitted broken HTML. It is the opposite: a claim about where something came from, and "I was told this before" and "I worked this out just now" are different things the reader could not otherwise tell apart. Each one becomes a card naming the files it came from, with the prose either side of it left as prose. Named rather than merely tinted, since a colour can say "this one is different" but not what kind of different. A tag still arriving is left alone: streaming means the closing half may be seconds away, and a half-written marker is not a marker yet. Verified on the emulator with two notes in one message, one of them citing two files, and prose before, between and after. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
c5dbd1535d |
Ask for permission on the call it is about, and read the input
A bash permission request arrived as a second card repeating the tool
call's input verbatim, so the same command appeared twice and the reader
had to work out it was one event. `Event::Question` now carries `about`:
the `tool_use_id` the CLI's `can_use_tool` request already names. That
makes the pairing a measured fact rather than a match on input text --
and it stays `Option`, because AskUserQuestion is not permission for
anything and an echo session's question is about no tool at all. Those
still draw as their own card, which is what every question did before.
The card also reads the input instead of dumping it. Every tool's input
is JSON, and showing it raw makes the reader parse `{"command":"…",
"timeout":5000}` to find the line they care about. A small table says
which field is the subject of which tool -- Bash's `command`, Read's
`file_path` -- and the rest is still listed, since dropping a field
would claim the tool had no other input when it might. The subject is
syntax-highlighted with dev.snipme:highlights, for the reason the
markdown renderer is a library: lexical rules are somebody else's
specification. Its theme is Catppuccin, mapped in Theme.kt beside the
rest of the palette rather than taken from the library's defaults.
The input shows whether or not the card is expanded. A row that says
only "Bash" says nothing anyone can act on, least of all when it is
asking to run something.
Verified on the emulator against a real haiku session: one card, the
description, `grep -rn "needle" /tmp | head -3` highlighted, `timeout:
5000` pulled out, and "Allow Bash?" with its buttons inside the card --
then Allow, which resolved in place and ran.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
206319045d |
Show usage per machine, and say why a machine has none
The server now reports limits per machine, so the screen has to as well: one card per machine that offers a paid service, named by the machine first, because these are one account's numbers and which account is decided by which box ran the session. It also has to say which of four things happened, and the reason for splitting them shows up here rather than in the data. A machine nobody has logged in on is working exactly as somebody set it up, so it reads as a plain statement in ordinary text -- marking it would be the interface nagging about a decision already made, and would dilute the marks that do mean something. Only "couldn't reach it" and "the endpoint refused" are coloured as faults, and they say different things because they need different things done. The old screen drew all three in the error colour. No machine offering a paid service is not an error either: it says so instead of drawing nothing. The app also stopped parsing: `available` no longer exists and `getBoolean` on a missing key throws, so this had to land with the server change rather than after it. An older backend sending no `state` is read as "failed" rather than "ok", since an empty card drawn as healthy is the worse failure. Looked at running, against five machines: local reporting notLoggedIn with the backend's HOME emptied, loopback-over-ssh returning real windows beside it, an unreachable host showing ssh's own message in red, and a machine with no Claude provider correctly absent. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01VETa8afmpWaYezLCqJhDB8 |
||
|
|
6166b1f626 |
State the deny-unknown-fields rule where it governs all the bodies
It had landed appended to `SshRequest`'s doc comment, so a rule about every request body in the module read as something about how to describe a machine. Moved to the module doc beside the route table, where the set it governs is what a reader is already looking at, and worded so a new request body knows it is expected to carry the attribute too. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
446fb92ba3 |
Refuse a request field the server does not know
A misspelled field was accepted and dropped. Sending `permission_mode` instead of `permissionMode` produced a 200 and a session running in the default permission mode -- so the caller's setting was gone, and nothing anywhere said so. That is the expensive shape: indistinguishable from success at the place you are looking. It cost an hour here, chasing a "startup race" that was a key serde had silently discarded; with the name spelled the way the API asks, a bypassPermissions session runs a `sleep` loop with no prompt at all. So every request body refuses unknown fields, not just the one that bit. Axum's message names the offending field and lists what was expected, which is the whole of what the caller needs. Query strings are deliberately left permissive: a stale link carrying an extra parameter is not a mistake worth failing a request over. 53 tests, clippy and rustfmt clean; verified against the running server that the misspelling is now a 422 naming the field and the correct spelling still spawns. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
d8570f4d5a |
Ask each machine about its own limits, not this one about all of them
`ClaudeUsage` read `~/.claude/.credentials.json` on the machine running the backend and asked the API about that account, once, globally. But a session runs wherever its setup says, so the numbers on the usage screen belonged to the backend's account rather than to the account that spent the tokens. That is not a rounding error in the layout this project is aiming at. `ai-server` belongs on the host; the host has no `claude` CLI at all and the VM is a remote. So the screen would have reported "is Claude Code logged in on this machine?" while every session ran fine on a machine whose limits nobody could see. It looked correct only because backend and CLI happen to be the same box today. Usage is now per machine, asked through the same `Transport` the sessions use -- `ssh host sh -c 'cat $HOME/...'` for a remote, unchanged for the local one. `$HOME` is left for the far shell to expand, since a path built here is this machine's home directory and over ssh that is somebody else's. Machines with no Claude provider are not asked and get no row: they have no Claude limits, and a row about them would be a fact about nothing. The snapshot gains the states it could not say. `available` plus an `error` string made three different situations look identical, and the one that suffered was the harmless one: a machine nobody has logged in on is a decision somebody made, with nothing to fix, and it read as broken. `notLoggedIn`, `unreachable` and `failed` are now distinct, and which one a failed read is gets decided in `why_no_credentials` rather than at the call site. Supporting changes: `ssh::command` builds a `std::process::Command` that tokio converts from, so a blocking caller can use the one place that knows what a correct ssh invocation is; `Transport::capture_blocking` is that caller's door. The cache is keyed by machine and service rather than by position, since the set is no longer fixed at startup -- a positional cache would hand one machine's numbers to another the moment a setup was added. A cached snapshot still picks up a rename immediately, because the name has nothing to do with the poll interval. Verified against a real ssh setup (loopback, per AGENTS.md) with five machines, all four states seen: local `ok`, loopback-over-ssh `ok`, unreachable host `unreachable` carrying ssh's own message, a claude-less machine correctly absent, and -- with the backend's HOME emptied -- local `notLoggedIn` while the ssh machine still reported real windows, which is the production shape. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01VETa8afmpWaYezLCqJhDB8 |
||
|
|
9f403cddab |
Render markdown, and stop wiping messages still waiting to be read
Two things, both about the transcript telling the truth about itself. Markdown is rendered rather than shown as its source. The parsing is mikepenz/multiplatform-markdown-renderer, not something written here: markdown is somebody else's specification, and a hand-written subset of one disagrees with it at the edges, which is where the bug reports come from. `Markdown.kt` is only the mapping onto this app's palette, so code, links and rules take the Catppuccin values the rest of the app uses rather than the renderer's defaults. The queued-message list was cleared wholesale whenever a turn ended. But the backend holds a queue of its own and takes one message per turn, so a turn ending is precisely the moment the *rest* are still waiting -- the bubbles vanished while the messages were on their way, which reads as everything after the first having been dropped. Now a held message leaves the list exactly two ways: the session reads it, which arrives as a UserMessage, or its send failed and there is nothing to wait for. Measured first, because the report was that the backend dropped them: three messages sent behind one long turn were all delivered in order (ONE, TWO, THREE) against current main, so the loss was in the display. Verified on the emulator: headings, emphasis, inline code, nested lists, a quote bar, a fenced block, a rule and a link all render, and the three queued messages sit through their turn. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
90a57ca7e9 |
Revert "Default a new session to bypassPermissions"
This reverts commit
|
||
|
|
d81c9a7d65 |
Default a new session to bypassPermissions
Auto already allows most of what a session does, so the prompts it did raise were mostly the interruption without the choice -- and on a phone each one is a round trip to a question card. Measured rather than assumed: a haiku session spawned this way ran a bare `echo` and a `sleep` loop with no prompt at all. The stricter modes stay one tap away in the same picker, which is where a session that warrants one picks it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
19e3525131 |
Say what continuing a session will cost, not just how big it is
The import list reported a file size, which predicts the wrong thing. Most
of a large transcript is history from before a compaction, and the model is
not given it again: of the 133 MB session behind the 2026-08-29 incident,
99% of the bytes sat before its last compaction summary.
So each row now carries the tokens the model was actually holding at the
last turn -- the input side of the most recent assistant message's usage,
prompt plus both cache figures, which the CLI records itself rather than
anything inferred from the file. The two disagree in exactly the way that
makes the size misleading. Measured on this machine: `ai-app` is an 80 MB
file with 150k of context, while `ai-app-backup` is 3 MB with 481k. The
smaller file is the more expensive one to continue.
Absent rather than zero when no turn has recorded usage, since a session
with no turns has no figure rather than a figure of none.
The row is three lines instead of one run of separators:
path cut at the head, keeping the tail, and the only thing here that
is cut -- one long value with no natural break, where the lines
below it are short enough to wrap
stats named, context, lines, size
warning only when there is one, in the warning colour
The warning gets its own line and its own colour because it differs in kind
from the stats rather than in degree: those describe the session, it says
whether taking it is safe at all. Colour makes it findable, the words make
it actionable -- "open somewhere else" and "we could not check" are not
distinguishable by shade.
Titles no longer ellipse either; they wrap.
Looked at on the emulator rather than reasoned about, including the states
that are not the default: a long path truncating, a row with no warning,
and a row with no context figure.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VETa8afmpWaYezLCqJhDB8
|
||
|
|
c1a468432d |
Read a session's transcript once at launch, not twice
Reading the last status back from the transcript -- added so a restart stops claiming an exited session is idle -- walked the whole file a second time, after `Transcript::open` had just walked it for the sequence number. Both answers are wanted at the same moment by the same caller, so the cost was paid per session at exactly the point a restart is trying to be quick. `Transcript::open` now finds both in its one pass and reports the status it saw. The free function goes; a transcript knowing what it last recorded is where that belongs anyway. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01VETa8afmpWaYezLCqJhDB8 |
||
|
|
3a74bd9c35 |
Make the process record survive a crash mid-write
Reviewing the reattach code found the fault it exists to prevent, sitting in its own save point. `process::write` used `fs::write`, which truncates before it fills. A crash inside that window leaves no readable record -- and a missing record reads as "nothing is running", which is the single answer that makes the next launch start a *second* CLI against a conversation that already has one. The window is not rare: the record is rewritten on every read that makes progress, so many times a second while a turn is producing output. Written to a neighbouring file and renamed over the real name now. The rename is atomic, so a reader sees the whole old record or the whole new one. That also makes the fixed-size padding pointless -- a rename replaces the file rather than overwriting part of it -- so it goes. Two more from the same pass: - A failed read of the stdout log was logged and nothing else. The session then went deaf with nothing on screen: no more output, no error, a status that stayed wherever it was. It now says so, closes the queue rather than stranding messages in it, and reports `Unknown` -- not `Exited`, because the process may well still be running; what failed is this server's ability to hear it. - Sizing the stderr log by reading it. `read_from` with a large offset answers "how long is it" by allocating the whole file first, which on a chatty process is a large pointless read on every reattach. `size_of` asks the filesystem. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01VETa8afmpWaYezLCqJhDB8 |
||
|
|
aff2cb90e2 |
Stop reporting the app being switched away from as a failure
Backgrounding the app left "Lost the event stream (SocketTimeoutException: null)" waiting at the top on return. Android stops the activity, the socket dies with it, and the reconnect loop -- which kept running on a phone nobody was looking at -- recorded the failure. Switching apps is a choice somebody made, not a fault to report. Worse, it could not clear. `streamError` was reset when an event arrived, so a session that reconnected and then sat idle displayed a connection error it had already recovered from, indefinitely. That is the expensive half: a stale failure is indistinguishable from a live one. So the stream now runs only while the screen is at least STARTED, which makes the drop a deliberate close rather than an error (EventStream already distinguishes them), and resuming reconnects from the same cursor. What takes a failure off the screen is `onOpen` -- the measured moment the server accepted the connection -- rather than the first event to follow it. The message that does get shown leads with what will happen next rather than with the exception's class name, which named nothing the reader could act on. lifecycle-runtime-compose is declared rather than inherited from activity-compose, for the reason core-ktx already is: this code calls repeatOnLifecycle and LocalLifecycleOwner directly now, and a transitive could change under it. 2.11.0, the current stable. Verified on the emulator against an idle session, which is the case the old code could never clear: backgrounded 35s, returned, no banner -- and a message sent afterwards arrived live, so the reconnect genuinely reattached rather than merely staying quiet. Build, lint and ktfmt clean. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
d7c692a4ec |
Show a replayed session's images instead of dropping them
An imported session showed no screenshots. `text_of` kept only `text` blocks, so every image in the replayed tail was silently discarded -- while the *live* translator has always saved them into the session's `files/` and referenced them. Two readings of the same records, and the one used for history was the lesser. `save_image` moves out of `Translator` to a free function both paths call, since the naming scheme for that directory should exist once. `events_from` now takes the session directory to write into, which means the conversion has to happen where that directory exists -- so `Seed` carries the raw JSONL and `launch` turns it into events, rather than `routes` doing it before the session is created. Costs nothing in tokens, which is the point worth recording: this writes into ai-app's own session directory and the phone fetches a reference only when it draws one. Nothing here is ever written to the CLI's stdin -- it reads its own session file, and the only things this app sends it are typed messages, control requests and `/compact`. Verified against the 133 MB session behind the 2026-08-29 incident: 45 images in the replayed tail, written as real PNGs and served over the files route, with the transcript itself staying at 756 KB of references. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01VETa8afmpWaYezLCqJhDB8 |
||
|
|
d96bc041a7 |
Say how big a session is before it is imported
The import list reported a line count, which is the wrong axis: these transcripts embed screenshots as base64, so one line can be a megabyte. On this machine a 69 MB session has 3,427 lines while a 44 MB one has 6,792 — the number on the row said nothing about what continuing the session would cost, and size is the only thing there that predicts it. The session behind the 2026-08-29 incident was 65 MB across 13,000 lines, a line count that looks unremarkable. Shown beside the line count rather than instead of it, since a short file of long lines is exactly the expensive case. Not warned about and not marked: importing a large session is a choice somebody is entitled to make, and flagging it would be the interface nagging about a decision already taken. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01VETa8afmpWaYezLCqJhDB8 |
||
|
|
362d436d4f |
Let sessions outlive the backend, and never resume one twice
Three `claude` processes ended up running against this checkout on 2026-08-29, and the account hit its session limit. One cause, several ways in. An agent imported the Claude Code session it was *itself* running in. That is an ordinary import, and importing runs `--resume` -- so a second CLI attached to a file the first was still writing. The whole 65 MB conversation, 154 embedded screenshots included, was re-appended to the transcript under a new prompt id; both copies then read each other's writes as work done elsewhere, and the adopted one was billed for re-reading all of it. Meanwhile `shutdown_all` asked each session to stop and the process exited immediately, so the SIGKILL timer died with the runtime, the stop was unreliable, and whatever survived was orphaned with nothing written down to find it by. The processes leaked either way. So leak them on purpose, and be able to pick them back up. A session's process now outlives the backend and is adopted again on the way up, which is worth having for its own sake: restarting the server no longer ends a turn somebody is waiting on. Its stdio lives in the session directory -- a fifo opened read-write so the process is its own last writer and never reads EOF, plus stdout/stderr logs read from a byte offset. `session::process` records the pid *and* the kernel's start time for it, because a pid alone is reused and adopting a stranger's would mean never resuming the real conversation. That makes the fix structural rather than a check: everything goes through `ClaudeDriver::launch`, which adopts if it can and starts if it cannot, and `--resume` is reachable only on the second path. `Driver` gains two ways out where it had one -- `detach` (coming back) and `stop` (the session is being deleted, so the process must not survive). Importing a session that is open is now refused outright. Claude Code keeps `~/.claude/sessions/<pid>.json` for every live session, so this is a measurement rather than a guess; it reports no/yes/unknown, because a machine that keeps no such record cannot answer and "could not check" is not "nobody is using it". `SessionStatus` gains `Unknown` for the same reason. Also here, found on the way: - A reconnecting phone was sent the entire backlog. Opening a session was bounded to a page but reconnecting was not, so a long disconnect delivered thousands of events one frame at a time. Past `CATCH_UP_LIMIT` the stream sends a `reset` frame and the newest window, and the client rebuilds from it as it does on open -- without the reset the window is spliced onto rows no longer adjacent to it. - A session's status was assumed idle at launch. Read from the transcript instead, so a restart stops claiming an exited session is waiting for you. - `llama-server`'s stdout was piped and never drained, so a chatty one blocked on a full pipe buffer mid-load. It goes to a log now. - A turn that exited or errored never emitted `Idle`, so the queue stayed "running" for good: every later message was held forever and, since a message is only recorded when taken, vanished with nothing on screen. - Two doc comments had drifted onto the wrong functions. Verified by killing the server mid-turn: the process survived, finished its turn unattended (12.8 KB of output nothing was reading), and the restarted server adopted it -- one process, all 700 lines in the transcript, no hole, and it still took a new message afterwards. Deleting a session stops its process; a 266-event backlog resets while a 16-event one streams. 46 tests, clippy and rustfmt clean, app compiles and lints. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01VETa8afmpWaYezLCqJhDB8 |
||
|
|
9791afcfd6 |
Hold a queued message below the indicator until the turn takes it
The backend records a message the moment it is sent, so one sent into a running turn landed in the transcript at the time *we* spoke -- above the working indicator, in among things the session had already read. It had not read it. Showing it there says otherwise. Messages sent while a turn is in flight are now held below the indicator, drawn quieter, and take their place in the conversation when the turn ends. That is an approximation and worth naming: the CLI injects a queued message at a tool boundary, and tells nobody when it does, so the end of the turn is the first moment anything here can honestly say the message was taken. It errs toward "not yet read", which is the direction that cannot mislead. Also in this change, from the same pass over the screen: the app draws above the gesture strip rather than under it, and the three status colours that were literals in two other files -- an amber, a green and a red off Material's defaults -- are now Catppuccin members in Theme.kt beside the rest of the scheme. The ordering is verified by construction rather than photographed: the list is bottom-anchored, so the first item emitted is the lowest on screen, and the queued block is emitted before the indicator. Staging a real long-running turn to photograph cost four model turns and never produced one, because the model kept declining to sleep -- which is its own finding, and the reason the next change is a test command in the echo driver. |
||
|
|
f402a1f6c8 |
Hold the bottom while typing, and put the status where the answer goes
**Typing moved the newest message out of sight.** The list re-pinned on a new item and nothing else, but the two things that shrink it while somebody writes a reply are the field growing from one line to four and the keyboard opening under it -- neither of which is a new item. So the message being replied to drifted upward, and the view only came back when the reply was finally sent, which is the one moment it did not matter. It now watches the viewport as well as the item count, and only acts while already pinned. **The status has moved out of the corner.** As a label up there it said something about the session; at the end of the transcript it says something about a place -- this is where the next thing appears -- and that is where the reader is already looking, because it is where the last message is. `exited` is still said, in the same place. Removing the corner label without it would have left a session whose process is gone looking exactly like one waiting for input, on a screen whose whole purpose is typing at it. Verified on the phone-sized case rather than reasoned about: three lines of text with the keyboard open keeps the newest message directly above the field, and sending shows "working" immediately under the sent message. |
||
|
|
e03b757dec |
Stop losing tool calls at the seam between transcript pages
A tool's start and end are two events folded into one row, and the fold only ever *updated* an existing row -- so an end whose start was not in the same fold changed nothing and vanished. Not a broken row: no row at all, which is indistinguishable from a tool that never ran, and is what Iris saw as gaps where she remembered work happening. Paging made it routine. Each page was folded on its own and prepended, so every seam split whatever spanned it: 30 tool ends in the first page of this conversation, one of them already orphaned before a single scroll. Two changes, because the two halves fail differently. Pages are re-folded from the events they came from rather than folded apart and stitched together. That needs the events kept beside the rows, since folding is one-way. One pass over everything loaded, paid only when somebody scrolls back, which is the moment they asked for it. And an end with no start now creates a row instead of disappearing. Its name is unknown from an end alone, so it says "tool" until the earlier page arrives and replaces it -- a row that admits what it does not know beats silence, because silence is a claim that nothing happened. Verified across a real seam: scrolling back through the 80-event boundary of an 864-event import is continuous, with no gaps where tool calls were. |
||
|
|
2fe34176c0 |
Give the transcript a way back, and stop losing an import's name
**Every imported session was called "claude-cli session".** The app has nothing to say about the title -- the server names it after the session it is continuing -- so it sent `""`. That is `Some`, which satisfied the `or_else` meant to catch "no title given", so the imported name was computed and then thrown away in favour of the `<provider> session` fallback. Normalised at the boundary instead: blank means absent, because that is what it means to the person who left it blank. Both the client's value and the imported one go through the same trim, so neither can be a string of spaces standing in for a name. **And a jump-to-latest button**, shown only while the newest message is off-screen. Reading back through a conversation is a place to be rather than a state to be rescued from, so it waits to be wanted and leaves once there is nowhere to jump to. It says where it goes instead of drawing an arrow. The list is laid out from the bottom, so "down" in the data is up on the screen, and an arrow would be asking the reader to hold that in their head to press a button. Verified on screen: an import now arrives titled "ai-app" rather than "claude-cli session", and the button appears on scrolling back, returns to the newest message, and disappears on arrival. |
||
|
|
dcb158ee44 |
Fetch the transcript instead of replaying it one event at a time
**The five seconds of loading top-down.** Opening a session subscribed to the event stream from sequence zero, so the backlog arrived as one SSE frame per event -- 864 of them for an imported conversation, rendered as they landed. That is not a slow list; it is a conversation being replayed at network speed, and it looks like loading from the top because it is. The newest page now comes as one request, and the stream starts from where that page ended, carrying live events only -- which is what a stream is good at. Scrolling back fetches the page before it, so history costs something only when somebody actually reads it. 80 events instead of 864, and the first frame is already the end of the conversation. I had called this fixed after anchoring the list at the bottom, on the strength of an emulator on the same machine as the server. That test could not have shown the problem: the whole backlog arrived in one frame's worth of time over loopback. Iris's phone, over a tunnel, took five seconds. **Send disappearing while running.** It was never conditional -- the row simply ran out of width. A Row hands out intrinsic widths in order and clips the overflow, so when Stop appeared the pickers I had added pushed Send off the screen: the app's central control, gone at exactly the moment the app is most in use. The settings now share what is left after the actions have taken what they need. While a turn is in flight the button says **Queue**, because that is what sending then does -- the message is injected at the next tool boundary rather than starting a turn of its own. The backend has always done this; the button was describing something else. **And the model picker no longer dismisses the keyboard**, which it did by taking focus. Changing the model mid-sentence is an aside, not a departure from what you were typing. Verified on the 864-event import: at the newest message within a second, history paging back continuously past the first page, and Stop beside Queue while running. |
||
|
|
ba25a5cacf |
Open the transcript at the bottom instead of travelling there
The list was built oldest-first and then scrolled to the end, so opening a session started at the top and raced downward through everything in it. On an imported conversation that is nine hundred items measured before a word is readable, and it was visible every single time. Laying the list out from the bottom removes the journey rather than hiding it. The newest message is index 0, so the first frame is already the right one, and older items are composed only when somebody scrolls back to them -- which is also what makes a long history cheap rather than something to load up front. Following the tail gets simpler as a result. There is no longer a moment where new content pushes the anchor away, so "am I pinned" is read straight from the scroll position instead of being remembered across scrolls, and a new message is one step back to index 0 rather than a jump across the transcript. Checked on the 863-event import of this very conversation: a screenshot one second after opening is already at the newest message, and scrolling back reaches older ones in the order they happened. |
||
|
|
3f8805a610 |
Say which kind of delete this is, and stop offering what is already open
**Deleting was one word for two different acts.** An imported session's real transcript belongs to Claude Code and outlives anything this app does, so removing it here undoes a view. A session started here has no copy anywhere, and removing it ends the conversation. The dialog warned "this can't be undone" of both, which makes the warning worthless on the one where it is true -- and frightening on the one where it is not, since what it actually deletes is a cache of a conversation still sitting on the machine. Sessions now report whether they were imported, and the dialog says which act this is. No new mechanism: the soft delete already existed, it was just indistinguishable from the hard one. **And a session already open here is no longer offered for import.** Importing one twice would leave two `--resume` processes appending to the same transcript, each seeing the other's writes as work done elsewhere and replaying them -- both sessions then showing a conversation neither is having. The route refuses it as well, so the rule holds for anything not going through the app. Left out of the list rather than shown and disabled. The usual argument says absence is ambiguous, and it is wrong here: an imported session has not disappeared, it has moved to the session list, which is where it now belongs. Absence means "already somewhere you can reach it". Deleting the app's copy puts it straight back -- verified: 68 offered, 67 after importing one, 68 again after the soft delete, which is also the clearest demonstration that a soft delete keeps the conversation. |
||
|
|
d2915c12fa |
Record how the import sync tells its own writes apart
The reasoning belongs where somebody would look before changing the poll: status is the obvious discriminator and is wrong, and the failure it produces reads as the model repeating itself rather than as a bug. |
||
|
|
a9ea84c96c |
Stop choosing a model, and keep an imported session up to date
**Why the model became fable.** `spawn_session` fell back to the
provider's first listed model when none was given. That list is a shortcut
for the spawn screen, written in whatever order somebody typed it, and its
first entry is `fable` -- so every session spawned without a model, which
is every import, silently became a fable session. It looked like a default
and was an artefact of list order. Absent now means absent: no `--model`
flag, and the CLI uses whatever the person configured for themselves.
**Model and permission mode are now visible and changeable** from the
session, as buttons that read as their current value rather than labels
beside one. The mode was spawn-only; the CLI turns out to accept
`control_request{subtype:set_permission_mode}` and echo the mode back,
probed against 2.1.237 the same way the rest of the protocol record was.
Both default to `auto` -- on a phone every ask is a round trip to a
question card, which is how "allow Bash?" became the most-answered
question in the app.
The mode is reported by the API so the picker shows what the session is
actually set to, and it is kept in the live session beside the model for
the reason the model already was: `meta` is the shape a session was
*launched* with, so reporting from it shows the value a change replaced.
**And an imported session keeps itself level with its source file**, so
work done at a terminal arrives without a button. `--resume` appends to
the same transcript rather than forking -- measured, not assumed -- so the
only hard question is which new lines came from here.
Answered by counting the events this session has recorded. Status is the
obvious signal and is wrong, which cost a round trip to find: a turn that
starts and finishes between two polls reads as idle at both, so its output
is replayed on top of itself. It showed up on screen as `donedone`, and
only because the reply was one word -- with a longer answer it would have
looked like the model repeating itself.
Verified against both halves: text appended to the source file the way a
terminal writes it appears within one interval, and a message sent through
the app appears exactly once, before and after a turn.
|
||
|
|
c3e7f07a5d |
Pin the transcript to its tail, and stop asking about every command
**The scroll.** The transcript scrolled on new items and nothing else, which missed the two cases that matter most. An imported session's history arrived and left the view wherever it landed; the keyboard opening shrank the viewport and slid the newest messages under the IME, so typing meant typing into a view showing the middle of something. The view is now pinned to the tail, and it is the reader's scroll that decides: settling anywhere above the bottom releases the pin, settling back at the bottom re-arms it. The pin is written only when a scroll *ends*, so it survives the moment when new content has just pushed the bottom away but the reader never moved -- deriving it continuously from "is the bottom visible" would release it on every append, which is the race that makes naive follow-the-tail implementations let go at random. New items and viewport resizes both re-scroll; the jump is instant rather than animated, because an imported session appends hundreds of items at once and animating through them is a light show. **The input field gets a row of its own**, above the buttons. Sharing one row put the full width behind three controls, so the thing being typed into was the narrowest thing on the row. **Permissions default to auto, and importing can choose.** The spawn screen defaulted to "manual" and imports passed no mode at all, so the CLI asked about everything -- and on a phone every ask is a round trip to a question card, which is how "allow Bash?" became the most-answered question in the app. Both paths now default to auto, with the other modes one tap away for a session that warrants caution. Looked at running, all three: an imported session opens at its bottom, the tail stays visible while typing with the keyboard open, and the mode picker shows auto selected. |
||
|
|
233689ced6 |
Name a session in the import list, and let one be deleted
Three things about finding a session in a list of seventy, and one about getting rid of it. **A name beats anything inferred.** `/rename` appends a `custom-title` record, so if somebody has said what a session is, that is the row. Eleven of the seventy here turned out to be named already and none of it showed. **Otherwise the last thing said, not the first.** The question this list answers is "which one was I just in", and a session's opening line is the least distinctive thing about it -- several of these begin with the same slash command. Finding that last message took three tries, and the two wrong ones are worth recording because they failed in opposite directions. Grepping the user record type caught tool results, which are *also* user records -- so a session that ended mid-tool showed a tail of empty records and a row saying nothing was said, when plenty had been. Narrowing to a string `content` fixed those two and broke twenty others, because a message carrying an attachment stores its text in a list. Excluding `tool_use_id` keeps both shapes of a real message and drops the one that is not: seven rows still have nothing to show, and those are sessions that really are empty. **Sorted by when it was last used**, and the time is on the row. Naming was tried as the first sort key and is a worse list -- it buries what somebody was just doing under everything they ever named. A name is for recognising a row, not for ordering it, so it stays as the title and as a word beside it. **And a session can be deleted**, which is asked before it is done. The transcript *is* the session, so this ends any chance of resuming that conversation, and the dialog says exactly that rather than "are you sure?". Deletion resolves the id against what the machine reported, like importing, so no path crosses the wire in either direction. Looked at on the emulator, including the dialog -- which is where the delete button turned out to be missing entirely after a patch that compiled fine, and where the row layout got its second look. |
||
|
|
6bbc829a3e |
Import a Claude Code session the machine already has
Claude Code keeps every session as JSONL under `~/.claude/projects/`, and the CLI continues one with `--resume <id>`. `claude.rs` already resumes whenever it finds a resume token in the session directory, for crash recovery -- so importing is that same path with the token written before the driver starts, and there is deliberately no second way to begin a session. The seed goes through `launch` with the ordinary spawn, so the driver never learns which kind it got. Two things the machine answers and the phone does not. **Which sessions exist.** One command per setup rather than one per file, for the reason discovery already gives: over ssh each would be its own connection. Titles come from the first few user records rather than the first, because a session opens with records the CLI injected -- slash commands, caveats around local command output -- which are stored as ordinary user records without the meta flag, so titling by "first user record" produced a list where most rows read `<command-name>/clear`. **Which file an id names.** The phone sends an id and never a path; the server looks it up again among the sessions it enumerated. An enrolled token must not be able to turn a spawn into "read me this file", which is the same rule that keeps a provider's command out of `POST /setups`. Only the tail is replayed. The imported conversation is for reading -- continuing it is the CLI's job, and it reads the whole file itself -- so this is a display budget, and it has to be one: the session this was written in is 39 MB, and all of it would otherwise cross a tunnel to a phone. A recorded working directory can outlive itself, which this found immediately: every session from before the checkouts moved to `~/repos` still records `~/host/repos/...`. Resuming into one fails at `cd` before the CLI starts -- a confusing way to meet a feature whose promise is "carry on where you left off" -- so the directory is checked, and a missing one is dropped with a log line naming it rather than being passed on to fail. Verified against this very session: 905 events replayed from the tail (351 tool calls, 350 results, 185 assistant messages, 19 mine), the resume token pointing at its id, and the stale directory reported and dropped. The list was read on the emulator, where the top row is that session under its opening sentence. |