A command, a tool's output and a code block in a reply are the one thing on
this screen that is not somebody's prose, and they now say so: Mocha's
Crust, which sits below Base, so the same colour is one clear step down both
on the page where a reply is drawn and on the card where a tool call is.
The renderer's code background was `surfaceVariant`, which is exactly a
card's own fill -- a fenced block inside a tool call had no background at
all, and one in a reply read as a step *up* out of the page.
Tool output takes the monospace face with it. It is column-aligned far more
often than it is prose -- a listing, a diff, a table of numbers -- and a
proportional font silently destroys the alignment that carried the meaning.
`RawBlock` is a composable rather than a modifier because the inset is part
of it: monospace text against the edge of a tinted block reads as clipping.
A call with neither a subject nor any other field draws nothing at all
rather than an empty tinted rectangle.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Five things Iris asked for, all about the transcript screen holding still
around whoever is reading it.
A tool call opened on its own stayed open when a second call in the same
run turns it into a group. Watching a Bash call and having the session
make another one used to shut the card being read and fold it behind
"Called 2 tools" -- the reader lost their place because something else
happened. The transition is noticed once, at the moment a run first
becomes a group; after that the group's own toggle owns it, so shutting a
group whose inner call is still expanded does not re-open it.
The compaction clock is taken from the `compacting` status event's own
timestamp rather than from this device noticing one, so it survives
leaving the session and coming back -- it used to disappear, because the
only thing that knew when the compaction started was a screen that had
been disposed. The server timestamps every transcript line, so this is
still a measurement; it is compared against the phone's wall clock, which
is the same comparison a session's "last active" already makes.
Session settings are a dialog over the session instead of a screen below
it. Two controls did not warrant a page transition and a back stack, and
the thing they change was hidden while they were on screen. Captions are
gone -- each control is a labelled noun -- and "Notify me" is
"Notifications" with a bell beside it (`md-bell`, added to the committed
Nerd Fonts subset). Failures keep their words, since those are what a
reader cannot work out by looking.
Tool groups are rounded like every other card, their foot bar is the same
height as their heading (both derived from the heading's own line height,
so the pair cannot drift), and the calls inside are a connected stack:
square where they face a neighbour, rounded on the outside, with a small
gap so the join reads as a join.
Scroll position is persistent on the device, per session, keyed by the
row rather than by an index -- an index means nothing across a reopen,
where the transcript is fetched newest-first. Reopening pages backwards
until that row is loaded *and* has something older behind it, because the
oldest loaded row is a half-row that grows when the page behind it
arrives; anchoring into one landed a screen and a half out. The list
draws nothing until the position lands, so there is no frame in which the
transcript is somewhere other than where it was left.
Two things found on the way. `snapshotFlow`'s first emission is the state
before anybody has touched the list, and reading it as a scroll that had
just ended at the newest end wiped every saved position on the way in.
And backwards pages now ask for 800 events rather than 80: ai-app-2
measured a real transcript at 2,426 events for seven assistant messages,
so a page of eighty is a fifth of one row and filling the lookahead took
about thirty sequential round trips -- seconds of a list that will not
move, over the tunnel.
`/tools [n] [gap]` in the echo driver takes seconds between calls, which
is what makes a run grow slowly enough for somebody to have opened one of
its calls first.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Everything that opens now behaves alike. Touch a row's upper half and its
top edge holds, so it opens and closes downwards; touch the lower half and
the bottom edge holds, which is what the list does on its own. A group's
heading and the bar at its foot fall in the halves they already occupy, so
they keep the behaviour they had, and a single tool call -- one card, with
no bar -- gets the same choice for the first time: tapping low on an open
Bash card now shuts it downwards exactly as a group's bar does.
That makes the position of the tap the one mechanism, and RowEdge goes
away with the pair of hardcoded ends it existed to name. Controls report
where they were touched in root coordinates, which is all a control can
know -- a group is one row with a control at each end and calls in the
middle, and only the row knows where its own ends are -- and the row turns
that into an edge.
`clickableAt` is built on `clickable` rather than replacing it, so the
ripple and the click action assistive technology reads are unchanged; the
down position is observed on the initial pointer pass and nothing is
consumed.
Verified with ui-trace: on a collapsed group, a tap at y=1370 holds the
heading and one at y=1450 lets the row grow upward instead. On the same
nested call inside an open group, opening it from the group's upper half
holds the heading at 565 and from the lower half moves it to 296. ktfmt,
lint and 85 tests clean.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Reinstates the reverted downward-opening rows with the two defects that
made the first attempt worse than what it replaced.
The correction waited on the row's top edge and could wait up to half a
second for it to move. A top edge also moves when the reader scrolls, so a
correction still pending would wake on their drag, read the scroll distance
as the row's growth, and undo it -- the transcript jumping on every expand
and refusing to scroll back at all. It now waits on the row's *size*, which
nothing but a resize changes.
The second is why collapsing a group taller than the screen did nothing at
all. Such a group is the list's own anchor item, so as it shrinks it slides
down behind its anchored bottom edge and out of the viewport, and its size
reads as null -- which `withTimeoutOrNull` cannot tell from the null that
means the wait expired. The case most needing the correction was the one
silently skipped. The wait now answers a value that a timeout cannot, and
the distance is read off any row from the pressed one upwards, all of which
move by exactly the row's growth.
Verified on the emulator with ui-trace (~/.local/bin), which samples the
accessibility tree at 60Hz and reports node bounds in device pixels:
expanding and collapsing from a heading hold it to the pixel, collapsing
from the foot bar holds all four rows below it, a drag 120ms after a tap is
left alone, and scrolling back stays put for six seconds. ktfmt, lint and
85 tests clean.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This reverts commit f4d4c82. The anchoring it added made the transcript
jump on every expand and collapse, and left the list snapping back to the
bottom when somebody scrolled up, which is worse than the upward-opening
it was meant to fix.
Two things to look at when this is retried. `LazyListItemInfo.offset` in a
`reverseLayout` list is not obviously the coordinate space this assumed,
so `offset + size` may have been measuring the bottom edge -- the one the
list already holds -- rather than the top. And the anchoring scroll ran in
a coroutine that could still be pending when the reader started dragging;
`scrollBy` takes the default mutation priority, so it cancels that drag.
Verify the next attempt with `uiautomator dump` -- node bounds in device
pixels, before and after a toggle -- rather than by eye from screenshots.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Tapping a group's heading used to send that heading up off the top of the
screen and fill the space above it, so the calls appeared on the far side
of the control that produced them. The transcript is laid out from the
bottom, so every row's bottom edge is what the list holds still and all
growth goes upward.
The rule now is that the end the reader pressed is the end that must not
move. A heading anchors the top, so the row opens downwards under it; the
bar at the foot of an open group anchors the bottom, so shutting it from
there leaves what follows the group where it is -- which is what already
happened, but by accident of the layout rather than on purpose, and would
have been lost the moment anything else changed.
Bottom is the list's own behaviour and costs nothing. Top is measured
rather than calculated: only the layout knows how tall an open group is,
so `toggleAnchored` reads where the top edge was, lets the change land,
and scrolls by however far it moved.
Applied to every row that opens, not just groups -- a lone tool call and a
peer message are the same gesture, and one of them opening the other way
would be the odder for it.
An AskUserQuestion arrived in the middle of a run of tool calls and was
folded into the collapsed card with them, so the one row where somebody
was asked something -- and the answer they gave -- was hidden behind
"Called 6 tools" like any other grep.
It now starts a run of its own and ends the one before it, which needs no
change to the grouping: a run of one is drawn as itself. The calls around
it become a group before and a group after, so where the work stopped to
ask is legible from the shape of the transcript without opening anything.
Echo's `/ask` now runs three ordinary calls on each side of the question,
because that is the shape this has to be looked at in and there was no
way to produce it.
A group of adjacent tool calls was identified by its first call, and the
list is keyed by that identity. But a run can gain members at *either*
end -- a new call arriving beside it, or a page of history arriving in
front of it -- so its first member is not a name, it is a description
that changes. Every time it changed, the row was a different row as far
as the list was concerned: the anchor went with it, and the transcript
stepped under whoever was reading.
Each call now carries the run it belongs to, decided once when it is
folded in and never recomputed, and the row keys on that. A lone call
that gains a neighbour becomes a group *without* changing identity,
which the old key got wrong in the other direction too -- one row was
replaced by another rather than updated.
`joinPages` hands the arriving older calls the name of the run they are
joining, rather than renaming that run after them. The obvious way round
is the wrong one: the newer half is the part already on screen, so
naming the joined run after the older half renames the row the reader is
looking at, which is the whole failure this is meant to remove.
Checked against the same twelve-`/tools 8` rig, whose page boundary falls
inside the second group: every group still reads eight, so the grouping
is unchanged -- what changed is that none of their identities move.
Toward the standing rule for this screen, which is that it may only move
when the reader is at the newest end and something new arrives.
A question is now fully described by the event that reports it: the tag
it was asked under, each option's label, what it means, and the sample of
what picking it would produce, plus whether several may be picked at
once. The app renders from that alone.
It had been reading Claude Code's tool input to find the parts the event
dropped -- that dialect's schema, written out a second time in Kotlin,
where no other provider could reach it and where it would drift the
first time the schema moved. Echo could not describe an option at all,
and llama never will.
Answers travel as a list for the same reason. A question that takes one
answer sends a list of one rather than being a different shape, and the
one place that flattens it is where the CLI is spoken to: its answers
map holds a string, so several choices are joined there. That join was
in the phone.
Also here because it is the same rule: the permission ask reuses the
question body rather than owning a second one, so Allow/Deny renders and
resolves through exactly the code an AskUserQuestion does.
Verified against both, since a refactor that only satisfies the case it
was written for has been tried on the half that cannot fail: a two
question `/ask` answered from the phone, one option and then two, and a
real sonnet session's `rm -f` permission asked, allowed, and run.
Reported by Iris through the dev-updater session: a two-question
AskUserQuestion arrived with only one option visible per question, so
the answer she sent was the only one she had been offered.
The cause was a `Row`. It hands out intrinsic widths in order and clips
whatever runs past the edge, so the first option or two drew and the
rest went off the side of the screen -- which does not read as a bug, it
reads as those having been the only choices. The same Row was in the
permission ask beside it; both wrap now. That pairing is the reason to
look: a rule stated on one member of a set is usually missing from the
others.
The rest of what she asked for, and what each was:
- It drew twice, as the tool call and again as loose question cards,
because the backend marked these questions as belonging to no call.
They belong to the call that asked, and now say so.
- So it renders like any other tool: one card, its own heading, opened
because a decision cannot be made from a closed row.
- Each option shows its description and its `preview` block, which is
the part a reader is deciding on and none of which was reaching them.
- "Other" is a field on every question. The harness always offers it, so
leaving it out narrowed a question that was never that narrow.
- A multi-select sends the labels it collected as one string, which is
the tool's own schema rather than a guess -- its answers map is
string-valued.
- No spinner while it waits. A spinner says the machine is working; here
the machine is idle and the turn is stopped on the person, so the card
says "your turn" in the colour this app already uses for that.
Verified against a real session as well as the echo fixture: haiku asked
two questions with three described options each, both were answered from
the phone, and the model carried on with the answers. Echo grew `/ask`
so the shape can be looked at without paying a model to produce one, and
its option cards are outlined rather than tinted -- as one surface step
up they were three paragraphs where three things to press should be.
Scrolled back through a conversation, every new message dragged the view
with it -- which reads as the screen scrolling down on its own, at the
exact moment somebody is trying to read something else.
The list had no keys, so its rows were identified by position. The
transcript is drawn newest-first, so a new message is an insertion at
index 0: every existing row shifts up one index, the viewport stays on
the index it was on, and the content slides through it. The working
indicator appearing and disappearing did the same thing at the same end.
So rows now carry the identity they always had in the data. Every
transcript event has a seq, which is what the transcript is ordered by
and never changes, and every row keeps the seq of the first event behind
it -- a streaming message keeps the seq of its first delta, so it holds
still for the whole answer rather than becoming a new row on every
frame, and a tool call keeps its start's.
Paging older history is the same insertion from the other end, and it is
the thing this could plausibly have broken. Checked on the emulator:
scrolled back mid-turn, the view sat still through twenty seconds of
streamed deltas, and scrolling to the far end still fetched earlier
pages and stayed where it was while they arrived.
**The queue was holding messages the CLI would have taken.** Two claims
in this file contradicted each other: the module header said a mid-turn
message is injected at the next tool boundary -- "the behavior this app
exists for" -- and `Queue`'s own doc said a line written mid-turn simply
becomes the next turn. The code followed the second, parking every
message until `Status::Idle`, which is the end of the whole turn.
Measured rather than argued, twice. Writing a line straight into a live
session's stdin fifo mid-turn produced one `result` for the whole thing,
so it was consumed inside that turn, not as a new one. The header was
right and the queue was built on the wrong claim.
The cost was exactly what Bryan reported: he steered after the second
tool call and it sat unread until every remaining call had finished.
Measured before and after on the same three-step turn -- steer sent at
+13s, recorded at +24.7s before this change and at +14.1s after, which
is the next tool boundary.
So the line goes out immediately. What stays behind is the
*announcement*: the CLI says nothing on stdout about having read a
message, so `MessageTaken` now waits for the next assistant text or tool
call, which is proof another model call happened and the steer was in
it. That keeps a held message drawn below the working indicator until
the session has actually taken it -- the thing that mattered when this
was last changed -- without delaying the message to get it. Idle counts
too, and is the case that must not be missed: a message written after a
turn's last model call has no later output to prove anything.
`closed` is untouched, and `Queue::close` still reports held messages by
name rather than dropping them.
**An image now names the call that produced it.** `Event::Image` gains
`about`, the `tool_use_id` from the tool result it came out of, so a
screenshot is drawn inside that call's card instead of floating beside
it -- pairing them by position is what a page boundary breaks. `None`
for a person's own attachment, which belongs to no call. The import path
threads it through as well, so replayed history reads the same as live.
Images show whether the card is open or closed: a call whose result *is*
a picture says less closed than the one line it replaced.
Verified on the emulator against a real haiku turn: the checkerboard sits
inside `Read /tmp/tiny.png`, and the steer sits between that call and the
next, where it was taken. 53 tests, clippy, rustfmt, lint and ktfmt clean.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Committed unformatted again: I piped the check through grep, so the
task's failure never reached the shell's exit status and the chained
commit ran anyway. Check exit codes, not output.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
It was a labelled button beside a tool group that collapses with a
drawn chevron -- two controls doing the same kind of thing in two
visual languages. Now one `Chevron` composable serves both directions,
parameterised rather than copied, since a pair that differs by a minus
sign drifts and the drift is a bug in exactly one direction.
The comment it replaces argued against an arrow here, on the grounds
that the list is laid out upside down. That reasoning was about the
code: nobody reading the screen knows the list is reversed, and on
screen the newest message is at the bottom, which is where this goes.
It draws no text, so the name lives in its content description -- the
whole of what a screen reader has, and the answer to "what was that
arrow for" later.
Verified on the emulator: scrolled up, the chevron appears bottom
centre matching the group's; tapped, it returns to the newest message
and takes itself away.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Five changes to how a transcript reads.
**Images no longer move the page.** The row was as tall as whatever had
loaded, so it grew when the bytes arrived and pushed everything below it
-- and in a bottom-anchored list, an image loading above the viewport
moved the text under the reader's eyes. The height is now decided before
the fetch and never changes: four lines of the body style, measured from
the type so it stays four lines when the reader has scaled their fonts.
Nothing to see when loading finishes, which is the point.
**A small image is enlarged with nearest neighbour**, a large one shrunk
smoothly -- decided per image from its actual size rather than set once,
since blowing a 16px sprite up with interpolation turns it into a blur
of exactly the thing being looked at.
**Tapping one opens it full screen**, fitted so the whole image is
visible first, with two-finger zoom to 8x and pan once zoomed. A dialog
rather than a screen, so back returns to the transcript.
**A tool call is one line closed**: the tool's name and what the call is
for. The command is not on it, because a wrapped command turns one row
into four. Open, it shows the command, the rest of the input and the
output, with the timeout at the top right -- a limit on the call rather
than part of what it does, worth seeing beside the command it constrains.
A call waiting on permission is shown open regardless, since the command
is the thing being decided.
**Adjacent calls fold into "Called n tools"**, closed by default, and it
closes again from either end -- a long group's heading scrolls away while
its last call is still on screen, and the reader who wants it shut is
looking at the bottom. The calls keep their full width; what says they
belong together is the surface behind them, one cue rather than two
half-cues. Grouping happens at display time, not in the fold: the
transcript's own order is what paging and the stream depend on.
Echo gains `/tools [n]` so a run of calls can be produced without paying
for one.
Verified on the emulator: four calls folded and expanded, one opened
inside the group showing `timeout 5000` top right, a 16px checkerboard
enlarged with hard pixel edges beside a shrunk screenshot at the same
height, the screen byte-identical between one second and six after
opening, full screen fitted, and back returning to the same scroll
position. Pinch itself is the one thing not verified here -- `adb input`
cannot inject a two-finger gesture.
53 tests, clippy, rustfmt, Android lint and ktfmt clean.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>