Two things the measurements for the previous commit turned up. Highlighting a fence cost 174ms for a two-hundred-line Kotlin block, and the lazy list charged it again every time that block scrolled back into composition -- six times in one bench run, with the scroll's draw phase at 1.29ms per frame. So it is warmed and cached where parses already are: `highlight` is a plain function taking no colour from the theme, `warm` fills `ParsedReplies.highlighted` from `fences(parse)` off the drawing thread, and `fenceContent` extracts the code here rather than through the library's composable, so the string warmed is the string drawn. And the anchor restore asked for a span counted in events, which goes negative when the anchor's row is the oldest half-row and was coerced to one -- a request per delta, six hundred round trips walking one reply back a word at a time with the spinner up. It asks for a page of rows now. Clean pairs, fresh sessions each side, same gestures. Fence scroll (transcript-bench.sh): draw phase 0.74ms per frame before highlighting existed, 0.77ms after, no lexing in the window either side. Streaming forty linked items (stream-bench.sh): 2412ms of reparsing before, 674ms after; mean 5.0ms to 1.4ms, worst 8.9ms to 7.9ms. Bullet glyphs, fence colours, the image links and the reference link checked on the emulator; lint clean on AGP 9.4.0. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
315 lines
18 KiB
Markdown
315 lines
18 KiB
Markdown
# Transcript rendering: what was learned, and what is next
|
|
|
|
Written 2026-09-03 at the end of a week of work on the session screen's
|
|
transcript, so the next session can start from here rather than from a
|
|
compacted context. `AGENTS.md` holds the one-paragraph conventions; this is
|
|
the longer record: the measurements that drove each decision, the
|
|
techniques that worked, the ones that did not, and the order to do the rest
|
|
in. `PLAN.md` remains the design source of truth; nothing here contradicts
|
|
it.
|
|
|
|
## The goal, and where it stands
|
|
|
|
A reply of any length must scroll at the phone's 120Hz without a bump, and
|
|
must keep doing so while the reply is still streaming in. Measured on the
|
|
Pixel 9 Pro XL by Bryan, the transcript went from visible stalls at long
|
|
replies and at lists of links to "I have to actually try to feel any
|
|
bumps". The remaining work is finish and extensibility rather than
|
|
performance.
|
|
|
|
## The architecture, as built
|
|
|
|
Everything below lives under `app/androidApp/src/main/kotlin/com/example/aiapp/`.
|
|
|
|
**Rows become units, and units are bounded.** `TranscriptUnits.kt` turns a
|
|
transcript row into the things the lazy list actually holds. An assistant
|
|
reply is not one unit: it is one unit per piece of its markdown, so the
|
|
list composes and draws a paragraph, a fence, a table or one bullet at a
|
|
time. The reason is the draw phase: a row's display list holds every glyph
|
|
of it and is re-recorded whenever drawing is invalidated, and the lazy list
|
|
composes an item whole in the frame it scrolls into. The tallest single
|
|
row still being drawn before this was 36,982px, twenty-five screens in one
|
|
message. Long user messages are sliced the same way (`UserChunk`), through
|
|
the shared `cardPiece` modifier that draws one card in lazy-list pieces.
|
|
|
|
**One parse per message, addressed by piece.** `MarkdownPieces.kt`'s
|
|
`Piece(block, item)` is an address into the message's single parse tree,
|
|
not a substring: `block` indexes the root's children and `item` one
|
|
`LIST_ITEM` of a top-level list. Cutting was originally done by
|
|
re-parsing substrings, which cost a parse per piece and broke reference
|
|
links defined at the foot of a message. `ParsedReplies` caches the parse
|
|
and the piece list per text (`of`, `piecesOf`), warmed off the composing
|
|
thread by `TranscriptItems.warm`. The parser is still intellij-markdown via
|
|
the mikepenz renderer; the renderer's `Markdown()` composable is kept only
|
|
as the provider of its `Local*` environment (`MarkdownRoot` in
|
|
`Markdown.kt`), and `MarkdownElement` dispatches a whole block through our
|
|
component table.
|
|
|
|
**Lists are drawn an item at a time, by us.** The renderer has no element
|
|
for a single list item, so `MarkdownListItem` draws one: marker, then the
|
|
item's children, nested lists recursing through `MarkdownList`. The
|
|
marker is drawn in one place on purpose; styled bullets per depth go
|
|
there.
|
|
|
|
**Links are spans, not nodes.** `MarkdownLinks.kt`. Compose turns every
|
|
`LinkAnnotation` into a layout node (clipped, focusable, hoverable,
|
|
clickable, outline recomputed from the text layout). A paragraph of eight
|
|
links was nine nodes, and measured against the same paragraphs with each
|
|
link replaced by plain words it cost 26.3ms worst measure against 5.2ms,
|
|
1.7x the place time. That was the bump at a reply's list of sources.
|
|
`LinkedText` builds the annotated string with the renderer's own inline
|
|
builder but answers links itself: colour, underline, a string annotation
|
|
carrying the URL, and one tap detector for the whole text that asks the
|
|
layout which glyph is under the finger. Hit-testing must check the glyph
|
|
on either side of the returned caret, because `getOffsetForPosition`
|
|
returns the nearest boundary; taps on the right half of a glyph otherwise
|
|
open nothing. Headings need the `ATX_CONTENT`/`SETEXT_CONTENT` child, since
|
|
the inline builder draws nothing for a node type it does not know (a week
|
|
of blank headings). Tables go through `LinkedTable`/`LinkedTableRow` so
|
|
cells get the same treatment.
|
|
|
|
**Text draws on the platform directly.** A paragraph without an image
|
|
skips the renderer's `MarkdownText`, which charges every paragraph for the
|
|
possibility of inline images (placement callback, derived inline-content
|
|
map, semantics group, size animation). Paragraphs that contain an image
|
|
still take the renderer's path.
|
|
|
|
**Tables spread or scroll without subcomposition.** The renderer used
|
|
`BoxWithConstraints` to decide; `LinkedTable` uses
|
|
`fillMaxWidth().horizontalScroll().layout { }` -- `horizontalScroll`
|
|
passes `minWidth` through and lifts `maxWidth` to infinity, so the inner
|
|
layout reads `minWidth` as the room available and takes
|
|
`max(minWidth, columns * cellWidth)`.
|
|
|
|
**A streaming reply is reparsed one block at a time.** `LiveParse` in
|
|
`Markdown.kt` freezes every finished top-level block with its parse and
|
|
reparses only the tail block per delta. Markdown's block rules make later
|
|
text unable to alter an earlier block, with the single exception of a
|
|
late reference definition, which is accepted. Measured on a 58-word stream
|
|
of list, fence, table and quote: 47 tail reparses at 1.7ms mean. A
|
|
single-list stream still reparses per delta because the whole list is one
|
|
tail block.
|
|
|
|
**Expansion anchors the edge that was tapped, and the list never moves
|
|
under the reader** except when pinned to the bottom with new content
|
|
arriving. Those two rules are in `ScrollAnchor.kt` and `TranscriptList.kt`
|
|
and are the reason several tempting simplifications were rejected.
|
|
|
|
## Techniques and harness
|
|
|
|
- **`app/ui-sandbox.sh`** starts a second `ai-server` against a sandbox
|
|
home with the echo driver, so nothing touches real sessions.
|
|
`spawn [title]` makes an echo session and prints its id; `send SID text`
|
|
or `send SID @file` sends into it; `api /path [curl args]` is an
|
|
authenticated request. Restarting it regenerates the config but keeps
|
|
enrolled tokens.
|
|
- **The echo driver is the test rig** (`server/src/session/echo.rs`, the
|
|
list at the top of the file). `/stream N`, `/mixed N`, `/table N`,
|
|
`/tools N gap`, `/ask`, `/peer`, `/compact`, `/slow`, `/bash command`
|
|
each produce a shape the real CLI produces only when it feels like it.
|
|
Build what a UI test needs into it rather than spending model turns.
|
|
- **`app/transcript-bench.sh`** is the standard measurement: restart, open
|
|
the first session, scroll, print the render report. The report is what
|
|
the "Copy render timings" button copies and also logs
|
|
(`adb logcat -d -s ai-app:I`), and it includes the last crash's stack
|
|
(`CrashLog.kt`), which is how a crash on the phone reaches a session
|
|
here.
|
|
- **`DebugStats`/`FrameStats`** time our own phases (`record: one block`,
|
|
`measure: the app root`) and count events (`markdown reparsed while
|
|
streaming`, `markdown cut into pieces`). Add a counter before guessing.
|
|
- **`app/trace-draw.sh`** names what a scrolling frame spends inside the
|
|
framework, via `atrace` text output, no trace processor needed. It is
|
|
how the link-node cost was attributed.
|
|
- **`app/debug-transcript.sh`** loads a real Claude Code conversation onto
|
|
the emulator; two faults were invisible on fixtures and obvious on it.
|
|
Real transcripts are private: fixtures stay in `/tmp`, never in the repo.
|
|
- **`ui-trace`** reads the screen as text. Bounds print as
|
|
`x1,y1..x2,y2`; unanchored `-m` patterns match labels, anchored ones do
|
|
not. A row taller than the viewport reports clipped bounds, so compare
|
|
screenshots for that case.
|
|
- **Emulator frame times are not app measurements.** Software rendering
|
|
puts the stock Settings app at 60ms of UI-thread traversal per frame.
|
|
Costs of operations in milliseconds rank correctly; smoothness itself is
|
|
judged on the phone.
|
|
- **System Tracing on the phone does not work on GrapheneOS.** Its
|
|
Categories list is empty because the tracing daemon builds it by running
|
|
`atrace --list_categories`, which returns nothing there, and a recorded
|
|
trace contains zero ftrace events: no app sections, no frames, no
|
|
scheduling. Callstack sampling records, but the app's profiler config
|
|
unwinds one process shard in four. GrapheneOS issues 2206 and 6094 are
|
|
open on exactly this. Until they close, phone numbers come from the
|
|
render report and from Bryan noticing.
|
|
- **Compose `DropdownMenu` in an edge-to-edge activity** needs
|
|
`PopupProperties(clippingEnabled = false)` or it opens a status bar's
|
|
height away from its anchor (`~/.claude/TOOLCHAIN.md`).
|
|
- **highlights 1.1.0's shell lexer** returns a span whose end precedes its
|
|
start for `x '*/a/*'`; `ToolInput.kt` drops such spans. The fix belongs
|
|
upstream and has not been filed.
|
|
|
|
## Rejected, and why
|
|
|
|
- **Writing our own markdown renderer.** Rejected in favour of keeping
|
|
the intellij-markdown parser and the library's inline builder while
|
|
owning block dispatch and the leaf composables. The parser is the hard
|
|
part and is not the slow part; everything that was slow lived in the
|
|
composables, which are now ours.
|
|
- **Re-parsing substrings per piece.** Cost a parse per piece and broke
|
|
foot-of-message reference links. Replaced by addressed pieces of one
|
|
parse.
|
|
- **Animated or timing-dependent corrections.** Anything the reader could
|
|
catch at 120Hz is a bug; corrections must be structurally impossible to
|
|
see.
|
|
|
|
## Session of 2026-09-03: the list above, worked through
|
|
|
|
Items 1-5 of the previous list are done and checked on the emulator, and so
|
|
is the AGP bump from item 7. Two things were found while measuring them: a
|
|
174ms stall this work introduced and then removed, and a restore bug that
|
|
predates it. Both are described below with their numbers.
|
|
|
|
**1. Styled markers per depth.** `MarkdownListItem`'s `Marker` draws `•`,
|
|
`◦`, `▪` by depth, cycling past the third, in `listMarkerColor` (Theme.kt,
|
|
Lavender -- the scheme's secondary accent, which nothing else used, so it
|
|
now means "structure"). Ordered numbers take the same colour. All three
|
|
glyphs were checked on screen at four depths: they render from the system
|
|
fonts, no missing-glyph boxes. The colour is the same at every depth on
|
|
purpose -- depth is said by the glyph and the indent, and a colour per depth
|
|
would make a difference in degree look like one in kind.
|
|
|
|
**2. Syntax highlighting inside fences.** `CodeFence.kt` holds it:
|
|
`highlight` (moved out of `ToolInput.kt`, so a reply's code and a tool
|
|
call's command are the same colours), the `fenceLanguage` alias table
|
|
(extension or name to the highlights lexer; a word not in it stays plain,
|
|
because a fence coloured by the wrong language's rules looks highlighted and
|
|
is wrong in a way the reader cannot see), `fenceContent`, and the
|
|
`codeFence`/`codeBlock` entries of the component table.
|
|
|
|
The measurement is the reason there is a cache. Highlighted naively, with
|
|
the answer held by a `remember` inside the fence, a two-hundred-line Kotlin
|
|
fence cost **174ms to lex** and the lazy list charged it again every time
|
|
the block scrolled back into composition -- six times in one bench run,
|
|
1043ms of lexing, and the scroll's draw phase up at 1.29ms per frame. So
|
|
highlighting is now warmed and cached exactly as parsing is
|
|
(`ParsedReplies.highlighted`, filled by `warm` from `fences(parse)`), and
|
|
`highlight` is a plain function taking no colour from the theme, which is
|
|
what lets it run off the drawing thread.
|
|
|
|
Because the warming has to ask for the same string the drawing does,
|
|
`fenceContent` extracts the code and the language word itself -- the rule
|
|
copied from the library's `MarkdownCodeFence`, which is a composable and so
|
|
cannot be called from `warm`. Two extractions would be two keys, and the
|
|
warmed answer would be missed at every fence with nothing saying so.
|
|
|
|
Clean pair, two fresh sessions of the same 200-line Kotlin fence, no saved
|
|
anchor, same gestures (`transcript-bench.sh`):
|
|
|
|
| | before (no highlighting) | after (warmed) |
|
|
|---|---|---|
|
|
| draw phase per frame | 0.74ms | 0.77ms |
|
|
| the transcript's share | 0.33ms | 0.33ms |
|
|
| lexing during the scroll | none | none |
|
|
|
|
**3. `MarkdownRoot` no longer calls the library's `Markdown()`.** It
|
|
provides the locals itself -- reference links from the parse, padding,
|
|
dimens, colours, typography, a no-op image transformer, animations,
|
|
components. Nothing between a piece and the screen is the library's now
|
|
except the leaf composables named in the component table.
|
|
|
|
**4. Paragraphs with images -- done differently from the plan.** The plan
|
|
said draw the image as its own piece; measuring first showed the app has no
|
|
image loader and the renderer's transformer was the no-op one, so an image
|
|
in a reply drew as *nothing at all*. An `IMAGE` node is now appended by
|
|
`appendPlainLink` as a link carrying its alt text (the address when there is
|
|
none), which says what was there and opens it. With no image to place, the
|
|
`hasImage` branch and the renderer's `MarkdownText` are gone: every
|
|
paragraph is the platform `BasicText`.
|
|
|
|
**5. Per-item units for a streaming list.** `LiveParse.advanceTo` now cuts
|
|
at the last item of a multi-item list (`openPiece`), provided that item has
|
|
content beyond its marker -- a bare `-` is an empty item now and the first
|
|
character of a paragraph line once `-x` arrives, so cutting on it would draw
|
|
that line as a new item. The cut is at the start of the item's line, so the
|
|
indentation the reparse reads its nesting from survives.
|
|
`Segment.continues` marks a tail that carries on a list, and
|
|
`MarkdownPiece`'s `continuesList`/`listContinues` keep an inner item's
|
|
padding at the seam, so nothing moves when the seam does.
|
|
|
|
Clean pair, two fresh sessions, forty linked bullet items streamed word at a
|
|
time (`stream-bench.sh /tmp/longlist.md`):
|
|
|
|
| | before | after |
|
|
|---|---|---|
|
|
| reparses while streaming | 482 | 483 |
|
|
| total time in them | 2412ms | 674ms |
|
|
| mean / worst | 5.0ms / 8.9ms | 1.4ms / 7.9ms |
|
|
| `record: one block` worst | 1.8ms | 0.7ms |
|
|
|
|
The worst case moves least, which is the shape to expect: the first reparse
|
|
of a tail still covers whatever has arrived, and the last item can be long.
|
|
What changes is that every reparse after it covers one item instead of the
|
|
whole list.
|
|
|
|
**The restore walked back one event per request.** Found while benching, and
|
|
older than this work. `savedAnchor`'s loop asked for
|
|
`oldestSeq - anchor.seq + RESTORE_PAGE_CUSHION` events; when the anchor's row
|
|
is already loaded but is the oldest half-row (which `anchorRow` refuses,
|
|
correctly -- it grows when the page behind it lands), that span is negative
|
|
and was coerced to 1. So the restore fetched one event, then one more, at a
|
|
round trip each: six hundred requests walking a long reply back a word at a
|
|
time, with the screen on its spinner the whole way and the sandbox log
|
|
printing `limit=1` once a second. It now asks for a page counted in rows,
|
|
which is the only kind that can promise to reach the row behind the anchor.
|
|
|
|
**Harness.** `app/stream-bench.sh [-k] FILE` is `transcript-bench.sh` for a
|
|
reply still arriving: opens the first session, taps the app's own "Jump to
|
|
latest" so the list is pinned to the newest end, resets the report, sends
|
|
FILE, waits for the transcript to stop growing, prints the report. Both of
|
|
those last two are corrections to a first version that measured nothing:
|
|
a transcript parked further back never redraws while a reply streams into it
|
|
(the list must not move under a reader), and a session is idle at *both*
|
|
ends of a turn, so polling for idle answers before the turn has started.
|
|
Fixtures live in `/tmp` and are regenerated from the shapes named here:
|
|
`fixture.md` (lists four deep, ordered and nested, fences in
|
|
kotlin/rust/sh/none, a table with a link, a quote with a list, an inline and
|
|
a standalone image, a reference link), `longfence.md` (200-line Kotlin
|
|
fence), `longlist.md` (40 linked items).
|
|
|
|
**Two traps in the emulator loop**, both of which cost a bench run here.
|
|
`adb shell pm clear` removes the enrolment and the notification permission
|
|
along with the saved anchors, so the next run measures a permission dialog;
|
|
re-enrol with the command `ui-sandbox.sh` prints and
|
|
`pm grant … POST_NOTIFICATIONS`. And a saved anchor is per session id, so
|
|
the only way two builds start a scroll from the same place is a *fresh
|
|
session for each*.
|
|
|
|
## What is next, in order
|
|
|
|
1. **A fence still arriving is lexed per delta.** `warm` covers settled
|
|
messages; the live tail's fence is highlighted on the composing thread by
|
|
`ParsedReplies.highlighted`'s inline miss, once per delta, and at 174ms
|
|
for a long one that is the stall above wearing a different hat. The tail
|
|
is short while it is being written, so this may already be cheap -- but
|
|
it is unmeasured, and `stream-bench.sh /tmp/longfence.md` is the run that
|
|
says. Cheapest fix if it bites: highlight the live tail only when it is
|
|
under some length, or move the miss off-thread the way `LiveParse` moved
|
|
parsing, keeping the plain text until the answer lands.
|
|
2. **The reconnect loop.** Restarting the app onto a session with a saved
|
|
anchor while a long reply was streaming left it reconnecting every 1.5s
|
|
(`RECONNECT_DELAY_MS`), spinner up, until the server was restarted.
|
|
`events?after=N` more than `CATCH_UP_LIMIT` (200) behind answers `reset`
|
|
plus the newest 200 *raw* deltas -- a window starting mid-message -- and
|
|
the reset clears `items`, which is the state the restore loop then pages
|
|
against. The one-event-per-request bug above was part of what made it so
|
|
visible; whether it survives that fix is the first thing to find out.
|
|
3. **File the highlights range bug upstream.** There is no `gh` and no
|
|
GitHub credential in this VM, so it needs Bryan or a token. One-line
|
|
repro: lexing `x '*/a/*'` as `SyntaxLanguage.SHELL` in highlights 1.1.0
|
|
returns a highlight whose `location.end` precedes its `location.start`.
|
|
Until then `highlight` drops such spans, which is why the shell fence in
|
|
`fixture.md` -- whose command contains `'*/.git/*'` -- draws plain while
|
|
an ordinary shell fence colours.
|
|
4. **Regression runs.** `transcript-bench.sh` and `stream-bench.sh` before
|
|
and after any change to the files above, with the report in the commit.
|
|
The numbers to watch are the worst `record: one block`, the reparse mean
|
|
while streaming, and the draw phase's accounting line.
|