Files
ai-app/TRANSCRIPT_RENDERING.md
T
irisandClaude Opus 5 b0b91ad28f Draw a fence plain while it is still being written
Warming covers a settled message, so it did not reach the block a reply is
still writing: that one was re-lexed at every delta, on the composing
thread. Measured streaming a two-hundred-line Kotlin fence -- 211 lexes,
13.7 seconds across the turn, the worst 177ms -- for colours on text being
replaced as fast as they were computed.

A live reply's last segment now draws its code plain and takes its colours
when the block freezes, which for a finished fence is as soon as the next
block starts. The settle-time lex then happens in warm, off the drawing
thread. Same fixture after: one lex of 374ms in warm, draw phase 1.51ms to
1.01ms per frame while streaming, and the fence coloured on screen once the
reply settles.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-03 18:58:55 -04:00

329 lines
18 KiB
Markdown

# Transcript rendering: what was learned, and what is next
Written 2026-09-03 at the end of a week of work on the session screen's
transcript, so the next session can start from here rather than from a
compacted context. `AGENTS.md` holds the one-paragraph conventions; this is
the longer record: the measurements that drove each decision, the
techniques that worked, the ones that did not, and the order to do the rest
in. `PLAN.md` remains the design source of truth; nothing here contradicts
it.
## The goal, and where it stands
A reply of any length must scroll at the phone's 120Hz without a bump, and
must keep doing so while the reply is still streaming in. Measured on the
Pixel 9 Pro XL by Bryan, the transcript went from visible stalls at long
replies and at lists of links to "I have to actually try to feel any
bumps". The remaining work is finish and extensibility rather than
performance.
## The architecture, as built
Everything below lives under `app/androidApp/src/main/kotlin/com/example/aiapp/`.
**Rows become units, and units are bounded.** `TranscriptUnits.kt` turns a
transcript row into the things the lazy list actually holds. An assistant
reply is not one unit: it is one unit per piece of its markdown, so the
list composes and draws a paragraph, a fence, a table or one bullet at a
time. The reason is the draw phase: a row's display list holds every glyph
of it and is re-recorded whenever drawing is invalidated, and the lazy list
composes an item whole in the frame it scrolls into. The tallest single
row still being drawn before this was 36,982px, twenty-five screens in one
message. Long user messages are sliced the same way (`UserChunk`), through
the shared `cardPiece` modifier that draws one card in lazy-list pieces.
**One parse per message, addressed by piece.** `MarkdownPieces.kt`'s
`Piece(block, item)` is an address into the message's single parse tree,
not a substring: `block` indexes the root's children and `item` one
`LIST_ITEM` of a top-level list. Cutting was originally done by
re-parsing substrings, which cost a parse per piece and broke reference
links defined at the foot of a message. `ParsedReplies` caches the parse
and the piece list per text (`of`, `piecesOf`), warmed off the composing
thread by `TranscriptItems.warm`. The parser is still intellij-markdown via
the mikepenz renderer; the renderer's `Markdown()` composable is kept only
as the provider of its `Local*` environment (`MarkdownRoot` in
`Markdown.kt`), and `MarkdownElement` dispatches a whole block through our
component table.
**Lists are drawn an item at a time, by us.** The renderer has no element
for a single list item, so `MarkdownListItem` draws one: marker, then the
item's children, nested lists recursing through `MarkdownList`. The
marker is drawn in one place on purpose; styled bullets per depth go
there.
**Links are spans, not nodes.** `MarkdownLinks.kt`. Compose turns every
`LinkAnnotation` into a layout node (clipped, focusable, hoverable,
clickable, outline recomputed from the text layout). A paragraph of eight
links was nine nodes, and measured against the same paragraphs with each
link replaced by plain words it cost 26.3ms worst measure against 5.2ms,
1.7x the place time. That was the bump at a reply's list of sources.
`LinkedText` builds the annotated string with the renderer's own inline
builder but answers links itself: colour, underline, a string annotation
carrying the URL, and one tap detector for the whole text that asks the
layout which glyph is under the finger. Hit-testing must check the glyph
on either side of the returned caret, because `getOffsetForPosition`
returns the nearest boundary; taps on the right half of a glyph otherwise
open nothing. Headings need the `ATX_CONTENT`/`SETEXT_CONTENT` child, since
the inline builder draws nothing for a node type it does not know (a week
of blank headings). Tables go through `LinkedTable`/`LinkedTableRow` so
cells get the same treatment.
**Text draws on the platform directly.** A paragraph without an image
skips the renderer's `MarkdownText`, which charges every paragraph for the
possibility of inline images (placement callback, derived inline-content
map, semantics group, size animation). Paragraphs that contain an image
still take the renderer's path.
**Tables spread or scroll without subcomposition.** The renderer used
`BoxWithConstraints` to decide; `LinkedTable` uses
`fillMaxWidth().horizontalScroll().layout { }` -- `horizontalScroll`
passes `minWidth` through and lifts `maxWidth` to infinity, so the inner
layout reads `minWidth` as the room available and takes
`max(minWidth, columns * cellWidth)`.
**A streaming reply is reparsed one block at a time.** `LiveParse` in
`Markdown.kt` freezes every finished top-level block with its parse and
reparses only the tail block per delta. Markdown's block rules make later
text unable to alter an earlier block, with the single exception of a
late reference definition, which is accepted. Measured on a 58-word stream
of list, fence, table and quote: 47 tail reparses at 1.7ms mean. A
single-list stream still reparses per delta because the whole list is one
tail block.
**Expansion anchors the edge that was tapped, and the list never moves
under the reader** except when pinned to the bottom with new content
arriving. Those two rules are in `ScrollAnchor.kt` and `TranscriptList.kt`
and are the reason several tempting simplifications were rejected.
## Techniques and harness
- **`app/ui-sandbox.sh`** starts a second `ai-server` against a sandbox
home with the echo driver, so nothing touches real sessions.
`spawn [title]` makes an echo session and prints its id; `send SID text`
or `send SID @file` sends into it; `api /path [curl args]` is an
authenticated request. Restarting it regenerates the config but keeps
enrolled tokens.
- **The echo driver is the test rig** (`server/src/session/echo.rs`, the
list at the top of the file). `/stream N`, `/mixed N`, `/table N`,
`/tools N gap`, `/ask`, `/peer`, `/compact`, `/slow`, `/bash command`
each produce a shape the real CLI produces only when it feels like it.
Build what a UI test needs into it rather than spending model turns.
- **`app/transcript-bench.sh`** is the standard measurement: restart, open
the first session, scroll, print the render report. The report is what
the "Copy render timings" button copies and also logs
(`adb logcat -d -s ai-app:I`), and it includes the last crash's stack
(`CrashLog.kt`), which is how a crash on the phone reaches a session
here.
- **`DebugStats`/`FrameStats`** time our own phases (`record: one block`,
`measure: the app root`) and count events (`markdown reparsed while
streaming`, `markdown cut into pieces`). Add a counter before guessing.
- **`app/trace-draw.sh`** names what a scrolling frame spends inside the
framework, via `atrace` text output, no trace processor needed. It is
how the link-node cost was attributed.
- **`app/debug-transcript.sh`** loads a real Claude Code conversation onto
the emulator; two faults were invisible on fixtures and obvious on it.
Real transcripts are private: fixtures stay in `/tmp`, never in the repo.
- **`ui-trace`** reads the screen as text. Bounds print as
`x1,y1..x2,y2`; unanchored `-m` patterns match labels, anchored ones do
not. A row taller than the viewport reports clipped bounds, so compare
screenshots for that case.
- **Emulator frame times are not app measurements.** Software rendering
puts the stock Settings app at 60ms of UI-thread traversal per frame.
Costs of operations in milliseconds rank correctly; smoothness itself is
judged on the phone.
- **System Tracing on the phone does not work on GrapheneOS.** Its
Categories list is empty because the tracing daemon builds it by running
`atrace --list_categories`, which returns nothing there, and a recorded
trace contains zero ftrace events: no app sections, no frames, no
scheduling. Callstack sampling records, but the app's profiler config
unwinds one process shard in four. GrapheneOS issues 2206 and 6094 are
open on exactly this. Until they close, phone numbers come from the
render report and from Bryan noticing.
- **Compose `DropdownMenu` in an edge-to-edge activity** needs
`PopupProperties(clippingEnabled = false)` or it opens a status bar's
height away from its anchor (`~/.claude/TOOLCHAIN.md`).
- **highlights 1.1.0's shell lexer** returns a span whose end precedes its
start for `x '*/a/*'`; `ToolInput.kt` drops such spans. The fix belongs
upstream and has not been filed.
## Rejected, and why
- **Writing our own markdown renderer.** Rejected in favour of keeping
the intellij-markdown parser and the library's inline builder while
owning block dispatch and the leaf composables. The parser is the hard
part and is not the slow part; everything that was slow lived in the
composables, which are now ours.
- **Re-parsing substrings per piece.** Cost a parse per piece and broke
foot-of-message reference links. Replaced by addressed pieces of one
parse.
- **Animated or timing-dependent corrections.** Anything the reader could
catch at 120Hz is a bug; corrections must be structurally impossible to
see.
## Session of 2026-09-03: the list above, worked through
Items 1-5 of the previous list are done and checked on the emulator, and so
is the AGP bump from item 7. Two things were found while measuring them: a
174ms stall this work introduced and then removed, and a restore bug that
predates it. Both are described below with their numbers.
**1. Styled markers per depth.** `MarkdownListItem`'s `Marker` draws `•`,
`◦`, `▪` by depth, cycling past the third, in `listMarkerColor` (Theme.kt,
Lavender -- the scheme's secondary accent, which nothing else used, so it
now means "structure"). Ordered numbers take the same colour. All three
glyphs were checked on screen at four depths: they render from the system
fonts, no missing-glyph boxes. The colour is the same at every depth on
purpose -- depth is said by the glyph and the indent, and a colour per depth
would make a difference in degree look like one in kind.
**2. Syntax highlighting inside fences.** `CodeFence.kt` holds it:
`highlight` (moved out of `ToolInput.kt`, so a reply's code and a tool
call's command are the same colours), the `fenceLanguage` alias table
(extension or name to the highlights lexer; a word not in it stays plain,
because a fence coloured by the wrong language's rules looks highlighted and
is wrong in a way the reader cannot see), `fenceContent`, and the
`codeFence`/`codeBlock` entries of the component table.
The measurement is the reason there is a cache. Highlighted naively, with
the answer held by a `remember` inside the fence, a two-hundred-line Kotlin
fence cost **174ms to lex** and the lazy list charged it again every time
the block scrolled back into composition -- six times in one bench run,
1043ms of lexing, and the scroll's draw phase up at 1.29ms per frame. So
highlighting is now warmed and cached exactly as parsing is
(`ParsedReplies.highlighted`, filled by `warm` from `fences(parse)`), and
`highlight` is a plain function taking no colour from the theme, which is
what lets it run off the drawing thread.
Because the warming has to ask for the same string the drawing does,
`fenceContent` extracts the code and the language word itself -- the rule
copied from the library's `MarkdownCodeFence`, which is a composable and so
cannot be called from `warm`. Two extractions would be two keys, and the
warmed answer would be missed at every fence with nothing saying so.
A fence *still arriving* was the same stall in a second place, and the
warming does not reach it: the tail is re-lexed at every delta, on the
composing thread, for colours on text that is being replaced as fast as they
are computed. Measured streaming the same fence: **211 lexes, 13.7 seconds**
across the turn, the worst 177ms. So a block that is still being written is
drawn plain and takes its colours when it freezes -- `MarkdownRoot`'s
`streaming`, which is true only for a live reply's *last* segment, so a
finished fence colours as soon as the next block starts. It is the bargain
[LiveParse] already makes for a reference link defined at the foot of a
message, and the settle-time lex then happens in `warm`, off the drawing
thread: one lex, 374ms, and the row redraws coloured on the tick that
follows it.
Clean pair, two fresh sessions of the same 200-line Kotlin fence, no saved
anchor, same gestures (`transcript-bench.sh`):
| | before (no highlighting) | after (warmed) |
|---|---|---|
| draw phase per frame | 0.74ms | 0.77ms |
| the transcript's share | 0.33ms | 0.33ms |
| lexing during the scroll | none | none |
Streaming that fence in (`stream-bench.sh /tmp/longfence.md`), before and
after the plain-while-writing rule:
| | before | after |
|---|---|---|
| lexes during the turn | 211 | 1 (in `warm`, off-thread) |
| time in them | 13,665ms | 374ms |
| worst single lex | 177.1ms | -- |
| draw phase per frame | 1.51ms | 1.01ms |
**3. `MarkdownRoot` no longer calls the library's `Markdown()`.** It
provides the locals itself -- reference links from the parse, padding,
dimens, colours, typography, a no-op image transformer, animations,
components. Nothing between a piece and the screen is the library's now
except the leaf composables named in the component table.
**4. Paragraphs with images -- done differently from the plan.** The plan
said draw the image as its own piece; measuring first showed the app has no
image loader and the renderer's transformer was the no-op one, so an image
in a reply drew as *nothing at all*. An `IMAGE` node is now appended by
`appendPlainLink` as a link carrying its alt text (the address when there is
none), which says what was there and opens it. With no image to place, the
`hasImage` branch and the renderer's `MarkdownText` are gone: every
paragraph is the platform `BasicText`.
**5. Per-item units for a streaming list.** `LiveParse.advanceTo` now cuts
at the last item of a multi-item list (`openPiece`), provided that item has
content beyond its marker -- a bare `-` is an empty item now and the first
character of a paragraph line once `-x` arrives, so cutting on it would draw
that line as a new item. The cut is at the start of the item's line, so the
indentation the reparse reads its nesting from survives.
`Segment.continues` marks a tail that carries on a list, and
`MarkdownPiece`'s `continuesList`/`listContinues` keep an inner item's
padding at the seam, so nothing moves when the seam does.
Clean pair, two fresh sessions, forty linked bullet items streamed word at a
time (`stream-bench.sh /tmp/longlist.md`):
| | before | after |
|---|---|---|
| reparses while streaming | 482 | 483 |
| total time in them | 2412ms | 674ms |
| mean / worst | 5.0ms / 8.9ms | 1.4ms / 7.9ms |
| `record: one block` worst | 1.8ms | 0.7ms |
The worst case moves least, which is the shape to expect: the first reparse
of a tail still covers whatever has arrived, and the last item can be long.
What changes is that every reparse after it covers one item instead of the
whole list.
**The restore walked back one event per request.** Found while benching, and
older than this work. `savedAnchor`'s loop asked for
`oldestSeq - anchor.seq + RESTORE_PAGE_CUSHION` events; when the anchor's row
is already loaded but is the oldest half-row (which `anchorRow` refuses,
correctly -- it grows when the page behind it lands), that span is negative
and was coerced to 1. So the restore fetched one event, then one more, at a
round trip each: six hundred requests walking a long reply back a word at a
time, with the screen on its spinner the whole way and the sandbox log
printing `limit=1` once a second. It now asks for a page counted in rows,
which is the only kind that can promise to reach the row behind the anchor.
**Harness.** `app/stream-bench.sh [-k] FILE` is `transcript-bench.sh` for a
reply still arriving: opens the first session, taps the app's own "Jump to
latest" so the list is pinned to the newest end, resets the report, sends
FILE, waits for the transcript to stop growing, prints the report. Both of
those last two are corrections to a first version that measured nothing:
a transcript parked further back never redraws while a reply streams into it
(the list must not move under a reader), and a session is idle at *both*
ends of a turn, so polling for idle answers before the turn has started.
Fixtures live in `/tmp` and are regenerated from the shapes named here:
`fixture.md` (lists four deep, ordered and nested, fences in
kotlin/rust/sh/none, a table with a link, a quote with a list, an inline and
a standalone image, a reference link), `longfence.md` (200-line Kotlin
fence), `longlist.md` (40 linked items).
**Two traps in the emulator loop**, both of which cost a bench run here.
`adb shell pm clear` removes the enrolment and the notification permission
along with the saved anchors, so the next run measures a permission dialog;
re-enrol with the command `ui-sandbox.sh` prints and
`pm grant … POST_NOTIFICATIONS`. And a saved anchor is per session id, so
the only way two builds start a scroll from the same place is a *fresh
session for each*.
## What is next, in order
1. **The reconnect loop.** Restarting the app onto a session with a saved
anchor while a long reply was streaming left it reconnecting every 1.5s
(`RECONNECT_DELAY_MS`), spinner up, until the server was restarted.
`events?after=N` more than `CATCH_UP_LIMIT` (200) behind answers `reset`
plus the newest 200 *raw* deltas -- a window starting mid-message -- and
the reset clears `items`, which is the state the restore loop then pages
against. The one-event-per-request bug above was part of what made it so
visible; whether it survives that fix is the first thing to find out.
2. **File the highlights range bug upstream.** There is no `gh` and no
GitHub credential in this VM, so it needs Bryan or a token. One-line
repro: lexing `x '*/a/*'` as `SyntaxLanguage.SHELL` in highlights 1.1.0
returns a highlight whose `location.end` precedes its `location.start`.
Until then `highlight` drops such spans, which is why the shell fence in
`fixture.md` -- whose command contains `'*/.git/*'` -- draws plain while
an ordinary shell fence colours.
3. **Regression runs.** `transcript-bench.sh` and `stream-bench.sh` before
and after any change to the files above, with the report in the commit.
The numbers to watch are the worst `record: one block`, the reparse mean
while streaming, and the draw phase's accounting line.