dev.snipme:highlights 1.1.0 found comments before it knew the language and paired /* with */ by ordinal, so `//` in any URL commented out the rest of its line, every Rust `#[derive(...)]` greyed out as a comment, a `#` inside a Kotlin string swallowed the line, and `x '*/a/*'` in shell produced a span whose end preceded its start -- the one that crashed a card holding `-path '*/.git/*'`. None of that could be post-processed away, because comments won over strings before the language was known. Highlighter.kt is one left-to-right scanner: at each position it is in a line comment, a block comment, a string, or ordinary code, and every span is emitted by advancing an index, so spans cannot overlap, arrive out of order or run backwards. Languages.kt is a `Rules` row per language -- comment tokens, block comment and whether it nests, the string forms, what opens an attribute, and the keyword set -- so a new language is a table entry. The keyword lists came from the library's SyntaxTokens.kt (Apache-2.0, noted at the table) so nothing that is coloured today turns plain, and RON, TOML, fish and JSON are coloured for the first time. HighlighterTest.kt is a new JVM unit test source set -- 24 cases, the library's mistakes kept as regressions, plus a sweep asserting no span escapes the code for any language on unterminated and empty input. AGENTS.md's app line now runs :androidApp:testDebugUnitTest. Measured on the ai-app emulator, debug build, a ~200-line Kotlin fence sent into a sandbox session: before code highlighted: 1, 101.9ms total, 101.9ms mean, 101.9ms worst after code highlighted: 1, 15.0ms total, 15.0ms mean, 15.0ms worst and a second fence in the same run took 13.9ms, so that is the steady cost rather than class loading. stream-bench.sh after the change: code highlighted: 1, 12.1ms total, 12.1ms mean, 12.1ms worst markdown reparsed while streaming: 1329, 2130.7ms total, 1.6ms mean, 8.7ms worst record: one block: 131, 11.4ms total, 0.1ms mean, 0.4ms worst draw phase 1.21ms per frame, the transcript 0.23ms of it transcript-bench.sh after: draw phase 1.10ms per frame, the transcript 0.49ms (place 0.48), worst place 4.3ms -- unchanged within run-to-run noise, as expected, since the scan happens in `warm` and not while drawing. Looked at on the emulator: a URL inside a Kotlin string, a Rust attribute with a lifetime and a raw string, a shell line with globs and `$#`, a RON fence and a TOML fence all colour correctly; a Bash tool card still colours its command; a plain Python fence -- which this change had no reason to touch -- looks as it did; an unknown language stays plain; and a fence is plain while it streams and colours when it freezes. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
266 lines
16 KiB
Markdown
266 lines
16 KiB
Markdown
# Transcript rendering: what was learned, and what is next
|
|
|
|
Written 2026-09-03 at the end of a week of work on the session screen's
|
|
transcript, so the next session can start from here rather than from a
|
|
compacted context. Work that is finished lives in "the architecture, as
|
|
built"; the running log of how each piece got there has been dropped. `AGENTS.md` holds the one-paragraph conventions; this is
|
|
the longer record: the measurements that drove each decision, the
|
|
techniques that worked, the ones that did not, and the order to do the rest
|
|
in. `PLAN.md` remains the design source of truth; nothing here contradicts
|
|
it.
|
|
|
|
## The goal, and where it stands
|
|
|
|
A reply of any length must scroll at the phone's 120Hz without a bump, and
|
|
must keep doing so while the reply is still streaming in. Measured on the
|
|
Pixel 9 Pro XL by Bryan, the transcript went from visible stalls at long
|
|
replies and at lists of links to "I have to actually try to feel any
|
|
bumps". The remaining work is finish and extensibility rather than
|
|
performance.
|
|
|
|
## The architecture, as built
|
|
|
|
Everything below lives under `app/androidApp/src/main/kotlin/com/example/aiapp/`.
|
|
|
|
**Rows become units, and units are bounded.** `TranscriptUnits.kt` turns a
|
|
transcript row into the things the lazy list actually holds. An assistant
|
|
reply is not one unit: it is one unit per piece of its markdown, so the
|
|
list composes and draws a paragraph, a fence, a table or one bullet at a
|
|
time. The reason is the draw phase: a row's display list holds every glyph
|
|
of it and is re-recorded whenever drawing is invalidated, and the lazy list
|
|
composes an item whole in the frame it scrolls into. The tallest single
|
|
row still being drawn before this was 36,982px, twenty-five screens in one
|
|
message. Long user messages are sliced the same way (`UserChunk`), through
|
|
the shared `cardPiece` modifier that draws one card in lazy-list pieces.
|
|
|
|
**One parse per message, addressed by piece.** `MarkdownPieces.kt`'s
|
|
`Piece(block, item)` is an address into the message's single parse tree,
|
|
not a substring: `block` indexes the root's children and `item` one
|
|
`LIST_ITEM` of a top-level list. Cutting was originally done by
|
|
re-parsing substrings, which cost a parse per piece and broke reference
|
|
links defined at the foot of a message. `ParsedReplies` caches the parse
|
|
and the piece list per text (`of`, `piecesOf`), warmed off the composing
|
|
thread by `TranscriptItems.warm`. The parser is still intellij-markdown via
|
|
the mikepenz renderer, but its `Markdown()` composable is not called at all:
|
|
`MarkdownRoot` in `Markdown.kt` provides the `Local*` environment itself --
|
|
reference links from the parse, padding, dimens, colours, typography, a
|
|
no-op image transformer, animations, components -- and `MarkdownElement`
|
|
dispatches a whole block through our component table. Nothing between a
|
|
piece and the screen is the library's now except the leaf composables that
|
|
table names.
|
|
|
|
**Lists are drawn an item at a time, by us.** The renderer has no element
|
|
for a single list item, so `MarkdownListItem` draws one: marker, then the
|
|
item's children, nested lists recursing through `MarkdownList`. The
|
|
marker is drawn in one place on purpose; styled bullets per depth go
|
|
there.
|
|
|
|
**Links are spans, not nodes.** `MarkdownLinks.kt`. Compose turns every
|
|
`LinkAnnotation` into a layout node (clipped, focusable, hoverable,
|
|
clickable, outline recomputed from the text layout). A paragraph of eight
|
|
links was nine nodes, and measured against the same paragraphs with each
|
|
link replaced by plain words it cost 26.3ms worst measure against 5.2ms,
|
|
1.7x the place time. That was the bump at a reply's list of sources.
|
|
`LinkedText` builds the annotated string with the renderer's own inline
|
|
builder but answers links itself: colour, underline, a string annotation
|
|
carrying the URL, and one tap detector for the whole text that asks the
|
|
layout which glyph is under the finger. Hit-testing must check the glyph
|
|
on either side of the returned caret, because `getOffsetForPosition`
|
|
returns the nearest boundary; taps on the right half of a glyph otherwise
|
|
open nothing. Headings need the `ATX_CONTENT`/`SETEXT_CONTENT` child, since
|
|
the inline builder draws nothing for a node type it does not know (a week
|
|
of blank headings). Tables go through `LinkedTable`/`LinkedTableRow` so
|
|
cells get the same treatment.
|
|
|
|
**Text draws on the platform directly.** A paragraph without an image
|
|
skips the renderer's `MarkdownText`, which charges every paragraph for the
|
|
possibility of inline images (placement callback, derived inline-content
|
|
map, semantics group, size animation). Paragraphs that contain an image
|
|
still take the renderer's path.
|
|
|
|
**Tables spread or scroll without subcomposition.** The renderer used
|
|
`BoxWithConstraints` to decide; `LinkedTable` uses
|
|
`fillMaxWidth().horizontalScroll().layout { }` -- `horizontalScroll`
|
|
passes `minWidth` through and lifts `maxWidth` to infinity, so the inner
|
|
layout reads `minWidth` as the room available and takes
|
|
`max(minWidth, columns * cellWidth)`.
|
|
|
|
**A streaming reply is reparsed one block at a time.** `LiveParse` in
|
|
`Markdown.kt` freezes every finished top-level block with its parse and
|
|
reparses only the tail block per delta. Markdown's block rules make later
|
|
text unable to alter an earlier block, with the single exception of a
|
|
late reference definition, which is accepted. Measured on a 58-word stream
|
|
of list, fence, table and quote: 47 tail reparses at 1.7ms mean. A
|
|
single-list stream would reparse the whole list per delta, since it is one
|
|
tail block; that is what the rule below cuts.
|
|
|
|
**A streaming list becomes a unit per item.** `LiveParse.advanceTo` cuts at
|
|
the last item of a multi-item list (`openPiece`), provided that item has
|
|
content beyond its marker -- a bare `-` is an empty item now and the first
|
|
character of a paragraph line once `-x` arrives, so cutting on it would draw
|
|
that line as a new item. The cut is at the start of the item's line, so the
|
|
indentation the reparse reads its nesting from survives. `Segment.continues`
|
|
marks a tail that carries on a list, and `MarkdownPiece`'s
|
|
`continuesList`/`listContinues` keep an inner item's padding at the seam, so
|
|
nothing moves when the seam does. Forty linked bullets streamed a word at a
|
|
time went from 2412ms of reparsing to 674ms, and `record: one block` from
|
|
1.8ms worst to 0.7ms.
|
|
|
|
**Fences are highlighted off the drawing thread, and a fence still being
|
|
written is drawn plain.** `Highlighter.kt` holds `highlight` and the scanner
|
|
behind it (shared with a tool call's input, so the same code is the same
|
|
colours wherever it appears); `CodeFence.kt` holds the `fenceLanguage` alias
|
|
table and `fenceContent`. A word not in the table stays plain, because a
|
|
fence coloured by the wrong language's rules looks highlighted and is wrong
|
|
in a way the reader cannot see. Highlighting is warmed and cached exactly as
|
|
parsing is (`ParsedReplies.highlighted`, filled by `warm` from
|
|
`fences(parse)`), and `highlight` takes no colour from the theme, which is
|
|
what lets it run off the drawing thread: a two-hundred-line Kotlin fence
|
|
costs 15ms to scan on the emulator's debug build -- it cost 102ms through
|
|
the library that used to do this -- and a `remember` inside the fence was
|
|
charged that again every time the block scrolled back into composition. Because the warming has to ask for the same
|
|
string the drawing does, `fenceContent` extracts the code and the language
|
|
word itself -- two extractions would be two keys, and the warmed answer
|
|
would be missed at every fence with nothing saying so. A fence still
|
|
arriving is the same stall in a second place, and warming cannot reach it:
|
|
the tail was re-lexed at every delta, on the composing thread, for colours
|
|
on text being replaced as fast as they were computed -- 211 lexes and 13.7
|
|
seconds across one turn. So `MarkdownRoot`'s `streaming`, true only for a
|
|
live reply's last segment, draws the block plain until it freezes; a
|
|
finished fence colours as soon as the next block starts, and the settling
|
|
lex happens once, in `warm`.
|
|
|
|
**Markers, and images.** `MarkdownListItem`'s `Marker` draws the bullet by
|
|
depth, cycling past the third, in `listMarkerColor` (Theme.kt). The colour
|
|
is the same at every depth on purpose: depth is said by the glyph and the
|
|
indent, and a colour per depth would make a difference in degree look like
|
|
one in kind. The app has no image loader and the renderer's transformer was
|
|
the no-op one, so an image in a reply drew as *nothing at all*; an `IMAGE`
|
|
node is now appended by `appendPlainLink` as a link carrying its alt text
|
|
(the address when there is none), which says what was there and opens it.
|
|
|
|
**Expansion anchors the edge that was tapped, and the list never moves
|
|
under the reader** except when pinned to the bottom with new content
|
|
arriving. Those two rules are in `ScrollAnchor.kt` and `TranscriptList.kt`
|
|
and are the reason several tempting simplifications were rejected.
|
|
|
|
## Techniques and harness
|
|
|
|
- **`app/ui-sandbox.sh`** starts a second `ai-server` against a sandbox
|
|
home with the echo driver, so nothing touches real sessions.
|
|
`spawn [title]` makes an echo session and prints its id; `send SID text`
|
|
or `send SID @file` sends into it; `api /path [curl args]` is an
|
|
authenticated request. Restarting it regenerates the config but keeps
|
|
enrolled tokens.
|
|
- **The echo driver is the test rig** (`server/src/session/echo.rs`, the
|
|
list at the top of the file). `/stream N`, `/mixed N`, `/table N`,
|
|
`/tools N gap`, `/ask`, `/peer`, `/compact`, `/slow`, `/bash command`
|
|
each produce a shape the real CLI produces only when it feels like it.
|
|
Build what a UI test needs into it rather than spending model turns.
|
|
- **`app/transcript-bench.sh`** is the standard measurement: restart, open
|
|
the first session, scroll, print the render report. The report is what
|
|
the "Copy render timings" button copies and also logs
|
|
(`adb logcat -d -s ai-app:I`), and it includes the last crash's stack
|
|
(`CrashLog.kt`), which is how a crash on the phone reaches a session
|
|
here.
|
|
- **`app/stream-bench.sh [-k] FILE`** is `transcript-bench.sh` for a reply
|
|
still arriving: opens the first session, taps "Jump to latest" so the list
|
|
is pinned to the newest end, resets the report, sends FILE, waits for the
|
|
transcript to stop growing, prints the report. Both of those last two are
|
|
corrections to a first version that measured nothing -- a transcript parked
|
|
further back never redraws while a reply streams into it, and a session is
|
|
idle at *both* ends of a turn, so polling for idle answers before the turn
|
|
has started. Fixtures live in `/tmp` and are regenerated from the shapes
|
|
named here: `fixture.md` (lists four deep, ordered and nested, fences in
|
|
kotlin/rust/sh/none, a table with a link, a quote with a list, an inline
|
|
and a standalone image, a reference link), `longfence.md` (200-line Kotlin
|
|
fence), `longlist.md` (40 linked items).
|
|
- **Two traps in the emulator loop**, each of which cost a bench run.
|
|
`adb shell pm clear` removes the enrolment and the notification permission
|
|
along with the saved anchors, so the next run measures a permission
|
|
dialog; re-enrol with the command `ui-sandbox.sh` prints and
|
|
`pm grant ... POST_NOTIFICATIONS`. And a saved anchor is per session id,
|
|
so the only way two builds start a scroll from the same place is a *fresh
|
|
session for each*.
|
|
- **`DebugStats`/`FrameStats`** time our own phases (`record: one block`,
|
|
`measure: the app root`) and count events (`markdown reparsed while
|
|
streaming`, `markdown cut into pieces`). Add a counter before guessing.
|
|
- **`app/trace-draw.sh`** names what a scrolling frame spends inside the
|
|
framework, via `atrace` text output, no trace processor needed. It is
|
|
how the link-node cost was attributed.
|
|
- **`app/debug-transcript.sh`** loads a real Claude Code conversation onto
|
|
the emulator; two faults were invisible on fixtures and obvious on it.
|
|
Real transcripts are private: fixtures stay in `/tmp`, never in the repo.
|
|
- **`ui-trace`** reads the screen as text. Bounds print as
|
|
`x1,y1..x2,y2`; unanchored `-m` patterns match labels, anchored ones do
|
|
not. A row taller than the viewport reports clipped bounds, so compare
|
|
screenshots for that case.
|
|
- **Emulator frame times are not app measurements.** Software rendering
|
|
puts the stock Settings app at 60ms of UI-thread traversal per frame.
|
|
Costs of operations in milliseconds rank correctly; smoothness itself is
|
|
judged on the phone.
|
|
- **System Tracing on the phone does not work on GrapheneOS.** Its
|
|
Categories list is empty because the tracing daemon builds it by running
|
|
`atrace --list_categories`, which returns nothing there, and a recorded
|
|
trace contains zero ftrace events: no app sections, no frames, no
|
|
scheduling. Callstack sampling records, but the app's profiler config
|
|
unwinds one process shard in four. GrapheneOS issues 2206 and 6094 are
|
|
open on exactly this. Until they close, phone numbers come from the
|
|
render report and from Bryan noticing.
|
|
- **Compose `DropdownMenu` in an edge-to-edge activity** needs
|
|
`PopupProperties(clippingEnabled = false)` or it opens a status bar's
|
|
height away from its anchor (`~/.claude/TOOLCHAIN.md`).
|
|
- **The syntax highlighter is ours: `Highlighter.kt` and `Languages.kt`.**
|
|
One left-to-right scanner with a small state -- in a line comment, in a
|
|
block comment, in a string, or in ordinary code -- and a `Rules` row per
|
|
language, so a new language is a table entry rather than code. Every span
|
|
is emitted by advancing an index, so spans cannot overlap, arrive out of
|
|
order or run backwards, and an unterminated string or comment simply runs
|
|
to the end of the code. `HighlighterTest.kt` is the JVM unit test
|
|
(`./gradlew :androidApp:testDebugUnitTest`); the cases in it are the
|
|
library's mistakes, kept as regressions.
|
|
It replaced dev.snipme:highlights 1.1.0 on 2026-09-03, which found
|
|
comments before it knew the language and paired `/*` with `*/` by
|
|
ordinal. That library used one set of delimiters for every language, so
|
|
`//` in any URL commented out the rest of its line (in `curl
|
|
https://example.com/x && echo done` the comment ran to the end and took
|
|
`echo` with it, and in Kotlin `val url = "https://..."` the string
|
|
disappeared inside it), every Rust `#[derive(...)]` greyed out as a
|
|
comment, a `#` inside a Kotlin string swallowed the line, and `x '*/a/*'`
|
|
in shell yielded `start=6, end=5` -- a range `AnnotatedString` rejects,
|
|
which crashed a card holding `-path '*/.git/*'`. Comments were located
|
|
before strings and won over them, so post-processing could not recover
|
|
what a wrong comment range had already suppressed. The scanner is also
|
|
about seven times faster on the same fixture, and it colours RON, TOML,
|
|
fish and JSON, which the library did not know at all.
|
|
|
|
## Rejected, and why
|
|
|
|
- **Writing our own markdown renderer.** Rejected in favour of keeping
|
|
the intellij-markdown parser and the library's inline builder while
|
|
owning block dispatch and the leaf composables. The parser is the hard
|
|
part and is not the slow part; everything that was slow lived in the
|
|
composables, which are now ours.
|
|
- **Re-parsing substrings per piece.** Cost a parse per piece and broke
|
|
foot-of-message reference links. Replaced by addressed pieces of one
|
|
parse.
|
|
- **Animated or timing-dependent corrections.** Anything the reader could
|
|
catch at 120Hz is a bug; corrections must be structurally impossible to
|
|
see.
|
|
|
|
## What is next, in order
|
|
|
|
1. **The reconnect loop.** Restarting the app onto a session with a saved
|
|
anchor while a long reply was streaming left it reconnecting every 1.5s
|
|
(`RECONNECT_DELAY_MS`), spinner up, until the server was restarted.
|
|
`events?after=N` more than `CATCH_UP_LIMIT` (200) behind answers `reset`
|
|
plus the newest 200 *raw* deltas -- a window starting mid-message -- and
|
|
the reset clears `items`, which is the state the restore loop then pages
|
|
against. The restore's one-event-per-request bug was part of what made it
|
|
so visible and has been fixed; whether this survives that fix is the
|
|
first thing to find out.
|
|
2. **Regression runs.** `transcript-bench.sh` and `stream-bench.sh` before
|
|
and after any change to the files above, with the report in the commit.
|
|
The numbers to watch are the worst `record: one block`, the reparse mean
|
|
while streaming, and the draw phase's accounting line.
|