Split the newest reply once its turn ends, and make the UI harness reusable

The transcript's remaining lag was the newest assistant reply: transcriptUnits
kept the last row whole -- right while it streams (splitting a changing text
is a parse per delta), wrong forever after, so a session that ends on a long
reply drew it as one lazy-list item with every node alive. On a Pixel 9 Pro
XL that was 13.8ms of draw phase a frame, 79% of it the framework's own
per-node bookkeeping, against a 34,996px item.

An AssistantMsg now carries `settled`, folded from the status event that ends
its turn (status changes are transcript events with seqs, so replay settles
the same way), and cleared if a delta ever grows the message again. A settled
newest reply splits like every other. Folding it -- rather than reading the
screen's status -- routes the resplit through the held-events gate, so it can
only happen at the newest end while pinned, never under a reader. The
"session is working" predicate now lives once, in sessionWorking().

Measured on the emulator, same session and gestures, a 43KB reply as the
last row: draw phase 3.92ms -> 1.20ms per frame, framework share 3.07ms
(78%) -> 0.54ms (45%), worst single measure 82.5ms -> 9.1ms. The report's
"on screen" line went from one 60,674px AssistantMsg to five blocks of
95-846px. A live streamed turn settles and splits the moment it goes idle.

The harness half, asked for by Bryan: ui-sandbox.sh now derives its port and
root from the checkout name (two checkouts' sandboxes cannot reach each
other), keeps its token in ~/.config/ai-app/sandbox-token and salvages
enrolled device tokens across restarts (enrol the emulator once, ever), and
gained the driving verbs every UI session was re-inventing in /tmp: spawn,
send (text or @file), api. transcript-bench.sh is the standard
scroll-and-report measurement. AGENTS.md documents all of it.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
irisandClaude Fable 5 committed 2026-09-01 17:19:19 -04:00
1 parent 917eb9a7b3
commit 7a48f8ff1f
8 files changed
+250 -23

No files matched your search

@@ -297,6 +297,16 @@ fun parseSeqEvent(json: String): SeqEvent {
* compaction that finished without saying how much it recovered, or a clear nobody has run a turn
* since.
*/
/**
* Whether [state] is one the session is doing work in -- the states a turn is still open under.
*
* One predicate because two readers have to agree on the list: the session screen's working
* indicator, and the fold's decision that the newest reply is finished
* ([TranscriptItem.AssistantMsg.settled]). Two copies would drift the first time the server grows a
* state, and the drift would be a reply that never splits or one split mid-stream.
*/
fun sessionWorking(state: String): Boolean = state == "running" || state == "compacting"
fun contextAfter(current: Long?, event: SessionEvent): Long? =
when (event) {
// Falls back to what we had, so a turn the dialect reported no usage for is stale by a
@@ -352,7 +352,7 @@ fun StatusText(status: String) {
else -> status to MaterialTheme.colorScheme.onSurfaceVariant
}
Row(verticalAlignment = Alignment.CenterVertically) {
if (status == "running" || status == "compacting") {
if (sessionWorking(status)) {
// The same colour as the word beside it: the two are one signal, and a spinner in
// the theme's accent says the state is something other than what the label says.
CircularProgressIndicator(
@@ -319,7 +319,7 @@ fun SessionScreen(settings: ServerSettings, summary: SessionSummary, onBack: ()
// them. From the server rather than from this screen, so a rename sent from the settings
// screen -- or from another device -- is drawn waiting here too.
var waitingCommands by remember { mutableStateOf(listOf<Pair<String, String>>()) }
val running = status == "running" || status == "compacting"
val running = sessionWorking(status)
var moreHistory by remember { mutableStateOf(true) }
var loadingHistory by remember { mutableStateOf(false) }
var ready by remember { mutableStateOf(false) }
@@ -45,7 +45,27 @@ sealed class TranscriptItem {
val images: List<String> = emptyList(),
) : TranscriptItem()
data class AssistantMsg(override val seq: Long, val text: String) : TranscriptItem()
data class AssistantMsg(
override val seq: Long,
val text: String,
/**
* Whether this reply is finished: the session has stopped working since its last delta.
*
* What it buys is the split. [transcriptUnits] keeps the newest reply whole because a
* streaming reply's text changes per delta and splitting a changing text is a parse per
* delta -- but "newest" outlives the turn, so a session that ends on a long reply was
* drawing it as one item indefinitely, with every node of it alive. Measured on a Pixel 9
* Pro XL: one 34,996px reply on screen put the frame's draw phase at 13.8ms, 79% of it the
* framework's own bookkeeping, which grows with alive nodes.
*
* Folded from the status event that ended the turn, rather than read off the screen's
* status, because rows only change through the held-events gate: the split changes the
* newest row's list identity, and doing that from a status flip while somebody is reading
* inside that reply would step the list under them. An event has to wait for the reader to
* be at the newest end; a screen state does not.
*/
val settled: Boolean = false,
) : TranscriptItem()
data class ToolRun(
override val seq: Long,
@@ -362,7 +382,8 @@ fun foldEvent(items: List<TranscriptItem>, entry: SeqEvent): List<TranscriptItem
// every frame, and the list would jump for the whole of a streamed answer.
val last = items.lastOrNull()
if (last is TranscriptItem.AssistantMsg) {
items.dropLast(1) + last.copy(text = last.text + event.delta)
// A message growing again is not finished, whatever a status said in between.
items.dropLast(1) + last.copy(text = last.text + event.delta, settled = false)
} else {
items + TranscriptItem.AssistantMsg(entry.seq, event.delta)
}
@@ -456,7 +477,7 @@ fun foldEvent(items: List<TranscriptItem>, entry: SeqEvent): List<TranscriptItem
// is nothing it belongs above.
is SessionEvent.MessageDropped -> items
is SessionEvent.Settings -> items
is SessionEvent.Status -> items
is SessionEvent.Status -> settleReply(items, event.state)
is SessionEvent.Error -> items + TranscriptItem.ErrorMsg(entry.seq, event.message)
is SessionEvent.Image ->
// Under the call that produced it when there is one, and a row of
@@ -479,6 +500,20 @@ fun foldEvent(items: List<TranscriptItem>, entry: SeqEvent): List<TranscriptItem
is SessionEvent.UsageDelta -> items
}
/**
* A status saying the session stopped working is the moment its newest reply is finished.
*
* See [TranscriptItem.AssistantMsg.settled] for what the mark buys and why it is made here in the
* fold. Status changes are transcript events with seqs of their own, so a replayed session settles
* its replies the same way a live one does.
*/
private fun settleReply(items: List<TranscriptItem>, state: String): List<TranscriptItem> {
if (sessionWorking(state)) return items
val last = items.lastOrNull() as? TranscriptItem.AssistantMsg ?: return items
if (last.settled) return items
return items.dropLast(1) + last.copy(settled = true)
}
private fun updateTool(
items: List<TranscriptItem>,
id: String,
@@ -126,10 +126,12 @@ sealed class TranscriptUnit {
* Every settled reply is cut into its blocks ([markdownBlocks], via the caches on [replies] so a
* message is only ever split once), and so is an *opened* peer message -- [openNotes] is which ones
* those are, which is why the flatten needs it. A shut one is a single heading and cannot be worth
* splitting. The reply still arriving -- the last row -- stays whole: its text changes with every
* delta, and splitting it here would parse the whole message per delta on whichever thread is
* composing. [AssistantMessage]'s own streaming path already parses deltas off the main thread and
* gives the live message a layer per block.
* splitting. The reply still arriving -- the newest row, until the status event that ends its turn
* marks it [TranscriptItem.AssistantMsg.settled] -- stays whole: its text changes with every delta,
* and splitting it here would parse the whole message per delta on whichever thread is composing.
* [AssistantMessage]'s own streaming path already parses deltas off the main thread and gives the
* live message a layer per block. Once settled it splits like every other reply, which is what
* bounds the newest row's cost after a session ends on a long one.
*
* Runs per fold, so it must stay proportional to what is loaded with no parsing in it on the warm
* path: [ParsedReplies.partsOf] and [ParsedReplies.blocksOf] are lookups for any text [warm] has
@@ -163,7 +165,9 @@ fun transcriptUnits(
)
}
}
} else if (item is TranscriptItem.AssistantMsg && index != rows.lastIndex) {
} else if (
item is TranscriptItem.AssistantMsg && (item.settled || index != rows.lastIndex)
) {
var ordinal = 0
fun gap() = if (ordinal == 0) rowGap else BLOCK_SPACING
replies.partsOf(item.text).forEach { part ->