fd2e1d079836553704ea49bd15083faf71dc591f
163
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
fd2e1d0798 |
Delete and import sessions in batches, and say which rows are busy
Clearing out imported sessions was one confirmation dialog per row, which is why it was not worth doing. Holding a row on the import screen now selects it and plain taps add more; Delete and Import act on the whole selection from a bar along the bottom. Submitting hands the work over and puts the screen back as it was: the selection clears, the bar goes, and what says the work is happening is the rows it is happening to -- the one in flight marked with its operation, the rest marked "waiting". Both are inert, so a queued row cannot be tapped into starting a second CLI behind the batch already coming for it. Rows leave as each one lands rather than all at the end, because a finished row still sitting there looks exactly like one that was never imported; the rows below it therefore move, so a row that has just moved ignores taps for half a second. That busy appearance is one composable shared with the session list, which had its own dimmed row and its own word for it. It is a word rather than a bare spinner because deleting and importing differ in kind. Deleting a session can now take the machine's own transcript with it, as a switch in the confirmation and only where the driver keeps a record this app's delete cannot otherwise reach. Off by default, since leaving that copy is what makes an ordinary delete recoverable -- and the paragraph is rewritten rather than appended to when it is on, because the sentence promising the conversation is still there to import again is exactly the one the switch makes false. The server removes the machine's copy first, so a machine it cannot reach leaves the session where it was. `app/ui-sandbox.sh` is how all of this was driven: a second server with its own $HOME, invented transcripts and a two-line `claude`. Against the ordinary server, testing delete deletes somebody's conversation and testing import spends a turn on a real account. |
||
|
|
d07325401d |
Clean up the app: one fix per rule already written down somewhere else
A pass over every Kotlin file, fixing each place a rule the codebase had already learned was applied to only one member of its set: - Notifications: checkSelfPermission(POST_NOTIFICATIONS) on Android 12 and below answers "denied" for a permission that does not exist there, so every notification on API 24-32 was silently dropped. Version-guarded; before 13 the app-level switch is the whole answer. - The permission-mode list was written three times and had drifted: the import screen was missing "plan". One list in Api.kt now. - The import screen's delete refetched the whole list through a loading spinner -- the exact fault the session list's delete already fixed and documented. It now removes only the deleted row. - The import screen truncated paths at a hardcoded 40 characters; it now uses StartEllipsis against the row's real width, like the models screen. - warm() bypassed the partsOf cache built for it, re-scanning every loaded message per page, and warmed a multi-block memory note under keys no row looks up. It now mirrors transcriptUnits through the same caches. - The transcript's data model (TranscriptItem, foldEvent, joinPages, warm) moved out of SessionScreen.kt into TranscriptItems.kt: pure folding with no screen in it, changing for unrelated reasons in the same file. - One image fetch/decode/failed block was written twice; it is rememberSessionBitmap in SessionImage.kt now. - Lint is fully clean: android.media.ExifInterface replaced with the androidx one (the framework copy lacks the hostile-image parsing fixes, and these images arrive from outside the phone), highlights bumped to 1.1.0, and the notification fix above closed InlinedApi. - Dead weight out: an unused act() onFailure parameter, and five orphaned or misattached doc comments (UserBubble carried SessionImage's doc). Verified: ktfmt, compileDebugKotlin, lintDebug (0 issues), server suite (93 passed), clippy and rustfmt clean; exercised on the emulator against a real imported transcript -- paging, tool groups, block rendering, no crashes. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
a63ae8347f |
Load history before the reader arrives, and estimate room from real sizes
The first loading boundary sat barely off-screen: the opening page is deliberately small (sized for time-to-first-frame), so one short scroll met a spinner and a round trip. And the prefetch that should have hidden every later boundary estimated room-ahead from the average size of the units currently visible -- the worst possible sample, since two tall blocks filling a viewport multiply out over dozens of unseen one-line rows and report screens of room when the end is one swipe away. Iris's report showed exactly that: viewport 676px, two units visible. Three changes, one per cause: - The opening page is followed by a full page fetched in the background as soon as the screen is up -- the same reasoning that keeps the opening page off the critical path puts the first real page right behind it. A restore skips this; it just paged as deep as its anchor needed. - Room ahead is now added up from the real size of every unit the list has ever laid out, recorded by key as units pass through the viewport, with the running average standing in for units never seen. The walk early-outs at the cushion, so a frame's check stops after a few units in the common case. - The cushion is six screens (was three) and a page is 400 events (was 800). The cushion grew because firing early costs a page nobody may read while firing late is a spinner under a finger for a tunnel round trip; the page shrank because pages now load before anybody waits on them, so a page only has to stay above the fold floor -- the median run of deltas folding into one reply is ~400 events, and below that a page can add no visible room at all. Smaller pages cross the tunnel faster and land in smaller frame spikes, which is what makes the boundary hard to catch. Verified on the emulator against the biggest transcript on the machine at --delay 300: opening a session loads the first background page unscrolled, twenty-five hard flings paged the rest in with zero frames in which the spinner was on screen, and at --delay 2000 -- outrunning the chain deliberately -- the boundary shows the spinner and then heals with the paragraph being read holding its position. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
9477cd288a |
Stop the keyboard re-laying-out the whole screen, and float the composer
Opening the keyboard was late on 82% of frames on the Pixel while the transcript itself cost 0.25ms of each -- the cost was everywhere else. imePadding() sat on the activity's root box, so every frame of the IME animation resized the entire tree: measured on the emulator (with the new app-root timers) as a full re-measure (4.1ms), re-place (2.5ms) and re-record (1.0ms) of everything, ~34 frames per open, while the newly added recomposition counters read zero -- pure layout traversal, no recomposition to fix. So the keyboard now touches only what actually moves. The session screen's composer (status row, suggestions, attachments, text field, buttons) is a bottom-aligned overlay on its own layer, translated by the IME inset read inside the graphicsLayer block -- a keyboard frame invalidates layer properties only. The transcript box reserves the overlay's measured height plus imePadding, and that modifier is the whole of what the keyboard re-measures: the box's own size never changes, so the header and everything above it are untouched. The overlay is opaque for the one frame between it growing and the reservation catching up. imePadding moved off the activity root onto AppRoot's other screens, which keep the old arrangement -- none of them has a keyboard open over anything that scrolls at 120Hz. Same five-open protocol on the emulator, before and after: the app root is now measured zero times (was 168), per-frame app work 7.6ms -> 1.9ms (transcript measure 1.2 + place 0.5 + record 0.2), draw-phase p90 9.6ms -> 5.0ms, waited p50 2.6ms -> 0.4ms. What is left per frame is the transcript's own one-box remeasure, whose children skip measurement because their width is unchanged. Verified the states the overlay could have broken: keyboard over a long and a two-message conversation (content hangs from the composer in both), a three-line draft growing the composer upward with the reserve following, slash suggestions stacking above the field, and the closed state identical to before. The app-root timers and the recomposition counters stay in: they are the difference between this report saying "draw is high" and saying where. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
acb59adcf6 |
Replace the hand-built transcript window with a reversed lazy list of blocks
The plain Column held every loaded row and reimplemented what a lazy
list is -- windowing (retained ranges, stand-in spacers), height
bookkeeping (a heights map and prefix-summed tops), scroll anchoring,
and restore -- and every constraint kept intersecting that surface
somewhere new: the retained window was a fresh way to flicker, the
spacers hit the Constraints height limit, the IME re-measured the
world, and the restore needed pending anchors threaded through layout.
The reason a lazy list was abandoned no longer holds. It was dropped
when an item was a whole message, and a message can be twenty-five
screens tall -- composing one mid-fling is a hundred-millisecond frame.
The block splitting built later (for draw granularity) is the missing
piece: with one item per markdown *block*, entering composition costs
laying out a paragraph, and the parse is already cached by the same
warm() that always ran. So the transcript is now a
LazyColumn(reverseLayout) over TranscriptUnits -- settled replies
flattened to block items, everything else one item, the live reply kept
whole because its text changes per delta.
What each constraint rests on now, all verified on the emulator against
a real 1200-event imported transcript at --delay 120:
- Following the newest message, paging in history, and the keyboard are
the reversed layout itself: item zero is the bottom, an arriving
message extends the pinned end, a page lands past every visible
index, and an IME resize keeps the anchored item against the
composer. A short conversation stacks from the bottom.
- No item enters unready: pages are warmed before the fold lands (the
opening page now folds a scratch copy off-thread first), so heights
are real on first measure and scrolling back is cache hits -- zero
markdown parses on the composing thread across a full page-back
through all 1200 events.
- Restore resolves the saved seq to a unit index and snaps before the
draw gate lifts; anchors gained a unit ordinal (ScrollAnchor grew a
third field, read compatibly) so a position inside a forty-block
reply survives.
- Unloaded history is a spinner item at the far end while moreHistory
holds; paging triggers on estimated pixels ahead (units remaining at
the typical visible unit size), since a lazy list has not measured
what it never composed.
- The tap-half expansion anchoring survives, but through
requestScrollToItem: a raw dispatchRawDelta from onSizeChanged forces
remeasure inside the measure pass and crashes
("performMeasureAndLayout called during measure").
Deleted: TranscriptScroll.kt entirely (1000 lines of window, spacers,
tops, anchors, seeding). ParsedReplies gained a parts cache so the
per-fold flatten never re-scans a settled message, and clear() now
empties all three maps.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
||
|
|
0a0949ee7f |
Search the row tops, tell the truth about the window, report the residual
Three things, and the middle one is why the last commit's "zero flickers" was worth nothing. **The detector was measuring the wrong thing.** `covered()` compared what is on screen against `retained` -- which is the *intent*. Those came apart the moment the window started being recomputed during layout, because it is then updated before the frame draws: the check passed while the screen was still showing the spacers from the composition before. It now compares against what the list actually composed, recorded as it composes it. The first honest run found the open path still failing, and the counters named a branch of my own that did nothing at all -- position lost, scroll container reporting a measured position from a *different* session, so neither arm ran and the window was left empty. The whole transcript then composed as a single spacer. The test is no longer "is there a position to trust" but "is the window empty", which is the question that was actually being asked. **The scan was the lag.** Putting the recompute where the layout is made it correct and put a linear pass over every row into every frame with it -- 0.5ms to 1.1ms of placement at 190 rows, the same O(rows)-per-frame mistake this file exists to undo, reintroduced by its own fix. The tops are ascending by construction, so it is a binary search. **The button now does the subtraction.** The draw phase carries Compose's measurement as well as its recording, so "draw is high" never said which of three things was high, and the split was being worked out by hand in a conversation every time a report arrived. It reports the per-frame split directly: the transcript's own measure, placement and recording, and what is left, which is the framework's bookkeeping after a layout. The immediate window went from two screens to three, and from six rows to twelve. What the list has composed is always one frame behind what the window says -- the window is recomputed in layout and the rows it names are built by the next composition -- so the margin has to cover a frame of movement as well as the gap between recomputations. Measured down from misses of one and two rows. Verified: three cycles of open, fling up through a page load, fling back, close, on a 92-row session, plus reopens of a short one. Zero, with the detector that no longer flatters itself. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
c9e9ddb62f |
Split a stood-down run across spacers Constraints can hold
`Can't represent a width of 0 and height of 273238 in Constraints`, from `Modifier.height` during measure, which took the app down on opening one particular session. Compose packs a Constraints into a single Long, and at a width of zero that leaves eighteen bits for the height -- so 262,143px is the ceiling and a taller spacer throws. Collapsing every stood-down row into one spacer is what made that reachable. It is the change that stopped the per-frame cost growing with the conversation, and it also turned "the history above the window" into a single fixed height -- which for a long conversation read at the newest end is the whole transcript. The one that crashed was 273,238px. Split rather than clamped: the height is load-bearing. It is what keeps the transcript's total the same as the rows it stands in for, so shortening it would move everything under the reader -- a silent wrong answer in place of a loud one. The gaps the arrangement inserts between the pieces come out of the total for the same reason. The cap is 100,000px rather than the 262,143 that would just fit, because the limit depends on how many bits the width took: a spacer sized against today's screen width is a crash waiting for a wider one. Verified by forcing the split -- built with the cap at 5,000px, a run came out as ten spacers with the transcript intact and nothing reporting an unbuilt row -- then restored. Testing it only at the real cap would have tested the branch that cannot fail on any conversation reachable here. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
c2ceaf01d6 |
Count the flicker, then remove the four things causing it
"It still flickers sometimes" is not something a fix can be tested against,
so the first change is a counter. Drawing is the one place that knows both
what is on screen and what was built, so it asks: `covered()` compares the
rows in the viewport against the window and records how far short it fell
and which way. That turned a symptom into a number, and the number found
four separate causes -- three of which I would not have guessed, and one of
which I had already "fixed" twice.
- Opening a session recomputed the window before the scroll container had
measured anything. `maxValue` is zero then, which reads as the reader
being at the oldest end, so the window landed a whole transcript away
from where the session was about to open.
- A page of history recomputed it from a scroll position one layout out
of date -- stale by exactly the height of the page that had just
arrived. The reader's position is now carried across the change as a
row rather than as a pixel, and resolved through the row that *holds*
that seq, because a regroup can fold the row it named into another.
- The window was widened to cover the screen only when it was recomputed,
which was every two screens of movement. A fling covers a screen in a
frame or two, so it outran the window and arrived at rows that were
still spacers. Standing rows up a few per frame made that worse rather
than causing it: after a seed the built window is two screens wide and
grows two rows a frame. What is near the screen is now widened every
frame; only the outer bound is lazy.
- Both were measured in pixels, and a pixel budget cannot know how many
rows it covers until the rows have been measured -- which is never, on
the frame a session opens. The margins are now a number of screens *or*
a number of rows, whichever is larger.
The recompute moved to the transcript's placement, which is the one moment
both halves are current: the rows have just been measured and so has the
scroll container. Everywhere else it ran could be right about one and stale
about the other. It writes only when the answer changes, so a frame where
nothing moved costs one scan and no recomposition.
Two supporting fixes. `covered()` also repairs, so however the window goes
stale the damage is one frame rather than until the next scroll. And the
scroll position is now keyed per session -- `rememberScrollState()` is not
keyed, so a second session opened without leaving the screen inherited the
first one's offset and, worse for the window, its `maxValue`.
Verified on the emulator: flings up and down, three scroll-and-restore
cycles on a 92-row session and three reopens of a short one -- zero, where
before each reopen cost one to three. Placement stayed at 0.8ms.
Also here: crashes are recorded and travel out through the debug button, so
"it crashes opening that chat" arrives with a stack next time; the report
goes to the log as well as the clipboard, so a session driving the app over
adb can read it; and `trace-draw.sh` captures a labelled frame breakdown,
with a note that it must be run against the phone -- on this emulator two
thirds of a frame is `dequeueBuffer` and Compose's own draw is 0.40ms.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
41df8bb76a |
Stop re-measuring the transcript at the keyboard, and seed the window
Three things, all of them the same shape: work done because something was asked a question at the wrong moment. The keyboard. The layout modifier passed `minHeight = viewportSize` down to its child, so every frame of the IME animation changed the child's constraints -- and changed constraints are exactly what defeats the early-return in `MeasurePassDelegate.remeasure`. The whole transcript was re-measured on the way up, measured on a Pixel 9 Pro XL as 250 measurements averaging 2.8ms with a 54.5ms worst. The minimum is applied to what this node reports now, not to what it asks its child for, so the child early-returns; the content is placed against the bottom to keep the anchoring the minimum existed for. It only ever bites on a conversation shorter than the screen, which was never the case paying for it. The flicker. `retained` starts empty, so on the composition that introduces the rows every one of them is outside the window, the whole transcript collapses to a single spacer, and it draws blank for a frame. Seeded now -- at the newest end on open, and around the anchor on a restore. The restore case is the one that bites: `placed()` jumps the view during placement, and a window recomputed there only schedules a recomposition, so the destination would draw as spacer for a frame or two after drawing ungates. Seeded by pixels rather than by a row count, because a run of tool calls is eight rows and less than half a screen. The block layers. Every paragraph of every reply had a layer, which was right when whole rows were re-recording constantly and one row's display list was 36,982px tall. Re-recording is rare now -- 65 whole rows in fifty seconds of reading -- and a layer costs a layout node and a display list held for the life of the row, against the node count the per-frame cost scales with. They go to the message still arriving, which is the only one whose drawing is invalidated often enough to want the granularity. Verified on the emulator against the case none of this was written for: a two-message conversation far shorter than the viewport still hangs from the composer, with the keyboard both up (messages at y1756-1940, composer 2109) and down (936-1120 against 1289). Mechanism for the block layers and the hole in the restore seeding both from the ai-app-2-6d session; the layers are not re-positioned per frame as I had it, they are baked into the row's display list. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
f497d1f3ad |
Actually run the amortization, and stop paying for history twice
Two things committed in |
||
|
|
94b1509733 |
Stand rows up a few per frame instead of two screens at once
The counters now say which half is left. Re-recording is rare -- 157 row display lists in a minute of scrolling, 407ms in total -- while the frame's draw phase sits at 7.9ms. Since that phase also carries Compose's own measurement, what is left is laying text out: shaping glyphs, on the thread drawing the frame, and none of it timed by anything here. The window moved in one step, so every step boundary shaped two screens of markdown inside a single frame, a page landing did the same, and opening a session did seventeen screens of it in one. It now walks towards its target a couple of rows a frame. That does not make the work cheaper and is not meant to: it stops it arriving together, which is the difference between a frame that is late and a frame that is missed by ten. What is on screen is never amortised. The visible range goes up in the frame it is needed whatever else is pending, and only the margin being read *towards* is spread out -- so this cannot show anybody a gap, which is the failure the last two changes in this area both had. Diagnosis and the ordering from the ai-app-2-6d session, whose test this follows: records small and draw high means shaping rather than recording. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
9052e5f55e |
Notice a row arriving, and stop subcomposing the list at the keyboard
Two faults from making the retained window lazy, both of them mine. A screen of blank between the last message and the box it was typed in. The window was recomputed when the view had moved far enough, and a message arriving does not move the view -- so the new row fell outside the window and stood in as a spacer of its guessed height. The version number the check compares against is only bumped by the recompute it guards, so asking before refreshing meant never noticing. It refreshes first now, and the window also watches how many rows there are, because a row arriving is the case it exists to catch and the one that does not announce itself through the scroll position. And the keyboard, which was the most expensive thing on the screen. The visible height came from a `BoxWithConstraints` wrapped around the transcript -- that is a `SubcomposeLayout`, and the IME animation changes the visible height on every frame of its slide, so the entire transcript was being subcomposed again for each of them. The same number read in the layout phase, from the scroll container's own measurement, makes it a relayout instead, and the rows keep the measurements they already have. Checked on the emulator: the newest message sits against the composer with the keyboard up, and there is no gap. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
2b24362cc4 |
Work out the retained window once a frame, not once a row
The same mistake as the draw-phase culling, moved into another phase. Each row held a `derivedStateOf` over the scroll position to decide whether it stayed built. That reads correctly and costs the same shape: every one of those derived states is invalidated by every scroll frame and has to be re-evaluated to learn whether its answer changed, so the per-frame work grew with the number of loaded rows again -- this time landing in the recomposition pass, which the platform reports as the frame's animation phase. On a Pixel 9 Pro XL at 153 rows that was 11.6ms at the median, against an 8.3ms frame, and it was the largest term left. The window is now one range that every row reads, recomputed once a frame from one observer. And recomputed lazily: every row reads it, so every change to it disturbs all of them, which is affordable once every couple of screens and is not affordable at row boundaries, where a fling would cross one every few frames. The window is eight screens either side and it moves in steps of two, so the margin absorbs the staleness. Worth naming as a pattern, since this is the third time: a per-row answer to a question about the scroll position is O(rows) per frame wherever it is evaluated -- in draw, in a derived state, anywhere. The question has one answer and it belongs in one place. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
718fb5320c |
Stop the transcript being able to go blank
Scrolling up emptied the screen and only reopening the session brought it back. The guess for an unmeasured row's height was re-derived from a running average each time it was wanted, and that average moves as rows are measured -- so the height a spacer had been *built* at stopped matching the height the running totals were adding up, and the two drifted apart. Once they differed by more than the retain window every row failed the distance test at once, and a transcript of nothing but spacers has nothing left to measure and so nothing to correct itself with. Two changes, and the second is the one that matters. A guess is now made once per row and kept, so it cannot drift from what was built with it. And which rows stay built is decided as a *range* rather than by each row testing itself: the row nearest the viewport is in that range by construction, whatever the arithmetic says, so the worst a mistake here can do is build too few rows or too many. A blank transcript is no longer a state this can reach. That is the shape worth keeping from the bug. Every row answering independently meant one wrong number could stand all of them down together, and the failure was silent, self-sustaining, and looked exactly like the screen having nothing to show. Checked by scrolling sixty swipes to the oldest loaded end and back -- the content holds throughout. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
53735f8df4 |
Let a paged-in row wait its turn instead of standing up with the page
The layers took care of the steady state -- the transcript's content re-recorded fifteen times in a thirty-second scroll, and the frame's draw phase sits at 3.5ms. What is left is entirely spikes, and they are pages landing: one of those fifteen recordings took 99.9ms on its own, with the frame after it unable to start. The cause was a rule that read as caution and was not. A row with no measured height was retained whatever its distance, on the grounds that it had just been paged in and was about to be looked at -- but "no measured height" describes the *whole page*, not the near edge of it, so eight hundred events' worth of markdown was shaped inside the frame the page arrived in. A row that has not been measured now stands in at the average of those that have, so it is placed and judged by distance like every other row and is built when the reader comes near it. The average rather than a constant because these run from a one-line note to a screenful, and the average row in a conversation is a fair guess at the next one. Being wrong is cheap here and self-correcting: an estimate is only ever used above the viewport, and this layout hangs from its far end, so a correction up there moves nothing on screen. Checked by scrolling twenty-five swipes into history and back on a real transcript -- rows stand up as they are reached, and the position does not shift as the guesses are replaced by measurements. Second half of the diagnosis from the ai-app-2-6d session. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
e65c961e1e |
Give every row its own layer, and stop culling during draw
The piece of a lazy list this had not rebuilt was the render node per item. Without one, a row's glyphs are recorded into its parent's display list, and that list is re-recorded every frame the parent is invalidated -- which, while the list is scrolling, is every frame. With one, the row is recorded once and afterwards moved by a transform, and the render thread culls the ones off screen itself. The draw-phase culling it replaces could not have won, and the two are mutually exclusive rather than complementary: `onScreen` and `RowWindow` read the scroll position from inside every row's and every block's drawing, and reading a scroll position during draw invalidates the drawing it is in. So the machinery for deciding which rows need not be drawn was re-recording all of them, every frame, to decide it. Keeping both would have bought the cost of the first and none of the second. Measured on the emulator over a thirty-swipe scroll: **one** row display list recorded across 481 frames, against 1,544 recordings for the same gesture before, and the frame's draw phase down from 4.2ms to 2.9ms at the median. Two things this also corrects. `LAYOUT_MEASURE_DURATION` is the *View* hierarchy's pass, and Compose is one view -- `AndroidComposeView. dispatchDraw` calls `measureAndLayout()` before it records, so all of Compose's own measurement, text shaping above all, is reported inside DRAW_DURATION. Every "layout is 0.0ms, so nothing is being re-measured" reading in this file's history was reading a bucket that never contained it. And the retain window is now load-bearing for a second reason: a retained display list costs memory, so bounding what is retained bounds that too. Found by the ai-app-2-6d session reading the render reports against this code; the diagnosis and the ordering are theirs. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
5b9941cfaa |
Tell Compose the transcript's rows are values, so it can leave them alone
A page of history landing recomposed every loaded row and everything inside it -- 701 compositions for 148 rows in one of Iris's scrolls, which is four rebuilds of the whole visible transcript, markdown and all. That is what her `waited` at the 90th percentile was: the frame could not start because the thread was rebuilding rows whose content had not changed. Nothing was stopping Compose skipping them except that it could not prove it was safe to. Stability is inferred from a class's fields and a `List` field makes it assume the worst, so `TranscriptItem`, `TranscriptRow` and `ParsedReplies` were all treated as things that might change underneath a composable at any moment. They are not: a row is rebuilt from the transcript rather than edited, two rows describing the same events are equal, and the parse cache is keyed on the text it parsed. Saying so is the whole change. Measured on the emulator over the same scroll: 139 row compositions, and **15** of them rebuilt the message inside. The other 124 skipped straight past the markdown, which is where the cost was. The transcript's own draw came down with it, from 3.1ms mean to 1.2ms. The promise these annotations make has to stay true -- nothing described by them is mutated after it is built. It is not today, and the note on `TranscriptRow` says so where somebody adding a field will read it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
76c51117d6 |
Parse ahead on two threads, not on every core
The frame is now failing to *start* rather than taking too long once it has: 21ms of `waited` at the 90th percentile against a median frame of 8.8ms that is inside budget. What is holding it up arrives with a page of history -- 1.5 seconds of markdown parsed in a twelve second scroll, all of it work nobody is waiting for. Off the composing thread was the right call and it is not the same as free. The default dispatcher sizes itself to the machine, which is right for work somebody is waiting on: a page's worth of parses takes every core, and the thread that draws the frame queues behind one of them. Two threads, and a yield between messages, leaves the phone somewhere to run the frame. Nothing here makes the parsing faster, and it should not: the whole point of doing it ahead is that its duration does not matter. What matters is that it stops being in the way. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
cc41a03746 |
Keep eight screens of rows built, and stand the rest down
Iris's readings made the shape unarguable: smooth with one page loaded, a step worse at the next, worse again at the one after, and 62.7ms at the median by 390 rows -- with layout at 0.1ms and the GPU at 1.8ms the whole time. The cost was following what was *loaded* rather than what was on screen, and keeping every row built is the only thing in here that does that. So the window is bounded. Eight screens either side stay fully built, which is about sixteen times what a lazy list keeps: everything somebody has just read is still there, and only a deliberate journey back through the conversation pays to rebuild anything. That was the point of retaining rows and it survives; what does not survive is retaining all of them. A row outside the window is replaced by a spacer of the height it was last measured at, so the transcript's total height is unchanged and nothing under the reader moves. A row that has never been measured has no height to stand in for it and is kept whatever its distance, which is exactly the row that has just been paged in. Each row decides for itself, in a composable of its own, from a derived state -- so a row hears about the scroll only when its own answer changes. Read from the list's body instead, every row would recompose whenever any row crossed the edge. On the emulator, same transcript and same gesture, the frame's draw phase goes from 2.7ms at the median to 0.4ms. Iris's own timing reading from before this says where the rest of her frame goes: the transcript's whole draw is 3.0ms mean against a 4.2ms median draw phase, so the median frame was already close to budget and what is left is the tail -- 20ms of waiting and 2.3 seconds of parsing in bursts, both of which arrive with a page. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
bdc62b6442 |
Time the draw phase, since counting it stopped being informative
The counters did their job and then ran out: 4,090 rows drawn against 94,815 skipped, the tallest drawn block down from 36,982px to 1,765px, and drawing still took 30.9ms at the median. Almost nothing is being recorded and recording is still the expensive phase, so the next question is not "how much" but "where" -- and there are only two answers left. Either the little that is drawn is somehow costly, or the time is not inside the transcript at all and every count above is beside the point. So the draw is timed at three levels that nest: the whole transcript, one row, one block. If the transcript's own figure is most of the frame's draw, the cost is ours and the rows and blocks say which. If it is a fraction of it, the frame is being spent somewhere this has not been looking. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
981856e046 |
Draw a reply one block at a time, and skip the blocks off screen
A reply's display list holds every glyph of it and is re-recorded whenever drawing is invalidated, so one long message costs as much to draw as a hundred short ones and skipping the rows around it cannot help while it is the one on screen. That was the whole of the remaining draw cost: 97% of rows correctly skipped, and the tallest one still drawn was 36,982px -- about twenty-five screens in a single message. So a message is cut into its top-level blocks and each is drawn, or not, on its own. The cut comes from the parser's own boundaries rather than from a line scanner looking for blank lines, which is what makes it safe: a heading, a table, a fenced block and a list are each one node whatever is inside them, so a loose list does not become five one-item lists and a fence is never split down the middle. Checked against a real reply whose list items are separated by blank lines -- it still draws as one list with its bullets aligned, which is the case a blank-line split gets wrong. It bounds parsing too, which was the other symptom in the same reading: one message took 1.4 seconds to parse as a single unit, and a block is a paragraph. The row and the message each know half of where a block is, so they meet at an interface declared where it is used: the list supplies the row's position, the message supplies the block's offset inside it, and drawing a message does not have to know it is inside a transcript. What this cannot divide is a single node, and a long fenced code block is one -- so the report now also carries the tallest *drawn block*, which is the number that says whether splitting bounded anything. This emulator's tallest row is one such fence, which is why its own figures do not move. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
dffa365e54 |
Count what the transcript draws, not just how long drawing took
"Drawing costs too much" and "drawing was skipped and still costs too much" need opposite fixes, and a millisecond figure cannot tell them apart. So the report now carries how many rows were drawn, how many were skipped, and the height of the tallest one that was drawn. The first reading answers it. Skipping works -- 854 rows drawn against 24,078 skipped, so 97% of the transcript is correctly not being recorded -- and the tallest row that *is* drawn is 36,982px. One assistant message about twenty-five screens tall, whose display list holds every glyph of it, re-recorded whenever drawing is invalidated. A row that size is a hundred short rows as far as the draw phase is concerned, and no amount of skipping its neighbours helps while it is the one on screen. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
f9432b69cc |
Keep every row laid out, but only draw the ones near the screen
Iris's Pixel 9 Pro XL said which half was costing, and it was not the half either of us was looking at. With 138 rows loaded: layout 0.0ms at the median, GPU 1.9ms, and **draw 13.6ms** against a 120Hz budget of 8.3ms. Nothing was being re-measured and the phone's rasteriser was idle; the UI thread was recording draw commands for a transcript that was almost entirely off screen. Retaining rows was right and stays. What does not follow the same rule is drawing: the draw pass walks the whole tree, so a list that keeps every row alive records every row every frame, and that cost grows with each page of history -- which is exactly what "worse afterwards" was. Composition and measurement are what must not be thrown away, because they are what has to be rebuilt from nothing when the reader comes back. A display list is rebuilt from a layout that is still there. So each row skips its own draw when it is off screen, by more than a screen's margin either side. The check reads the scroll position from the draw phase, so moving the list invalidates drawing and nothing else, and it is a lookup rather than a sum -- the running totals are rebuilt at most once a frame and only after something has actually changed height, since adding them up per row per lookup would have made the fix quadratic in the thing it was fixing. Also reports the three frame phases that were missing, which is why the phases on that reading did not add up to the total: the time a frame spent waiting for the UI thread to be free, handling input, and running animations. About 15ms of the 30.6 was in that gap and unattributed. On the emulator, same transcript, before and after: draw p50 3.9ms -> 2.5ms, p90 7.3ms -> 4.2ms, p99 48.6ms -> 5.2ms. The saving there is small because only 60 rows were loaded; it is proportional to how much is off screen, and on the phone that was 95,000px of content against a 1,474px viewport. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
3e653c4797 |
Add a button that copies what this session costs to draw
A speedometer left of the usage chart, because it is the same kind of thing about the app that the chart is about the conversation. It copies rather than opens: a screenful of timings read on the phone is a screenful nobody can act on, and what it produces is a message to whoever is looking at the code. It carries both halves of the question, because "the scroll is laggy" has two causes and one appearance. From the platform, each frame's cost split into its phases -- if measuring and drawing are small and the total is large, the time is going into rasterising and no amount of doing less work per row will move it. From the app, counts and timings of the work the transcript actually does: rows composed, replies parsed on the composing thread rather than ahead of it, runs grouped. Counts rather than frame times are the point of the second half. This emulator scrolls at the same 21ms median as the stock Settings app, so every app-level cost here is under the floor of what it can measure and no number taken in it says anything about a 120Hz phone. How many times a row was composed is the same number on any machine, and it is the one that says whether the work follows what is on screen or everything ever loaded. Pressing it empties both, so two presses measure two separate stretches of scrolling rather than one and then the same one again. The first reading from the emulator already says something: measure and layout are 0.0ms at the median, so nothing is being re-measured, and the cost is 3.9ms of recording the draw against 10.9ms of GPU. It also shows 60 loaded rows composing 177 times across three page loads -- every row recomposing whenever a page lands -- which is the next thing to look at if the phone says the work is ours. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
f0f6d2b099 | Merge branch 'main' of git.arirex.me:iris/ai-app | ||
|
|
fd250378a0 |
Attach the restore hook only while there is a restore
`onPlaced` on the transcript content is the one callback left that would run every frame, and on all but the two frames of a restore it looks at a null and returns. Attaching it only while a position is waiting says that in the modifier chain rather than in a branch inside it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
d829a77ddf | Merge branch 'main' of git.arirex.me:iris/ai-app | ||
|
|
c6bb337c60 |
Take the icon square up to the touch target, and let it fill the row
Asked for: more space between the header icons, and a box tall enough to fill the header rather than sitting inside it. 48dp does both -- the marks stand 31dp apart and 31dp from the screen edge, measured on the emulator, where the 40dp square had them 23dp apart and 27dp from it. It is also the platform's minimum touch target, which the previous size was short of, and it is taller than any header's text: the session header's row now takes its height from the button and needs no vertical padding of its own, for the same reason the rows add no gap between two buttons. The ring around the mark is still the only spacing rule; every number here moved because it did. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
19072a7e96 | Merge branch 'main' of git.arirex.me:iris/ai-app | ||
|
|
f22c96abe2 |
Stop paying a callback per row per frame, and drop the reply cap
Holding every row alive turned two per-row callbacks from a cost paid by
whatever was on screen into a cost paid by everything ever loaded, which
is why loading one more page of history made the whole transcript lag
and each page after made it worse. Both fire for every registered node
on every layout pass:
- an `onGloballyPositioned` per row, kept so a tap could be told which
half of the row it landed in
- an `onGloballyPositioned` per clickable inside a row, for the same
reason, so a tool group paid several
Neither is needed. One gesture detector on the whole list records where
a touch went down in the content's own coordinates, and the row's own
start is added up from heights when somebody actually taps -- so the
transcript keeps a height per row, which changes when the row does, in
place of a position per row, which is wrong the moment anything scrolls.
Positions for the saved-scroll anchor come from the same sum, and the
anchor is applied once per layout from the content rather than once per
row.
The reply cap is gone at Iris's request; she wasn't seeing it and does
not want it for now. It was doing measurable work -- p50 rises from 16ms
to 20ms on the emulator with it removed -- so it is worth knowing where
to look if long history feels heavy.
That 4ms is the only figure here I trust. This emulator's own noise
between identical repeats is larger than the difference the callback
change makes (identical builds measured 3.5% and 9.3% of frames over
budget), and its p50 of 16ms is already past a 120Hz budget, so it
cannot rank any of this the way the phone will. The callbacks are gone
because they are O(rows) per frame by construction and the list no
longer bounds how many rows there are -- not because a number here says
so.
Checked on the emulator against a real 1,200-event transcript: tapping a
collapsed row expands it with its top edge held exactly still, and a
scroll position comes back pixel-identical after leaving the session and
returning.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
4dd1e6ed8c |
Drop the import the icon row no longer needs
ktfmt's, after the merge: nothing in the session screen arranges a row by hand any more. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
5737c06654 | Merge branch 'main' of git.arirex.me:iris/ai-app | ||
|
|
c5fecb296f |
Give an icon button its own ring of space, and let that be the spacing
Every gap around a header icon now comes out of the button's own padding: one ring to the screen edge, two where a button meets its neighbour. The box was the size of the mark (28dp) and the separation was bolted on beside it, so on the session header the two marks stood 31dp apart while the outer one was 14dp from the edge of the screen -- a pair that acts on one screen reading as two unrelated marks, one of them falling off it. Measured on the emulator at 40dp: 23dp between the marks, 27dp to the edge, and the same ring above and below. The mark was also not square, which is why an arrow looked taller than it was wide. `Glyph` is a `Text`, and it inherited the body style's 24sp line height around a 17sp mark; this font's ascent and descent add up to exactly one em, so a line height of the point size is the square the glyph draws in. That leading is also what a button's padding had to be measured through. 40dp is Material's own icon-button state layer, and it is what the pressed-state ripple draws -- at 28dp that circle was inscribed in the mark's corners, and beside a title it arrived at the first letter. The touch target comes up with it, from well under the platform's 48dp minimum to within 8dp of it; going the rest of the way would put the marks back 31dp apart. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
ee1c493559 |
Hold every loaded row, and page by pixels rather than by rows
The transcript is a plain Column scrolled in reverse instead of a LazyColumn. Nothing about Compose was re-measuring text that had already been drawn -- a node that is still alive and whose constraints have not changed skips measurement outright, in MeasurePassDelegate.remeasure. What was throwing that away was disposal: a lazy list drops a row the moment it leaves the viewport, and the markdown tree, the measured lines and the cached paragraph go with it. Keeping the rows is the fix, and it is what a browser-based client does that we were not. Two scroll corrections go with it, each of which had a comment explaining a way it had been seen to fire at the wrong moment. The content now hangs from its newest end, so a page of older history extends the far end and moves nothing on screen, and an arriving message extends the end the viewport is already pinned to. Following the newest message is no longer an effect that notices and corrects; it is where the content is. The same goes for the keyboard opening, which was the case that used to get missed. Paging asks its question in pixels of scroll -- how far can the reader keep going before they run out -- which is what it was always about. Rows were the wrong unit twice: a fixed count of them is a distance only by accident, and counting screenfuls of rows fixed the size of that mistake without fixing its kind. A saved position is resolved to the row that now holds that seq before the layout is asked to put it back. The events behind a row regroup between the save and the reopen, so the seq that was a row's first is often no longer any row's first, and handing the layout the saved seq named a row that did not exist -- the position was never applied and the session opened at the newest end. Verified on the emulator against a real 1,200-event transcript: the position survives leaving and reopening, jump-to-latest arrives and stays followed, and the newest message holds its place against the composer as the keyboard opens and the draft grows. Scrolling measures the same as the lazy version did (2.6% vs 2.0% janky, p50 16ms, p90 21ms, no slow UI-thread frames either way) -- the cap and the off-thread parse had already taken that cost out, so this change is about what the list can no longer do wrong rather than about frames. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
3095052df0 | Merge branch 'main' of git.arirex.me:iris/ai-app | ||
|
|
592038aa54 |
Draw a long reply to a couple of screens, with the rest a press away
Iris approved the cap and set the one exception: never the latest
response. That one is being read as it arrives, often still arriving, and
putting "show the rest" under the turn somebody is waiting for hides the
answer they are waiting for. Everything behind it is history, which is
what this is for.
The limit is on the **source**, not on the height, and that is the whole
point. Clipping a laid-out row to a height saves nothing -- Compose
measures the text and throws the overflow away, so the line-breaking has
already happened -- while cutting the string before it is parsed stops
the work being done at all, the parse included. Four thousand characters
is about two screenfuls, with a thousand of slack so that a reply barely
over the limit is not given a control worth nothing to anybody.
The cut is at a line ending, because a markdown source cut mid-line is a
different document: half a heading marker, a list item with no bullet, a
link whose closing bracket was dropped. A fence left open is closed,
which is the one case a line boundary does not cover -- an unterminated
``` swallows the rest of the reply into a code block, so the truncation
would change how the part still on screen is *drawn* rather than only how
much of it there is, and a reader cannot tell that from the reply
genuinely having been code. Verified against a reply whose cut lands
inside a 400-line Rust block: the block renders as a block, closes, and
the control sits below it in ordinary text.
`markdownIn` now names both forms, because which one a row draws is not
decided there -- so pressing the control is a cache hit rather than a
50ms parse on the frame it happens in, and neither form is a key that no
row ever looks up.
Measured on the emulator, now that it renders on the GPU and its frame
numbers mean something -- two flings over replies of the same size in the
same session, one capped and one not:
uncapped 62,310 capped 54,990 (4,000 drawn)
janky (legacy) 58.20% 11.33%
slow UI-thread frames 5 1
p50 23ms 16ms
p90 30ms 24ms
The modern "Janky frames" figure is 4.04% against 3.59% -- near identical,
because both stay inside the compositor's deadline on this emulator. The
win is UI-thread work, which is what was aimed at, and the legacy metric
is the one that counts it.
Expanded is remembered by the screen, beside the other expansion sets, so
scrolling away and back does not shut something deliberately opened.
Known and not fixed: the floating jump-to-latest control is centred at the
bottom of the list and overlaps this one's label whenever a capped reply's
foot lands in that band. It is the overlay's pre-existing disregard for
content -- long replies have always had text under it -- but a control
there is worse than prose, and it is now common rather than incidental.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
4edc96de09 |
Tell "this window is not running" from "we could not read the time"
The five-hour window's reset time is absent between blocks, because the window is anchored to the block it started in and there is nothing to reset until one is running. Measured against a live response: the reset came back as exactly five hours after work resumed, and the weekly windows in the same response carried the identical microsecond, so both are computed from one now() at request time. The weeklies always have a reset because a week is always running -- which is why only the five-hour row looked wrong. `remainingUntil` returned null for that and for a timestamp it could not parse, so the session bar announced "reset time unknown" about a machine behaving perfectly, on the one row somebody reads before starting something big. The usage dialog, looking at the same field, drew nothing at all and printed a raw ISO string when a parse did fail. One missing value, two rules, and neither of them right. `WindowEnd` names the three answers and both callers go through it. A window that is not running shows its percentage and no countdown, in the bar as well as the dialog; an unreadable timestamp says so in words rather than showing itself. Looked at both on the emulator, the second by making the server drop the field: 15% with "4h 43m left" when a block is running, "13%" alone when none is, and the dialog's five-hour row with no reset line beside weeklies that have one. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
3da0f2e2f6 |
Replace a claim measured through the software rasteriser
Two places said the history cushion's evidence was a swipe that moved nothing for 689ms. That reading came from `ui-trace show` on a row taller than the viewport, which reports clipped bounds and so prints "nothing moved" for a list that is scrolling fine -- and it was taken on the software renderer besides. The finding it supported is sound and has a better witness: counted at the server, ten swipes asked for ten pages before the change and three after, which does not depend on how anything renders. Also records what the screen actually costs now, measured on the GPU emulator: 5.2-5.9% janky frames and 0-2 slow UI-thread frames flinging fast, against 3.3% for the stock Settings app on the same device. With the note that matters for repeating it -- let the screen settle before resetting gfxinfo, because the first seconds after opening a session are every row being composed for the first time and read three times worse. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
0dc248b9e4 |
Fetch a restore's history in one request, and let RUST_LOG through
Now that a scroll anchor is a seq, the restore knows exactly how far back
it has to reach, so it asks for that span in one request instead of
walking there a page at a time. ai-app-2 suggested it; the arithmetic is
theirs. `read_window` counts lines and a transcript numbers them one per
seq, so the distance to the anchor is the number of events to ask for --
and were seqs ever sparse, that difference is larger than the count,
which overshoots into older history rather than stopping short.
Capped at [RESTORE_PAGE_MAX], and the loop already there is what makes
the cap safe: a span past it comes back in several requests rather than
one, which is what every restore did until now. The bytes are the same
either way -- every row between the anchor and the newest end has to be
loaded for the list to be able to count to it -- so this only trades
round trips against response size.
Measured on a 16,133-event session, restoring to seq 2000 (14,133 events
back, the extreme case): **19 requests before, 5 after.** The realistic
case, a couple of thousand events back, is 4 before and 2 after. At
`--delay 150` the deep one puts the row on screen at 5.7s and the
moderate one at 2.5s, and in both the row does not move once it lands.
Also: `RUST_LOG` did nothing. `with_env_filter("info")` is a fixed
directive that never reads the environment, so the per-request page
diagnostics AGENTS.md tells you to turn on with `RUST_LOG=ai_server=debug`
printed nothing at all -- which reads as the code you are instrumenting
being wrong rather than as the switch being disconnected. It is a
fallback now, so the default is still `info`. Those diagnostics are what
the request counts above were measured with.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
c5ecc63995 | Merge branch 'main' of git.arirex.me:iris/ai-app | ||
|
|
b3070f16ff |
Keep the transcript ahead of the reader, and off the thread that draws
Three things, all of them the same complaint: scrolling back through a long session stalls. **The fold was on the main thread.** Only `fetchTranscript` was inside `withContext(Dispatchers.IO)`; the fold loop that turns a page into rows ran on the caller's dispatcher, which is Main. `foldEvent` returns a new list per event, so a page is that many copies of a list growing to that length -- about three hundred thousand element copies -- run in the middle of the scroll that asked for it. Affordable at 80 events per page and not at 800. **`warm` scanned the whole transcript on the calling thread.** Only `replies.warm` was off it; the `markdownIn` split that decides *what* to parse ran before the hop, over every assistant message loaded, on every page. The scan grew with the conversation while the work it found stayed one page's worth. **The cushion was eight rows, which is not a distance.** A row is anything from one line to a page: on a tool-heavy transcript eight rows is less than one screen, so the reader reached the end of what was loaded on every swipe and waited a round trip standing there. It is three screenfuls now, measured from what is actually on screen. On the emulator against a 24,000-event transcript that is 3 page fetches for 10 swipes rather than 10. Also a spinner while a chat loads. Nothing is drawn while the newest page is in flight or a saved position is being put back, and a blank page is what this screen otherwise means by "there is nothing here" -- so the state that does not know needed its own appearance. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
e76617e210 |
Draw verbatim text on its own dark surface
A command, a tool's output and a code block in a reply are the one thing on this screen that is not somebody's prose, and they now say so: Mocha's Crust, which sits below Base, so the same colour is one clear step down both on the page where a reply is drawn and on the card where a tool call is. The renderer's code background was `surfaceVariant`, which is exactly a card's own fill -- a fenced block inside a tool call had no background at all, and one in a reply read as a step *up* out of the page. Tool output takes the monospace face with it. It is column-aligned far more often than it is prose -- a listing, a diff, a table of numbers -- and a proportional font silently destroys the alignment that carried the meaning. `RawBlock` is a composable rather than a modifier because the inset is part of it: monospace text against the edge of a tinted block reads as clipping. A call with neither a subject nor any other field draws nothing at all rather than an empty tinted rectangle. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
75e02b3857 |
Anchor a saved scroll position on a seq, not on a row's name
ai-app-2 found the reason a session Iris had scrolled in took a while to open, and it is a defect in what I shipped this morning. A `ScrollAnchor` stored the list's own row key. That key looks stable and is not: a tool row is named after its run, `joinPages` gives a run the name its *newest* half carries, and the newest half is whatever the newest page happened to start with. The newest page is the last eighty events, so the name is fixed for an idle session -- which is why this passed on the emulator -- and changes the moment the session says anything. An active session therefore renamed its tool runs on every reopen, the anchor was never found, and the restore paged backwards until there was no history left, i.e. to the first event of the conversation. So the anchor is a sequence number now, taken from the oldest event behind the row. That is the server's own numbering: assigned once, never moved, the same on every device. The restore finds the last row starting at or before it, so it still lands correctly when grouping has changed underneath it -- a run folding differently, two halves of a reply becoming one message -- and the page-back loop gains an exact stop, because `oldestSeq` walks strictly backwards and a row always matches once the window passes the anchor. It cannot run to seq 1 any more. `TranscriptRow` now carries `startSeq` beside `key`, and the two are documented against each other: `key` is the list's identity and a display decision, `startSeq` is a place in the conversation. Anything that has to point at a place and find it later uses the second. Two more found while testing this, both the same shape -- a listener waiting for a scroll to *end* never runs when the list moves inside one frame: - The anchor was read from `rows`, which a `LaunchedEffect(listState)` captures from the first composition, where it is empty. Every scroll saved a null anchor, indistinguishable from being left at the newest end, so the position was silently never recorded at all. It reads the live state now, as the paging code beside it already warned it must. The anchor is also written from the settled *position* rather than from the scroll flag, because Jump to latest snaps within one frame: it left the old position recorded, so pressing the control that means "take me to the end" and coming back put the reader where they had been. - Jump to latest did not set `followTail` either, which is the same hole in the sibling and predates today: the view landed at the bottom with following switched off, and the next message did not bring it along. The button states what it means now. `followTail` itself stays on the scroll settle, deliberately -- a keyed list moves its own anchor when a row arrives, so the position reports itself as scrolled back for a frame every time a message lands, which is the whole reason that value is remembered rather than read. Verified against an echo session grown to 1,327 events between saving the position and reopening, so the newest window slid and the runs were renamed: the same row is on screen before and after, it appears 419ms in and never moves, and a 200-piece streamed reply still follows the bottom after a jump. The threading of `loadOlderPage`'s fold is ai-app-2's, not touched here. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
61d2c78afe |
Keep the reader's place: in a group, in a compaction, and in the list
Five things Iris asked for, all about the transcript screen holding still around whoever is reading it. A tool call opened on its own stayed open when a second call in the same run turns it into a group. Watching a Bash call and having the session make another one used to shut the card being read and fold it behind "Called 2 tools" -- the reader lost their place because something else happened. The transition is noticed once, at the moment a run first becomes a group; after that the group's own toggle owns it, so shutting a group whose inner call is still expanded does not re-open it. The compaction clock is taken from the `compacting` status event's own timestamp rather than from this device noticing one, so it survives leaving the session and coming back -- it used to disappear, because the only thing that knew when the compaction started was a screen that had been disposed. The server timestamps every transcript line, so this is still a measurement; it is compared against the phone's wall clock, which is the same comparison a session's "last active" already makes. Session settings are a dialog over the session instead of a screen below it. Two controls did not warrant a page transition and a back stack, and the thing they change was hidden while they were on screen. Captions are gone -- each control is a labelled noun -- and "Notify me" is "Notifications" with a bell beside it (`md-bell`, added to the committed Nerd Fonts subset). Failures keep their words, since those are what a reader cannot work out by looking. Tool groups are rounded like every other card, their foot bar is the same height as their heading (both derived from the heading's own line height, so the pair cannot drift), and the calls inside are a connected stack: square where they face a neighbour, rounded on the outside, with a small gap so the join reads as a join. Scroll position is persistent on the device, per session, keyed by the row rather than by an index -- an index means nothing across a reopen, where the transcript is fetched newest-first. Reopening pages backwards until that row is loaded *and* has something older behind it, because the oldest loaded row is a half-row that grows when the page behind it arrives; anchoring into one landed a screen and a half out. The list draws nothing until the position lands, so there is no frame in which the transcript is somewhere other than where it was left. Two things found on the way. `snapshotFlow`'s first emission is the state before anybody has touched the list, and reading it as a scroll that had just ended at the newest end wiped every saved position on the way in. And backwards pages now ask for 800 events rather than 80: ai-app-2 measured a real transcript at 2,426 events for seven assistant messages, so a page of eighty is a fifth of one row and filling the lookahead took about thirty sequential round trips -- seconds of a list that will not move, over the tunnel. `/tools [n] [gap]` in the echo driver takes seconds between calls, which is what makes a run grow slowly enough for somebody to have opened one of its calls first. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
d981d63bcc | Merge branch 'main' of git.arirex.me:iris/ai-app | ||
|
|
fd616361eb |
Hold the transcript still under a scroll, and parse replies before drawing
Both faults needed a real conversation to see, so `app/debug-transcript.sh`
now puts one on the emulator: it copies a Claude Code transcript into /tmp,
gives an ai-server a HOME of its own so the import can only see the copy, and
enrols the app against it. The transcript itself never enters this repository
-- those files hold whatever was said, read and written in a session. Beside
it, `ai-server --delay MS` holds every response back, because a phone's
requests take tens to hundreds of milliseconds over the tunnel and several
faults live entirely in what the app does while one is outstanding.
**Scrolling up threw the reader back to the newest end, once.** `followTail`
is deliberately a remembered answer, rewritten only when a scroll settles, so
for the whole of a fling it still reports the newest end -- where the reader
was when they threw it. A page of history landing during that fling is a
change in the item count, and the correction written for an insertion at the
newest end fired for one at the oldest. Captured on the emulator:
scrolling=true atNewest=false followTail=true
history: START last=20 total=28
history: page of 80 events -> rows now 36
countChanged count=36 followTail=true scrolling=true
>>> scrollToItem(0) SNAP
It could happen only once, which is what made it look arbitrary rather than
mechanical: the snap settles the scroll at the newest end, so the next fling
gets far enough to settle away from it, and from then on `followTail` is
false. So the list is no longer moved while a scroll is running, which is a
rule of its own rather than a refinement of that condition -- and skipping
the correction outright is right rather than merely safe, because the count
can only grow at the newest end while the reader is already there, `record`
holding everything else until they come back.
**A page of history stalled the frame it appeared in.** Parsing is the
expensive half of drawing a reply and costs in proportion to what was
written: against this transcript one message took 51ms and several took
10-25ms, where the synthetic replies this was tuned on took 4.6ms. So each
page's replies are parsed on a background thread as the page arrives --
after the join, since a boundary falling through a reply leaves a message
made of both halves whose text has existed for no time at all, and warming
the page alone warmed the two halves and missed the one thing drawn. A row
with no answer waiting still parses inline: a row measured at nothing before
it is measured at its real height collapses the transcript above it. Misses
are not stored, so a reply still streaming cannot fill the map with copies
of itself on the way to being finished.
Measured over the same twelve flings: 13.5ms average per composed reply
before, 7us after, the remaining parse being one message at session open.
Verified with ui-trace at 1kHz: with a page landing mid-drag the suppression
fires and the row the reader is on moves monotonically down, 266 -> 1063,
with no step backwards; at rest 0 of 65 elements move. 86 server tests pass,
ktfmt/lint/clippy clean.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
257f4c85c1 | Merge branch 'main' of git.arirex.me:iris/ai-app | ||
|
|
50f9b956d9 |
Check "exited" against the process record before believing it
A session adopted at a backend start keeps the transcript's last status, so one whose process had been reported gone and was then found again read as `exited` while its CLI was running. `exited` is the word that draws the phone's Start button and lets `start_session` build a driver, so Start was accepted every time it was pressed -- and since starting replaces the driver without retiring the old one, each press left another reader on the same process. Every line the CLI wrote was then translated once per reader: three presses put three interleaved copies of one reply on screen, which is what it was reported as. So `exited` is now checked against `session::process`, the one authority on whether a process exists, in `launch` and again in `start_session`. A record that is not known to be dead makes it false, and what replaces it is `unknown` -- there is a process, and nothing here has heard from it, which is the answer `status_of_unlaunched` already gave to the same question. The correction goes out through the sink rather than into the manager's view alone, or the list and the session screen would disagree about it in the way this same button did a commit ago. A driver that `start_session` replaces now gets `Driver::detach`, which already existed for the backend going away and is the whole of what a driver whose process has exited is owed. On the phone the process button is disabled while its own request is in flight, so a second press cannot be decided against a status the first has not changed yet. That is a courtesy rather than the fix; the server refuses it either way, because a phone that has lost the stream cannot be relied on to know. Verified against a stand-in CLI, with the state forced by hand: before, three Starts returned 204 and left four readers on one process and the status still `exited`; after, the session reports `unknown` on both surfaces and all three are refused. Then driven on the emulator -- Stop, Start, Stop, Start alternated correctly with one process at a time, and the list, the transcript and the record all agree. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
1ed6e29bc6 | Merge branch 'main' of git.arirex.me:iris/ai-app | ||
|
|
62c22e38b7 |
Parse a reply's markdown once per composition, not once per delta
Scrolling was laggy, and measuring said where the time went. Instrumenting the transcript's main-thread work on the emulator, against `/stream 200`: markdown parsing ran fifty-eight times in three seconds -- once per streamed delta, each one re-parsing the whole message the reply had grown into -- for 49-78ms of main-thread work per three seconds, with single parses reaching 7.3ms. That is most of a frame at 60Hz and more than a whole one at 120. Everything else the transcript does per event was under a tenth of it. So only the first parse stays on the composing thread. That one has to: the renderer's asynchronous path draws an empty loading slot until its result arrives, which measures a row at nothing before it is measured at its real height, and the whole transcript above it collapses and springs back. Every parse after the first is the same row growing, and there is a previous parse to keep drawing until the new one lands -- so those go to a background thread and no frame is ever without a height. What is on screen stays a real prefix of the reply rather than a guess at it; it is simply one parse behind. The same measurement found `loaded` costing an ArrayList copy per event, and nothing reading it. It recorded every event the screen had ever seen against the possibility that a page arriving in front of them would need the events themselves to stitch on -- but `joinPages` heals the boundary from the folded rows and has since it was written, so this was a list that only ever grew. Checked on the emulator with ui-trace at 1kHz. Streaming at the newest end: the row's bottom edge holds at y=1940 while it grows upward, and the header, status row and composer do not move for six seconds. Scrolled back with a reply streaming: nothing moves at all, 0 of 45 elements over five seconds. Scrolling a mixed transcript: rows keep a constant height as they translate, so none of them arrives blank and fills in afterwards. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |