Files
ai-app/AGENTS.md
irisandClaude Fable 5.1 46d3a6fd41 docs: record the streaming-rebuild fix, its numbers, and the new scripts
RUST.md's P0 box gets the fix, the before/after streaming-phase numbers
(with their caveats), the build-apk.sh/run-bench.sh scripts, and what the
dropout-fix pass's three remaining verifications are blocked on (the
sandbox ai-server currently fails to build, unrelated to this change).
IRIS.md gets the List::replace_back/clear and TranscriptScreen::apply
API entries. AGENTS.md's rigs section gets one sentence on each script.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-05 22:14:35 -04:00

35 KiB
Raw Permalink Blame History

ai-app

A phone interface to AI coding sessions (Claude Code and llama.cpp), replacing the Claude app for daily use. Rust/Axum backend on the desktop, Kotlin/Compose Android app, WireGuard + pinned self-signed TLS + bearer token between them.

docs/PLAN.md is the design source of truth — every decision with its date, its rationale, and what was rejected. Read it before changing anything structural, and update it in place when a decision changes rather than letting this file and the plan become two versions of the truth. This file is the working notes layer: layout, commands, rigs, and things that have bitten. The design and working documents live under docs/ — everything except this file and CLAUDE.md, which stay at the root because that is where Claude Code and other agent harnesses look for them.

The central design point, worth not undoing by accident: a session is a child process, translated into one common event model. A new session type is a new driver — never a session-type branch in shared code (routes, transcript, app screens).

Layout

Mirrors ../dev-updater deliberately: same stack (axum 0.8 + axum-server/rustls, tokio, clap; Kotlin 2.4.x + Compose Multiplatform, single :androidApp module), same cert scheme, same registry pattern. Read dev-updater's README.md and AGENTS.md before diverging from them. Module-by-module intent is in docs/PLAN.md's "Backend layout".

  • server/ — the Rust backend (ai-server). routes.rs's module doc comment is the HTTP table and the surface's source of truth.
  • app/ — the Compose app, package com.example.aiapp, label "AI Sessions". AppRoot.kt is the navigation when; MainScreen.kt the root's four tabs (sessions, import, models, setups); Api.kt/EventStream.kt the REST + SSE clients; Events.kt the event model mirror; ServerConfig.kt settings and the Keystore-sealed token.
  • wg-app-link/ — a git submodule shared with dev-updater: the pinned CA and leaf (certs), QR enrollment and the bearer token (enroll), wg0 binding and the certificate's SANs (netif), owner-only files (private), and the RON house rules (format). Clone with --recurse-submodules, or git submodule update --init in an existing checkout — server/ will not build without it, since it is a path dependency, which is what keeps the two projects version-locked to the commit this repo pins. What deliberately did not move is the API surface and the config schema: routes, drivers, sessions and setups are what makes this project itself.
  • docs/ — every design and working document except this file and CLAUDE.md:
    • docs/EXPLORER.md — the file explorer's design (server/src/files.rs and FilesScreen.kt / FileViewer.kt / FileEditor.kt).
    • docs/TRANSCRIPT_CACHE.md — the phone's copy of what it has been sent. Read it before touching TranscriptCache.kt, TranscriptSource.kt, or the opening and stream effects in SessionScreen.kt.
    • docs/TODO.md — the working list.
    • docs/RUST.md — the plan for moving the app to Rust (on the rustify branch of the ai-app-2 clone): what has to be reproduced, the framework decision, and the ordered experiments with their pass conditions. Read it before touching anything under that branch.
    • docs/IRIS.md, docs/IRIS_TODO.md, docs/DECISIONS.md, docs/LAYOUT.md, docs/TEXTURES.md, docs/CLIENT_CORE.md — iris's own public API log, working list, decisions log, layout/render design, and texture-atlas design, and the client-core crate's design, respectively.
  • .dev-updater.ron — what Dev Updater builds here: the server (run as service: Managed(…), supervised by Dev Updater's own implementation rather than a script kept here) and the APK, in parallel. It points at resources.ron, which is ours rather than Dev Updater's — it names ~/.local/share/ai-app and ~/.config/ai-app so the Uninstall dialog can offer them. Note what deleting the config directory takes with it: the CA under certs, which is the one-way door. Stop on the server card stops the server a phone reaches through the tunnel, so on that phone it stays down until somebody starts it again; Dev Updater reaches it over its own port and is unaffected, which is what makes the button safe to press and easy to regret.

Icons

Nerd Fonts glyphs from a committed subset, not vector assets and not ordinary Unicode. NerdIcons.kt declares each codepoint and app/build-icon-font.sh subsets the font; the two lists have to agree, because a codepoint in the Kotlin that the script did not subset is a glyph that silently isn't there. Rerun the script and commit its output when adding one — it needs network access. md-cog and md-refresh are deliberately the same codepoints dev-updater uses and must not drift from it. The subset is the Mono face, where every glyph is one em square, which is what makes two icon buttons the same width without either being given one — and why GLYPH_SIZE is smaller than it looks like it should be.

Checking your work

  • Server: ./run-tests.sh from the repo root (or cargo test from server/), plus cargo clippy --all-targets and cargo fmt. The build stays warning-clean and rustfmt-clean at the defaults — there is no rustfmt.toml and there should not be one.
  • App: from app/, . ./android-env.sh && ./gradlew :androidApp:ktfmtFormat :androidApp:compileDebugKotlin :androidApp:lintDebug :androidApp:testDebugUnitTest. The unit tests are JVM-only and cover the syntax highlighter, the ANSI parser and the transcript cache — the app's pure logic with no Android in it. Touching anything under BenchFixture.kt, BenchNetwork.kt, BenchRun.kt or the bench build type also needs :androidApp:compileBenchKotlin :androidApp:lintBench — a second build type compiles separately and lint has caught real bugs debug alone never would (see "Android Lint" below).
  • Android Lint is not optional and is not run by a build. It found a crash that had been shipping (java.time on a minSdk-24 app with desugaring off) and later a permission check that silently dropped every notification on Android 12 and below. Fully clean as of 2026-08-31; keep it that way, and suppress with tools:ignore plus a written reason rather than by lowering the bar.
  • Then ./build-apk.sh for the APK to install on a phone through Dev Updater, or ./run-android.sh to build, install and launch on the emulator. The phone gets the release build, signed with a key the script generates once under ~/.config/ai-app/release.jks (never in the repo); ./build-apk.sh debug builds the other variant, and Dev Updater's build modes call the script with exactly that word. Dev Updater lists every variant under build/outputs/apk, so pick release there; a phone still holding the debug build has to uninstall it first, since the two are signed differently.
  • The emulator scripts stay on the debug build. Never read a frame time from one as the app's — a debuggable build runs Compose at a fraction of release speed; the render report says which build it came from.

Running it here

  • Run the server for development with --bind 127.0.0.1. Without it the server binds wg0, which exists here but is unreachable from the emulator (it dials 10.0.2.2). First run prints the enrollment QR/URI with the token. ai-server --enroll-link mints one more device's link while the server keeps running; the server adopts that token on its first use. It is what Dev Updater's Enroll button runs.
  • Point development at a scratch state directory rather than the real one: --config /tmp/…/config.ron --data-dir /tmp/…/sessions --port 8444.
  • The APK pins the CA of the machine that builds it, read at build time from $XDG_CONFIG_HOME/ai-app/certs/ca.pem (AI_APP_CA overrides). So the server must have started once on that machine first — the build stops with that instruction otherwise — and an APK built in this VM only works against a server in this VM.
  • Prefer exercising the server directly over going through the UI: curl --cacert ~/.config/ai-app/certs/ca.pem -H "Authorization: Bearer …" https://127.0.0.1:8443/sessions. The CA is wherever --certs put it — by default under $XDG_CONFIG_HOME, never in the checkout, so a relative certs/ca.pem finds nothing. The emulator app reaches it at https://10.0.2.2:8443; enroll with adb shell "am start -a android.intent.action.VIEW -d 'aiapp://enroll?host=10.0.2.2&port=8443&token=…'".
  • ai-server --delay MS holds every response back. Over the tunnel a phone's requests take tens to hundreds of milliseconds, and several faults live entirely in what the app does while one is outstanding. On a loopback server those windows close before anything can be observed, so the bug looks like it is not there.
  • RUST_LOG=ai_server=debug logs every transcript page with its before, after and what came back, and logs each SSE subscriber's cursor and whether it was continued or reset (stream backlog:). That is the only place "how far had this phone fallen behind" is answerable — the app sees a window arrive and cannot tell.
  • ./test-wg-tunnel.sh up|test|down builds a real tunnel between two network namespaces inside one machine and drives the server through it — a genuine handshake against 10.66.0.1 with pinned TLS, no router or phone involved. That is how to verify the wg0-only posture.

The rigs

Each exists because something was invisible without it.

  • The bench build type and app/bench-fixture/ exist for P0 (RUST.md and DECISIONS.md's 2026-09-05 entries), the phone benchmark gate Iris asked for before porting continues: a deterministic, checked-in synthetic transcript (app/bench-fixture/generate.py, never a real one) that both this app and iris open with no server, so a frame-time comparison measures the renderer rather than the data. ./build-apk.sh bench builds it — own application id (com.example.aiapp.bench) and label ("AI Sessions bench") so it installs beside a real enrollment rather than replacing it. Opening it goes straight to a session screen holding the fixture (no enrollment, no permission prompts) with a "Run benchmark" control beside "Copy" in session settings: it drives the same scroll loop and streaming phase transcript-bench.sh/stream-bench.sh drive over ui-trace, but in-process, since a real phone has no usable system tracing and no agent can drive one (this-machine-android's skill). BenchFixture.kt/BenchNetwork.kt fake the backend by installing a URLStreamHandlerFactory that answers TranscriptSource/EventStream's requests from an in-memory copy of the fixture instead of opening a socket — so the fold, the paging and uniqueItems under test are the screen's real ones, never a shortcut built just for this. The report gains a bench: section (process CPU time, peak RSS, battery current) on every build, empty except when BenchRun.kt filled it in.
  • app/ui-sandbox.sh — a second ai-server with its own $HOME, config and data directory, holding eight invented Claude Code transcripts and a claude that is two lines of shell. That isolation is the point: the import screen lists whatever is in ~/.claude/projects, which in this VM is real agent transcripts, so exercising delete against the ordinary server deletes somebody's conversation and exercising import starts a real --resume on the owner's account. Its port and root derive from the checkout's name, so two checkouts' sandboxes cannot reach each other, and its token is generated once into ~/.config/ai-app/sandbox-token and carried across restarts along with any the enrolment flow appended — so the emulator app is enrolled once (the start banner prints the command) and stays enrolled. It shares the real TLS certificates, because the installed APK pins that CA. Driving verbs, so none of this is re-derived per session: ./ui-sandbox.sh spawn [title] (an echo session, prints its id), ./ui-sandbox.sh send SID text|@file, and ./ui-sandbox.sh api /path [curl args]. ./ui-sandbox.sh keep restarts the server without wiping the sessions and enrolment already there — for when the fixture under test was expensive to build; plain start wipes them, which is right for the list-screen fixtures and wrong for that. It passes --delay by default, and AI_SANDBOX_BIG_MB puts one large transcript among the small ones while AI_SANDBOX_SPAWN_DELAY makes the fake CLI slow to start. Both exist because operations that finish in milliseconds have states on the way that nothing can observe, and an unobservable state is one where broken and working look identical. It also builds a fixture tree at the sandbox home's ~/files for the explorer, holding the states otherwise only reachable by finding a real machine in one: an empty directory, a name with a tab and one with an apostrophe, a binary file, one over FILE_LIMIT, one chmod 000, a symlink to a directory and a broken one, a source file per language, and the three sizes the limits were measured against (edit-32k.rs, edit-128k.rs, big-source.rs). Point a session at it with ./ui-sandbox.sh api /sessions/<id>/cwd -X POST -H 'content-type: application/json' -d '{"cwd":"~/files"}'. The explorer's 409 is produced by editing the file on the machine (printf … > file) between pressing the pencil and pressing save.
  • app/debug-transcript.sh — a real conversation on the emulator. The echo driver is the right rig for most things and the wrong one for anything whose cost scales with what was actually written: a real reply is longer, is real markdown, and carries tool calls whose input and output are kilobytes. Two faults were invisible until a real transcript was loaded — a page of history landing mid-fling threw the reader back to the newest end, and parsing one real reply took 51ms against 4.6ms for a synthetic one. -b takes the biggest conversation on the machine rather than the newest, which is what a scrolling test wants; --stop takes it down. It copies the transcript into /tmp and gives the server a HOME of its own, so the import can only see the copy — importing spawns claude --resume, and against the real file that is a second CLI writing to a conversation somebody may still be in. A transcript never goes in this repository: they hold whatever was said, read and written in that session, and ~/repos is shared with the host besides.
  • A fake CLI exercises the process lifecycle without a token. Point a claude_cli provider's command at a two-line script — #!/bin/sh and cat > /dev/null — and it behaves the way the lifecycle code cares about: it holds the fifo open, records a real pid, writes nothing, and dies on a signal. So adopt, stop, restart and start are all drivable without a real --resume and without spending a turn on somebody's account. Reach for this when what is under test is whether a process is running, and for debug-transcript.sh when it is what the transcript draws.
  • app/transcript-bench.sh is the standard scroll measurement: it opens the first session (or -k keeps the current screen), scrolls a fixed gesture loop, and prints the app's render report — the same one the in-app copy button produces, whose on screen: line names what the viewport was holding. Compare two runs with the same gestures; the emulator's absolute frame times transfer nothing, the report's accounting does. Run it either side of any change under Markdown*.kt, Transcript*.kt or SessionScreen.kt's list, and put the report in the commit. The numbers that move first are the worst record: one block, the reparse mean while streaming, and the draw phase's accounting line.
  • app/stream-bench.sh [-k] FILE is that measurement for a reply still arriving. It taps "Jump to latest" so the list is pinned to the newest end, resets the report, sends FILE, waits for the transcript to stop growing, and prints. Both of those are corrections to a first version that measured nothing: a transcript parked further back never redraws while a reply streams into it, and a session is idle at both ends of a turn, so polling for idle answers before the turn has started.
  • app/trace-draw.sh names what a scrolling frame spends inside the framework, from atrace text output with no trace processor needed. It is how the cost of a layout node per link was attributed to the framework rather than guessed at.
  • iris/android-app/build-apk.sh [debug|release] [--abi ...] [--features ...] builds iris-android-app's cdylib (cargo ndk) and its APK (Gradle) in one step and verifies the result (aapt2/apksigner), and iris/android-app/run-bench.sh [--apk PATH] installs it on this checkout's own emulator, taps "Run benchmark" by label, and prints the report -- written so the P0 build/install/tap/read-report cycle stops being retyped by hand each time (docs/RUST.md's P0 box).

Driving the UI

No script that drives this app's UI presses a coordinate. Every control is found by the name it already carries for assistive technology — ui-trace record --do "tap 'Session settings'" — which resolves the label against the screen at the moment of the gesture and fails the whole run when it is not there. app/bench-lib.sh is what the bench scripts share for it. A coordinate is a position measured once by hand, and anything that moves the control makes the tap land on whatever now sits there — the bench then reports a number that was never measured, which reads exactly like a result. Both bench scripts pressed the render report at tap 723 205 until that button moved into the session settings dialog on 2026-09-03. The check that none has crept back:

grep -n "tap [0-9]" app/*.sh

Swipes are still coordinates, deliberately: a gesture across a scrolling area is a distance rather than a control.

Two traps in the emulator bench loop, each of which cost a run. adb shell pm clear removes the enrolment and the notification permission along with the saved anchors, so the next run measures a permission dialog — re-enrol with the command ui-sandbox.sh prints, and pm grant … POST_NOTIFICATIONS. And a saved scroll anchor is per session id, so the only way two builds start a scroll from the same place is a fresh session for each.

The emulator is ~/repos/emulator-tools' business, not this repo's. emu up creates and boots the AVD named after this checkout — whatever emu name prints, never a name typed out here, since this file is the same in every clone. run-android.sh is that plus a build and an install. The adb on PATH after sourcing android-env.sh is that repo's wrapper, which fills in -s from the same rule. Gradle does not go through it, so a Gradle init script from emulator-tools runs emu check before installDebug, uninstallDebug and connectedAndroidTest and fails rather than fanning out to every attached device; when it refuses, say which device you mean at the moment you use it — ANDROID_SERIAL=$(emu serial) ./gradlew ….

Testing llama.cpp and ssh here

The prebuilt CPU llama.cpp lives outside the repo at ~/.local/opt/llama.cpp (the 15 MB ubuntu-x64 release asset). It needs its own directory on LD_LIBRARY_PATH, so start the server as LD_LIBRARY_PATH=~/.local/opt/llama.cpp ai-server … and point a provider's command at ~/.local/opt/llama.cpp/llama-server. A 0.6B Q8_0 answers at usable speed on this VM's 8 cores. Do not test with a 2-bit quant: the IQ2_XXS of that model produces fluent nonsense, which reads exactly like a broken driver — llama-cli produces the same from the file directly, which is how to tell the two apart in a hurry.

There is no second machine, so ssh this VM to itself: generate a throwaway key, append the public half to ~/.ssh/authorized_keys, and configure a host of bob@127.0.0.1 with identityFile pointing at it plus options: ["StrictHostKeyChecking=no", "UserKnownHostsFile=…"] so it touches nothing real. Point a provider's command at something harmless like /bin/echo rather than at claude: the transport is what is under test, the process exiting immediately is the signal, and it costs no tokens. Take the key back out afterwards. The remote login shell here is fish; the remote script and ssh.rs's POSIX quoting happen to mean the same thing in both, but that is luck rather than design, and a shell that is neither is the thing to suspect first if a remote spawn ever mangles an argument.

Where things run (host vs this VM)

The machine itself — the two boxes, the shared ~/repos mount, and why the VM is untrusted — is described once in ~/.claude/MACHINE.md. What that means here:

  • ai-server belongs on the host in production. That is where the LAN address the phone can reach is, and where WireGuard terminates. wg-setup-host.sh sets that up (keys, wg0.conf, the phone's QR); run it there with sudo WG_ENDPOINT=<ddns name>.
  • The tunnel and the real phone can never terminate in the VM, because nothing outside can open a connection into it. Phone bring-up is host work.
  • wg0 (10.66.0.1) exists in this VM too, so the production path is exercisable during development. It has no reachable peer and does not need one — but with no --bind the emulator cannot reach the server.
  • The claude CLI is only in the VM, so from the host it is a remote. The backend reaches it as it would any other machine.
  • Starting the server in the VM makes a separate throwaway dev CA. Never install a build pinning that on the real phone.

Sessions outlive the backend

Since 2026-08-29 a session's process is deliberately left running when ai-server stops, and adopted again when it starts. docs/PLAN.md has the design; day to day:

  • Stopping the server no longer stops the sessions. After pkill ai-server the claude processes are still there, on purpose (reattaching to the claude-cli it left running in the log). To end one, POST /sessions/{id}/stop — which keeps the session and its transcript, and /start brings the process back on the same conversation — or delete the session, which ends the conversation too.
  • A message or a command sent to a stopped session starts it, so the Start button is for when you want a process and nothing to say to it yet.
  • A backend start adopts and starts nothing. If you are looking for a stopped session's process after a restart, there is deliberately none.
  • A session spawned while testing cleans itself up: --throwaway-sessions, which a debug build defaults to on. Pass --throwaway-sessions=false to keep what a development server spawns. The flag decides only what new sessions are marked as; what happens on the way out is decided by the mark.
  • Each session directory holds process.json, stdin.fifo, stdout.log and stderr.log. stdout.log is the driver's input, read from the byte offset in process.json; removing either by hand while the session is live loses output or replays it.

Importing

The import list reports each session's size as well as its line count, because the two disagree in the way that matters: these transcripts embed screenshots as base64, so one line can be a megabyte. On this machine a 69 MB session has 3,427 lines and a 44 MB one has 6,792 — nothing about a line count tells you what continuing a session will cost. Shown, not warned about; importing a large session is a choice somebody is entitled to make.

Never import a Claude Code session that is open in a terminal. The app refuses it — see docs/PLAN.md for the incident that made that a refusal rather than a warning.

One Claude Code session id can name two files, and the listing offers it once. Resuming from a different working directory makes the CLI write a second transcript with the same id under that directory's project folder — an ordinary state of a machine, not corruption. Everything downstream addresses a session by id, and the phone keyed its list on it, so two rows sharing one closed the app on a Compose duplicate-key throw. parse_listing keeps the copy with the most lines, because the other is usually a few-hundred-byte stub and is often the newer of the two, so recency is the wrong key. Deleting removes every copy rather than the first, or the row came back after a delete that reported success. The phone's half is uniqueItems, which every list keyed on a server-chosen id goes through: a repeat there must never be able to close the app, whatever produced it.

Deleting a session offers to take the machine's own transcript with itDELETE /sessions/{id}?deleteForeign=true, behind a switch in the confirmation, and only where the driver keeps a record of its own (keepsOwnTranscript, which today means Claude Code). Off by default, because leaving that copy is what makes an ordinary delete recoverable — and the dialog's paragraph is rewritten when it is on rather than appended to, since the sentence promising the conversation "should still be there to import again" is exactly the one the switch makes false. The server deletes the machine's copy first, so a machine it cannot reach leaves the session where it was instead of half-deleted.

Shared appearance

  • A row something is happening to is dimmed, drained of colour, and says which operation in a wordBusyItem, used by both the session list and the import list so the appearance is learned once. The word rather than a bare spinner because "deleting" and "importing" differ in kind. It does not make the row inert: the caller disables its own click handler while it passes a label. An overlay consuming pointer events was tried and swallowed the drag along with the tap, so a list could not be scrolled while anything in it was busy.

Things that have bitten

Project-specific only — a lesson that would bite any project on this machine belongs in ~/.claude/TOOLCHAIN.md or ~/.claude/MACHINE.md instead.

  • tracing caches callsite interest process-wide. A test that hits a tracing::warn! with no subscriber installed can poison the interest cache for a concurrent test that captures logs (flaky "nothing was logged" failures). Keep every exercise of a logging code path under the one capturing subscriber — that is why the auth middleware has a single combined gating+logging test.
  • The composer can get stuck floating above the bottom of the screen after the keyboard closes, while a reply is streaming. The composer's position and the transcript's bottom padding are both driven by the raw, animated WindowInsets.ime value read inside a graphicsLayer block, to avoid recomposing the whole screen every frame of the keyboard's animation. That animation is carried by a WindowInsetsAnimationCallback, and a callback interrupted mid-flight leaves whatever it was carrying frozen at its last value with nothing left to correct it. A streaming reply invalidates the view every frame, which is exactly the condition known to starve that callback of its onEnd. WindowInsets.isImeVisible does not share the failure mode — it is set once, from the platform's own start/end of the transition over a different path — so it is read once per keyboard toggle and used to force both places back to zero. The guard is a boolean; the inset itself must never be read in the composable body. That correction first shipped as a padding(bottom = … imeInsets.getBottom(this) …), which subscribes the whole screen to a value that changes every frame: measured at 16 full recompositions of SessionScreen per keyboard open, against 1. It is .then(if (imeVisible) Modifier.imePadding() else Modifier) instead — imePadding reads the inset in the layout phase, and dropping the modifier is the same coercion to zero the boolean was added for. The counter to check is session screen recomposed in the debug report, which should move by one across a keyboard open, not by the number of frames it took.
  • The keyboard pans the window unless the activity opts into resize. Without android:windowSoftInputMode="adjustResize", opening the IME slides the whole window up (top bar off screen) instead of resizing — imePadding() alone does not fix it and the transcript looks empty.
  • A PEM constant must start at the opening quotes. A generated """\n-----BEGIN CERTIFICATE----- costs Android's CertificateFactory its preamble sniff, so it tries DER instead and fails at runtime with ASN.1 … DECODE_ERROR — nowhere near the code that produced it.
  • ZXing only looks for a dark code on a light ground. The enrollment QR is block characters in the terminal's foreground colour, so a dark-themed terminal renders it as a negative and the in-app scanner silently never matches — while the phone's own camera app, which tries both, does. The scanner asks for Intents.Scan.MIXED_SCAN, which alternates normal and inverted frames; keep it that way rather than making the server dictate the colours.
  • serde_json's default float parser is not correctly rounded, so the server handed out the same transcript line two different ways: a ts of 1788546972.6030757 came back from /transcript as …0755 while the SSE stream sent the original. Nothing on screen could show it — a ts is drawn as a relative time — and what found it was the phone's cache comparing a line it held against the server's answer. The float_roundtrip feature in server/Cargo.toml is the fix and a_line_read_back_is_the_line_that_was_written is what keeps it; that test fails within a second of the feature being dropped.
  • Resolving one importable session used to list every one of them. import::delete and the import seed both called list, which reads every transcript Claude Code has ever written — measured at 3.7 seconds against the 867 MB in this VM, paid once per session in a batch. import::find takes the same script with one glob narrower: 78ms. Ids are checked (is_session_id) before they reach that glob, since a / or .. walks it out of the projects directory.
  • A transcript page used to cost the whole transcript. read_window read and parsed every line and then kept the last limit of them, so the work was the size of the conversation rather than the size of the answer: one page of a 21 MB, 24,000-event transcript took ~500ms to return 620 KB, and took the same 500ms whichever page was asked for. It is a bisection now (Indexed in transcript.rs) — sequence numbers only increase, so the edge of a range is found by parsing one line per halving. Same page, ~110ms, of which ~20ms is the file scan. The file is still read whole; that is where the remaining cost is, and going further means a chunked backwards reader.
  • Paging back has two failures that look like "there is simply no more history", and neither says anything on screen. Both invisible on a loopback server and reproducible at --delay 150. The pager fires on the first layout, before any event has arrived — moreHistory starts true, so the spinner is in the list and visibleItemsInfo is not empty — and before = 0 asks for the events before the first one, which is none, which is exactly how this code is told it has reached the start. loadOlderPage refuses oldestSeq == 0 now. And joinPages only ran adoptRun on the path where a split call had been found, so a boundary landing cleanly between two calls — most of them — left one run of tool calls drawn as two groups with the seam wherever the reader happened to have paged. Reproducing either takes a boundary placed on purpose: the opening page is 80 events, so arrange the transcript so that event counts back from the newest.
  • A page is 800 events and a screen is a handful of rows, and the two have no fixed ratio. A run of thirty-five tool calls is one row; a reply is hundreds of text deltas folded into one. So anything that budgets in rows has to measure a screen rather than name a number: the history cushion was eight rows, which on a tool-heavy transcript is less than one screenful, so the reader hit the end of what was loaded on every swipe and stood there for a round trip. It is HISTORY_SCREENS viewports now, counted from what is actually on screen.
  • Only fetchTranscript was off the main thread; the fold was not. foldEvent returns a new list per event, so a page is that many copies of a growing list — fine at 80 events and about 300,000 element copies at 800, run in the middle of the scroll that asked for it. warm had the same shape: the markdownIn scan that decides what to parse ran before the hop to Dispatchers.Default. The shape to watch for is a withContext that wraps the fetch and leaves the work done with the result outside it.

Measurements worth not re-taking

  • What the transcript screen costs to scroll. Taken 2026-08-30 on the GPU emulator against a real imported transcript with the server at --delay 120. Settled and flinging fast, both into fresh history and back through rows already drawn: 5.25.9% janky frames, 99th percentile 2932ms, 02 slow UI-thread frames. The stock Settings app on the same device is 3.3% and 38ms, so this is at the platform floor. The number that is not at the floor is the first few seconds after opening a session, where every row on the way is being composed for the first time; that is inherent to a lazy list and it is why a measurement taken before the screen settles reads three times worse. Settle first, then reset gfxinfo.
  • The reset path is not reachable by reopening a session. Measured 2026-09-04 against a session streaming at 20 events a second: reopening one with an anchor 1,800 events back connects 87119 events behind, well under CATCH_UP_LIMIT's 200, because the restore is two requests — the opening page, then one span covering the whole distance. To exercise the reset at all you have to lower CATCH_UP_LIMIT in a throwaway build; at 5 the app takes the reset on a live connection, clears, refills and carries on without reconnecting.
  • The session screen's stream survives backgrounding here — 20 seconds at the launcher while 415 events were produced brought no reconnect at all, which is not what the comment above that loop expects, and is most likely this emulator being headless rather than the phone's behaviour.
  • Reopening a cached session costs one request for one event (the probe), and scrolling the whole conversation back costs nothing more; a cold open of the same 500-event session is two pages, 100 events. Measured 2026-09-04 on the emulator against the sandbox.
  • Reading is cheap and editing is not. The viewer handles a 1 MiB, 28,000-line file because it draws one row per line; the editor is one BasicTextField, which costs two seconds a frame at 128 kB and stops the app at 1 MiB, so EDIT_LIMIT caps it at 32 kB with the reason said on screen. If you make the editor faster, that number is what to move. docs/EXPLORER.md's "What the measurements said" has the rest.