Author SHA1 Message Date
iris 800da46188 Merge remote-tracking branch 'origin/rustify' into worktree-agent-a27094a7db775552a
# Conflicts:
#	docs/IRIS.md
2026-09-05 21:37:04 -04:00
irisandClaude Fable 5.1 00767eed4d docs: P0's iris half done -- bench feature, emulator smoke run, APK
RUST.md's P0 box gets the iris-half account: the fixture, the scroll/stream
mechanism, the report fields, build commands (all clean), packaging (no
cargo xtask apk yet, so a new Gradle release build type on top of cargo
ndk), and the emulator smoke run's report next to Compose's own. Used a
second, differently-named AVD rather than contend with the session already
on this checkout's own emulator.

DECISIONS.md's P0 entry gets a matching summary bullet. IRIS.md records
AndroidAppState::platform_ready. IRIS_TODO.md notes the one gap found:
no read-only selectable text primitive, so the bench report's TextEdit
picks up a keyboard on tap it has nothing to type into.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-05 21:35:51 -04:00
irisandClaude Fable 5.1 683db4908a iris-android-app: a bench feature, P0's iris half
A third AndroidAppState (BenchClient) on top of transcript-screen: embeds
app/bench-fixture/assets/transcript.jsonl with include_str! (no server, no
enrollment), folds the first 3,200 lines through client_core's real
fold_page as the opening backlog, and holds the rest back as a streaming
tail. "Run benchmark" resets FrameReport, animates the same 24-swipe/
6-cycle scroll BenchRun.kt drives (List::scroll in ~60Hz steps, since iris
has no built-in tween), then replays the tail at 20/s through fold_event --
the same fold path a live SSE reply takes -- and shows a report in a
selectable TextEdit. "Copy report" puts it on the clipboard.

The report adds process CPU time (libc::getrusage), peak RSS (/proc/self/
status's VmHWM) and battery current (BatteryManager.getIntProperty via
direct JNI, bench_jni.rs's PlatformHandle) to FrameStats's existing
frames/janky%/percentiles/CPU-GPU-split line -- "unavailable" rather than a
fabricated number wherever the platform can't answer.

build.rs now exits early under the bench feature before requiring a live
server's host/port/token/CA: BenchClient never calls build_transport().
app/build.gradle gains a signed `release` build type (previously only
debug) so the cdylib cargo ndk builds can be packaged for a phone, the same
key app/build-apk.sh generates.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-05 21:35:44 -04:00
irisandClaude Fable 5.1 8d23a20792 iris: AndroidAppState::platform_ready, a JavaVM+View handle for later JNI calls
Default no-op lifecycle hook, called once from new_peer right after new.
P0's bench build needs to call BatteryManager/ClipboardManager through the
view's own Context from a background thread as well as the UI thread, and
neither a JavaVM nor a GlobalRef to the view was reachable from
AndroidAppState::new before this. Existing implementors (Client,
TranscriptClient) are unaffected.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-05 21:35:33 -04:00
irisandClaude Fable 5.1 d01c105037 iris: stop requesting compute-shader limits nothing uses
adapter.request_device asked for Limits::default(), which requests
desktop-tier compute-shader limits unconditionally even though nothing in
iris/iris-core creates a ComputePipeline or writes a @compute stage. That
crashed device creation outright on a downlevel GL adapter reporting
OpenGL ES 3.0 (no compute at all) -- the Android emulator's
EMU_GPU=software/force-gles path, and any real GLES-3.0-only device.

New iris_core::device_limits(), shared by both platform backends, zeros
exactly the six max_compute_* fields rather than switching to a downlevel
Limits preset -- downlevel_webgl2_defaults() also zeros
max_storage_buffers_per_shader_stage, which shader.wgsl's vertex stage
needs. rigs/gpu-probe's own mirrored limits were updated to match.

Not verified against the actual SwiftShader-ES-3.0 crash on-device this
pass: the cold boot needed would have force-restarted this checkout's
emulator while another session had its own app running on it.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-05 21:18:30 -04:00
iris 88631f5e8b Merge remote-tracking branch 'origin/rustify' into worktree-agent-a27094a7db775552a
# Conflicts:
#	AGENTS.md
#	server/src/session/driver.rs
2026-09-05 21:10:19 -04:00
irisandClaude Fable 5.1 e6924298bc iris: fix the intermittent touch-scroll dropout (missed ACTION_DOWN hit-test)
Root-caused via temporary logcat tracing (touch events, DragArbiter state,
Selection::drag dispatch), reproduced against a real sandbox session: a
gesture's ACTION_DOWN can land on a row's own padding/gap or its header,
which CursorSense has no sensor over, so the widget that ends up handling
the gesture only ever sees Pressing frames and DragArbiter never gets
press_start -- leaving it stuck in Idle (answers Undecided forever) for the
rest of that gesture. Not the previously-suspected coalesced first
ACTION_MOVE, which is now ruled out.

DragArbiter::is_idle() lets Selection::drag notice a Pressing frame with
no matching press_start and recover the press there instead. Four new unit
tests, one of which fails on the pre-fix code.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-05 21:05:03 -04:00
irisandClaude Fable 5.1 68b48cfd14 docs: record P0's Compose half (bench build, fixture, smoke run)
RUST.md's P0 box gets the emulator smoke run's report and what's done vs.
left; DECISIONS.md gets a dated summary entry; AGENTS.md's "Checking your
work" and "The rigs" get one paragraph each on the bench build type and
app/bench-fixture/.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-05 21:04:07 -04:00
irisandClaude Fable 5.1 e6c884a0cd app: fixture-mode session screen and a "Run benchmark" control
BenchFixture.kt/BenchNetwork.kt fake the backend for the bench build: a
URLStreamHandlerFactory installed only under BuildConfig.FIXTURE_MODE
answers TranscriptSource/EventStream's requests from an in-memory copy of
the bundled fixture instead of opening a socket, so the fold, the paging
and uniqueItems under test are the screen's real ones rather than a
shortcut built for this. MainActivity opens straight onto that session
when FIXTURE_MODE is set, with no enrollment and no permission prompts.

BenchRun.kt drives the same scroll loop and streaming phase
transcript-bench.sh/stream-bench.sh drive over ui-trace, but in-process
(24 swipes through the real LazyListState, then 400 fixture events
appended at 20/s through the real live-fold path), and adds process CPU
time, peak RSS and battery current to the render report -- "unavailable"
rather than a fabricated number where the device can't answer.

"Run benchmark" sits beside the existing "Copy" in session settings,
found by that exact label the way every other control here is
(SessionSettingsDialog's onRunBenchmark, null on every build but bench).
debugReport gained an optional `extra` section for this; empty and
invisible on every other build's report.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-05 21:03:57 -04:00
irisandClaude Fable 5.1 a6cb9a9082 app: a bench build type for P0's benchmark gate
Own application id (.bench suffix) and label ("AI Sessions bench" via a
build-type resValue over the new @string/app_name), release
optimisations, signed with the same key build-apk.sh already generates,
FIXTURE_MODE=true wired through BuildConfig. Its asset source set points
straight at app/bench-fixture/assets rather than a copy under androidApp,
so there is one file to keep in sync with the generator, not two.

build-apk.sh bench builds it; the CA-pinning step is untouched and still
requires a real ca.pem to exist, even though this build never connects --
simplest to let it pin whatever is there rather than special-casing it.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-05 21:03:44 -04:00
irisandClaude Fable 5.1 0be6a571c4 app/bench-fixture: the synthetic transcript P0's benchmark opens in both apps
Deterministic (seeded), in the app's own event model rather than a real
transcript: 3,601 events split into a 3,200-event opening backlog and a
400-event tail both bench harnesses replay as the streaming phase, with
headings, inline markdown, fenced code in six languages, a table, tool
calls with kilobyte-scale input/output, and two embedded PNGs.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-05 21:03:34 -04:00
irisandClaude Fable 5.1 bfe93c4188 RUST.md, DECISIONS.md: P0, the phone benchmark gate Iris asked for before P1
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-05 20:31:20 -04:00
irisandClaude Fable 5.1 5b7dc0e4e2 RUST.md, DECISIONS.md, IRIS_TODO.md: the port plan, P1-P7, after iris-over-Masonry
Adds "The port, in order (decided 2026-09-05)" to RUST.md: seven ordered
steps building the app on iris now that the framework is decided, each
naming the Kotlin files it replaces, the client-core pieces it needs
(and which are not yet covered and must be ported first), the missing
iris widgets it needs (recorded in IRIS_TODO.md's new "Build (for the
port)" section), and a pass condition a later agent can run. Ordered by
risk to the daily-use path: session screen parity, then the shell merge
and a real phone install, then root tabs, the file explorer,
settings/enrolment, desktop parity, and the cutover itself.

Crate-shape decision recorded in DECISIONS.md: one UI crate, app-ui,
grown out of transcript-ui rather than started beside it, with
desktop-app/android-app as thin entry points over it and platform-only
code staying in the E3/E5 Java shell.

Updates RUST.md's "Where things stand" and "For the next agent" to point
at P1 rather than the now-closed framework decision.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-05 20:30:02 -04:00
irisandClaude Fable 5.1 621f08d725 DECISIONS.md, RUST.md: Iris decided iris over Masonry, 2026-09-05
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-05 20:25:43 -04:00
irisandClaude Fable 5.1 e49d0e606f RUST.md, DECISIONS.md, IRIS.md: iris's host-GPU frame time, 2026-09-05
Takes the -gpu host pair the earlier software-mode comparison flagged as
missing. Under real GPU rendering (--features force-gles: the default
Vulkan backend has no adapter at all under plain host-GPU boot, confirmed
by the exact wgpu error), iris's median frame (15.0ms) is faster than
Compose's (20.0ms) on the same session content -- the opposite shape from
the software-mode table. The new redraw-to-submit/submit-to-present split
shows iris's own CPU work is a median 0.2ms per frame; almost the whole
frame is time handing off to the driver, consistent with (but not proof
of) the software-mode gap being mostly SwiftShader's CPU rasterisation
cost rather than iris-specific slowness.

A same-mode software force-gles run, meant to isolate the backend, hit a
third distinct crash instead (SwiftShader's GL path reports itself as
OpenGL ES 3.0, which has no compute shaders, and iris's device request
assumes them unconditionally) -- real scope to fix, not done here, so the
software-mode question stays open. A real intermittent touch-scroll
dropout was also reproduced (six consecutive swipes produced zero
redraws while taps kept working; an identical retry then succeeded) and
is not explained. The idle-redraw and virtualised-culling findings from
the software-mode pass were confirmed to hold under real GPU rendering
too.

DECISIONS.md's DEFERRED item carries the updated table; the iris-vs-
Masonry choice itself is still Iris's to make. IRIS.md records the
FrameReport::record_split/FrameStats::cpu_p50/gpu_wait_p50 API from the
prior commit (e2a1fad), which this pass's measurement used.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-05 20:07:20 -04:00
irisandClaude Fable 5.1 e2a1fadbec iris: FrameReport CPU/GPU split, force-gles backend switch, iris-scroll.sh rig
Splits each frame sample at queue.submit into redraw-to-submit (iris's own
CPU work) and submit-to-after-present (driver/GPU wait), so RUST.md's I5
"where does iris's frame time go" question can be answered with a number
per half instead of a single total. Adds a force-gles Cargo feature that
switches the Android wgpu::Instance from Backends::PRIMARY to Backends::GL
at compile time (no runtime env-var path exists into an already-launched
Android process on this machine), for isolating SwiftShader-Vulkan vs.
GLES/virgl as the software-mode gap's cause. app/iris-scroll.sh extracts
transcript-bench.sh's exact 24-swipe/6-cycle gesture loop for iris's own
demo app, which transcript-bench.sh cannot drive directly since it opens a
session through the Compose app's own UI.

Verification (this pass, on a disk-pressure-limited host running low on
space): cargo fmt --all clean, no diff. cargo clippy --workspace
--all-targets: no warnings from this diff (pre-existing future-incompat
notices from wgpu/winit/naga only). cargo test --workspace and cargo ndk
for iris-android-app --features transcript-screen were verified clean by
the previous pass on this identical diff (fmt/clippy/test/ndk all clean,
per that pass's own report); not re-run here because the host's disk was
93% full and a concurrent ai-server rebuild (stable toolchain moved to
1.98.1, rebuilding aws-lc-sys from scratch) had driven I/O pressure to
~60%, so a repeat cargo test --workspace sat 50+ minutes doing no useful
work and was stopped rather than left to make the disk situation worse.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-05 18:47:01 -04:00
irisandClaude Fable 5.1 0e4629361b docs: I5's clean scroll comparison between Compose and iris, one session
Same sandbox session content, same emulator, EMU_GPU=software: Compose
(debug, in-app report) 1102 frames/99.0% late/p50 33.8ms/p99 79.5ms vs
iris (release -- debug SIGSEGVs on this emulator) FrameReport 299
frames/94.65% janky/p50 79.1ms/p99 117.8ms (repeat: 233/94.42%/p50
109.3ms). Ticks I5 [x]; states plainly what's not comparable (build
profile forced asymmetric, three different jank definitions, both are
software-rasterised emulator numbers). The two "zero frames" attempts
that preceded the clean runs traced to this session's own script bug
(a cd into /tmp changed which emulator ui-trace targeted), not a
reproduction of the previously-suspected touch-delivery dropout; a
sampler ran the whole session and saw load rise during the gesture
without correlating to any failure. DECISIONS.md's DEFERRED item gets
the same table so Iris can decide iris-vs-Masonry from it -- that
choice is left to her.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-05 15:21:45 -04:00
irisandClaude Fable 5.1 1e7b1cddb7 RUST.md, IRIS.md, IRIS_TODO.md, DECISIONS.md: record I5's frame report and holddrag results
FrameReport gave a real, measured on-device number (frames=34,
janky%=61.76, p50=26.5ms p90=48.0ms p99=98.1ms worst=98.1ms) and
long-press-then-drag-to-select is now confirmed on-device (logcat plus a
screenshot of the highlighted selection). Neither closes I5's box to [x]
yet: the frame number is real but not the clean single 24-swipe loop
comparable to Compose's, because gestures against this checkout's
EMU_GPU=software emulator intermittently delivered zero touch input this
session -- a new, separately named finding (candidate cause: the
emulator's own software rasterisation measured at ~78% of a CPU core
continuously), not yet root-caused. DECISIONS.md's DEFERRED item is
updated with these numbers rather than a decision made here.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-05 15:00:29 -04:00
irisandClaude Fable 5.1 470f8e5019 transcript-ui: log selection begin/extend, for on-device verification
Selection has no accessibility label of its own yet, so a logcat line at
begin/extend is the smallest way to confirm a real long-press-then-drag
reached DragArbiter/Selection on-device. Driven with the new ui-trace
holddrag action against iris-android-app's transcript screen: produced
"iris selection: begin at row ..." then a sequence of "... extend to row
..." lines, and a screenshot right after shows the expected highlighted
selection spanning multiple rows.

New `log = "0.4.28"` dependency (matching iris-android-app's own pin) --
transcript-ui had no logging facility before this.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-05 15:00:18 -04:00
irisandClaude Fable 5.1 cf10b17c5b End a subagent on its own end_turn, not the parent's tool_result, and give the expander a touch-sized row
The Agent tool runs subagents in the background, so the parent's result
arrives at launch while the subagent works on for minutes; finishing on it
read a running agent as finished with a transcript cut off at launch. A
subagent now ends on its own message_delta end_turn, and a later line for a
finished one reopens it, since a background agent can be messaged again.

The card's expander row was only the chevron's height, so a tap for it
landed on the first subcard; it is the platform's 48dp minimum now.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-05 14:44:00 -04:00
irisandClaude Fable 5.1 7ae53ad797 iris: FrameReport, a per-frame wall-time report of iris's own render path
dumpsys gfxinfo cannot see a SurfaceView's own GPU-drawn frames at all
(RUST.md's I5 box), so iris needs its own equivalent of Compose's
render-report button before item 3 of the recommendation can be decided
by a number. FrameReport (iris/core/src/render/frame_report.rs) records
each frame's wall time -- from render()'s redraw start to after
queue.submit + present() -- into a fixed 4096-entry ring, and reports
total frames, janky % (>16.7ms, gfxinfo's own budget), P50/P90/P99 and
the worst. Wired into AndroidUiState and android/view.rs's render(), and
exposed as two named controls ("Frame report", "Reset frame report") on
iris-android-app's transcript screen, logged under the crate's fixed tag
so a script can grep "iris frame report" the way transcript-bench.sh
greps "ai-app render report".

6 new unit tests for the ring/percentile math. cargo fmt/clippy/test
--workspace clean; cargo ndk (iris, transcript-ui, and
iris-android-app --features transcript-screen) all clean.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-05 14:23:51 -04:00
irisandClaude Fable 5.1 d17040b601 RUST.md, IRIS.md, IRIS_TODO.md, DECISIONS.md: record I5's Android integration and measurements
I5's transcript screen now runs on-device against a real ai-server on
iris-android-app's new transcript-screen feature (extends I2's shell
rather than a third one), with real scrolling, real touch-drag panning
and tap-by-name accessibility all confirmed by screenshot/log evidence.
I4's own emulator-side check (tap-by-name on the tabs demo) closed the
same session, so its box ticks [x] now.

Still [~], not [x]: the render-time number RUST.md's recommendation
wants for iris couldn't be produced this pass, for a precise and
recorded reason rather than a vague one -- dumpsys gfxinfo cannot see a
SurfaceView's own GPU-drawn frames at all (0 frames reported across a
gesture loop that visibly scrolled), and a SurfaceFlinger --latency
fallback gave no per-frame history either on this Android version. The
Compose side of the same loop did produce a real number under identical
conditions (8.96% janky, 99th percentile 150ms), so this is now a
one-sided number rather than a missing one on both sides.

Also found and recorded: the AVD's saved snapshot carries a GPU config
across restarts, so switching between the documented Vulkan boot
recipes needs a cold boot (clearing snapshots/) that the emu wrapper
does not force -- cost three different-looking crashes before the
pattern was the snapshot, not the code.

DECISIONS.md's DEFERRED item is updated with the numbers Iris needs to
weigh the iris-vs-Masonry call; the call itself stays hers.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-05 14:09:04 -04:00
irisandClaude Fable 5.1 bf5087a598 iris/android: fix background-thread redraw crash, add missing INTERNET permission
Two real bugs found bringing up I5's Android transcript client, neither
specific to that screen -- any future caller of Tasks::redraw_handle()
from a background thread would hit the first one.

AndroidRedrawHandle::request_redraw called View::post_frame_callback from
a tokio worker thread; its Java side calls Choreographer.getInstance(),
which throws IllegalStateException unless the *calling* thread already
has a Looper, and a JNI-attached background thread has none. That crashed
the whole process (SIGABRT, unwrap() on a JavaException) the first time a
background fetch asked for a second frame. Fixed by routing through
View::post_delayed(0) instead, Android's own thread-safe way to queue
work onto a View's UI thread, landing on a new
IrisViewPeer::delayed_callback override that drains tasks and renders --
same body as do_frame, now running safely on the UI thread.

iris-android-app's manifest never needed INTERNET before (the tabs demo
makes no network call); its absence read as EPERM ("Operation not
permitted") from UreqTransport::new's connect, not the
ECONNREFUSED/ENETUNREACH a dead server would give.

Full account in RUST.md's I5 box and IRIS.md's Tasks::redraw_handle entry.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-05 14:08:40 -04:00
irisandClaude Fable 5.1 9fa09b0af1 Show a session's subagents as subcards, each with a read-only transcript
A subagent is a second transcript owned by a session, in the same event
model, with no process and no controls. The claude translator routes lines
carrying parent_tool_use_id to a per-subagent translator and transcript
under <session>/subagents/<tool_use_id>; three routes expose the list, a
transcript page and the SSE stream. Echo grows /subagent [n] as the rig.

On the phone a card with subagents ends in a chevron expander, collapsed by
default, opening to outlined subcards styled like dev-updater's components;
a subcard opens SessionScreen in read-only form, addressed through
TranscriptAddress so paging, cache and stream are shared.

Design in SUBAGENTS.md; choices awaiting review in DECISIONS.md.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-05 13:41:15 -04:00
irisandClaude Fable 5.1 aa3d11471f iris-android-app: transcript-screen feature -- I5's Android integration
Extends the existing tabs demo shell (I2/E5's Gradle project, JNI
registration) with a second, mutually-exclusive AndroidAppState rather than
building a third shell -- it already has the working IrisView/MainActivity
Java and the register_view_class wiring, and the only thing a transcript
screen needs on top is a different Client type (the same axis
tabs_ui::build vs. transcript_ui::build already varies along on winit).

`--no-default-features --features transcript-screen` builds
transcript_client::TranscriptClient instead of the plain tabs Client:
fetches the sandbox's session list, opens the first one, and follows it
live, reusing desktop-app's app.rs shape (fold_event/group_tool_runs/
fold_page/raw_seq, a generation counter) almost verbatim. The one real
difference is the redraw path -- android-view has no winit::EventLoopProxy,
so Tasks gained redraw_handle() (iris/src/task.rs) to let a caller request
a frame after each TaskCtx::update from inside a still-running task, not
just once when the whole future completes.

Deliberate simplification, not a template: there is no session list or
enrollment UI here. build.rs bakes the sandbox's host/port/token plus the
pinned CA in at build time from AI_APP_TRANSCRIPT_HOST/_PORT/_TOKEN and
AI_APP_CA, the same trust-boundary reasoning as the Compose app's
GeneratePinnedCert Gradle task, extended to also bake the enrollment since
building a real one (Keystore-sealed storage, a QR/link scanner) is E3's
scope, not this box's. Recorded in RUST.md's I5 box.

tabs-ui and the transcript-screen deps are now both optional, gated by
mutually exclusive tabs-screen (default) / transcript-screen features --
building one screen with the other's default deps still active tripped
Cargo's unused_dependencies lint.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-05 13:26:10 -04:00
irisandClaude Fable 5.1 78aff64844 client-core: hoist transcript_fold::{fold_page,raw_seq} out of desktop-app
Both desktop-app's app.rs and the new Android transcript client (RUST.md's
I5) need the same page-fold and live-stream resume-cursor logic; per
CODE_RULES's "write the logic once" it now lives in client-core alongside
fold_event/group_tool_runs instead of being duplicated. desktop-app calls
the shared functions; its own copies and their tests moved with them.

Also fixes a clippy::collapsible_if in config.rs's percent_decode, found
while re-running clippy after this change (let-chains are stable now).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-05 13:25:46 -04:00
irisandClaude Fable 5.1 9a33cb5384 docs/: move the design and working documents out of the repo root (CLAUDE.md and AGENTS.md stay, harnesses read them there)
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-05 13:03:42 -04:00
iris 45ced405f3 Merge branch 'worktree-agent-afe80868604fef704' into tmp-merge 2026-09-05 12:59:54 -04:00
iris 62199aa3a7 Merge remote-tracking branch 'origin/rustify' into worktree-agent-afe80868604fef704
# Conflicts:
#	IRIS.md
#	RUST.md
2026-09-05 12:59:39 -04:00
irisandClaude Fable 5.1 b133d85943 RUST.md, IRIS.md, CLIENT_CORE.md: record E4 done
RUST.md: E4 ticked with the screenshot path, the exact commands against
app/ui-sandbox.sh, and the streaming-duplication bug the screenshot found;
"Where things stand" moved E4 out of "in flight" into its own done bullet.
IRIS.md: transcript_ui::build_tree, the public API change transcript-ui
gained for this. CLIENT_CORE.md: client_core::config's table row and its
correspondence note.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-05 12:55:59 -04:00
irisandClaude Fable 5.1 ba6817fee5 iris: run-headless.sh --bin, for screenshotting a real binary not just an example
desktop-app (RUST.md's E4) is a real crate binary a person runs, not a
demo under examples/, and it needs its own argv (--ca, --link) to start
at all -- neither of which the script had a way to express. --bin swaps
`cargo build --example`/`target/debug/examples/NAME` for the `--bin`
equivalents; $RUN_HEADLESS_ARGS is word-split into the launched binary's
own argv, since no example ever needed one before.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-05 12:55:52 -04:00
irisandClaude Fable 5.1 73ee63bc1b iris: desktop-app, a winit window for the transcript screen (RUST.md's E4)
The pass condition was the same screen, from the same crate, running in a
window with only the layout differing. desktop-app is a new workspace
member: a session list (plain iris::widget::Span, rebuilt on selection)
beside transcript_ui::build_tree's screen, talking to a real ai-server
through client-core's ApiClient/UreqTransport/follow_session_events, with
background network I/O on plain std::threads reporting back through
winit's EventLoopProxy rather than iris's Tasks (which only redraws once
per async closure, not once per SSE event).

Both pass-condition proofs held against app/ui-sandbox.sh's real server:
the list showed a spawned session, selecting it loaded its transcript, and
a message sent from the composer streamed its reply back live. Along the
way, a real bug: resuming the SSE stream from a folded item's seq (which
for a still-open assistant message is its *first* delta's seq by design)
replayed already-folded deltas and duplicated the tail of the reply --
found by a run-headless.sh screenshot, fixed by resuming from the raw wire
seq instead, and covered by a regression test.

Deliberately simple and said so in app.rs's module doc: every SSE event
refolds the whole transcript and rebuilds the right-hand tree from
scratch rather than reaching for TranscriptScreen::push_row's incremental
append, since a streaming reply is a row whose text keeps changing after
it appears and push_row can only add a new one. Fine at a desktop
session's scale; the real fix needs transcript-ui to expose updating a
row in place. Android is untouched by this step.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-05 12:55:41 -04:00
irisandClaude Fable 5.1 8f0aec449a transcript-ui: build_tree, the screen without claiming the window root (RUST.md's E4)
build() always finished by calling ui_state.set_root(), which is right for
a window that *is* the transcript screen and wrong for a caller embedding
it beside something else (the desktop app's session list). build_tree()
is build() minus that last step, returning the widget tree instead of
planting it; build() is now one line on top of it, so nothing else
changes for existing callers.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-05 12:55:29 -04:00
irisandClaude Fable 5.1 6d5fd64bb0 client-core: EnrolledServer, the aiapp:// enrol-link parser (RUST.md's E4)
A Rust client needs the same host/port/token an Android phone gets from
scanning an aiapp://enroll?... QR, so a desktop build can enrol from the
identical text pasted rather than a second format invented for it.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-05 12:55:22 -04:00
irisandClaude Fable 5.1 e5880c33f4 iris: DragArbiter closes I5's touch-drag pan-vs-select gap
A row's own click_or_drag() selection handler always won the same
gesture a list-level pan wanted, since run_sensors gives the inner
layer first refusal every frame it's pressed. DragArbiter
(iris/src/sense.rs) decides pan vs. select the way Android does:
vertical drag pans immediately, a held stationary press starts a
selection after LONG_PRESS, and a horizontal drag on already-selected
text extends immediately. transcript-ui's Selection::drag routes
every row's drag through one arbiter per list, driving List::scroll
for a pan instead of a second scroll mechanism.

8 new unit tests (iris::sense::drag_arbiter_tests); cargo
fmt/clippy/test --workspace and cargo ndk (iris, transcript-ui) all
clean; run-headless.sh screenshot byte-identical to before the change.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-05 12:26:54 -04:00
irisandClaude Fable 5.1 a853eb5a4d DECISIONS.md: the summary file for choices made without Iris; RUST.md: note the two in-flight pieces
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-05 12:18:26 -04:00
iris eff5c8b0c0 Let the machine's own CLI refresh an expired token, and retry once
A 401 from the usage endpoint means the stored access token has expired.
Refreshing it here is not an option: Anthropic's OAuth rotates the refresh
token, so a second refresher invalidates the CLI's copy and forces a
re-login on a machine that usually has a live session on it. So run the CLI
there instead and re-read what it wrote.

`doctor` rather than `auth status`: probed against 2.1.258 with an invalid
token, `auth status` answers loggedIn:true from the file alone and never
reaches the network. The same probe showed a failed refresh blanks both
tokens, which is why this stays on the 401 path.

Also gives ProviderConfig one program() so the CLI's default path is not
written down twice.
2026-09-05 12:07:34 -04:00
iris 7b63330aaa Say when a usage 401 is an expired login, not an unreachable endpoint
A 401 is the endpoint answering and refusing the stored OAuth token, which
Claude Code refreshes as it runs -- so a machine whose CLI has been idle
hands us a stale one. Reporting it as "usage endpoint unreachable" pointed
at the network instead of at the one thing that fixes it.
2026-09-05 11:58:17 -04:00
iris 22d5c6585a Merge branch 'worktree-agent-a23aa965694b9eaf5' into rustify (I5: transcript-ui, SpanStyle) 2026-09-05 07:56:09 -04:00
irisandClaude Sonnet b063fbd7f9 RUST.md, IRIS.md, IRIS_TODO.md: record I5 -- transcript-ui built and
partial, the recommendation's numbers still missing

I5's own box: the seven "hard to get back" behaviours each shown or
given a sourced reason, the exact verification commands and results,
and what remains (Android integration, touch-drag-vs-selection
arbitration, row accessibility names, a tappable link, code-span chip,
selection's anchor-row shortcut, code-fence syntax highlighting) --
each also a dated IRIS_TODO.md item so it is not silently dropped.
Ticked [~] rather than [x]: the widget-tree half is built and tested,
the emulator half is not.

"Where things stand" and the Recommendation's item 3 updated in place
to say plainly that neither Masonry (E2) nor iris (I5) has produced a
render number yet, and why -- not a bad measurement, no measurement
obtainable yet on either side -- with the structural findings that do
exist (iris now does cross-row selection and per-span inline rich text,
neither of which exists in masonry/masonry_core/xilem today) recorded
as what currently favours iris absent a number.

IRIS.md gets SpanStyle's own entry: what changed, why, and the one
thing a future TextBuilderOutput impl must remember (both TextOutput
and TextEditOutput apply .spans() -- this box shipped the bug of
missing one of the pair once already).

Co-Authored-By: Claude Sonnet <noreply@anthropic.com>
2026-09-05 07:55:04 -04:00
irisandClaude Sonnet 3f25e7ebca iris: transcript-ui, the transcript screen (RUST.md's I5)
A new workspace member, iris/transcript-ui/, built the same way
tabs-ui is: generic over Rsc: HasEvents + Rsc::State: FocusHost, on
client-core/event-model by path (real code, matching E2's precedent).
Four modules:

- markdown.rs: CommonMark (pulldown-cmark) -> one plain string plus a
  Vec<SpanStyle>, so a row's headings/bold/italic/inline-code/links
  render inline inside one wrapped TextEdit rather than one widget per
  block -- the actual proof that iris can do what E2 found Masonry
  structurally unable to (masonry/src/widgets/text_area.rs's
  "TODO: RichTextInput").
- row.rs: one iris::widget::List row per folded TranscriptRow. A
  TranscriptRow::Tools group collapses to a summary and expands to
  every call's own tool/input/output on tap, using List::extent +
  note_tap for hold-the-edge exactly as list.rs's module doc describes.
- selection.rs: cross-row selection -- a drag that starts in one row's
  TextEdit and crosses into another's, coordinating each visible row's
  own select/select_all/deselect from one pointer gesture. The one
  Masonry's own text_area.rs cites as impossible (no
  SelectionContainer-shaped type anywhere in masonry/masonry_core/
  xilem).
- composer.rs: a growing multi-line composer with no fixed height,
  wired beside the list with .height(rest(1)) -- the real screen for
  IRIS_TODO.md's "input box" benchmark case.

9 new tests (5 pure markdown, 4 selection), all passing. Screenshotted
via run-headless.sh: real inline rich text visible (bold, italic,
inline code, a bigger bold heading, a coloured link, a monospaced
fenced block, a collapsed tool-call row).

What this box does not close, each recorded at its own point (RUST.md's
I5 box, IRIS_TODO.md's dated entries): no Android integration exists
yet for this screen (no cdylib/Gradle shell the way iris-android-app
wraps tabs-ui), so the emulator-side render-number pass condition was
not attempted; a touch-drag pan over a row's own text currently loses
gesture arbitration to that row's own drag-select (diagnosed and named,
not silently broken); row-level accessibility names, a tappable link, a
code-span background chip, and Selection's anchor-row shortcut are
scoped shortcuts recorded in place.

cargo fmt/build/clippy/test --workspace clean; cargo ndk -t x86_64
-P 26 build/clippy clean for both transcript-ui and iris.

Co-Authored-By: Claude Sonnet <noreply@anthropic.com>
2026-09-05 07:54:51 -04:00
irisandClaude Sonnet 0af4c88d08 iris: SpanStyle, per-range text styling (RUST.md's I5)
A TextBuffer used to have exactly one style for its whole string,
applied via parley's push_default. SpanStyle adds a second, optional
layer -- a byte range plus whichever of colour/family/font size/
bold/italic/underline it overrides, pushed with parley's own
push(property, range) -- so a heading, bold, inline code and a link can
each carry their own look inside one wrapped, selectable TextEdit. This
is the actual answer to RUST.md's E2 finding against Masonry
(TextArea::edit_styles() is one StyleSet for the whole editor).

PlacedGlyph gains a color field, read from parley's own per-run
Style::brush, and Painter::glyphs draws each glyph in its own colour
instead of one colour for the whole RenderedText.

Real bug found while wiring this into a live screen (not caught by any
test, since markdown's own tests only check string/range logic): spans
were threaded through TextOutput::run but not the sibling
TextEditOutput::run, so every editable field silently dropped them.
Fixed in build.rs; see IRIS.md's entry for why both call sites are a
pair to keep in sync.

cargo fmt/build/clippy/test --workspace and cargo ndk (iris,
iris-android excluded per its own workspace exclusion) all clean; 28
existing iris tests unaffected.

Co-Authored-By: Claude Sonnet <noreply@anthropic.com>
2026-09-05 07:54:28 -04:00
irisandClaude Opus 5 6bdec6e785 Let a session resume itself when its usage limit lifts
Off by default and per session: it spends quota the moment quota exists,
with nobody watching, which is not a thing a default may decide. Switched
on from the session settings dialog, with the message it sends editable
("continue" unless something else is typed).

Running out of quota becomes a state rather than an error. The Claude
driver recognises its dialect's sentence -- `Claude AI usage limit
reached|1788546972` -- and reports `LimitReached` with the reset time it
gave; nothing above a driver matches on a string. The transcript draws it
as a divider, like a clear or a compaction.

The schedule is a plan to *ask*, never a plan to send. Both reset times
available are untrustworthy in the direction that matters -- the dialect's
is written when the turn fails, the endpoint's moves when the window does
-- so the wait ends in a question to the usage meter, and only `ok` with
no window at 100% sends anything. A window still spent reschedules to its
own reset time, which is what makes a limit that lifts late wait longer
and one that lifts early resume sooner. A meter that cannot be asked is a
longer wait too, never a send. A day after the limit was hit the wait
gives up and says so in the transcript, so a machine that can never be
asked is not retried for ever.

The schedule is persisted on the session: a five-hour window outlasts a
backend restart, and a wait forgotten across one never comes back.

Driven end to end with echo, never a real account: `/limit [minutes]`
reports the same event a real driver does and `/usage` sets what the meter
answers, deliberately separate so the two can disagree. The wait moved
from the dialect's two minutes to the meter's seven when the meter changed
its mind, and the message went out on the first check after the meter came
back under the limit.

Also makes the settings dialog scrollable, which these two controls made
necessary: at a 1.5x system font it clipped the last of them with nothing
on screen to say so.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-05 05:21:43 -04:00
irisandClaude Opus 5 4821a02bd3 Default thinking level for new sessions, and move the rigs out of AGENTS.md
`Config::default_effort` is what a session starts at when nothing chose one,
applied in `spawn_session` rather than filled in by the spawn screen so it
holds for an import and a bare API call too. It is set by the spawn screen's
own picker, whose label says so: one control, where new sessions are made,
rather than a settings page for a single value. Not on a provider, because
providers are discovered and the next rediscovery would erase it; not on the
phone, because a second device would then spawn at a level nobody there
chose. `GET`/`POST /defaults` carry it as a struct, so the permission mode --
still hardcoded to `auto` on the spawn screen -- can move there later without
a second route.

Only drivers that read a level are given one: an echo session was storing a
`--effort` it never passes to anything, which is a config file answering a
question about itself wrongly.

Separately, `AGENTS.md` is 35 KB sent with every request in this repo, and 12
KB of it was rigs and reference measurements that only matter once you are
running one. Those are the `ai-app-rigs` skill now -- the same text, still the
only copy, read when the work touches it. 35,198 -> 20,813 chars.

Verified on the emulator against the sandbox: the spawn screen pre-fills from
the server, picking `low` spawned a session at `low` and left `/defaults` set
to it, and an echo session spawned afterwards took no level at all.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-04 21:42:05 -04:00
irisandClaude Opus 5 1ff662c7c3 Let a session choose how hard it thinks
Output is about an eighth of what a session costs and thinking is nearly
all of it -- prose is ~1.5% of output tokens, measured over 27,015 requests
of this account's own transcripts -- so the level is the largest saving
available short of shortening the conversation itself.

Shaped like the working directory rather than like the model: the CLI's
only two setting control requests are `set_model` and `set_permission_mode`
(checked against the 2.1.258 binary), so `--effort` is read when the process
launches and cannot be asked of a running one. `set_session_effort` records
the level and stops the process; the next message or Start launches one that
has it. That is also why the picker is in the session settings dialog beside
Move, and not on the bar beside the model and the mode, which take effect
mid-turn.

`None` is a level in its own right -- the CLI's own default -- so the picker
can return to it, and a blank is normalized to it at the boundary rather
than stored as a level the CLI would reject.

Offered only where it means something: `DriverKind::takes_effort` reports
the capability and the phone leaves the row out entirely, rather than the
session-type branch this app does not have anywhere else. A llama session
would otherwise get a control whose only effect is stopping its process.

Verified on the emulator against the sandbox's fake CLI: the picker sets it,
the server reports it, and an echo session's dialog is unchanged.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-04 21:23:58 -04:00
irisandClaude Opus 5 e4f0935f98 Keep the second auth test under a subscriber, so the tripwire is not flaky
`gates_every_route_and_never_logs_the_token` failed about one full-suite
run in ten, on the assertion that a rejection *was* logged. Its sibling
ends with an unauthenticated request of its own, made with no subscriber
on that thread -- and tracing caches a callsite's interest process-wide
the first time it is reached, so whichever test got there first decided
whether the warning would ever be recorded.

That is the rule already written at the top of "Things that have bitten",
applied to one member of a set: the combined gating+logging test exists
because of it, and the enrollment test added later did not get it.
Twenty runs clean since.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-04 17:58:11 -04:00
iris 3c0214ece8 Merge branch 'main' of git.arirex.me:iris/ai-app
# Conflicts:
#	AGENTS.md
#	PLAN.md
#	app/androidApp/src/main/kotlin/com/example/aiapp/SessionUsageBar.kt
#	app/androidApp/src/main/kotlin/com/example/aiapp/SpawnScreen.kt
#	server/src/config.rs
#	server/src/main.rs
#	server/src/routes.rs
#	server/src/session/echo.rs
#	server/src/session/llama.rs
#	server/src/session/transport.rs
#	server/src/ssh.rs
#	server/src/usage.rs
2026-09-04 17:56:50 -04:00
irisandClaude Opus 5 127b25e60a Meter a session by its provider, and let llama.cpp run over ssh
The rate-limit bar answered a question about an account, and picked the
answer by machine. One machine runs echo, the Claude CLI and a local
model side by side, so every echo session on it drew the CLI's five-hour
window: a quota that session cannot spend and could never run down. A
session now names its meter (`usageProvider`, from
`DriverKind::usage_provider`, which `usage::providers_for` reads too so
the two lists cannot disagree), and the phone matches on machine *and*
provider. Nothing meters echo or llama, and nothing at all is drawn --
including while the first fetch is out, since "checking" under a session
that turns out to meter nothing is a row the screen then withdraws.

Echo gets a meter it can be *told* about instead: `/usage 42`,
`/usage 95 20`, `/usage 42 never`, `/usage notloggedin`,
`/usage unreachable`, `/usage failed`, `/usage off`. Those states cost
real quota to arrange, which is why none of them had been looked at.

And llama.cpp runs wherever a setup says, which was the last of phase 5.
`Transport::reserve_port` is the second half of what a transport is --
"run this" plus "reach this port" -- returning the port the server binds
there and the port that reaches it here, and `Launch::reaching` puts the
`-L` tunnel on the connection that already carries the command. Three
things that came out of building it:

- A forwarded launch gets a pty and every other one keeps `-T`. Killing
  the ssh client ends a CLI by closing the stdin it reads; llama-server
  never reads its stdin, so the same kill left it running on the far
  machine with the model loaded -- one orphan per stopped session.
- The model is looked for on the machine that will serve it, at that
  machine's own models directory, so `GET /setups/{id}/models` is what
  the spawn screen offers rather than the backend's own downloads.
- The readiness poll watches the process, not only the port: a model
  that will not load exits in a second and would otherwise have been
  reported as "gave up after 300s". The failure carries the log's tail.

Exercised end to end against this VM over ssh to itself: spawn, load,
answer, outlive a backend restart, be adopted, answer again, and stop --
with both the ssh client and the far llama-server gone afterwards. The
local path, the Claude bar and the spawn screen checked on the emulator.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-04 17:45:32 -04:00
108 changed files with 17907 additions and 789 deletions

No files matched your search

+244
View File
@@ -0,0 +1,244 @@
---
name: ai-app-rigs
description: ai-app's test rigs, harness scripts and reference measurements - ui-sandbox.sh, debug-transcript.sh, transcript-bench.sh, stream-bench.sh, trace-draw.sh, the /usage fixture vocabulary, the fake CLI, the rule that no UI-driving script may tap a coordinate, how to test llama.cpp and ssh on this machine, how importing behaves, and the scroll/stream/explorer numbers not worth re-measuring. Read before running or writing a benchmark, driving the app's UI from a script, exercising the session lifecycle, testing a llama or remote session, or touching the import screen.
---
# ai-app: rigs, harnesses and measurements
Moved out of `AGENTS.md` on 2026-09-04 so it is read when it is relevant
rather than sent with every request in this repo -- it was 12 KB of the 35 KB
that file cost on every one. Unchanged in the move, and still the only copy.
## The rigs
Each exists because something was invisible without it.
- **`app/ui-sandbox.sh`** — a second `ai-server` with its own `$HOME`, config
and data directory, holding eight invented Claude Code transcripts and a
`claude` that is two lines of shell. **That isolation is the point**: the
import screen lists whatever is in `~/.claude/projects`, which in this VM is
real agent transcripts, so exercising *delete* against the ordinary server
deletes somebody's conversation and exercising *import* starts a real
`--resume` on the owner's account.
Its port and root derive from the checkout's name, so two checkouts'
sandboxes cannot reach each other, and its token is generated once into
`~/.config/ai-app/sandbox-token` and carried across restarts along with any
the enrolment flow appended — so the emulator app is enrolled **once** (the
start banner prints the command) and stays enrolled. It shares the real TLS
certificates, because the installed APK pins that CA.
Driving verbs, so none of this is re-derived per session:
`./ui-sandbox.sh spawn [title]` (an echo session, prints its id),
`./ui-sandbox.sh send SID text|@file`, and
`./ui-sandbox.sh api /path [curl args]`.
`./ui-sandbox.sh keep` restarts the server without wiping the sessions and
enrolment already there — for when the fixture under test was expensive to
build; plain `start` wipes them, which is right for the list-screen
fixtures and wrong for that.
It passes `--delay` by default, and `AI_SANDBOX_BIG_MB` puts one large
transcript among the small ones while `AI_SANDBOX_SPAWN_DELAY` makes the
fake CLI slow to start. Both exist because operations that finish in
milliseconds have states on the way that nothing can observe, and an
unobservable state is one where broken and working look identical.
It also builds a fixture tree at the sandbox home's `~/files` for the
explorer, holding the states otherwise only reachable by finding a real
machine in one: an empty directory, a name with a tab and one with an
apostrophe, a binary file, one over `FILE_LIMIT`, one `chmod 000`, a
symlink to a directory and a broken one, a source file per language, and
the three sizes the limits were measured against (`edit-32k.rs`,
`edit-128k.rs`, `big-source.rs`). Point a session at it with
`./ui-sandbox.sh api /sessions/<id>/cwd -X POST -H 'content-type: application/json' -d '{"cwd":"~/files"}'`.
The explorer's 409 is produced by editing the file on the machine
(`printf … > file`) between pressing the pencil and pressing save.
- **`app/debug-transcript.sh`** — a real conversation on the emulator. The
echo driver is the right rig for most things and the wrong one for anything
whose cost scales with what was actually written: a real reply is longer,
is real markdown, and carries tool calls whose input and output are
kilobytes. Two faults were invisible until a real transcript was loaded — a
page of history landing mid-fling threw the reader back to the newest end,
and parsing one real reply took 51ms against 4.6ms for a synthetic one.
`-b` takes the biggest conversation on the machine rather than the newest,
which is what a scrolling test wants; `--stop` takes it down.
It copies the transcript into `/tmp` and gives the server a `HOME` of its
own, so the import can only see the copy — importing spawns `claude
--resume`, and against the real file that is a second CLI writing to a
conversation somebody may still be in. **A transcript never goes in this
repository**: they hold whatever was said, read and written in that
session, and `~/repos` is shared with the host besides.
- **`/usage` in an echo session puts up an invented meter**, which is how the
rate-limit screens' states are reached without spending quota: `/usage 42`,
`/usage 95 20` (minutes left), `/usage 42 never` (the between-blocks window
with no reset time), `/usage 42 unreadable`, `/usage notloggedin`,
`/usage unreachable`, `/usage failed`, `/usage off`. The vocabulary is
`usage::Fixture`'s, since those are its states. With none set an echo
session meters nothing, which is the ordinary case and draws no bar.
- **A fake CLI exercises the process lifecycle without a token.** Point a
`claude_cli` provider's `command` at a two-line script — `#!/bin/sh` and
`cat > /dev/null` — and it behaves the way the lifecycle code cares about:
it holds the fifo open, records a real pid, writes nothing, and dies on a
signal. So adopt, stop, restart and start are all drivable without a real
`--resume` and without spending a turn on somebody's account. Reach for
this when what is under test is *whether a process is running*, and for
`debug-transcript.sh` when it is *what the transcript draws*.
- **`app/transcript-bench.sh`** is the standard scroll measurement: it opens
the first session (or `-k` keeps the current screen), scrolls a fixed
gesture loop, and prints the app's render report — the same one the in-app
copy button produces, whose `on screen:` line names what the viewport was
holding. Compare two runs with the same gestures; the emulator's absolute
frame times transfer nothing, the report's accounting does. Run it either
side of any change under `Markdown*.kt`, `Transcript*.kt` or
`SessionScreen.kt`'s list, and put the report in the commit. The numbers
that move first are the worst `record: one block`, the reparse mean while
streaming, and the draw phase's accounting line.
- **`app/stream-bench.sh [-k] FILE`** is that measurement for a reply still
arriving. It taps "Jump to latest" so the list is pinned to the newest end,
resets the report, sends FILE, waits for the transcript to stop growing,
and prints. Both of those are corrections to a first version that measured
nothing: a transcript parked further back never redraws while a reply
streams into it, and a session is idle at *both* ends of a turn, so polling
for idle answers before the turn has started.
- **`app/trace-draw.sh`** names what a scrolling frame spends inside the
framework, from `atrace` text output with no trace processor needed. It is
how the cost of a layout node per link was attributed to the framework
rather than guessed at.
### Driving the UI
**No script that drives this app's UI presses a coordinate.** Every control
is found by the name it already carries for assistive technology —
`ui-trace record --do "tap 'Session settings'"` — which resolves the label
against the screen at the moment of the gesture and fails the whole run when
it is not there. `app/bench-lib.sh` is what the bench scripts share for it. A
coordinate is a position measured once by hand, and anything that moves the
control makes the tap land on whatever now sits there — the bench then
reports a number that was never measured, which reads exactly like a result.
Both bench scripts pressed the render report at `tap 723 205` until that
button moved into the session settings dialog on 2026-09-03. The check that
none has crept back:
grep -n "tap [0-9]" app/*.sh
Swipes are still coordinates, deliberately: a gesture across a scrolling area
is a distance rather than a control.
**Two traps in the emulator bench loop**, each of which cost a run.
`adb shell pm clear` removes the enrolment and the notification permission
along with the saved anchors, so the next run measures a permission dialog —
re-enrol with the command `ui-sandbox.sh` prints, and
`pm grant … POST_NOTIFICATIONS`. And a saved scroll anchor is per session id,
so the only way two builds start a scroll from the same place is a *fresh
session for each*.
**The emulator is `~/repos/emulator-tools`' business, not this repo's.**
`emu up` creates and boots the AVD named after this checkout — whatever `emu
name` prints, never a name typed out here, since this file is the same in
every clone. `run-android.sh` is that plus a build and an install. The `adb`
on `PATH` after sourcing `android-env.sh` is that repo's wrapper, which fills
in `-s` from the same rule. Gradle does not go through it, so a Gradle init
script from `emulator-tools` runs `emu check` before `installDebug`,
`uninstallDebug` and `connectedAndroidTest` and fails rather than fanning out
to every attached device; when it refuses, say which device you mean at the
moment you use it — `ANDROID_SERIAL=$(emu serial) ./gradlew …`.
### Testing llama.cpp and ssh here
**Both are set up here as of 2026-09-04** and need nothing typed. The
prebuilt CPU llama.cpp lives outside the repo at `~/.local/opt/llama.cpp`
(the 15 MB `ubuntu-x64` release asset) and is symlinked as
`/usr/local/bin/llama-server`, which is what makes **discovery find it over
ssh**: `~/.local/bin` is not on the PATH a non-interactive ssh session gets.
It resolves its own libraries through `$ORIGIN`, so no `LD_LIBRARY_PATH` is
needed. One model is downloaded — `unsloth/Qwen3-0.6B-GGUF/Qwen3-0.6B-Q8_0.gguf`,
639 MB under `~/.local/share/ai-app/models` — and answers at usable speed on
this VM's 8 cores. **Do not test with a 2-bit quant**: the
IQ2_XXS of that model produces fluent nonsense, which reads exactly like a
broken driver — `llama-cli` produces the same from the file directly, which
is how to tell the two apart in a hurry.
There is no second machine, so **ssh this VM to itself**. That is set up
too: the key is `~/.config/ai-app/ssh-self` (its public half is in
`~/.ssh/authorized_keys`, labelled removable), and the real config carries a
setup called **"this vm over ssh"** — `bob@127.0.0.1` with that
`identityFile` plus
`options: ["StrictHostKeyChecking=no", "UserKnownHostsFile=/tmp/ai-app-known-hosts"]`
so it touches nothing real — offering `claude-cli` and `llama-cpp`. It is the
whole rig for "does a remote llama session work", since the far machine is
this one and the model file is the same file. For a throwaway setup of your
own, point a provider's `command` at something harmless like `/bin/echo`
rather than at `claude`: the transport is what is under test, the process
exiting immediately is the signal, and it costs no tokens. The remote login
shell here is **fish**; the
remote script and `ssh.rs`'s POSIX quoting happen to mean the same thing in
both, but that is luck rather than design, and a shell that is neither is the
thing to suspect first if a remote spawn ever mangles an argument.
## Importing
The import list reports each session's **size as well as its line count**,
because the two disagree in the way that matters: these transcripts embed
screenshots as base64, so one line can be a megabyte. On this machine a 69 MB
session has 3,427 lines and a 44 MB one has 6,792 — nothing about a line
count tells you what continuing a session will cost. Shown, not warned about;
importing a large session is a choice somebody is entitled to make.
**Never import a Claude Code session that is open in a terminal.** The app
refuses it — see PLAN.md for the incident that made that a refusal rather
than a warning.
**One Claude Code session id can name two files, and the listing offers it
once.** Resuming from a different working directory makes the CLI write a
second transcript with the same id under that directory's project folder — an
ordinary state of a machine, not corruption. Everything downstream addresses
a session by id, and the phone keyed its list on it, so two rows sharing one
**closed the app** on a Compose duplicate-key throw. `parse_listing` keeps
the copy with the most lines, because the other is usually a few-hundred-byte
stub and is often the *newer* of the two, so recency is the wrong key.
Deleting removes every copy rather than the first, or the row came back after
a delete that reported success. The phone's half is `uniqueItems`, which
every list keyed on a server-chosen id goes through: a repeat there must
never be able to close the app, whatever produced it.
**Deleting a session offers to take the machine's own transcript with it**
`DELETE /sessions/{id}?deleteForeign=true`, behind a switch in the
confirmation, and only where the driver keeps a record of its own
(`keepsOwnTranscript`, which today means Claude Code). Off by default,
because leaving that copy is what makes an ordinary delete recoverable — and
the dialog's paragraph is rewritten when it is on rather than appended to,
since the sentence promising the conversation "should still be there to
import again" is exactly the one the switch makes false. The server deletes
the machine's copy *first*, so a machine it cannot reach leaves the session
where it was instead of half-deleted.
## Measurements worth not re-taking
- **What the transcript screen costs to scroll.** Taken 2026-08-30 on the GPU
emulator against a real imported transcript with the server at
`--delay 120`. Settled and flinging fast, both into fresh history and back
through rows already drawn: **5.25.9% janky frames, 99th percentile
2932ms, 02 slow UI-thread frames.** The stock Settings app on the same
device is 3.3% and 38ms, so this is at the platform floor. The number that
is *not* at the floor is the first few seconds after opening a session,
where every row on the way is being composed for the first time; that is
inherent to a lazy list and it is why a measurement taken before the screen
settles reads three times worse. **Settle first, then reset `gfxinfo`.**
- **The reset path is not reachable by reopening a session.** Measured
2026-09-04 against a session streaming at 20 events a second: reopening one
with an anchor 1,800 events back connects **87119 events behind**, well
under `CATCH_UP_LIMIT`'s 200, because the restore is two requests — the
opening page, then one span covering the whole distance. To exercise the
reset at all you have to lower `CATCH_UP_LIMIT` in a throwaway build; at 5
the app takes the reset on a live connection, clears, refills and carries
on without reconnecting.
- **The session screen's stream survives backgrounding here** — 20 seconds at
the launcher while 415 events were produced brought no reconnect at all,
which is not what the comment above that loop expects, and is most likely
this emulator being headless rather than the phone's behaviour.
- **Reopening a cached session costs one request for one event** (the probe),
and scrolling the whole conversation back costs nothing more; a cold open
of the same 500-event session is two pages, 100 events. Measured
2026-09-04 on the emulator against the sandbox.
- **Reading is cheap and editing is not.** The viewer handles a 1 MiB,
28,000-line file because it draws one row per line; the editor is one
`BasicTextField`, which costs two seconds a frame at 128 kB and stops the
app at 1 MiB, so `EDIT_LIMIT` caps it at 32 kB with the reason said on
screen. If you make the editor faster, that number is what to move.
EXPLORER.md's "What the measurements said" has the rest.
+49 -14
View File
@@ -5,11 +5,14 @@ replacing the Claude app for daily use. Rust/Axum backend on the desktop,
Kotlin/Compose Android app, WireGuard + pinned self-signed TLS + bearer token
between them.
**`PLAN.md` is the design source of truth** — every decision with its date,
its rationale, and what was rejected. Read it before changing anything
**`docs/PLAN.md` is the design source of truth** — every decision with its
date, its rationale, and what was rejected. Read it before changing anything
structural, and update it in place when a decision changes rather than
letting this file and the plan become two versions of the truth. This file is
the working notes layer: layout, commands, rigs, and things that have bitten.
The design and working documents live under `docs/` — everything except this
file and `CLAUDE.md`, which stay at the root because that is where Claude
Code and other agent harnesses look for them.
The central design point, worth not undoing by accident: **a session is a
child process, translated into one common event model.** A new session type
@@ -22,7 +25,7 @@ Mirrors `../dev-updater` deliberately: same stack (axum 0.8 +
axum-server/rustls, tokio, clap; Kotlin 2.4.x + Compose Multiplatform, single
`:androidApp` module), same cert scheme, same registry pattern. Read
dev-updater's `README.md` and `AGENTS.md` before diverging from them.
Module-by-module intent is in PLAN.md's "Backend layout".
Module-by-module intent is in `docs/PLAN.md`'s "Backend layout".
- `server/` — the Rust backend (`ai-server`). `routes.rs`'s module doc
comment is the HTTP table and the surface's source of truth.
@@ -40,16 +43,23 @@ Module-by-module intent is in PLAN.md's "Backend layout".
projects version-locked to the commit this repo pins. What deliberately did
**not** move is the API surface and the config *schema*: routes, drivers,
sessions and setups are what makes this project itself.
- `EXPLORER.md` — the file explorer's design (`server/src/files.rs` and
`FilesScreen.kt` / `FileViewer.kt` / `FileEditor.kt`).
- `TRANSCRIPT_CACHE.md` — the phone's copy of what it has been sent. Read it
before touching `TranscriptCache.kt`, `TranscriptSource.kt`, or the opening
and stream effects in `SessionScreen.kt`.
- `TODO.md` — the working list.
- `RUST.md` — the plan for moving the app to Rust (on the `rustify`
- `docs/` — every design and working document except this file and
`CLAUDE.md`:
- `docs/EXPLORER.md` — the file explorer's design (`server/src/files.rs`
and `FilesScreen.kt` / `FileViewer.kt` / `FileEditor.kt`).
- `docs/TRANSCRIPT_CACHE.md` — the phone's copy of what it has been sent.
Read it before touching `TranscriptCache.kt`, `TranscriptSource.kt`, or
the opening and stream effects in `SessionScreen.kt`.
- `docs/TODO.md` — the working list.
- `docs/RUST.md` — the plan for moving the app to Rust (on the `rustify`
branch of the `ai-app-2` clone): what has to be reproduced, the
framework decision, and the ordered experiments with their pass
conditions. Read it before touching anything under that branch.
- `docs/IRIS.md`, `docs/IRIS_TODO.md`, `docs/DECISIONS.md`,
`docs/LAYOUT.md`, `docs/TEXTURES.md`, `docs/CLIENT_CORE.md` — iris's
own public API log, working list, decisions log, layout/render design,
and texture-atlas design, and the client-core crate's design,
respectively.
- `.dev-updater.ron` — what Dev Updater builds here: the server (run as
`service: Managed(…)`, supervised by Dev Updater's own implementation
rather than a script kept here) and the APK, in parallel. It points at
@@ -86,7 +96,11 @@ two icon buttons the same width without either being given one — and why
:androidApp:compileDebugKotlin :androidApp:lintDebug
:androidApp:testDebugUnitTest`. The unit tests are JVM-only and cover the
syntax highlighter, the ANSI parser and the transcript cache — the app's
pure logic with no Android in it.
pure logic with no Android in it. Touching anything under `BenchFixture.kt`,
`BenchNetwork.kt`, `BenchRun.kt` or the `bench` build type also needs
`:androidApp:compileBenchKotlin :androidApp:lintBench` — a second build
type compiles separately and lint has caught real bugs debug alone never
would (see "Android Lint" below).
- **Android Lint is not optional and is not run by a build.** It found a
crash that had been shipping (`java.time` on a minSdk-24 app with
desugaring off) and later a permission check that silently dropped every
@@ -146,6 +160,27 @@ two icon buttons the same width without either being given one — and why
Each exists because something was invisible without it.
- **The `bench` build type and `app/bench-fixture/`** exist for P0 (RUST.md
and DECISIONS.md's 2026-09-05 entries), the phone benchmark gate Iris
asked for before porting continues: a deterministic, checked-in synthetic
transcript (`app/bench-fixture/generate.py`, never a real one) that both
this app and iris open with no server, so a frame-time comparison
measures the renderer rather than the data. `./build-apk.sh bench` builds
it — own application id (`com.example.aiapp.bench`) and label ("AI
Sessions bench") so it installs beside a real enrollment rather than
replacing it. Opening it goes straight to a session screen holding the
fixture (no enrollment, no permission prompts) with a "Run benchmark"
control beside "Copy" in session settings: it drives the same scroll loop
and streaming phase `transcript-bench.sh`/`stream-bench.sh` drive over
`ui-trace`, but in-process, since a real phone has no usable system
tracing and no agent can drive one (this-machine-android's skill).
`BenchFixture.kt`/`BenchNetwork.kt` fake the backend by installing a
`URLStreamHandlerFactory` that answers `TranscriptSource`/`EventStream`'s
requests from an in-memory copy of the fixture instead of opening a
socket — so the fold, the paging and `uniqueItems` under test are the
screen's real ones, never a shortcut built just for this. The report
gains a `bench:` section (process CPU time, peak RSS, battery current) on
every build, empty except when `BenchRun.kt` filled it in.
- **`app/ui-sandbox.sh`** — a second `ai-server` with its own `$HOME`, config
and data directory, holding eight invented Claude Code transcripts and a
`claude` that is two lines of shell. **That isolation is the point**: the
@@ -312,7 +347,7 @@ means here:
## Sessions outlive the backend
Since 2026-08-29 a session's process is deliberately left running when
`ai-server` stops, and adopted again when it starts. PLAN.md has the design;
`ai-server` stops, and adopted again when it starts. docs/PLAN.md has the design;
day to day:
- **Stopping the server no longer stops the sessions.** After `pkill
@@ -345,7 +380,7 @@ count tells you what continuing a session will cost. Shown, not warned about;
importing a large session is a choice somebody is entitled to make.
**Never import a Claude Code session that is open in a terminal.** The app
refuses it — see PLAN.md for the incident that made that a refusal rather
refuses it — see docs/PLAN.md for the incident that made that a refusal rather
than a warning.
**One Claude Code session id can name two files, and the listing offers it
@@ -521,4 +556,4 @@ belongs in `~/.claude/TOOLCHAIN.md` or `~/.claude/MACHINE.md` instead.
`BasicTextField`, which costs two seconds a frame at 128 kB and stops the
app at 1 MiB, so `EDIT_LIMIT` caps it at 32 kB with the reason said on
screen. If you make the editor faster, that number is what to move.
EXPLORER.md's "What the measurements said" has the rest.
docs/EXPLORER.md's "What the measurements said" has the rest.
+44
View File
@@ -0,0 +1,44 @@
# Decisions awaiting review
Choices made while working autonomously, for Bryan to keep or change. Each
says what was picked and why; the detail is in the design doc it names.
Delete an entry once it has been looked at.
## Subagent views (2026-09-05, `SUBAGENTS.md`)
Made on my own judgement, limited blast radius:
1. **A subagent is a transcript, not a session.** It has no process,
controls or settings; it is addressed as `/sessions/{id}/subagents/{sub}`
and stored under the session's directory, so deleting the session takes
it. Alternative rejected: registering it as a session of its own, which
would give it a card in the main list and a driver that can do nothing.
2. **Read-only view is the session screen minus its controls**, rather than
a second, simpler transcript screen. Keeps paging, caching, selection
and rendering in one place. Cost: a `readOnly` mode threaded through
`SessionScreen`.
3. **The list only carries a count.** Each session row says how many
subagents it has; their titles and statuses are fetched when the card is
expanded. Keeps `GET /sessions` from reading every subagent transcript.
Consequence: an expanded card's statuses refresh with the list, not live.
4. **Expanded/collapsed is remembered per session on the phone**, not on
the server. Collapsed by default, per the transcript convention that new
things arrive collapsed.
5. **Subagents of imported sessions are not shown.** The import path still
skips `isSidechain` records; the CLI's own `subagents/agent-*.jsonl` files
are not read. Only subagents run while this backend was watching exist.
6. **Echo grows `/subagent [n]`** as the test rig, so nothing here needs a
paid turn to exercise.
Deferred, because they reach further than this feature:
- **Live status on the list.** Whether the session list should follow a
stream at all (it refreshes on demand today) decides whether subagent
status can ever be live there. Not changed.
- **Nested subagents.** A subagent's own Task calls are shown as tool calls
in its transcript and are not given transcripts of their own. Supporting
that is the same mechanism one level down, but the UI would need nested
expanders.
- **The subagent status row says "context unknown".** Nothing measures a
subagent's context; the row could leave it out rather than admit it.
-171
View File
@@ -1,171 +0,0 @@
# iris: notable public API changes
For Iris to read on her own time. Each entry is a change to iris's public
surface that a widget author or app author would notice: a trait method
added, removed or re-shaped; a type that callers construct differently; a
capability that moved. Small and trivial changes do not go here.
An entry gives the date, what changed, why, and a short before/after where
it helps judge the change without the session that made it. Newest first.
## 2026-09-05: accessibility names via AccessKit (RUST.md's I4)
`.label()` (already in `trait_fns.rs`, previously unused anywhere in-tree)
is now load-bearing: it's the one thing that puts a widget in the AccessKit
tree `iris_core::ui::access::AccessTree` builds and both backends push
out. A widget author who wants a control to be findable by name (and
tappable by name, through `ui-trace`/a real screen reader) calls `.label()`
on it; nothing else is required, and a widget nobody labels is invisible
to this system at zero cost, not just zero UI.
```rust
let button = rect(Color::LIME)
.on(CursorSense::click(), move |_, rsc| { ... })
.label("Add task"); // now findable by uiautomator/AccessKit as "Add task"
```
Two new things a widget author might touch directly:
- **`Widget::access_role(&self) -> accesskit::Role`**, default `Unknown`.
Override it if your widget has a real platform equivalent —
`TextEdit` now returns `TextInput`/`MultilineTextInput` by `EditMode`.
Only consulted for a widget that also has a `.label()`; an unlabelled
widget's `access_role` is never called.
- **`Widgets::named() -> impl Iterator<Item = WidgetId>`** — every widget
with an explicit label, for anything else that wants to walk the same
set `AccessTree` does.
Nothing about `Painter`, `draw`, or the layout/move machinery changed —
this sits entirely beside them, reading `resolved_region`'s output rather
than participating in producing it.
## 2026-09-05: `List`, a virtualised bottom-anchored list (RUST.md's I3)
A new widget, `iris::widget::List` (`iris/src/widget/list.rs` -- read its
module doc first), for the transcript's kind of screen: variable-height
rows, keyed by a `u64`, composed only while visible, moved rather than
re-laid-out on scroll, a scroll anchor that survives a row inserted above
it, "more" sentinels at each end, and "hold the edge nearest the tap" when
a row's height changes (`note_tap`, resolved in the layout pass).
```rust
let mut list = List::new(Axis::Y);
list.push_back(ListRow::new(key, row_widget)); // O(1)
list.push_front(ListRow::new(older_key, row)); // O(1), anchor unaffected
list.set_more_before(Some(spinner_widget)); // sentinel, drawn at the edge
list.note_tap(viewport_y); // before mutating a row's height
let (top, bottom) = list.extent(key).unwrap(); // last frame's on-screen box, if visible
```
Built entirely out of existing primitives (`Painter::widget`/`widget_within`/
`reposition`/`draw_twice`, and `draw_inner`'s own old-children diffing) --
no new mechanism was added to the render core for it. One correctness
lesson worth reading even for other widgets: a row that fills whatever
region it is offered (`Rect`, `is_size_independent`) cannot be measured at
a throwaway oversized region and then merely `reposition`ed into place --
`reposition` only ever writes an offset, never a size, so the oversized
primitive stays oversized. `List` fixes this by caching each row's real
height once measured and placing an already-known row directly at its
exact box; see `list.rs`'s `place` for the full reasoning and
`a_fill_shaped_background_is_not_left_oversized` for the regression test.
## 2026-09-05: a second backend (android-view), and what moved to make room for it
RUST.md's I2. Three changes a widget or app author would notice, all in
service of the same thing: `default` (winit) and the new `android`
(android-view) backends sharing what does not depend on windowing.
- **`Selector`/`Selectable`'s bound changed from `Rsc::State:
HasDefaultUiState` to `Rsc::State: FocusHost`** (new trait, `attr.rs`).
`HasDefaultUiState` still exists and still works — `default/attr.rs` now
implements `FocusHost` for anything that has it — so a winit app's
existing code is unaffected. An Android app implements `FocusHost` via
`HasAndroidUiState` instead. Affects only an app that referenced
`HasDefaultUiState` directly at a `Selectable`/`Selector` call site
rather than through `.attr::<Selectable>(())`, which nothing in-tree
does.
- **`Tasks::init` takes `Arc<dyn RequestRedraw>` instead of
`Arc<winit::window::Window>`.** `RequestRedraw` (`task.rs`) is one method,
`fn request_redraw(&self)`; `winit::window::Window` implements it
(`default/render.rs`), so `Tasks::init(window)` at a call site is
unchanged by inference. Only matters if something constructed a `Tasks`
directly rather than through `DefaultRsc`/`AndroidRsc`.
- **`TextEdit::apply_event`/`TextInputResult` are `#[cfg(not(target_os =
"android"))]`** — they take a `winit::event::KeyEvent`, which does not
exist on Android; `android/input.rs` drives the same primitives
(`backspace`/`delete`/`motion`/`insert`, all still unconditional) from
`ndk::event::Keycode` directly instead. New unconditional getters on the
way: `TextEdit::text()`/`selection_range()`/`caret()`, and
`TextEditCtx::delete_byte_range`/`set_cursor_byte` — the primitives
`android/ime.rs`'s `InputConnection` bridge needed and that were not
previously exposed publicly.
## 2026-09-04: `Widget::draw` reports the size it used; `desired_width`/`desired_height` are gone
A widget used to implement three methods (`draw`, `desired_width`,
`desired_height`); it now implements one, `fn draw(&mut self, painter: &mut
Painter) -> Size`, which draws into `painter.region()` and returns how much
of it was used. Why: the two extra methods routinely re-simulated what
`draw` was about to do anyway (`Span::desired_ortho` copied its own draw
loop to get cross-axis sizing right) — one visit per widget per frame
instead of up to three. A container that needs a child's size before
placing it (alignment, centering) draws the child once at a provisional
region, reads the returned `Size`, and calls the new `Painter::reposition`
to move it into its final spot — an O(1) offset write, not a second draw. A
widget whose drawn output never depends on the size it's given (a
fixed-size `Rect`, a decoded `Image`) overrides the new `fn
is_size_independent(&self) -> bool { false }` to `true`, which skips
redrawing it when only its offered region changes shape.
```rust
// before
fn draw(&mut self, painter: &mut Painter) { /* ... */ }
fn desired_width(&mut self, ctx: &mut SizeCtx) -> Len { /* ... */ }
fn desired_height(&mut self, ctx: &mut SizeCtx) -> Len { /* ... */ }
// after
fn draw(&mut self, painter: &mut Painter) -> Size { /* ... */ }
```
`SizeCtx` and `Cache` are gone with it — see `LAYOUT.md` for the full
design, the move-offset mechanism this shipped alongside, and the file
list.
## 2026-09-04: texture pipeline rebuilt off the binding array
`Textures`/`TextureHandle`, `GlyphPrimitive`, and `UiRenderNode::new` all
changed shape. Why: the old pipeline bound every texture ever drawn in one
`binding_array<texture_2d<f32>>` and asked every device, unconditionally,
for `VK_EXT_descriptor_indexing` — a real share of Android GPUs lack it,
and it failed outright on the Android emulator's software Vulkan. See
TEXTURES.md's "Recommended shape" and "Implemented, 2026-09-04".
- **`UiRenderNode::new` drops its `limits: UiLimits` parameter, and
`UiLimits` is gone.** Before: `UiRenderNode::new(&device, &queue,
&config, UiLimits::default())`. After: `UiRenderNode::new(&device,
&queue, &config)`. Nothing replaces it — there are no more
binding-array limits to size.
- **`src/default/render.rs`'s device request asks for no features and no
binding-array limits.** Before: `required_features:
Features::TEXTURE_BINDING_ARRAY | Features::PARTIALLY_BOUND_BINDING_ARRAY
| Features::SAMPLED_TEXTURE_AND_STORAGE_BUFFER_ARRAY_NON_UNIFORM_INDEXING`
plus two `max_binding_array_*` limits. After: `Features::empty()` (the
`DeviceDescriptor` default) and only `max_buffer_size` set, which was
never about the binding array.
- **`TextureHandle` has no `primitive()` method any more**; a caller
outside `iris` shouldn't have been calling it (it fed the old renderer's
internals), but if something did: use `image_index()` for a standalone
image's bind-group index. There is no equivalent for a page — a page has
no bind group of its own now, see below.
- **`GlyphPrimitive` has no public constructor from a struct literal.**
Before: `GlyphPrimitive { uv_min, uv_max, view_idx, sampler_idx, color,
flags }`. After: `GlyphPrimitive::new(uv_min, uv_max, layer, color,
flags)` — one `layer` (the shared atlas array's layer) instead of a
`view_idx`/`sampler_idx` pair, since a page is now a layer of one array
texture rather than its own bound texture.
- **A widget author drawing images is unaffected**: `Painter::texture`/
`texture_at`/`texture_within` and `Textures::add` keep their signatures.
What changed underneath is that each standalone image now gets its own
`wgpu::BindGroup` and draw call instead of a slot in the shared array —
invisible from the widget API, visible only in `UiRenderNode`'s internals
and in `iris`'s device requirements.
-198
View File
@@ -1,198 +0,0 @@
# iris: known problems and things still to build
Iris's own list for the library, recorded 2026-09-04 in her words where it
matters, so the agents working through RUST.md pick these up in a sensible
order rather than rediscovering them. Each item says where it sits in the
order and what "done" looks like. Tick and date them in place.
## Fix
- [x] **Input does not fall through by input type (2026-09-04).**
`SensorUi::run_sensors` (`src/default/sense.rs`) used to set "consumed,
stop checking lower layers" from mere hover — a widget registered for
nothing but `click()` blocked a `Scroll` meant for whatever was behind
it, since "the cursor is over this widget" and "this widget handled the
event" were the same check. Fixed by judging consumption per input
kind: with no button transition and no scroll happening this frame
("momentary" activity), the topmost hovered widget still wins, same as
before; when something momentary *is* happening, only a widget whose
registered senses actually include a matching non-hover one (checked
via a new `TypeEventManager::registered`, which lists what a widget
registered without running anything) consumes it, so a widget with only
`Hovering`/click handlers can no longer block a scroll from reaching a
list underneath. `iris/src/sense_tests.rs` builds a button-over-a-list
`Stack` with a plain `HasEvents` impl (no GPU or window) and checks both
directions: a scroll over the button reaches the list, and a real click
still reaches the button — confirmed to fail on the pre-fix code and
pass after.
- [x] **Appending one image to an already-loaded list rebuilds every other
image's bind group (2026-09-05, fixed 2026-09-05).** Found by the
benchmark below: `GpuTextures::update` (`core/src/render/texture.rs`)
triggered `rebuild_image_bind_groups` — a loop over *every live
standalone image*, rebuilding its `BindGroup` — whenever the shared
`masks` or `move_offsets` GPU buffer was resized (`masks_resized ||
moves_resized` in `UiRenderNode::update`, `core/src/render/mod.rs`), and
a widget getting its *first* move-offset slot (LAYOUT.md section 2 —
every widget gets one on first draw) could be exactly what grows that
buffer. So one new message with one new image, appended to a transcript
that already has N images loaded, did not cost O(1): it cost one
`create_image` for the new image plus one `make_image_bind_group` per
*existing* image, because the new widget's own move slot pushed the
arena past its capacity. Measured directly in
`iris/examples/bench_images.rs`: appending a 1,001st image to 1,000
already-settled ones reported **1,001** bind-group creates for that one
frame, not 1 (`./run-bench.sh images`, frame 5 in the transcript below).
**Fix**: `masks`/`move_offsets` never belonged in a standalone image's own
bind group (group 2) in the first place — the group also holds that
image's own texture view, which is the only thing that is genuinely
per-image, so a buffer shared by *everything* forced a rebuild of
*every* group the moment it moved. Gave masks/move_offsets their own
bind group (group 3 in `shader.wgsl` and `UiRenderNode`: `masks_layout`/
`masks_group`), bound once per frame in `UiRenderNode::draw` rather than
once per draw call, instead of duplicating them into every per-image
group. `GpuTextures` and its image bind groups now know nothing about
either buffer — `rebuild_image_bind_groups` is called only from
`grow_array` (the atlas array texture growing, which genuinely does
change what every image's own bind group must reference) — so a
masks/move_offsets resize now touches exactly one bind group, ever,
regardless of how many images are live. Numbers after the fix, same
benchmark and command:
./run-bench.sh images
frame=1 bind_group_creates=1000 (cold load, unchanged)
frame=2 bind_group_creates=0 (was 1000 -- see the item below)
frame=3 bind_group_creates=0
frame=4 bind_group_creates=0
(append one image here)
frame=5 bind_group_creates=1 (was 1001)
frame=6 bind_group_creates=0
`run-headless.sh tabs --shot` still 27266 bytes, byte-for-byte unchanged,
confirming the bind-group restructuring changed nothing about what is
drawn.
- [x] **Bind-group creation takes two frames to reach the steady state, not
one (2026-09-05, closed by the fix above, 2026-09-05).** Same benchmark:
loading 1,000 images cold used to report 1,000 creates on frame 1
(expected — `create_image`, one per new image) *and again* 1,000 on
frame 2, before settling to 0 from frame 3. This was `rebuild_image_bind_groups`
firing a second time for the same masks/move-offsets buffer-growth
reason as the item above, confirming the guess recorded here — the two
were exactly the same root cause measured two different ways. Frame 2
now reports 0 (see the numbers above); not a separate fix.
## Build
- [x] **Benchmarks**, not unit tests, run on demand (2026-09-05; a
`benches/` or a script under `iris/`, never in `cargo test`). The
scenario that matters most is a **message list** — chat apps and this
app's transcript alike — stressed with many messages and many images.
One case in particular: **resizing an input box** (typing enough text to
grow it) that pushes a long list of messages above it must stay very
fast and recalculate almost nothing — a move of everything above, not a
re-layout. That is exactly the O(1) move chain in LAYOUT.md; the
benchmark is what proves it. Done when the numbers are in this file with
the command, and the input-box case reports draws re-run, not just frame
time.
**Built as two rigs**, chosen per scenario by whether a real `wgpu`
device is needed (`UiRenderState`/`Widgets` touch no GPU or window, so
most of this runs as an ordinary binary — the same property
`layout_tests.rs` relies on):
- `iris/benches/message_list.rs` — a plain `Instant`-timed binary
(`[[bench]] harness = false` in `iris/Cargo.toml`), not criterion: see
the file's own header for why (short version — every scenario here
reduces to a *count* `UiRenderState::take_counters` already produces,
which criterion's statistical machinery adds nothing to and which a
new dependency is not worth pulling in for). Covers (a) first-frame
cost of a message list of N wrapped-text rows (one in 20 also carrying
a small in-memory image) for N = 100/1,000/10,000; (b) per-frame cost
of scrolling that list, 200 ticks; (c) the input-box case — a
fixed-height field at the bottom of the screen growing by a line 40
times, with the message list above it filling the rest of the screen.
Run: `cd iris && cargo bench --bench message_list` (always release —
`cargo bench` builds the `bench` profile, which is optimized).
- `iris/examples/bench_images.rs` — needs a real device, so it runs
through `iris/run-headless.sh bench_images`, printing
`UiRenderNode::take_image_bind_group_creates()` (a new counter, added
in `core/src/render/texture.rs` and `core/src/render/mod.rs`,
mirroring `UiRenderState::take_counters`) each frame. Covers (d): 1,000
image rows, checked both cold (does bind-group creation reach zero
once loaded) and after appending one more image once settled (does
*that* stay cheap) — the second question is what actually matters for
a live transcript and is what turned up the two Fix items above.
- `iris/run-bench.sh [list|images]` runs either or both and is what to
run before/after touching `Scroll`, `Span`, `Sized`, the move-offset
chain, or `GpuTextures`.
**Numbers (2026-09-05, release, `cargo bench`/`run-headless.sh`, this
VM: AMD Ryzen 7 3800X, 8 cores, rustc 1.98.0 nightly-2026-09-03):**
cd iris && cargo bench --bench message_list
(a) first frame, N=100: 30.30ms draws=227 rewrites=15 moves=0
(a) first frame, N=1000: 186.04ms draws=2252 rewrites=150 moves=0
(a) first frame, N=10000:1770.36ms draws=22502 rewrites=1500 moves=0
(b) scroll, N=100/1000/10000, 200 ticks each:
draws=200 rewrites=0 moves=200 (identical at every N)
per-tick average: 0.0002ms (identical at every N)
(c) input grows 40 lines, N=100/1000/10000 rows above it:
draws=320 rewrites=40 moves=160 (identical at every N)
per-line average: 0.0012-0.0013ms (identical at every N)
cd iris && ./run-bench.sh images (2026-09-05, before the fix)
frame=1 bind_group_creates=1000 (cold load)
frame=2 bind_group_creates=1000 (see Fix item above)
frame=3 bind_group_creates=0
frame=4 bind_group_creates=0
(append one image here)
frame=5 bind_group_creates=1001 (see Fix item above)
frame=6 bind_group_creates=0
cd iris && ./run-bench.sh images (2026-09-05, after the fix)
frame=1 bind_group_creates=1000 (cold load, unchanged -- genuine work)
frame=2 bind_group_creates=0
frame=3 bind_group_creates=0
frame=4 bind_group_creates=0
(append one image here)
frame=5 bind_group_creates=1 (one image's own create_image, O(1))
frame=6 bind_group_creates=0
**Reading it**: (a) is real, necessary work — shaping and laying out N
never-before-seen text rows — and scales with N as it must, ~10x cost
per 10x N. (b) and (c) are the pass conditions that matter: both are
**exactly flat across N = 100 to 10,000**, confirming LAYOUT.md's O(1)
move chain holds for both scrolling and for a growing input box pushing
the message list — draws/moves per tick or per line do not grow with
list size, and the per-operation cost (a fraction of a microsecond) is
nowhere near a frame budget. (d)'s cold-load and steady-state halves
behave as designed; its *append* half did not, until the fix above moved
masks/move_offsets out of the per-image bind group — now flat at O(1)
the same way (b) and (c) are.
- [ ] **Masks defined relative to each other.** Wanted: mask A multiplies
by something *and also* applies mask B — a mask can reference a parent
mask, the way the move chain references a parent offset. Today masks
are independent regions. Design it beside the move chain (same shape:
a parent index and a bounded walk in the shader); do it when a real
widget needs it, not before.
- [ ] **Positions as a single float per scroll.** Iris raised, and half
rejected, letting a scroll update one float rather than positions:
input handling cares about most elements in a list, so absolute
positions must be computed on the CPU anyway. LAYOUT.md's design
already lands here (GPU walks the chain, CPU resolves on demand for
hit tests). Keep the CPU resolution lazy and per query; do not
materialise every row's absolute position per frame.
- [ ] **Animations, last.** Cosmetic, so after everything above. Must be
**modular — a piece of the library rather than a core part forced into
everything, the same way input is**. Whatever the mechanism, a widget
that does not animate must pay nothing and import nothing for it.
## Reconsider
- [ ] **`WidgetView`.** Iris is unsure of it: what she wants is an easy way
to compose a widget from others (a button is the main case). With
sizing folded into `draw`, composing may be easy enough that `View` is
redundant. Decide after the layout change lands, by writing a button
both ways and keeping the one that is shorter to explain; delete the
other rather than keeping two ways.
+141
View File
@@ -0,0 +1,141 @@
# Subagents
A session's subagents -- the helpers a Claude Code session starts through its
Task tool -- each get a transcript of their own, listed under the session's
card and readable in the same transcript view the session has. Designed
2026-09-05; the decisions Bryan has not yet reviewed are in `DECISIONS.md`.
## What a subagent is here
**A subagent is a second transcript owned by a session, in the same event
model, with no process and no controls.** It is not a session: it cannot be
messaged, stopped or started, and it has no setup, model or usage of its
own. Everything it shares with a session -- the transcript file format, the
paging routes, the SSE stream, the phone's cache and rendering -- is reused
by addressing, not by copying.
The CLI reports a subagent's messages on the parent's own stream-json
output, each carrying `parent_tool_use_id` = the id of the Task `tool_use`
that started it. Before this the translator dropped those lines
(`subagent_events_are_not_duplicated_into_the_transcript`); now it routes
them to that subagent's own translator and transcript. The parent's
transcript still shows only the Task call itself.
## Storage
Under the session directory:
```
<session>/subagents/<tool_use_id>/meta.json {title, created}
<session>/subagents/<tool_use_id>/transcript.jsonl same SeqEvent lines as the session's
```
The id is the Task tool_use id (`toolu_…`), which is unique, stable across a
backend restart, and already the key everything on the parent side uses.
Only ids matching `[A-Za-z0-9_-]+` are ever created or looked up, since the
id becomes a path.
The transcript's sequence numbers are its own, starting at 1. `Transcript`,
`read_window`, `catch_up` and `read_after` work on it unchanged.
Its path out: deleting the session deletes its directory, subagents included.
There is no separate delete.
## Lifecycle, as events in the subagent's transcript
1. Created on the first child line for an unseen parent id (or, when the
parent Task call was seen, at that call). First lines written:
`Status Running`, then `UserMessage { text: <the Task's prompt> }` when
the prompt is known -- it genuinely is the subagent's first user turn.
2. Every child line is translated by that subagent's own `Translator`
(one per subagent: tool ids are unique but streaming deltas are by
content-block index, and parallel subagents interleave).
3. **The parent's `tool_result` never finishes a subagent.** The Task tool
runs in the background by default: the `tool_result` -- "Async agent
launched..." -- arrives the moment it *starts*, while the subagent goes
on working for however long its own turn takes, sometimes minutes. What
ends it is its own turn ending: the raw API's `message_delta` on its
stream carrying `stop_reason: "end_turn"` (a `stop_reason` of `tool_use`
is the model about to call one, not an end), or a `result` line for its
own turn if a future CLI version ever sends one. Either maps to
`Status Exited`; the subagent's vocabulary has no `Idle`, so the
equivalent event `dispatch` produces for an ordinary session is dropped
rather than written. A shipped version of this finished on the
`tool_result` instead, which read a running background agent as
"finished" with its transcript truncated at the moment it launched.
4. **A child line for a subagent that already finished reopens it**
(`Status Running`) rather than being dropped: a background Task can be
sent another message long after its first turn ended, and that is
exactly what a further line for it means. Same transcript, same child
`Translator`, just picking back up.
5. When the parent session's process exits (`Status Exited` on the
session), every subagent still `Running` gets `Status Exited` too: its
process was the parent's.
A subagent that was mid-flight when the backend restarted keeps working:
the registry reopens the existing transcript on the next child line, and
the file continues its sequence -- the same reopening #4 describes, whether
what closed it was a restart or its own `end_turn`. If its turn ended while
the backend was down nothing recorded that until the next line arrives, so
its last status stays `Running`, which the list reports as **unknown**
rather than as running (see the wire shape) until then.
Title: the Task call's `description` input, then ` (<subagent_type>)` when
one is given; falling back to the tool's name when the child arrives before
(or without) the parent call being seen.
## Server layout
- `session/subagent.rs` -- the registry: `Subagents` (per session, in
`Shared`), `Subagent` (its `Transcript` behind a mutex plus a
`broadcast::Sender<SeqEvent>`), `record(id, event)`, `start(id, title,
prompt)`, `finish(id)`, `reopen(id)`, `finish_all()`, `list()` from disk. Drivers get an
`Arc<Subagents>` beside their `EventSink`; llama ignores it.
- `session/claude/translate.rs` -- routes child lines by parent id, holds
one child `Translator` per subagent, remembers pending Task calls'
description/prompt/subagent_type.
- `session/echo.rs` -- `/subagent [n]`: the test rig. Starts *n* (default 1)
subagents at once, each named "helper k". Each writes the prompt as its
user message, streams a few words of text, runs one `Bash` tool call, then
finishes about three seconds after starting, and the parent's Task calls
end when their subagent does. Three seconds so the running state can be
seen on the phone.
- `routes.rs` -- three routes, in the doc table.
## Wire shape
```
GET /sessions/{id} SessionInfo gains `subagents: N` (count, 0 when none)
GET /sessions same field on each row
GET /sessions/{id}/subagents [{id, title, status, created, lastActivity}], oldest first
GET /sessions/{id}/subagents/{sub}/transcript exactly the session transcript's query and answer
GET /sessions/{id}/subagents/{sub}/events?after=N exactly the session events stream
```
`status` is the transcript's last `Status` event, serialised like a session's
(`running`, `exited`), except that a subagent whose session is not itself
running cannot be running: the list answers `unknown` for that one. The
phone words these as *running*, *finished* and *unknown* on the subcard.
The count on `SessionInfo` is a directory listing, so the list stays cheap.
The per-subagent status is only read when the list route is asked for.
## Phone
- `SessionSummary.subagents: Int`. A card with a non-zero count ends in an
expander row -- a full-width `Chevron(Pointing.Down)` row that flips to
`Pointing.Up` -- collapsed by default. Expanding fetches
`/sessions/{id}/subagents` and draws one `OutlinedCard` per subagent,
indented inside the session card, the way dev-updater draws a project's
components: title, then the status word and a relative time. The
expansion state is per session id and survives a refresh of the list.
- Tapping a subcard opens `Screen.Subagent`, which is `SessionScreen` in
**read-only** form: the same transcript, paging, cache, selection,
images and status row, with the composer, the process button, the model
picker, the files button, the settings cog and the usage bar left out.
The header shows the subagent's title with the session's title beneath
it. Back returns to the list.
- Addressing: `fetchTranscript`, `EventStream`, `TranscriptSource` and the
cache take a transcript address rather than a session id --
`sessions/{id}` or `sessions/{id}/subagents/{sub}` -- so the cache nests a
subagent's copy under its session's and the same code serves both.
+39
View File
@@ -102,6 +102,16 @@ android {
targetSdk = 37
versionCode = 1
versionName = "1.0"
// Read by MainActivity to decide, at startup, whether this is the P0 benchmark build
// (docs/RUST.md's P0 box) rather than the app somebody enrolled. False everywhere except
// the `bench` build type below, which overrides it.
buildConfigField("boolean", "FIXTURE_MODE", "false")
}
buildFeatures {
// Only for FIXTURE_MODE above; nothing else here reaches for generated BuildConfig fields.
buildConfig = true
// Only for the bench build type's resValue("string", "app_name", ...) below.
resValues = true
}
packaging {
resources { excludes += "/META-INF/{AL2.0,LGPL2.1}" }
@@ -131,6 +141,35 @@ android {
isMinifyEnabled = false
if (keystore != null) signingConfig = signingConfigs.getByName("release")
}
// P0's benchmark build (docs/RUST.md, docs/DECISIONS.md's 2026-09-05 entry): release
// optimisations so a frame time measured here means what release means everywhere else in
// this project, its own application id so it installs beside a real enrollment rather than
// replacing it, and FIXTURE_MODE so MainActivity opens straight onto the fixture session
// instead of asking to be enrolled. Signed with the same key as release -- it never talks
// to a real backend, so there is no CA of its own to mismatch, and a second keystore would
// be one more secret to keep off this machine's shared mount for no benefit.
create("bench") {
initWith(getByName("release"))
// :link (wg-app-link) has no "bench" build type of its own -- it is a library shared
// with dev-updater and has no reason to know this project invented one -- so this says
// which of its build types to link against instead.
matchingFallbacks += listOf("release")
applicationIdSuffix = ".bench"
// "AI Sessions bench" everywhere the OS shows the app's name (launcher, recents,
// Settings): this resValue overrides res/values/strings.xml's app_name for this
// build type alone, and AndroidManifest.xml's android:label reads @string/app_name
// rather than a literal so a build type can override it without touching the
// manifest.
resValue("string", "app_name", "AI Sessions bench")
buildConfigField("boolean", "FIXTURE_MODE", "true")
if (keystore != null) signingConfig = signingConfigs.getByName("release")
}
}
sourceSets {
// The fixture both bench builds (this one and iris's) open with; see
// app/bench-fixture/README.md. Read directly from its own directory rather than copied
// into androidApp/src -- one file to keep in sync with the generator, not two.
getByName("bench").assets.directories.add("../bench-fixture/assets")
}
compileOptions {
sourceCompatibility = JavaVersion.VERSION_21
+1 -1
View File
@@ -25,7 +25,7 @@
the fix is a judgement about how this app should look. Drop this
suppression when a real icon lands. -->
<application
android:label="AI Sessions"
android:label="@string/app_name"
android:allowBackup="true"
android:theme="@android:style/Theme.Material.Light.NoActionBar"
tools:ignore="MissingApplicationIcon">
@@ -142,6 +142,20 @@ data class SessionSummary(
val keepsOwnTranscript: Boolean,
/** How much the session asks before acting; null when it was never set. */
val permissionMode: String?,
/**
* How hard the model thinks, or null for the CLI's own default.
*
* Null is a level somebody can choose, not only one to start in -- see [EFFORT_LEVELS]. It is
* reported rather than assumed for the same reason [permissionMode] is.
*/
val effort: String?,
/**
* Whether a thinking level does anything here -- a Claude CLI session, not a llama or echo one.
*
* Asked of the server rather than worked out from the provider's name, because this is a
* property of the driver's *kind* and the phone only has the name.
*/
val takesEffort: Boolean,
/**
* Whether this continues a session the machine already had, which changes what deleting means.
*/
@@ -153,6 +167,24 @@ data class SessionSummary(
* itself from a default is one you can turn off while believing you are reading it.
*/
val notify: Boolean,
/**
* Whether this session sends itself a message once its account's usage limit lifts, and what
* that message says.
*
* The message is what the server would actually send, with its own default already filled in,
* so the field shows the words rather than an empty box standing for them.
*/
val autoResume: Boolean,
val autoResumeMessage: String,
/**
* When the server next intends to check whether the limit has lifted, in epoch seconds, or null
* when nothing is waiting.
*
* A time to *ask*, not a time to resume: the server checks the meter at that moment and waits
* again if the limit is still on. Worded that way wherever it is shown, because a promise this
* app cannot keep is worse than no time at all.
*/
val resumeAt: Double?,
/**
* The directory the session works in, or null where it was never given one.
*
@@ -178,8 +210,27 @@ data class SessionSummary(
* server because that is where a provider's kind is known.
*/
val maxImageEdge: Int?,
/**
* Which of `GET /usage`'s snapshots is about this session, and null where nothing meters it.
*
* The rate-limit bar answers a question about an *account*, and what decides which account --
* if any -- is the provider this session runs, not the machine it runs on. Pairing by machine
* alone drew the Claude CLI's five-hour window under every echo session on a machine that also
* has the CLI: a quota that session cannot spend and could never run down. Decided by the
* server for the same reason [maxImageEdge] is -- it is a fact about the provider's kind, and
* this app has only its name.
*/
val usageProvider: String?,
val status: String,
val lastActivity: Double,
/**
* How many subagents this session has, however their own status now reads.
*
* A directory listing on the server rather than a status read per subagent, so the list stays
* cheap; the per-subagent state is only fetched when the card is expanded. Zero on a server
* that predates subagents, so this app still opens against one.
*/
val subagents: Int,
)
private fun parseSession(session: JSONObject) =
@@ -192,14 +243,24 @@ private fun parseSession(session: JSONObject) =
title = session.getString("title"),
model = session.optString("model").ifEmpty { null },
permissionMode = session.optString("permissionMode").ifEmpty { null },
effort = session.optString("effort").ifEmpty { null },
takesEffort = session.optBoolean("takesEffort", false),
imported = session.optBoolean("imported", false),
notify = session.optBoolean("notify", true),
autoResume = session.optBoolean("autoResume", false),
// The server sends its own default rather than nothing, so an empty answer means an older
// server -- and this app's word for it is the same word.
autoResumeMessage =
session.optString("autoResumeMessage").ifEmpty { DEFAULT_RESUME_MESSAGE },
resumeAt = if (session.has("resumeAt")) session.getDouble("resumeAt") else null,
cwd = session.optString("cwd").ifEmpty { null },
contextTokens =
if (session.has("contextTokens")) session.getLong("contextTokens") else null,
maxImageEdge = session.optInt("maxImageEdge", 0).takeIf { it > 0 },
usageProvider = session.optString("usageProvider").ifEmpty { null },
status = session.getString("status"),
lastActivity = session.getDouble("lastActivity"),
subagents = session.optInt("subagents", 0),
)
fun fetchSessions(settings: ServerSettings): List<SessionSummary> =
@@ -215,6 +276,35 @@ fun fetchSessions(settings: ServerSettings): List<SessionSummary> =
fun fetchSession(settings: ServerSettings, sessionId: String): SessionSummary =
requestFromServer(settings, "/sessions/$sessionId") { parseSession(it.jsonObject()) }
/**
* One row of `GET /sessions/{id}/subagents`, oldest first.
*
* A subagent is a second transcript owned by a session -- no process, no controls of its own -- so
* this carries only what a card needs to draw and to open it; see SUBAGENTS.md. [status] is
* "running", "exited" or "unknown": a subagent whose session is not itself running cannot be
* running, and the list says so rather than reporting a state that cannot hold.
*/
data class SubagentSummary(
val id: String,
val title: String,
val status: String,
val created: Double,
val lastActivity: Double,
)
fun fetchSubagents(settings: ServerSettings, sessionId: String): List<SubagentSummary> =
requestFromServer(settings, "/sessions/$sessionId/subagents") {
it.jsonObjects { row ->
SubagentSummary(
id = row.getString("id"),
title = row.getString("title"),
status = row.getString("status"),
created = row.getDouble("created"),
lastActivity = row.getDouble("lastActivity"),
)
}
}
// What the server offers, so the spawn screen has no hardcoded lists: a setup added to the server's
// config.ron appears here with no app rebuild.
//
@@ -375,6 +465,12 @@ data class SshDetails(
* Where files attached from here land on that machine; null for the session's own directory.
*/
val attachmentsDir: String? = null,
/**
* Where that machine keeps its GGUF models; null for the same place the backend keeps its own
* (`~/.local/share/ai-app/models`, read on that machine). A llama.cpp session serves the file
* from the machine it runs on, so this is where its models are looked for and listed.
*/
val modelsDir: String? = null,
)
private fun SshDetails.toJson() =
@@ -382,6 +478,7 @@ private fun SshDetails.toJson() =
if (port != null) put("port", port)
if (!identityFile.isNullOrBlank()) put("identityFile", identityFile)
if (!attachmentsDir.isNullOrBlank()) put("attachmentsDir", attachmentsDir)
if (!modelsDir.isNullOrBlank()) put("modelsDir", modelsDir)
}
/** What a machine turns out to have, without saving anything. */
@@ -450,6 +547,8 @@ fun spawnSession(
model: String? = null,
cwd: String? = null,
permissionMode: String? = null,
/** Null for whatever the server's default is; see [fetchDefaultEffort]. */
effort: String? = null,
params: Map<String, String> = emptyMap(),
/** Continue this Claude Code session instead of starting an empty one. */
import: String? = null,
@@ -467,6 +566,7 @@ fun spawnSession(
if (!model.isNullOrBlank()) put("model", model)
if (!cwd.isNullOrBlank()) put("cwd", cwd)
if (!permissionMode.isNullOrBlank()) put("permissionMode", permissionMode)
if (!effort.isNullOrBlank()) put("effort", effort)
if (!import.isNullOrBlank()) put("import", import)
if (params.isNotEmpty()) {
put("params", JSONObject(params.toMap<String, Any>()))
@@ -906,7 +1006,7 @@ fun startImport(
*/
fun fetchTranscript(
settings: ServerSettings,
sessionId: String,
address: TranscriptAddress,
before: Long? = null,
limit: Int = 80,
// Count [limit] in rows, not events, joining a reply's streamed deltas into one -- so a page of
@@ -925,7 +1025,7 @@ fun fetchTranscript(
if (coalesce) append("&coalesce=true")
if (after != null) append("&after=").append(after)
}
return requestFromServer(settings, "/sessions/$sessionId/transcript$query") { connection ->
return requestFromServer(settings, "/${address.urlPath}/transcript$query") { connection ->
val body = JSONArray(connection.inputStream.bufferedReader().readText())
// The text as well as the event: the transcript cache stores the one and the fold needs the
// other, and they have to be the same line.
@@ -973,6 +1073,56 @@ fun setSessionModel(settings: ServerSettings, sessionId: String, model: String)
*/
val PERMISSION_MODES = listOf("manual", "acceptEdits", "auto", "bypassPermissions", "plan")
/**
* What a new session's thinking level is when nothing chose one, or null for the CLI's own.
*
* Held by the server rather than by this phone, because a second device would otherwise spawn
* sessions at a level the first one's owner never picked.
*/
fun fetchDefaultEffort(settings: ServerSettings): String? =
requestFromServer(settings, "/defaults") {
it.jsonObject().optString("effort").ifEmpty { null }
}
/** Sets what new sessions start at. Nothing already running changes. */
fun setDefaultEffort(settings: ServerSettings, level: String?) {
requestFromServer(
settings,
"/defaults",
method = "POST",
jsonBody = JSONObject().put("effort", level ?: JSONObject.NULL).toString(),
) {}
}
/**
* How hard the model thinks, as `claude --effort` takes them, cheapest first.
*
* Not offered alongside the model and the permission mode on the session's own bar, because it does
* not behave like them: the CLI has a control request for those two and none for this (checked
* against 2.1.258), so a level is settled when the process is launched. Changing it therefore stops
* the process, which is what the working directory beside it in this dialog does, and why it is
* here rather than on a bar whose other controls take effect mid-turn.
*/
val EFFORT_LEVELS = listOf("low", "medium", "high", "xhigh", "max")
/** What the picker shows, and sends as null, for a session that has chosen no level. */
const val DEFAULT_EFFORT = "default"
/**
* Records how hard a session thinks and **stops its process**, since the level is read when the
* process is launched. The next message, or Start, runs one that has it.
*
* [level] is null for the CLI's own default.
*/
fun setSessionEffort(settings: ServerSettings, sessionId: String, level: String?) {
requestFromServer(
settings,
"/sessions/$sessionId/effort",
method = "POST",
jsonBody = JSONObject().put("effort", level ?: JSONObject.NULL).toString(),
) {}
}
/** Switches how much a running session asks before acting, also in place. */
fun setSessionPermissionMode(settings: ServerSettings, sessionId: String, mode: String) {
requestFromServer(
@@ -984,6 +1134,36 @@ fun setSessionPermissionMode(settings: ServerSettings, sessionId: String, mode:
}
/** Turns this session's notifications on or off. Stored on the backend -- see `SessionConfig`. */
/**
* What an auto-resume says when nothing else was typed. Mirrors the server's own default, so a
* cleared field shows the word that would actually be sent instead of going blank.
*/
const val DEFAULT_RESUME_MESSAGE = "continue"
/**
* Turns auto-resume on or off and sets what it would say, in one request because they are one
* decision -- see the server's `/sessions/{id}/auto-resume`.
*/
fun setSessionAutoResume(
settings: ServerSettings,
sessionId: String,
autoResume: Boolean,
message: String?,
) {
requestFromServer(
settings,
"/sessions/$sessionId/auto-resume",
method = "POST",
jsonBody =
JSONObject()
.put("autoResume", autoResume)
// Empty means the server's default rather than a session poked with nothing to
// read, which is the same rule the server applies to the field.
.put("message", message?.trim()?.ifEmpty { null } ?: JSONObject.NULL)
.toString(),
) {}
}
fun setSessionNotify(settings: ServerSettings, sessionId: String, notify: Boolean) {
requestFromServer(
settings,
@@ -1074,6 +1254,26 @@ private fun parseDownload(o: JSONObject) =
error = if (o.has("error")) o.getString("error") else null,
)
/**
* The models on one machine, which is the list a llama.cpp session there can choose from.
*
* Not [fetchModels], which is what the *backend* has downloaded. A session serves its model from
* the machine it runs on, so for a machine reached over ssh those are two different lists -- and
* offering the backend's would name files that are not there, turning a choice that cannot work
* into a session that fails when it tries to load one.
*/
fun fetchSetupModels(settings: ServerSettings, setupId: String): List<LocalModel> =
requestFromServer(settings, "/setups/${setupId.urlEncoded()}/models") { connection ->
JSONArray(connection.inputStream.bufferedReader().readText()).mapObjects { m ->
LocalModel(
key = m.getString("key"),
repo = m.getString("repo"),
file = m.getString("file"),
bytes = m.getLong("bytes"),
)
}
}
fun fetchModels(settings: ServerSettings): Models =
requestFromServer(settings, "/models") { connection ->
val body = JSONObject(connection.inputStream.bufferedReader().readText())
@@ -1,7 +1,9 @@
package com.example.aiapp
import androidx.activity.compose.BackHandler
import androidx.compose.foundation.background
import androidx.compose.foundation.layout.Box
import androidx.compose.foundation.layout.fillMaxSize
import androidx.compose.foundation.layout.imePadding
import androidx.compose.foundation.layout.padding
import androidx.compose.material3.AlertDialog
@@ -34,7 +36,23 @@ import kotlinx.coroutines.withContext
* session, spawning one, and settings.
*/
private sealed class Screen {
data object Main : Screen()
/**
* The session list, with a subagent's own transcript over it when [subagent] is set.
*
* A layer on this screen rather than a screen of its own, for the same reason [Session.files]
* is: [SessionListScreen] owns which cards are expanded and what each expansion fetched, kept
* in `remember`, and a subagent is opened from a card's expander. As a sibling `Screen` it was
* disposed and recreated on every return, which lost that state -- an expanded card collapsed
* itself the moment its own subagent's view was closed.
*/
data class Main(val subagent: SubagentTarget? = null) : Screen()
/**
* One subagent's own transcript, read-only. See [SessionScreen]'s `subagent` parameter and
* SUBAGENTS.md's "Phone". Closing it returns to [Main] under it, not to [Session]: a subagent
* is opened from the session list's card rather than from inside the session it belongs to.
*/
data class SubagentTarget(val summary: SessionSummary, val subagent: SubagentSummary)
/**
* One session, with the file explorer over it when [files] is set.
@@ -81,7 +99,7 @@ fun AppRoot(
val context = LocalContext.current
val scope = rememberCoroutineScope()
var settings by remember(settingsVersion) { mutableStateOf(loadServerSettings(context)) }
var screen by remember { mutableStateOf<Screen>(Screen.Main) }
var screen by remember { mutableStateOf<Screen>(Screen.Main()) }
// A notification tap this could not follow, and why. Null both before one is asked for and
// after one succeeds, since success is a screen rather than a message.
var failedOpen by remember { mutableStateOf<FailedOpen?>(null) }
@@ -96,7 +114,7 @@ fun AppRoot(
share = shareRequest
// A session already open takes it. Otherwise the list is where the choice is made,
// whatever screen was showing: Spawn and Settings have nowhere to put a file.
if (screen !is Screen.Session) screen = Screen.Main
if (screen !is Screen.Session) screen = Screen.Main()
}
}
@@ -123,7 +141,7 @@ fun AppRoot(
existing = null,
onSaved = { saved ->
settings = saved
screen = Screen.Main
screen = Screen.Main()
},
onBack = null,
)
@@ -136,7 +154,7 @@ fun AppRoot(
// shows, so it always refetches.
val goToMain = {
reloadToken++
screen = Screen.Main
screen = Screen.Main()
}
if (screen !is Screen.Main) {
BackHandler(onBack = goToMain)
@@ -185,6 +203,9 @@ fun AppRoot(
reloadToken = reloadToken,
share = share,
onOpen = { screen = Screen.Session(it) },
onOpenSubagent = { summary, subagent ->
screen = here.copy(subagent = Screen.SubagentTarget(summary, subagent))
},
onSpawn = { screen = Screen.Spawn },
onImported = { imported ->
reloadToken++
@@ -192,6 +213,27 @@ fun AppRoot(
},
onSettings = { screen = Screen.Settings },
)
// Its own back handler is registered after MainScreen's, so it is the one the
// platform asks first while a subagent is open -- the same rule the files
// explorer's handler follows over its session, below.
here.subagent?.let { target ->
BackHandler { screen = here.copy(subagent = null) }
// Its own opaque background: this screen was always the sole content under
// the theme's own Surface before, so it never had to paint one -- stacked over
// the list here, the space between its own cards let the list underneath show
// through without this. The same fix FilesScreen needed over its session.
Box(Modifier.fillMaxSize().background(MaterialTheme.colorScheme.background)) {
key(target.summary.id, target.subagent.id) {
SessionScreen(
settings = current,
summary = target.summary,
onBack = { screen = here.copy(subagent = null) },
onFiles = {},
subagent = target.subagent,
)
}
}
}
}
is Screen.Session ->
// Keyed on the id, because a different session is a different screen rather than this
@@ -0,0 +1,108 @@
package com.example.aiapp
import android.content.Context
import java.util.concurrent.CopyOnWriteArrayList
/**
* P0's benchmark gate (see docs/RUST.md and docs/DECISIONS.md's 2026-09-05 entry): an in-process
* fake of the backend, so the `bench` build type can drive a real session screen -- the real
* [TranscriptSource], the real fold, the real paging -- with no server and no network permission.
*
* Only ever installed when [BuildConfig.FIXTURE_MODE] is true (see [MainActivity]); everything else
* in this build compiles it in but never calls it, since Kotlin has no per-build-type source set
* that both [MainActivity] (which every variant compiles) and this can share without one.
*
* The design: [requestFromServer] and [Sse] talk to `https://$FIXTURE_HOST:$FIXTURE_PORT` through
* ordinary `java.net.URL`, exactly as they would talk to a real server. A
* [java.net.URLStreamHandlerFactory] registered once for the whole process intercepts every
* `https://` connection to that host and answers from this object's in-memory event log instead of
* opening a socket -- see BenchNetwork.kt. Everything above that (TranscriptSource, SessionScreen,
* the fold, uniqueItems) never learns the difference.
*/
object BenchFixture {
const val FIXTURE_HOST = "bench.fixture.invalid"
const val FIXTURE_PORT = 1
/** How many of the fixture's events are the opening backlog; see bench-fixture/README.md. */
private const val BACKLOG_COUNT = 3200
val settings = ServerSettings(FIXTURE_HOST, FIXTURE_PORT, "bench")
/** The session id every bench run opens; nothing else in this build ever mints one. */
const val SESSION_ID = "bench-fixture-session"
/**
* The whole transcript, seq order, growing as [pushLive] is called during the streaming phase.
* Read by both the REST page handler and the SSE handler, so a page requested mid- stream and a
* live frame agree on what has "already happened" -- the same thing a real server's own
* transcript file guarantees.
*/
private val log = CopyOnWriteArrayList<Pair<String, SeqEvent>>()
/** The events not yet appended to [log] -- the streaming phase's own source. */
private var streamTail: List<Pair<String, SeqEvent>> = emptyList()
private val images = mutableMapOf<String, ByteArray>()
@Volatile private var loaded = false
/**
* Parses the bundled fixture once. Safe to call more than once; only the first does anything.
*/
@Synchronized
fun ensureLoaded(context: Context) {
if (loaded) return
val lines =
context.assets.open("transcript.jsonl").bufferedReader().readLines().filter {
it.isNotBlank()
}
val parsed = lines.map { it to parseSeqEvent(it) }
log.addAll(parsed.take(BACKLOG_COUNT))
streamTail = parsed.drop(BACKLOG_COUNT)
for (name in listOf("bench1.png", "bench2.png")) {
images[name] = context.assets.open(name).readBytes()
}
loaded = true
}
/** The events the streaming phase has left to send. */
fun remainingStreamEvents(): Int = streamTail.size
/** Sends the next fixture event onto the live log, as a real SSE frame would arrive. */
fun pushNextLiveEvent(): Boolean {
val next = streamTail.firstOrNull() ?: return false
streamTail = streamTail.drop(1)
log.add(next)
return true
}
/** Undoes [pushNextLiveEvent] and reloads the opening backlog, for running the bench twice. */
@Synchronized
fun resetToBacklog(context: Context) {
loaded = false
log.clear()
ensureLoaded(context)
}
fun fileBytes(name: String): ByteArray? = images[name]
/**
* Raw JSON lines with seq > [after], in order -- what an `/events?after=` connection replays.
*/
fun linesAfter(after: Long): List<String> =
log.filter { it.second.seq > after }.map { it.first }
/**
* One REST page: [fetchTranscript]'s `before`/`limit`/`after`, against the growing log. Ignores
* `coalesce` -- the fixture's own deltas are already split the way a real reply streams, and
* what the benchmark exercises is the fold and the paging, not the server's row-joining, which
* client-core's own port tracks separately (CLIENT_CORE.md).
*/
fun page(before: Long?, limit: Int, after: Long?): List<String> {
val upper = before ?: (log.lastOrNull()?.second?.seq?.plus(1) ?: 1L)
val candidates = log.filter {
it.second.seq < upper && (after == null || it.second.seq > after)
}
return candidates.takeLast(limit).map { it.first }
}
}
@@ -0,0 +1,181 @@
package com.example.aiapp
import java.io.ByteArrayInputStream
import java.io.IOException
import java.io.InputStream
import java.io.PipedInputStream
import java.io.PipedOutputStream
import java.net.HttpURLConnection
import java.net.URL
import java.net.URLStreamHandler
import java.net.URLStreamHandlerFactory
import java.security.Principal
import java.security.cert.Certificate
import javax.net.ssl.HttpsURLConnection
import javax.net.ssl.SSLPeerUnverifiedException
import org.json.JSONArray
/**
* Installs the process-wide interception [BenchFixture] needs. Idempotent and safe to call more
* than once; the JDK only allows [URL.setURLStreamHandlerFactory] to be called successfully once
* per process, and a second real call throws -- so this guards it rather than relying on every
* caller to remember.
*
* Scoped to [BenchFixture.FIXTURE_HOST]: any other `https://` URL falls through to the platform's
* ordinary handler, so this only ever changes behaviour for the one host the bench build invents.
*/
@Synchronized
fun installFixtureNetworkOnce() {
if (installed) return
installed = true
URL.setURLStreamHandlerFactory(
URLStreamHandlerFactory { protocol ->
if (protocol != "https") null
else
object : URLStreamHandler() {
override fun openConnection(url: URL): HttpURLConnection =
if (url.host == BenchFixture.FIXTURE_HOST) FixtureConnection(url)
else
// The bench build makes no other https call -- this factory is
// installed only in FIXTURE_MODE (MainActivity) -- so there is
// deliberately no delegate to a platform handler here: once a
// URLStreamHandlerFactory is installed there is no supported way to
// ask the JDK for its own default handler back, and re-entering this
// same factory for the fallback would recurse forever rather than
// reach one.
throw java.io.IOException(
"bench build's fixture network has no route to https host " +
"${url.host} -- only ${BenchFixture.FIXTURE_HOST} is served"
)
}
}
)
}
private var installed = false
/**
* Answers one request against [BenchFixture] instead of opening a socket. Implements just enough of
* [HttpsURLConnection] for [requestFromServer] and [Sse] to work unmodified: both only call
* `connect`/`disconnect`, set a handful of request properties they never need answered, and read
* `responseCode` and `inputStream`.
*/
private class FixtureConnection(url: URL) : HttpsURLConnection(url) {
private var input: InputStream? = null
private var writer: Thread? = null
override fun connect() {
if (input != null) return
input = route(url.path, url.query)
}
override fun disconnect() {
writer?.interrupt()
try {
input?.close()
} catch (_: IOException) {}
}
override fun usingProxy() = false
override fun getResponseCode(): Int {
connect()
return 200
}
override fun getInputStream(): InputStream {
connect()
return input!!
}
override fun getErrorStream(): InputStream? = null
// Nothing here reads any of these; implemented only because HttpsURLConnection declares them
// abstract. A fixture never negotiates real TLS, so each says exactly that rather than
// fabricating a plausible-looking certificate.
override fun getCipherSuite() = "none (bench fixture, no TLS)"
override fun getLocalCertificates(): Array<Certificate>? = null
override fun getServerCertificates(): Array<Certificate> =
throw SSLPeerUnverifiedException("bench fixture connection presents no certificate")
override fun getPeerPrincipal(): Principal =
throw SSLPeerUnverifiedException("bench fixture connection presents no certificate")
override fun getLocalPrincipal(): Principal? = null
/**
* [path] is `/sessions/{id}/...`; everything else this build's fixture is asked for is a bug.
*/
private fun route(path: String, query: String?): InputStream {
val params =
(query ?: "")
.split("&")
.filter { it.contains('=') }
.associate {
val (k, v) = it.split("=", limit = 2)
k to java.net.URLDecoder.decode(v, "UTF-8")
}
return when {
path.endsWith("/transcript") -> {
val lines =
BenchFixture.page(
before = params["before"]?.toLongOrNull(),
limit = params["limit"]?.toIntOrNull() ?: 80,
after = params["after"]?.toLongOrNull(),
)
val body = JSONArray(lines.map { org.json.JSONObject(it) })
ByteArrayInputStream(body.toString().toByteArray())
}
path.endsWith("/events") -> openEventsStream(params["after"]?.toLongOrNull() ?: 0L)
path.contains("/files/") -> {
val name = path.substringAfterLast("/files/")
val bytes =
BenchFixture.fileBytes(name)
?: throw IOException("bench fixture has no file named $name")
ByteArrayInputStream(bytes)
}
else -> throw IOException("bench fixture has no route for $path")
}
}
/**
* A live SSE body: [BenchFixture.linesAfter] replayed immediately, then polled every 50ms for
* anything [BenchFixture.pushNextLiveEvent] has added since -- the same shape a real backend's
* backlog-then-follow gives [Sse], just polled instead of woken, which is a fixture's business
* rather than something worth a condition variable for.
*/
private fun openEventsStream(after: Long): InputStream {
val pipeIn = PipedInputStream(1 shl 16)
val pipeOut = PipedOutputStream(pipeIn)
var sent = after
val thread = Thread {
try {
while (!Thread.currentThread().isInterrupted) {
val fresh = BenchFixture.linesAfter(sent)
for (line in fresh) {
pipeOut.write("data: $line\n\n".toByteArray())
pipeOut.flush()
sent = org.json.JSONObject(line).getLong("seq")
}
Thread.sleep(50)
}
} catch (_: InterruptedException) {
// disconnect() -- the ordinary way this ends.
} catch (_: IOException) {
// The reader side (Sse) closed its end.
} finally {
try {
pipeOut.close()
} catch (_: IOException) {}
}
}
.also {
it.isDaemon = true
it.start()
}
writer = thread
return pipeIn
}
}
@@ -0,0 +1,143 @@
package com.example.aiapp
import android.content.Context
import android.os.BatteryManager
import android.os.Process
import androidx.compose.animation.core.tween
import androidx.compose.foundation.gestures.animateScrollBy
import androidx.compose.foundation.lazy.LazyListState
import java.io.File
import kotlinx.coroutines.CoroutineScope
import kotlinx.coroutines.delay
import kotlinx.coroutines.isActive
import kotlinx.coroutines.launch
/**
* P0's scripted benchmark, run in-process instead of by a shell script: the phone has no usable
* system tracing (this-machine-android's skill) and no agent can drive it, so the same scroll loop
* and streaming phase `transcript-bench.sh`/`stream-bench.sh` drive over `ui-trace` are reproduced
* here against [LazyListState] and [BenchFixture] directly. Only reachable from the `bench` build
* (see [SessionSettingsDialog]'s `onRunBenchmark`), but compiled into every build for the reason
* [BenchFixture]'s doc comment gives.
*/
object BenchRun {
/** transcript-bench.sh's default: 6 cycles of 4 swipes each, 900px over 200ms, 500ms apart. */
private const val CYCLES = 6
private const val SWIPE_PX = 900f
private const val SWIPE_MS = 200
private const val SWIPE_PAUSE_MS = 500L
/** stream-bench.sh's shape: a real reply arrives as many small deltas, not one big write. */
private const val STREAM_EVENTS_PER_SEC = 20
private const val STREAM_SECONDS = 20
/**
* Scrolls, then streams, then returns the extra report lines P0 asked for (CPU time, peak RSS,
* battery current) -- [FrameStats] and [DebugStats] are reset first, exactly as
* `copyRenderReport` resets them, so the two accountings cover the same stretch of work.
*/
suspend fun run(
context: Context,
scope: CoroutineScope,
listState: LazyListState,
): List<String> {
FrameStats.reset()
DebugStats.reset()
val cpuStartMs = Process.getElapsedCpuTime()
val battery = BatterySampler(context)
// Launched in the caller's scope rather than a fresh coroutineScope{} here, which would
// suspend this function until the sampler job ended -- and it only ends when told to.
val samplerJob = scope.launch {
while (isActive) {
battery.sample()
delay(1000)
}
}
// The swipe loop: transcript-bench.sh's four swipes per cycle are two drags toward newer
// content and two back, so a cycle returns to where it started and the whole loop measures
// steady-state scrolling rather than travelling somewhere new each time.
repeat(CYCLES) {
repeat(2) {
listState.animateScrollBy(SWIPE_PX, tween(SWIPE_MS))
delay(SWIPE_PAUSE_MS)
}
repeat(2) {
listState.animateScrollBy(-SWIPE_PX, tween(SWIPE_MS))
delay(SWIPE_PAUSE_MS)
}
}
// Pinned to the newest end before streaming starts, the way stream-bench.sh's "Jump to
// latest" tap is -- a reply streamed into a list parked further back arrives off-screen and
// the report would show nothing happened.
listState.scrollToItem(0)
var sent = 0
val total = STREAM_EVENTS_PER_SEC * STREAM_SECONDS
while (sent < total && BenchFixture.remainingStreamEvents() > 0) {
BenchFixture.pushNextLiveEvent()
sent++
delay(1000L / STREAM_EVENTS_PER_SEC)
}
// Lets the last few deltas land and draw before the report is read.
delay(300)
samplerJob.cancel()
val cpuMs = Process.getElapsedCpuTime() - cpuStartMs
val rssLine = peakRssLine()
val batteryLine = battery.finish()
return listOf(
" scroll: $CYCLES cycles (${CYCLES * 4} swipes), streamed $sent/$total fixture events",
" process CPU time over this run: ${cpuMs}ms",
rssLine,
batteryLine,
)
}
/** VmHWM from /proc/self/status: the process's high-water mark, in kB, since it started. */
private fun peakRssLine(): String {
val kb =
try {
File("/proc/self/status")
.readLines()
.firstOrNull { it.startsWith("VmHWM:") }
?.trim()
?.removePrefix("VmHWM:")
?.trim()
?.removeSuffix("kB")
?.trim()
?.toLongOrNull()
} catch (_: Exception) {
null
}
return " peak RSS: " +
(kb?.let { "${it}kB" } ?: "unavailable (/proc/self/status unreadable)")
}
}
/**
* Samples [BatteryManager.BATTERY_PROPERTY_CURRENT_NOW] (microamps) once a second for the length of
* a run. The property returns `Int.MIN_VALUE` on hardware that does not support it -- most
* emulators -- and that is reported as "unavailable" rather than folded into an average with the
* real samples, which would silently understate every number after it. See UI_RULES: never present
* an inferred value as a measured one.
*/
private class BatterySampler(context: Context) {
private val manager = context.getSystemService(BatteryManager::class.java)
private val samples = mutableListOf<Int>()
fun sample() {
val value = manager?.getIntProperty(BatteryManager.BATTERY_PROPERTY_CURRENT_NOW)
if (value != null && value != Int.MIN_VALUE) samples.add(value)
}
fun finish(): String {
if (samples.isEmpty()) return " battery current: unavailable on this device"
val meanUa = samples.sum() / samples.size
return " battery current: mean ${meanUa}µA over ${samples.size} samples" +
" (min ${samples.min()}, max ${samples.max()})"
}
}
@@ -127,6 +127,13 @@ fun debugReport(
frames: List<String>,
accounting: List<String>,
crash: String?,
/**
* P0's benchmark-only measurements (process CPU time, peak RSS, battery current) -- empty on
* every path but [BenchRun.runP0Benchmark], which is the only caller that has them. A section
* heading only appears when there is something to put under it, so an ordinary copy from the
* render-report button reads exactly as it did before this existed.
*/
extra: List<String> = emptyList(),
): String = buildString {
appendLine("ai-app render report")
appendLine(device)
@@ -152,6 +159,11 @@ fun debugReport(
appendLine("work since this was last copied:")
val work = DebugStats.lines()
if (work.isEmpty()) appendLine(" nothing recorded") else work.forEach { appendLine(it) }
if (extra.isNotEmpty()) {
appendLine()
appendLine("bench:")
extra.forEach { appendLine(it) }
}
}
/** Puts [text] on the clipboard under [label], which is what the system offers as its name. */
@@ -12,6 +12,10 @@ import androidx.compose.ui.Alignment
import androidx.compose.ui.Modifier
import androidx.compose.ui.graphics.Color
import androidx.compose.ui.unit.dp
import java.time.Instant
import java.time.ZoneId
import java.time.format.DateTimeFormatter
import java.time.format.FormatStyle
/**
* A line across the transcript saying what left the session's context.
@@ -47,3 +51,40 @@ fun TranscriptDivider(text: String, color: Color, modifier: Modifier = Modifier)
fun ClearedRow(modifier: Modifier = Modifier) {
TranscriptDivider("Context cleared", clearedColor, modifier)
}
/**
* The mark running out of quota leaves.
*
* The same red the usage bar takes when a window is spent, because it is the same fact in a second
* place: colour by consequence, so "there is nothing left to spend" is learned once.
*
* A time rather than a countdown. The row is folded once and never re-measured, so a span would go
* stale on screen the moment it was drawn; and this is when the *account* said it would reset,
* which is not a promise about when the session picks back up. A limit the session was told no
* reset time for says nothing about one -- that state has its own words rather than a plausible
* number.
*/
@Composable
fun LimitRow(item: TranscriptItem.LimitNote, modifier: Modifier = Modifier) {
TranscriptDivider(limitSummary(item.resetsAt, ZoneId.systemDefault()), overLimitColor, modifier)
}
/**
* What the row says. Split out so the wording is testable without a screen, since the two states it
* has to keep apart -- a reset time that arrived and one that never did -- are exactly the pair
* that reads the same when it goes wrong.
*
* [zone] is a parameter rather than read here so a test says the same thing wherever it runs.
*/
fun limitSummary(resetsAt: Double?, zone: ZoneId): String {
val at = resetsAt?.let {
try {
DateTimeFormatter.ofLocalizedTime(FormatStyle.SHORT)
.withZone(zone)
.format(Instant.ofEpochSecond(it.toLong()))
} catch (_: Exception) {
null
}
}
return if (at == null) "Usage limit reached" else "Usage limit reached • resets $at"
}
@@ -14,7 +14,7 @@ private const val RESET_EVENT = "reset"
* mean. [close] from any thread ends it, and the caller owns reconnecting -- with the last seq it
* saw as the new cursor.
*/
class EventStream(settings: ServerSettings, private val sessionId: String) {
class EventStream(settings: ServerSettings, private val address: TranscriptAddress) {
private val stream = Sse(settings)
fun close() = stream.close()
@@ -35,7 +35,7 @@ class EventStream(settings: ServerSettings, private val sessionId: String) {
// one and the screen folds the other, and they have to be the same line.
onEvent: (raw: String, event: SeqEvent) -> Unit,
) {
stream.run("/sessions/$sessionId/events?after=$after", onOpen) { name, data ->
stream.run("/${address.urlPath}/events?after=$after", onOpen) { name, data ->
// A named frame carries no payload and a data frame has no name.
if (name == RESET_EVENT) onReset()
else if (data.isNotEmpty()) onEvent(data, parseSeqEvent(data))
@@ -157,6 +157,18 @@ sealed class SessionEvent {
*/
data object Cleared : SessionEvent()
/**
* The session stopped because its account's usage limit was reached.
*
* Its own event rather than an [Error] carrying the CLI's sentence, because it is a state
* rather than something that went wrong -- and because the raw sentence is `Claude AI usage
* limit reached|1788546972`, which is not readable by the person it is shown to.
*
* [resetsAt] is epoch seconds and null where the session was told nothing. Only the server acts
* on it; what this draws it as is a time, not a countdown, because nothing here re-measures it.
*/
data class LimitReached(val resetsAt: Double?) : SessionEvent()
data class Error(val message: String) : SessionEvent()
/**
@@ -261,6 +273,10 @@ fun parseSeqEvent(json: String): SeqEvent {
trigger = body.optString("trigger").ifEmpty { null },
)
"cleared" -> SessionEvent.Cleared
"limitReached" ->
SessionEvent.LimitReached(
if (body.has("resetsAt")) body.getDouble("resetsAt") else null
)
"error" -> SessionEvent.Error(body.getString("message"))
else -> SessionEvent.Unknown(type)
}
@@ -66,6 +66,35 @@ class MainActivity : ComponentActivity() {
// Transparent status bar on every version; the Surface below paints through underneath it
// and content insets itself. Same reasoning as dev-updater's MainActivity.
enableEdgeToEdge()
// The `bench` build's entire purpose (P0, docs/RUST.md): open straight onto the session
// screen against BenchFixture's in-process fake backend, with no enrollment, no network
// permission, and no notification prompt -- none of them mean anything with no server and
// no real device to notify. See BenchFixture.kt and BenchNetwork.kt for how a screen built
// to talk to a real backend is made to talk to this instead. Still needs the same
// status/navigation-bar padding the ordinary flow below applies: edge-to-edge is the
// platform's own default from Android 15 on this app's targetSdk, with or without the call
// above, so skipping the padding here put the header's own buttons under the status bar --
// there to look at, but not there for `ui-trace`'s tap-by-label to land on.
if (BuildConfig.FIXTURE_MODE) {
installFixtureNetworkOnce()
BenchFixture.ensureLoaded(this)
setContent {
MaterialTheme(colorScheme = AiAppColors) {
Surface(modifier = Modifier.fillMaxSize()) {
Box(Modifier.fillMaxSize().statusBarsPadding().navigationBarsPadding()) {
SessionScreen(
settings = BenchFixture.settings,
summary = benchSessionSummary(),
onBack = { finish() },
onFiles = {},
)
}
}
}
}
return
}
// Dark status-bar icons only over a light background, decided from the scheme rather than
// fixed. It was hardcoded to `true`, which was right against the default light surface and
// became unreadable the moment the app wore Catppuccin Mocha.
@@ -147,6 +176,26 @@ class MainActivity : ComponentActivity() {
}
}
/** The one session the `bench` build ever shows -- BenchFixture's session id, nothing else. */
private fun benchSessionSummary() =
SessionSummary(
id = BenchFixture.SESSION_ID,
setup = "bench",
setupName = "bench",
provider = "bench",
title = "P0 benchmark",
model = null,
keepsOwnTranscript = false,
permissionMode = null,
imported = false,
notify = false,
cwd = null,
contextTokens = null,
maxImageEdge = null,
status = "idle",
lastActivity = 0.0,
)
// launchMode="singleTop": an enrollment scan, or a notification tapped while the app is open,
// lands here rather than in a second activity instance.
override fun onNewIntent(intent: Intent) {
@@ -48,6 +48,8 @@ fun MainScreen(
/** What another app shared in and no session has taken yet; see [ShareRequest]. */
share: ShareRequest? = null,
onOpen: (SessionSummary) -> Unit,
/** Opens one session's subagent, from the expander under its card. */
onOpenSubagent: (SessionSummary, SubagentSummary) -> Unit,
onSpawn: () -> Unit,
onImported: (SessionSummary) -> Unit,
onSettings: () -> Unit,
@@ -139,6 +141,7 @@ fun MainScreen(
settings = settings,
reloadToken = token,
onOpen = onOpen,
onOpenSubagent = onOpenSubagent,
onSpawn = onSpawn,
)
MainTab.Import ->
@@ -1,7 +1,9 @@
package com.example.aiapp
import androidx.compose.foundation.ExperimentalFoundationApi
import androidx.compose.foundation.clickable
import androidx.compose.foundation.combinedClickable
import androidx.compose.foundation.layout.Arrangement
import androidx.compose.foundation.layout.Box
import androidx.compose.foundation.layout.Column
import androidx.compose.foundation.layout.Row
@@ -9,6 +11,7 @@ import androidx.compose.foundation.layout.Spacer
import androidx.compose.foundation.layout.fillMaxSize
import androidx.compose.foundation.layout.fillMaxWidth
import androidx.compose.foundation.layout.height
import androidx.compose.foundation.layout.heightIn
import androidx.compose.foundation.layout.padding
import androidx.compose.foundation.layout.width
import androidx.compose.foundation.lazy.LazyColumn
@@ -17,6 +20,7 @@ import androidx.compose.material3.Card
import androidx.compose.material3.CircularProgressIndicator
import androidx.compose.material3.FloatingActionButton
import androidx.compose.material3.MaterialTheme
import androidx.compose.material3.OutlinedCard
import androidx.compose.material3.Switch
import androidx.compose.material3.Text
import androidx.compose.material3.TextButton
@@ -30,6 +34,8 @@ import androidx.compose.runtime.setValue
import androidx.compose.ui.Alignment
import androidx.compose.ui.Modifier
import androidx.compose.ui.platform.LocalContext
import androidx.compose.ui.semantics.contentDescription
import androidx.compose.ui.semantics.semantics
import androidx.compose.ui.unit.dp
import kotlinx.coroutines.Dispatchers
import kotlinx.coroutines.launch
@@ -47,12 +53,40 @@ fun SessionListScreen(
settings: ServerSettings,
reloadToken: Int,
onOpen: (SessionSummary) -> Unit,
/** Opens one session's subagent, from the expander under its card. */
onOpenSubagent: (SessionSummary, SubagentSummary) -> Unit,
onSpawn: () -> Unit,
) {
val scope = rememberCoroutineScope()
var listState by remember { mutableStateOf<LoadState<List<SessionSummary>>>(LoadState.Loading) }
var confirmingDelete by remember { mutableStateOf<SessionSummary?>(null) }
// Which session cards are expanded to show their subagents, and what each expansion fetched.
// Ids rather than a flag on the row for the same reason `deleting` is: the rows are rebuilt
// from
// whatever the server last said, and this belongs to the reader's own choice, which survives a
// refresh.
var expandedSessions by remember { mutableStateOf(setOf<String>()) }
var subagentLoads by remember {
mutableStateOf(mapOf<String, LoadState<List<SubagentSummary>>>())
}
fun loadSubagents(sessionId: String) {
subagentLoads = subagentLoads + (sessionId to LoadState.Loading)
scope.launch {
subagentLoads =
subagentLoads +
(sessionId to
try {
LoadState.Loaded(
withContext(Dispatchers.IO) { fetchSubagents(settings, sessionId) }
)
} catch (e: ApiException) {
LoadState.failed(e)
})
}
}
// Failures that belong to one session rather than to the list, keyed by its id and shown on its
// own card. The two scopes are decided by whether the server answered: it answered and refused,
// so this says nothing about the other rows.
@@ -84,6 +118,14 @@ fun SessionListScreen(
withContext(Dispatchers.IO) {
transcriptCache.retainOnly(loaded.value.map { it.id }.toSet())
}
// A session gone from this answer cannot still be expanded, and an expanded one
// that is still here asks again -- its subagents may have changed since the
// last
// fetch.
val ids = loaded.value.map { it.id }.toSet()
expandedSessions = expandedSessions intersect ids
subagentLoads = subagentLoads.filterKeys { it in ids }
expandedSessions.forEach(::loadSubagents)
loaded
} catch (e: ApiException) {
LoadState.failed(e)
@@ -127,6 +169,17 @@ fun SessionListScreen(
deleting = session.id in deleting,
onOpen = { onOpen(session) },
onLongPress = { confirmingDelete = session },
expanded = session.id in expandedSessions,
subagents = subagentLoads[session.id],
onToggleSubagents = {
if (session.id in expandedSessions) {
expandedSessions = expandedSessions - session.id
} else {
expandedSessions = expandedSessions + session.id
loadSubagents(session.id)
}
},
onOpenSubagent = { subagent -> onOpenSubagent(session, subagent) },
)
Spacer(Modifier.height(12.dp))
}
@@ -225,7 +278,7 @@ fun SessionListScreen(
deleteSession(settings, session.id, alsoDeleteForeign)
// After it succeeded, not before: a refused delete leaves the
// session exactly as it was, and its transcript with it.
transcriptCache.session(session.id).purge()
transcriptCache.session(TranscriptAddress(session.id)).purge()
}
// Only this row, and only what changed. Refetching the list instead
// put every other session back through loading and handed the
@@ -276,6 +329,12 @@ private fun SessionCard(
deleting: Boolean,
onOpen: () -> Unit,
onLongPress: () -> Unit,
/** Whether the expander below is open. Collapsed by default; see [SessionListScreen]. */
expanded: Boolean,
/** What the expander's own fetch answered, or null before it has been asked. */
subagents: LoadState<List<SubagentSummary>>?,
onToggleSubagents: () -> Unit,
onOpenSubagent: (SubagentSummary) -> Unit,
) {
BusyItem(label = if (deleting) "deleting" else null) {
Card(
@@ -332,11 +391,105 @@ private fun SessionCard(
color = MaterialTheme.colorScheme.error,
)
}
// Nothing at all for a card with no subagents: a disabled expander here would be
// noise on every ordinary session's card. Its own row at the bottom rather than
// beside the title or the machine line, so opening it never displaces text that was
// already on screen -- see UI_RULES on a control not displacing the text beside it.
if (session.subagents > 0) {
Spacer(Modifier.height(8.dp))
// The platform's minimum touch height, not the chevron's own ten or so dp:
// at the chevron's height a tap meant for it landed on the first subcard
// beneath and opened a subagent instead.
Row(
horizontalArrangement = Arrangement.Center,
verticalAlignment = Alignment.CenterVertically,
modifier =
Modifier.fillMaxWidth()
.heightIn(min = 48.dp)
.clickable(enabled = !deleting, onClick = onToggleSubagents)
.semantics {
contentDescription =
if (expanded) "Collapse subagents" else "Expand subagents"
},
) {
Chevron(if (expanded) Pointing.Up else Pointing.Down)
}
if (expanded) {
Spacer(Modifier.height(4.dp))
Column(verticalArrangement = Arrangement.spacedBy(8.dp)) {
when (subagents) {
null,
is LoadState.Loading ->
CircularProgressIndicator(
modifier = Modifier.width(20.dp).height(20.dp),
strokeWidth = 2.dp,
)
is LoadState.Error ->
// Said here rather than left silent: a fetch that failed and an
// expander that simply found nothing must not look the same --
// see UI_RULES on designing the unknown state first.
Text(
subagents.message,
style = MaterialTheme.typography.bodySmall,
color = MaterialTheme.colorScheme.error,
)
is LoadState.Loaded ->
subagents.value.forEach { subagent ->
SubagentCard(
subagent,
onClick = { onOpenSubagent(subagent) },
)
}
}
}
}
}
}
}
}
}
/**
* One subagent, indented inside its session's card -- the way dev-updater draws a project's
* components (`ComponentCard`, `UpdaterScreen.kt`): an outlined card, not the session card's own
* filled one, so the nesting reads as one step rather than as another session.
*/
@Composable
private fun SubagentCard(subagent: SubagentSummary, onClick: () -> Unit) {
OutlinedCard(Modifier.fillMaxWidth().clickable(onClick = onClick)) {
Column(Modifier.padding(horizontal = 12.dp, vertical = 8.dp)) {
Text(subagent.title, style = MaterialTheme.typography.titleSmall)
Spacer(Modifier.height(2.dp))
Row(modifier = Modifier.fillMaxWidth()) {
Text(
subagentStatusLabel(subagent.status),
style = MaterialTheme.typography.bodySmall,
color = MaterialTheme.colorScheme.onSurfaceVariant,
modifier = Modifier.weight(1f),
)
Text(
relativeTime(subagent.lastActivity),
style = MaterialTheme.typography.bodySmall,
color = MaterialTheme.colorScheme.onSurfaceVariant,
)
}
}
}
}
/**
* The subcard's word for a subagent's status -- see SUBAGENTS.md's "Wire shape". Its own function
* rather than a branch inside [StatusText], because a subagent's three states are not that
* composable's five: "exited" reads as "finished" here, since its process was always its parent's
* and never something of its own to have merely stopped.
*/
private fun subagentStatusLabel(status: String) =
when (status) {
"running" -> "running"
"exited" -> "finished"
else -> "unknown"
}
@Composable
fun StatusText(status: String) {
val (label, color) =
@@ -216,16 +216,33 @@ fun SessionScreen(
share: ShareRequest? = null,
/** Said once [share] has been attached here, so it is not attached again. */
onShareTaken: () -> Unit = {},
/**
* Draws this screen read-only, on a subagent's own transcript instead of the session's.
*
* A subagent has no process and no controls of its own -- see SUBAGENTS.md's "Phone" -- so
* every gate below keyed on this switches off the composer, the files button, the settings cog,
* the usage bar and notifications, while everything that draws a transcript (paging, cache,
* selection, images, the status row, stream reconnects) is reused unchanged, pointed at
* [address] instead of the session's own.
*/
subagent: SubagentSummary? = null,
) {
DebugStats.count("session screen recomposed")
val isSubagent = subagent != null
val address = TranscriptAddress(summary.id, subagent?.id)
val scope = rememberCoroutineScope()
val topEdgeHeld = remember { TopEdgeHold() }
var items by remember { mutableStateOf(listOf<TranscriptItem>()) }
var status by remember { mutableStateOf(summary.status) }
var status by remember { mutableStateOf(subagent?.status ?: summary.status) }
// Seeded from the row this screen was opened from, so a conversation already under way says how
// much it is holding before any turn happens here. Null is "nobody has measured it", which is a
// different answer from an empty context and is drawn differently.
var contextTokens by remember(summary.id) { mutableStateOf(summary.contextTokens) }
//
// A subagent has no context measurement of its own, so it always starts unmeasured rather than
// borrowing the parent session's figure -- see UI_RULES on not showing an inferred value as one
// that was measured.
var contextTokens by
remember(address) { mutableStateOf(if (isSubagent) null else summary.contextTokens) }
// When the current compaction started. The moment comes off the `compacting` status event
// itself -- the server timestamps every transcript line -- rather than off this device noticing
// one, which is what makes it survive leaving the session and reopening it.
@@ -241,7 +258,13 @@ fun SessionScreen(
val context = LocalContext.current
// Seeded from what was left in the box last time and written back on every keystroke, so
// leaving the screen does not throw away a half-typed message. See `Drafts.kt`.
var input by remember(summary.id) { mutableStateOf(atEnd(loadDraft(context, summary.id))) }
//
// A subagent has no box to type into, so it never touches a draft at all -- not this session's,
// which is what reading one keyed only by `summary.id` would do here.
var input by
remember(summary.id) {
mutableStateOf(if (isSubagent) atEnd("") else atEnd(loadDraft(context, summary.id)))
}
// A model the reader has chosen and not yet confirmed. See [ModelSwitchWarning]: switching
// makes the session re-read the whole conversation.
var pendingModel by remember { mutableStateOf<String?>(null) }
@@ -294,27 +317,26 @@ fun SessionScreen(
// Reload throws away what it was reading from.
val cache = remember(settings) { TranscriptCache(cacheRoot(context, settings)) }
val source =
remember(summary.id, epoch) {
TranscriptSource(settings, summary.id, cache.session(summary.id))
}
remember(address, epoch) { TranscriptSource(settings, address, cache.session(address)) }
// Whether the cached tail has been shown to still be the server's own line. Nothing is resumed
// from a cached cursor until it has, and a probe that could not be made leaves this false for
// the stream loop to try again.
var probePassed by remember(summary.id, epoch) { mutableStateOf(false) }
var probePassed by remember(address, epoch) { mutableStateOf(false) }
// Whether the opening effect is still settling that question. It draws the cached rows and
// lifts [ready] before the answer arrives, which is the point of the cache -- so the stream
// below waits for this rather than for `ready`, or it asks the same question twice.
var probing by remember(summary.id, epoch) { mutableStateOf(true) }
var probing by remember(address, epoch) { mutableStateOf(true) }
// The oldest sequence number loaded, and whether there is more behind it. Paging backwards is
// what keeps opening a long session cheap.
var oldestSeq by remember { mutableLongStateOf(0L) }
// Where this session was last being read, from this device's own store. Read once, because the
// answer stops being interesting the moment the list is on screen.
val savedAnchor = remember(summary.id, epoch) { loadScrollAnchor(context, summary.id) }
// Where this transcript was last being read, from this device's own store, keyed by the address
// rather than the session id so a subagent's saved position cannot collide with its session's.
// Read once, because the answer stops being interesting the moment the list is on screen.
val savedAnchor = remember(address, epoch) { loadScrollAnchor(context, address.cachePath) }
// Whether the saved position is still being put back. Nothing is drawn while it is: opening at
// the newest end and then travelling to the anchor is exactly the journey a reader must never
// see.
var restoring by remember(summary.id, epoch) { mutableStateOf(savedAnchor != null) }
var restoring by remember(address, epoch) { mutableStateOf(savedAnchor != null) }
// Messages the server has taken and the session has not read yet, by the id that will resolve
// them. From the event stream rather than from what this screen sent, so they survive leaving
// the session -- and a message sent from another device is drawn waiting on this one too.
@@ -327,11 +349,11 @@ fun SessionScreen(
var loadingHistory by remember { mutableStateOf(false) }
var ready by remember { mutableStateOf(false) }
// Replies parsed ahead of the rows that draw them; see [ParsedReplies].
val replies = remember(summary.id) { ParsedReplies() }
// Keyed like everything else describing one session's transcript. `rememberLazyListState` saves
// through `rememberSaveable`, and this screen restores by its own anchor instead -- two
// restores would fight over the first frame.
val listState = remember(summary.id) { LazyListState() }
val replies = remember(address) { ParsedReplies() }
// Keyed like everything else describing one transcript. `rememberLazyListState` saves through
// `rememberSaveable`, and this screen restores by its own anchor instead -- two restores would
// fight over the first frame.
val listState = remember(address) { LazyListState() }
// Whether the newest message is on screen right now. The list is reversed, so the newest end is
// the scrolling start: nothing behind you is exactly being at the bottom. Asked of the scroll
// state rather than of item indices, because a zero-height first item makes an index ambiguous.
@@ -637,7 +659,7 @@ fun SessionScreen(
// ended and carries live events only. The window comes from this phone's own copy when there is
// one, and then costs a single request to check that the server's transcript is still the one
// it came from. See TRANSCRIPT_CACHE.md.
LaunchedEffect(summary.id, epoch) {
LaunchedEffect(address, epoch) {
/**
* One opening window onto the screen, whichever side it came from.
*
@@ -667,11 +689,16 @@ fun SessionScreen(
// A replay is as old as the last visit; the row this screen was opened from was
// fetched moments ago. So the transcript comes from the cache and everything that
// is not the transcript comes from the summary -- otherwise a session that finished
// an hour ago opens saying "working" until the stream connects.
status = summary.status
// an hour ago opens saying "working" until the stream connects. A subagent's status
// comes from its own summary, never the parent session's: they are two different
// things running or not, and the parent's model and permission mode do not apply to
// it at all.
status = subagent?.status ?: summary.status
if (!isSubagent) {
model = summary.model
permissionMode = summary.permissionMode ?: "auto"
if (summary.status != "compacting") compactingSince = null
}
if (status != "compacting") compactingSince = null
// Nothing to put back, so these rows are the screen and the probe can return under
// them. A restore still has history to fetch and is gated below.
if (savedAnchor == null) ready = true
@@ -798,7 +825,7 @@ fun SessionScreen(
// at the top on their return. Switching apps is a choice somebody made, not a fault to report.
// Stopping the stream deliberately makes the drop a close rather than an error, and resuming
// reconnects from the same cursor.
LaunchedEffect(summary.id, ready, epoch, lifecycleOwner) {
LaunchedEffect(address, ready, epoch, lifecycleOwner) {
if (!ready) return@LaunchedEffect
// The opening effect draws cached rows and lifts `ready` *before* it has checked that the
// cursor under them is still the server's, so `ready` is no longer the whole gate. Without
@@ -868,10 +895,14 @@ fun SessionScreen(
// The screen going away entirely, which the lifecycle scope above does not cover: a composable
// can leave the composition while the activity stays started. Keyed on the epoch as well, so
// Reload's replacement source is the one a later disposal closes.
DisposableEffect(summary.id, epoch) { onDispose { source.close() } }
DisposableEffect(address, epoch) { onDispose { source.close() } }
// Nothing gets announced about the session somebody is reading; see NotificationService.
// RESUMED rather than STARTED because "looking at it" means the foreground.
//
// Not for a subagent: it has no notifications of its own, and it is not the session this would
// otherwise mark as being read.
if (!isSubagent) {
LaunchedEffect(summary.id, lifecycleOwner) {
lifecycleOwner.repeatOnLifecycle(Lifecycle.State.RESUMED) {
NotificationService.showing(context, summary.id)
@@ -882,6 +913,7 @@ fun SessionScreen(
}
}
}
}
// Back at the newest end, so the backlog [apply] held can land. Everything at once rather than
// paced out: they are at the bottom, which is the one place the list is allowed to follow new
@@ -924,7 +956,7 @@ fun SessionScreen(
val (index, offset, awayFromNewest) = settled
saveScrollAnchor(
context,
summary.id,
address.cachePath,
// Nothing to restore at the newest end, which is where a session with no anchor
// opens anyway. One *before* the index, because item zero is the "below" slot.
if (!awayFromNewest) null
@@ -947,7 +979,7 @@ fun SessionScreen(
//
// There is no correction beside this one. Following the newest message is not an effect: the
// list is reversed, so an arriving message extends the end the viewport is pinned to.
val unitSizes = remember(summary.id) { HashMap<Any, Int>() }
val unitSizes = remember(address) { HashMap<Any, Int>() }
LaunchedEffect(listState, moreHistory) {
snapshotFlow { listState.layoutInfo }
.collect { info ->
@@ -983,6 +1015,8 @@ fun SessionScreen(
}
}
// Only for the model picker, which a subagent does not have.
if (!isSubagent) {
LaunchedEffect(summary.setupName, summary.provider) {
offeredModels =
try {
@@ -995,10 +1029,12 @@ fun SessionScreen(
.orEmpty()
}
} catch (_: Exception) {
// Not worth reporting: the picker simply has nothing to offer, which is visible.
// Not worth reporting: the picker simply has nothing to offer, which is
// visible.
emptyList()
}
}
}
/**
* Asks the server to take back a message the session has not read yet.
@@ -1148,8 +1184,9 @@ fun SessionScreen(
}
// One poll for the machines' limits, read by everything on this screen that reports them.
val usageFeed = rememberUsageFeed(settings)
val usage = usageFeed.forSetup(summary.setup)
// Nothing meters a subagent -- it has no account of its own -- so it never starts this poll.
val usageFeed = if (isSubagent) null else rememberUsageFeed(settings)
val usage = usageFeed?.forSession(summary) ?: SessionUsage.NotMetered
RecordFrames()
var usageOpen by remember { mutableStateOf(false) }
var settingsOpen by remember { mutableStateOf(false) }
@@ -1191,7 +1228,12 @@ fun SessionScreen(
// the bench scripts keep working when this moves again. They pressed it at a hand-measured
// coordinate until 2026-09-03, and anything that moved the header made that tap land on
// whatever now sat there -- reporting a number that was never measured.
val copyRenderReport = {
// Shared by the ordinary "Copy" button and (bench build only) "Run benchmark": what differs
// between them is only whether there is a [extra] section, built by BenchRun.run beforehand --
// everything about assembling, copying and logging the report is exactly the same act either
// way, and a second copy of it beside `onRunBenchmark` below would be the two silently
// disagreeing about what "the report" contains the first time either one changed.
fun buildAndCopyReport(extra: List<String> = emptyList()) {
val report =
debugReport(
device =
@@ -1213,6 +1255,7 @@ fun SessionScreen(
accounting =
FrameStats.drawPhase().let { (nanos, count) -> drawAccounting(nanos, count) },
crash = lastCrash(context),
extra = extra,
)
context.copyToClipboard("ai-app render report", report)
// Also to the log, so a session driving the app over adb can read the same report the
@@ -1226,6 +1269,20 @@ fun SessionScreen(
DebugStats.reset()
Toast.makeText(context, "Copied render report", Toast.LENGTH_SHORT).show()
}
val copyRenderReport = { buildAndCopyReport() }
// Bench build only: P0's scripted scroll-and-stream benchmark (BenchRun.kt), against the
// fixture session opened below instead of a real server. Null everywhere else -- see
// [SessionSettingsDialog]'s onRunBenchmark.
val runBenchmark: (() -> Unit)? =
if (BuildConfig.FIXTURE_MODE) {
{
settingsOpen = false
scope.launch {
val extra = BenchRun.run(context, scope, listState)
buildAndCopyReport(extra)
}
}
} else null
Box(Modifier.fillMaxSize()) {
Column(Modifier.fillMaxSize()) {
Row(
@@ -1236,20 +1293,35 @@ fun SessionScreen(
// A ring's worth, which is what the arrow already keeps on its other three sides.
Spacer(Modifier.width(GLYPH_BUTTON_MARGIN))
Column(Modifier.weight(1f)) {
// A subagent's own title, with the session's beneath it in a smaller style --
// the header says whose conversation this is as well as what it is. Otherwise
// just the session's title, as before.
if (subagent != null) {
Text(subagent.title, style = MaterialTheme.typography.titleMedium)
Text(
title,
style = MaterialTheme.typography.bodySmall,
color = MaterialTheme.colorScheme.onSurfaceVariant,
)
} else {
Text(title, style = MaterialTheme.typography.titleMedium)
// Machine first, then what runs on it -- the same order and the same wording
// everywhere this pair appears, so it reads as one fact rather than two
// sentences with different grammar.
// Machine first, then what runs on it -- the same order and the same
// wording everywhere this pair appears, so it reads as one fact rather than
// two sentences with different grammar.
//
// No model. The picker in the footer already shows what this session is set to,
// and showing it twice means two things to keep in step -- they disagreed for a
// moment on every model change.
// No model. The picker in the footer already shows what this session is set
// to, and showing it twice means two things to keep in step -- they
// disagreed for a moment on every model change.
Text(
"${summary.setupName} · ${summary.provider}",
style = MaterialTheme.typography.bodySmall,
color = MaterialTheme.colorScheme.onSurfaceVariant,
)
}
}
// None of this is a subagent's: it has no files of its own to browse, no settings,
// and nothing meters it -- see SUBAGENTS.md's "Phone".
//
// Beside the provider it reports on, which is the line directly to its left. Its
// real home is this provider's settings, which do not exist yet. A session on a
// provider with no such service gets an honest "unavailable" rather than a hidden
@@ -1264,6 +1336,7 @@ fun SessionScreen(
// Usage, files, settings -- widest scope first, narrowing to the right, so the cog
// stays at the end where every other screen keeps it. Asked for in this order by
// Iris on 2026-09-03.
if (!isSubagent) {
Row {
GlyphButton(
USAGE_GLYPH,
@@ -1281,25 +1354,29 @@ fun SessionScreen(
FilesTarget(
setup = summary.setup,
setupName = summary.setupName,
// Where this session works, and the machine's own home when it
// was never given a directory -- resolved there rather than
// guessed at here, since this app does not know that home.
// Where this session works, and the machine's own home when
// it was never given a directory -- resolved there rather
// than guessed at here, since this app does not know that
// home.
start = summary.cwd?.takeIf { it.isNotBlank() } ?: "~",
)
)
},
)
// What it opens is about this session, so it sits at the end of the session's
// own row. A cog and not a word because there will be more, and a bar of words
// has nowhere to put it.
// What it opens is about this session, so it sits at the end of the
// session's own row. A cog and not a word because there will be more, and a
// bar of words has nowhere to put it.
GlyphButton(SETTINGS_GLYPH, "Session settings", { settingsOpen = true })
}
}
}
// Under the header, above everything the session itself says: it is a fact about the
// machine rather than a turn in the conversation, and it is the number that decides
// whether to keep going.
// whether to keep going. Nothing meters a subagent.
if (!isSubagent) {
SessionUsageBar(usage)
}
(streamError ?: actionError)?.let { message ->
Text(
@@ -1545,6 +1622,7 @@ fun SessionScreen(
is TranscriptItem.ClearedNote -> ClearedRow()
is TranscriptItem.CompactedNote ->
CompactedRow(item)
is TranscriptItem.LimitNote -> LimitRow(item)
// Never reached: a peer message is flattened into
// its own units. Here because a `when` over the
// item kinds has to stay exhaustive.
@@ -1643,24 +1721,34 @@ fun SessionScreen(
)
}
// Kept for a subagent -- see SUBAGENTS.md's "Phone" -- with the wording that turns
// "exited" into "finished" for one, since it has no process to leave running or stop.
SessionStatusRow(
status = status,
compactingFor = compactingFor,
contextTokens = contextTokens,
subagent = isSubagent,
)
// Between the transcript and the box: above what is being typed, so the list does not
// cover the thing the command is about, and below everything that explains it.
// Everything from here down is the composer: a subagent cannot be messaged, so none of
// it applies -- see SUBAGENTS.md's "Phone".
if (!isSubagent) {
// Between the transcript and the box: above what is being typed, so the list does
// not cover the thing the command is about, and below everything that explains it.
CommandSuggestions(
// Nothing to suggest about a suggestion that was just taken. `/compact` is a whole
// command *and* a prefix of itself, so picking it left the list standing there with
// the one row already chosen. Held by what was picked rather than by a flag, so
// typing anything else brings the list back without a second thing to reset.
commands = if (input.text == picked) emptyList() else suggestedCommands(input.text),
// Nothing to suggest about a suggestion that was just taken. `/compact` is a
// whole command *and* a prefix of itself, so picking it left the list standing
// there with the one row already chosen. Held by what was picked rather than by
// a flag, so typing anything else brings the list back without a second thing
// to
// reset.
commands =
if (input.text == picked) emptyList() else suggestedCommands(input.text),
onPick = { command ->
// At the end of what was inserted, which is where the reader carries on typing:
// a command with an argument is put in the box half-written, and a cursor left
// at the front makes the next keystroke the first character of "/rename".
// At the end of what was inserted, which is where the reader carries on
// typing: a command with an argument is put in the box half-written, and a
// cursor left at the front makes the next keystroke the first character of
// "/rename".
input = atEnd(command.typed())
picked = command.typed()
},
@@ -1669,11 +1757,14 @@ fun SessionScreen(
// Always enabled -- a send while the session is running becomes a steering message
// injected at the next tool boundary, which is the point of the whole app.
//
// The field gets a row of its own, above the buttons: sharing one put the full width
// behind three controls, so the thing being typed into was the narrowest on the row.
// The field gets a row of its own, above the buttons: sharing one put the full
// width
// behind three controls, so the thing being typed into was the narrowest on the
// row.
Column(Modifier.fillMaxWidth().padding(8.dp)) {
// Directly above the box they will be sent from, so what is attached is visible
// rather than counted: the "+2" on the button below said how many and never which.
// rather than counted: the "+2" on the button below said how many and never
// which.
PendingAttachments(
settings = settings,
sessionId = summary.id,
@@ -1687,8 +1778,8 @@ fun SessionScreen(
saveDraft(context, summary.id, it.text)
},
modifier = Modifier.fillMaxWidth(),
// No longer "(+image)": the images are on screen above this, and a placeholder
// saying so said it in words beside the thing itself.
// No longer "(+image)": the images are on screen above this, and a
// placeholder saying so said it in words beside the thing itself.
placeholder = { Text("Message") },
maxLines = 4,
)
@@ -1696,17 +1787,19 @@ fun SessionScreen(
verticalAlignment = Alignment.CenterVertically,
modifier = Modifier.fillMaxWidth(),
) {
// Photo or file, asked here rather than by two buttons: the row is full, and
// Photo or file, asked here rather than by two buttons: the row is full,
// and
// attaching is one action whichever picker answers it.
var attaching by remember { mutableStateOf(false) }
Box {
// Just "+". The count it used to carry was standing in for showing them.
// Just "+". The count it used to carry was standing in for showing
// them.
BubbleButton(onClick = { attaching = true }) { Text("+") }
DropdownMenu(
expanded = attaching,
onDismissRequest = { attaching = false },
// See PickerButton: without this the menu opens a status bar's height
// away from the button in an edge-to-edge activity.
// See PickerButton: without this the menu opens a status bar's
// height away from the button in an edge-to-edge activity.
properties = PopupProperties(clippingEnabled = false),
shape = BubbleMenuShape,
) {
@@ -1730,11 +1823,11 @@ fun SessionScreen(
)
}
}
// The settings share what is left after the actions have taken what they need.
// A Row hands out intrinsic widths in order and clips whatever runs past the
// edge, so with these laid out first the arrival of Stop pushed Send off the
// screen entirely -- the app's central control, gone at the moment it is most
// in use.
// The settings share what is left after the actions have taken what they
// need. A Row hands out intrinsic widths in order and clips whatever runs
// past the edge, so with these laid out first the arrival of Stop pushed
// Send off the screen entirely -- the app's central control, gone at the
// moment it is most in use.
Row(
verticalAlignment = Alignment.CenterVertically,
modifier = Modifier.weight(1f),
@@ -1742,16 +1835,18 @@ fun SessionScreen(
if (offeredModels.isNotEmpty()) {
PickerButton(
current = modelLabel(model),
// What the machine offers, plus the state a session is in when it
// has chosen none of them. The button has always been able to say
// "default"; until this the list could not, so leaving it was a
// one-way trip.
// What the machine offers, plus the state a session is in when
// it has chosen none of them. The button has always been able
// to
// say "default"; until this the list could not, so leaving it
// was a one-way trip.
options = listOf(DEFAULT_MODEL) + offeredModels,
// Not set here. The button follows what the session reports it is
// set to, which arrives a moment later and is sometimes a different
// answer -- a name the CLI resolved, or no change at all on a
// provider whose model is fixed. Asked about first, unless there is
// nothing to lose by it -- see [ModelSwitchWarning].
// Not set here. The button follows what the session reports it
// is set to, which arrives a moment later and is sometimes a
// different answer -- a name the CLI resolved, or no change at
// all on a provider whose model is fixed. Asked about first,
// unless there is nothing to lose by it -- see
// [ModelSwitchWarning].
onPick = { chosen ->
if (
modelLabel(chosen) == modelLabel(model) ||
@@ -1772,14 +1867,16 @@ fun SessionScreen(
},
)
}
// The same filled shape as the button beside it, not an outlined one: these are
// two things you can do about the session, and weighting one as secondary said
// they were a primary action and its qualifier. What separates them is the
// colour and the mark, which is what they mean.
// The same filled shape as the button beside it, not an outlined one: these
// are two things you can do about the session, and weighting one as
// secondary said they were a primary action and its qualifier. What
// separates them is the colour and the mark, which is what they mean.
//
// Always here, rather than arriving with the turn as it used to. A control that
// comes and goes makes its own presence the signal, and a button always in the
// same place also cannot push Send off the end of the row by turning up.
// Always here, rather than arriving with the turn as it used to. A control
// that comes and goes makes its own presence the signal, and a button
// always
// in the same place also cannot push Send off the end of the row by turning
// up.
val process =
when {
running -> ProcessAction.Pause
@@ -1799,18 +1896,21 @@ fun SessionScreen(
Glyph(
process.glyph,
colour = LocalContentColor.current,
modifier = Modifier.semantics { contentDescription = process.label },
modifier =
Modifier.semantics { contentDescription = process.label },
)
}
Spacer(Modifier.width(8.dp))
// The paper plane, with a clock on it while a turn is in flight: sending then
// queues the message for the next tool boundary rather than starting a turn of
// its own, and the two have to be told apart at a glance. The label says the
// same thing to a screen reader.
// The paper plane, with a clock on it while a turn is in flight: sending
// then queues the message for the next tool boundary rather than starting a
// turn of its own, and the two have to be told apart at a glance. The label
// says the same thing to a screen reader.
//
// Disabled while there is nothing to send, rather than pressable and silent:
// `send` has always returned early on an empty composer, so the button promised
// something it would not do. Disabled and not hidden, for the reason above.
// Disabled while there is nothing to send, rather than pressable and
// silent:
// `send` has always returned early on an empty composer, so the button
// promised something it would not do. Disabled and not hidden, for the
// reason above.
Button(
onClick = { send() },
enabled = input.text.isNotBlank() || pendingAttachments.isNotEmpty(),
@@ -1827,12 +1927,13 @@ fun SessionScreen(
}
}
}
}
// Beside the other two dialogs, and outside the list for the same reason as them: what is open
// is the screen's business rather than any row's. See [SessionImageViewer].
fullImage?.let { ref -> SessionImageViewer(settings, summary.id, ref) { fullImage = null } }
if (usageOpen) {
UsageDialog(feed = usageFeed, onDismiss = { usageOpen = false })
usageFeed?.let { UsageDialog(feed = it, onDismiss = { usageOpen = false }) }
}
if (settingsOpen) {
// Measured when the dialog opens rather than kept up to date: what the reader is being told
@@ -1846,6 +1947,8 @@ fun SessionScreen(
settings = settings,
sessionId = summary.id,
title = title,
effort = summary.effort.takeIf { summary.takesEffort },
takesEffort = summary.takesEffort,
cachedBytes = cachedBytes,
// The purge finishes before the epoch moves, because the relaunched opening effect
// reads the same directory and would otherwise draw what is about to be deleted. The
@@ -1869,6 +1972,7 @@ fun SessionScreen(
},
onDismiss = { settingsOpen = false },
onCopyRenderReport = copyRenderReport,
onRunBenchmark = runBenchmark,
)
}
}
@@ -2122,6 +2226,13 @@ private fun SessionStatusRow(
/** Context the session is holding, or null where nothing has measured it. */
contextTokens: Long?,
modifier: Modifier = Modifier,
/**
* Whether this row is for a subagent rather than a session, which changes only one word:
* "exited" reads as "finished" there too, the same as the subagent list's own card -- a
* subagent's process was always its parent's, so "exited" would read as a fault rather than the
* ordinary way one of these ends.
*/
subagent: Boolean = false,
) {
DebugStats.count("status row recomposed")
Row(
@@ -2178,7 +2289,7 @@ private fun SessionStatusRow(
Text(
when (status) {
"idle" -> "idle"
"exited" -> "exited"
"exited" -> if (subagent) "finished" else "exited"
"awaitingInput" -> "your turn"
"unknown" -> "can't tell"
else -> status
@@ -2249,7 +2360,7 @@ private const val ONE_TAP_MS = 250L
* session is set to without spending a second line on saying it.
*/
@Composable
private fun PickerButton(current: String, options: List<String>, onPick: (String) -> Unit) {
fun PickerButton(current: String, options: List<String>, onPick: (String) -> Unit) {
var open by remember { mutableStateOf(false) }
// When an outside touch last closed the menu.
//
@@ -6,8 +6,10 @@ import androidx.compose.foundation.layout.Spacer
import androidx.compose.foundation.layout.fillMaxWidth
import androidx.compose.foundation.layout.height
import androidx.compose.foundation.layout.width
import androidx.compose.foundation.rememberScrollState
import androidx.compose.foundation.text.KeyboardActions
import androidx.compose.foundation.text.KeyboardOptions
import androidx.compose.foundation.verticalScroll
import androidx.compose.material3.AlertDialog
import androidx.compose.material3.CircularProgressIndicator
import androidx.compose.material3.MaterialTheme
@@ -26,6 +28,10 @@ import androidx.compose.ui.Alignment
import androidx.compose.ui.Modifier
import androidx.compose.ui.text.input.ImeAction
import androidx.compose.ui.unit.dp
import java.time.Instant
import java.time.ZoneId
import java.time.format.DateTimeFormatter
import java.time.format.FormatStyle
import kotlinx.coroutines.Dispatchers
import kotlinx.coroutines.launch
import kotlinx.coroutines.withContext
@@ -54,6 +60,16 @@ fun SessionSettingsDialog(
*/
title: String,
onRenamed: (String) -> Unit,
/**
* How hard the model thinks, as the session reports it, or null for the CLI's own default.
*
* Taken from the row this dialog was opened over rather than fetched, because unlike the
* notification switch there is nothing else that changes it: the level is this app's to set and
* the server does not resolve it into something else.
*/
effort: String?,
/** Whether a level does anything here; the row is left out entirely where it does not. */
takesEffort: Boolean,
/**
* What this phone is holding of the conversation, or null while that is being measured -- see
* the Reload row below, which is what would discard it.
@@ -66,9 +82,18 @@ fun SessionSettingsDialog(
* measures is that screen's own state.
*/
onCopyRenderReport: () -> Unit,
/**
* Runs P0's scripted scroll-and-stream benchmark and copies the extended report, or null on
* every build but `bench` -- see [BuildConfig.FIXTURE_MODE] and BenchRun.kt. Null rather than
* always-present-but-disabled: this has no meaning at all outside the bench build, and a
* control with nothing behind it on every other build is not a state worth drawing.
*/
onRunBenchmark: (() -> Unit)? = null,
) {
val scope = rememberCoroutineScope()
var name by remember(sessionId) { mutableStateOf(title) }
var level by remember(sessionId) { mutableStateOf(effort) }
var effortError by remember { mutableStateOf<String?>(null) }
var saving by remember { mutableStateOf(false) }
var error by remember { mutableStateOf<String?>(null) }
// Null until the server has been asked. The row this dialog was opened over is a snapshot of
@@ -77,6 +102,16 @@ fun SessionSettingsDialog(
// and a spinner sits beside it, which is what not knowing looks like.
var notify by remember(sessionId) { mutableStateOf<Boolean?>(null) }
var notifyError by remember { mutableStateOf<String?>(null) }
// The same three-state shape the notification switch has, for the same reason: until the
// server has answered, the switch is disabled rather than showing a position nothing confirmed.
var autoResume by remember(sessionId) { mutableStateOf<Boolean?>(null) }
var resumeMessage by remember(sessionId) { mutableStateOf(DEFAULT_RESUME_MESSAGE) }
// When the server next intends to ask whether the limit has lifted, or null when nothing is
// waiting. Read once with everything else: it moves on the server's schedule, not this
// screen's, and a figure that redrew itself here would be this app re-measuring what it was
// told.
var resumeAt by remember(sessionId) { mutableStateOf<Double?>(null) }
var resumeError by remember { mutableStateOf<String?>(null) }
// Where the session works. Null until the server has been asked, for the same reason the switch
// above is. An empty answer is a session that was never given a directory, which is not the
// same as one whose directory is unknown -- the field is only enabled once one of those is
@@ -90,6 +125,9 @@ fun SessionSettingsDialog(
try {
val fresh = withContext(Dispatchers.IO) { fetchSession(settings, sessionId) }
notify = fresh.notify
autoResume = fresh.autoResume
resumeMessage = fresh.autoResumeMessage
resumeAt = fresh.resumeAt
cwd = fresh.cwd.orEmpty()
typedCwd = fresh.cwd.orEmpty()
} catch (e: ApiException) {
@@ -97,6 +135,8 @@ fun SessionSettingsDialog(
// instead of offering a position nothing confirmed.
notifyError = e.message
notify = null
resumeError = e.message
autoResume = null
}
}
@@ -125,6 +165,26 @@ fun SessionSettingsDialog(
}
}
/**
* Chooses a thinking level, which ends the process the old level was launched with.
*
* Put back if the request is refused, for the reason the notification switch below gives: a
* control that stays where it was put after a refusal is stating something untrue.
*/
fun setEffort(chosen: String?) {
val was = level
level = chosen
effortError = null
scope.launch {
try {
withContext(Dispatchers.IO) { setSessionEffort(settings, sessionId, chosen) }
} catch (e: ApiException) {
level = was
effortError = e.message
}
}
}
// Moved optimistically so the switch answers the finger that moved it, and put back if the
// request is refused -- a switch that waits for a round trip reads as broken on a slow tunnel,
// and one that stays moved after a refusal lies.
@@ -142,6 +202,39 @@ fun SessionSettingsDialog(
}
}
/**
* Turns auto-resume on or off, or changes what it would say.
*
* One request for both, because the server takes one: switching it on and typing the message
* are two halves of the same decision, and sending them separately would leave a moment where
* the session is armed with the old words.
*
* Put back if refused, like the notification switch. Turning it off also clears what was
* scheduled -- said here rather than only on the server, or the row would go on naming a time
* that no longer exists.
*/
fun setAutoResume(on: Boolean, message: String) {
val wasOn = autoResume
val wasMessage = resumeMessage
val wasAt = resumeAt
autoResume = on
resumeMessage = message
if (!on) resumeAt = null
resumeError = null
scope.launch {
try {
withContext(Dispatchers.IO) {
setSessionAutoResume(settings, sessionId, on, message)
}
} catch (e: ApiException) {
autoResume = wasOn
resumeMessage = wasMessage
resumeAt = wasAt
resumeError = e.message
}
}
}
// Nothing to do when the name has not changed, so the button says so rather than sending a
// request whose success would look exactly like the failure of having typed nothing.
val changed = name.trim().isNotEmpty() && name.trim() != title
@@ -168,7 +261,10 @@ fun SessionSettingsDialog(
onDismissRequest = onDismiss,
title = { Text("Session settings") },
text = {
Column {
// Scrollable, because this dialog grew past a screenful: a Material dialog constrains
// its own height and clips what does not fit, so the last control on the list is one
// large system font away from being unreachable with nothing on screen to say so.
Column(Modifier.verticalScroll(rememberScrollState())) {
OutlinedTextField(
value = name,
onValueChange = { name = it },
@@ -212,6 +308,70 @@ fun SessionSettingsDialog(
)
}
Spacer(Modifier.height(8.dp))
Row(
verticalAlignment = Alignment.CenterVertically,
modifier = Modifier.fillMaxWidth(),
) {
Text("Resume after a usage limit", modifier = Modifier.weight(1f))
if (autoResume == null && resumeError == null) {
CircularProgressIndicator(
modifier = Modifier.width(16.dp).height(16.dp),
strokeWidth = 2.dp,
)
Spacer(Modifier.width(8.dp))
}
Switch(
checked = autoResume == true,
onCheckedChange = { setAutoResume(it, resumeMessage) },
enabled = autoResume != null,
)
}
// Disabled rather than hidden while the switch is off: a field that comes and goes
// makes its own presence the signal, and a visible one teaches what the switch will
// do. Committed on the keyboard's Done rather than on every keystroke, so typing a
// sentence is one request instead of one per letter.
OutlinedTextField(
value = resumeMessage,
onValueChange = { resumeMessage = it },
label = { Text("Message to send") },
// What an empty field means, in the field: the server's own word rather than a
// session poked with nothing to read.
placeholder = { Text(DEFAULT_RESUME_MESSAGE) },
singleLine = true,
enabled = autoResume == true,
modifier = Modifier.fillMaxWidth(),
keyboardOptions = KeyboardOptions(imeAction = ImeAction.Done),
keyboardActions =
KeyboardActions(onDone = { setAutoResume(true, resumeMessage) }),
)
// What it does and what it costs, in the order it happens. The last sentence is the
// one that matters: the time below is when the server will *ask*, not a promise
// about when the session speaks.
Text(
"When this session stops because the account is out of quota, the server " +
"checks the limit and sends this message once it has lifted. It checks " +
"again if the limit is still on.",
style = MaterialTheme.typography.bodySmall,
color = MaterialTheme.colorScheme.onSurfaceVariant,
)
// Only where something is actually waiting. Absent is not a state worth a row: a
// session that has not hit a limit has nothing scheduled, which the reader can see
// from the switch.
resumeAt?.let { at ->
Text(
"Waiting now -- next check ${formatCheckTime(at)}.",
style = MaterialTheme.typography.bodySmall,
color = MaterialTheme.colorScheme.onSurfaceVariant,
)
}
resumeError?.let {
Text(
it,
color = MaterialTheme.colorScheme.error,
style = MaterialTheme.typography.bodySmall,
)
}
Spacer(Modifier.height(8.dp))
Row(
verticalAlignment = Alignment.CenterVertically,
modifier = Modifier.fillMaxWidth(),
@@ -257,6 +417,44 @@ fun SessionSettingsDialog(
style = MaterialTheme.typography.bodySmall,
)
}
// Left out rather than disabled, the one place this dialog does that: a disabled
// control teaches what the thing can do, and a llama session cannot do this at all
// -- the row would be teaching something false about it.
if (takesEffort) {
Spacer(Modifier.height(8.dp))
Row(
verticalAlignment = Alignment.CenterVertically,
modifier = Modifier.fillMaxWidth(),
) {
Text("Thinking", modifier = Modifier.weight(1f))
PickerButton(
current = level ?: DEFAULT_EFFORT,
// The level the CLI picks for itself is in the list as well as in the
// button, so leaving a level is not a one-way trip -- the same
// correction the model picker carries.
options = listOf(DEFAULT_EFFORT) + EFFORT_LEVELS,
onPick = { chosen ->
setEffort(chosen.takeIf { it != DEFAULT_EFFORT })
},
)
}
// What it costs, said where it is about to be pressed, like Move above: the
// CLI reads the level when it launches and has no control request for
// changing one.
Text(
"Changing this stops the session's process. It starts again with the " +
"next message, or with Start.",
style = MaterialTheme.typography.bodySmall,
color = MaterialTheme.colorScheme.onSurfaceVariant,
)
effortError?.let {
Text(
it,
color = MaterialTheme.colorScheme.error,
style = MaterialTheme.typography.bodySmall,
)
}
}
Spacer(Modifier.height(8.dp))
Row(
verticalAlignment = Alignment.CenterVertically,
@@ -317,6 +515,21 @@ fun SessionSettingsDialog(
Text("Render timings", modifier = Modifier.weight(1f))
TextButton(onClick = onCopyRenderReport) { Text("Copy") }
}
// Bench-build only: see [onRunBenchmark]. Named exactly "Run benchmark" because
// ui-trace and the emulator smoke run find it by that label, the same way every
// other control here is found -- see AGENTS.md's "Driving the UI".
onRunBenchmark?.let { run ->
Spacer(Modifier.height(8.dp))
Row(
verticalAlignment = Alignment.CenterVertically,
modifier = Modifier.fillMaxWidth(),
) {
Glyph(SPEED_GLYPH, colour = MaterialTheme.colorScheme.onSurface)
Spacer(Modifier.width(8.dp))
Text("P0 benchmark", modifier = Modifier.weight(1f))
TextButton(onClick = run) { Text("Run benchmark") }
}
}
}
},
// Disabled rather than absent while there is nothing to save: a button that comes and goes
@@ -329,3 +542,21 @@ fun SessionSettingsDialog(
dismissButton = { TextButton(onClick = onDismiss) { Text("Close") } },
)
}
/**
* When the server will next look, as a local time.
*
* A time rather than a countdown, for the reason the transcript's own limit row gives: this screen
* reads the figure once, and a span drawn from a value nothing refreshes goes stale while somebody
* is looking at it.
*/
private fun formatCheckTime(epochSeconds: Double): String =
try {
DateTimeFormatter.ofLocalizedTime(FormatStyle.SHORT)
.withZone(ZoneId.systemDefault())
.format(Instant.ofEpochSecond(epochSeconds.toLong()))
} catch (_: Exception) {
// A time that cannot be read is not a time to show: the sentence above still says a check
// is coming, which is the part the reader can act on.
"soon"
}
@@ -70,12 +70,21 @@ class UsageFeed(
/** Ask the backend again now. The dialog's refresh button; the poll does it on its own. */
val refresh: () -> Unit,
) {
/** What [setup]'s own limits came back as. See [usageFor] for why the states are these. */
fun forSetup(setup: String): SessionUsage =
when (val state = snapshots) {
/**
* What meters [session], and what that meter came back as. See [usageFor] for the states.
*
* A session rather than a machine, because a machine is not what is metered: one machine runs
* the Claude CLI and an echo session side by side, and only the first of them spends anything.
*/
fun forSession(session: SessionSummary): SessionUsage {
// Settled without asking anybody: a session nothing meters has nothing to check, and
// "checking" is what the fetch's own states would say about it for as long as one is out.
val provider = session.usageProvider ?: return SessionUsage.NotMetered
return when (val state = snapshots) {
is LoadState.Loading -> SessionUsage.Waiting
is LoadState.Error -> SessionUsage.Unavailable(state.message)
is LoadState.Loaded -> usageFor(state.value, setup)
is LoadState.Loaded -> usageFor(state.value, session.setup, provider)
}
}
}
@@ -158,9 +167,15 @@ fun SessionUsageBar(usage: SessionUsage, modifier: Modifier = Modifier) {
}
}
// Nothing at all for a machine that meters nothing: a row saying "unknown" there would report a
// problem about a setup somebody chose, on every screen, forever.
if (usage is SessionUsage.NotMetered) {
// Nothing at all for a session that meters nothing: a row saying "unknown" there would report
// a problem about a setup somebody chose, on every screen, forever.
//
// And nothing while the first fetch is out, which is a different silence. A request in flight
// is not a state to report -- and the session that meters nothing is exactly the one this
// cannot yet tell apart, so "5-hour usage: checking" appeared under an echo session for half a
// second and was then taken away. A row that has to be withdrawn is worse than one that
// arrives late.
if (usage is SessionUsage.NotMetered || usage is SessionUsage.Waiting) {
return
}
@@ -171,9 +186,10 @@ fun SessionUsageBar(usage: SessionUsage, modifier: Modifier = Modifier) {
// Words, not a colour and not an empty bar: every one of these is a different kind of
// answer from "this much is used", and only words carry a difference in kind.
when (val state = usage) {
SessionUsage.NotMetered -> Unit
// Both handled above, before the row exists at all.
SessionUsage.NotMetered,
SessionUsage.Waiting -> Unit
is SessionUsage.Unavailable -> UsageNote("5-hour usage unknown -- ${state.why}")
SessionUsage.Waiting -> UsageNote("5-hour usage: checking")
is SessionUsage.Known -> {
val window = state.windows.firstOrNull { it.kind == "session" }
if (window == null) {
@@ -234,16 +250,22 @@ private fun fiveHourLabel(window: UsageWindow, now: OffsetDateTime): String {
}
/**
* One machine's snapshot, out of every machine's.
* One meter's snapshot, out of every machine's: [setup]'s row for [provider].
*
* Both halves are needed to pick it. A machine can hold more than one meter -- the Claude CLI's
* account and, while a test has one set, an echo session's invented one -- and a snapshot is one
* service on one machine.
*
* Every way of having *failed* to get numbers is [SessionUsage.Unavailable] with the reason in it.
* None of them may look like zero, and none may look like [SessionUsage.NotMetered], which is the
* machine having no quota rather than the question going unanswered.
*/
fun usageFor(snapshots: List<UsageSnapshot>, setup: String): SessionUsage {
// No snapshot at all means the backend never asked, which it only does for a machine with
// nothing metered on it. That is a different answer from having asked and failed.
val mine = snapshots.firstOrNull { it.setup == setup } ?: return SessionUsage.NotMetered
fun usageFor(snapshots: List<UsageSnapshot>, setup: String, provider: String): SessionUsage {
// No snapshot at all means the backend never asked, which it only does where there is nothing
// to ask about. That is a different answer from having asked and failed.
val mine =
snapshots.firstOrNull { it.setup == setup && it.provider == provider }
?: return SessionUsage.NotMetered
if (mine.state != "ok") {
return SessionUsage.Unavailable(mine.detail ?: mine.state)
}
@@ -236,6 +236,7 @@ private fun AddSetupDialog(
var address by remember { mutableStateOf("") }
var identity by remember { mutableStateOf("") }
var attachmentsDir by remember { mutableStateOf("") }
var modelsDir by remember { mutableStateOf("") }
var tested by remember { mutableStateOf<String?>(null) }
var testing by remember { mutableStateOf(false) }
@@ -250,6 +251,7 @@ private fun AddSetupDialog(
port = typedPort,
identityFile = identity.trim().ifEmpty { null },
attachmentsDir = attachmentsDir.trim().ifEmpty { null },
modelsDir = modelsDir.trim().ifEmpty { null },
)
}
@@ -293,6 +295,14 @@ private fun AddSetupDialog(
label = { Text("Folder for attached files (optional)") },
singleLine = true,
)
// Where that machine's GGUFs are, for a llama.cpp session on it. Blank means
// the same place this backend keeps its own downloads, read on that machine.
OutlinedTextField(
value = modelsDir,
onValueChange = { modelsDir = it },
label = { Text("Folder for models (optional)") },
singleLine = true,
)
tested?.let {
Spacer(Modifier.height(8.dp))
Text(it, style = MaterialTheme.typography.bodySmall)
@@ -61,18 +61,31 @@ fun SpawnScreen(
// "auto" rather than "manual": on a phone every ask is a round trip to a question card, and
// answering "allow Bash?" dozens of times per task is what this app exists to avoid.
var permissionMode by remember { mutableStateOf("auto") }
// Null until the server has been asked, and null again if it answers "no level chosen" -- the
// two are told apart by [defaultsAsked], because a picker that shows a level before the answer
// arrives is one you can spawn at without having chosen it.
var effort by remember { mutableStateOf<String?>(null) }
var defaultsAsked by remember { mutableStateOf(false) }
var busy by remember { mutableStateOf(false) }
// Only the spawn's own failure. The fetch's lives in `options`: this one leaves a filled-in
// form worth keeping, and that one leaves nothing to fill in.
var spawnError by remember { mutableStateOf<String?>(null) }
// Downloaded models, for a llama provider to choose between. Kept separate from the setups: a
// Claude session needs none, so failing to list them must not stop the screen rendering.
// The models on the *chosen machine*, for a llama provider to choose between. Kept separate
// from the setups: a Claude session needs none, so failing to list them must not stop the
// screen rendering. Refetched when the machine changes, because a model is a file on one
// machine -- see [fetchSetupModels].
var models by remember { mutableStateOf<List<LocalModel>>(emptyList()) }
var modelKey by remember { mutableStateOf<String?>(null) }
var contextSize by remember { mutableStateOf("") }
var temperature by remember { mutableStateOf("") }
LaunchedEffect(Unit) {
// Separate from the setups fetch below and deliberately not fatal: failing to learn the
// default must leave a screen you can still spawn from, so the picker stays on "default"
// and says so rather than the whole form refusing to draw.
runCatching { withContext(Dispatchers.IO) { fetchDefaultEffort(settings) } }
.onSuccess { effort = it }
defaultsAsked = true
options =
try {
val fetched = withContext(Dispatchers.IO) { fetchSetups(settings) }
@@ -83,9 +96,6 @@ fun SpawnScreen(
} catch (e: ApiException) {
LoadState.failed(e)
}
models =
runCatching { withContext(Dispatchers.IO) { fetchModels(settings).local } }
.getOrDefault(emptyList())
}
Column(Modifier.fillMaxSize().verticalScroll(rememberScrollState()).padding(16.dp)) {
@@ -115,6 +125,17 @@ fun SpawnScreen(
is LoadState.Loaded -> state.value
}
val setup = setups.firstOrNull { it.name == setupName }
// Whichever machine is chosen now, asked again when that changes. The old machine's list
// is dropped first rather than left on screen: a file name from another machine looks
// exactly like one from this one.
LaunchedEffect(setup?.id) {
models = emptyList()
modelKey = null
val id = setup?.id ?: return@LaunchedEffect
models =
runCatching { withContext(Dispatchers.IO) { fetchSetupModels(settings, id) } }
.getOrDefault(emptyList())
}
val current = setup?.providers?.firstOrNull { it.name == providerName }
// Only the Claude CLI has models, a working directory and permission modes; keying the
// extra fields on the kind rather than the provider name keeps a second Claude provider
@@ -175,12 +196,13 @@ fun SpawnScreen(
)
if (isLlama) {
// A llama session names one of the models this backend has downloaded, so the choice is
// that list rather than free text -- a name that is not on disk is a session that
// cannot start.
// A llama session names one of the models on the machine it will run on, so the
// choice is that list rather than free text -- a name that is not on that machine's
// disk is a session that cannot start.
if (models.isEmpty()) {
Text(
"No models downloaded yet. Get one from the Models screen first.",
"No models on ${setup?.name ?: "this machine"}. The Models screen downloads " +
"to the backend; another machine needs the file put there itself.",
style = MaterialTheme.typography.bodyMedium,
color = MaterialTheme.colorScheme.onSurfaceVariant,
)
@@ -251,6 +273,19 @@ fun SpawnScreen(
selected = permissionMode,
onSelect = { permissionMode = it },
)
Spacer(Modifier.height(16.dp))
// Says what it does to *later* spawns as well, because it does: the level chosen here
// is stored as the default, which is the whole way that default is set. A picker that
// quietly changed a global would be the same control with the fact left out.
ChipGroup(
label = "Thinking (kept as the default for new sessions)",
options = listOf(DEFAULT_EFFORT) + EFFORT_LEVELS,
// The CLI's own default is a level in the list, so this cannot be a one-way trip.
// Disabled-looking until the server has answered, for the reason above.
selected = if (defaultsAsked) effort ?: DEFAULT_EFFORT else null,
onSelect = { chosen -> effort = chosen.takeIf { it != DEFAULT_EFFORT } },
)
}
Spacer(Modifier.height(24.dp))
@@ -268,6 +303,13 @@ fun SpawnScreen(
try {
val spawned =
withContext(Dispatchers.IO) {
// Stored before the spawn and not after it: choosing a level is
// an intent about new sessions in general, so a spawn that then
// fails must not also lose the choice. Non-fatal for the same
// reason the fetch above is -- the session is what was asked for.
if (isClaude) {
runCatching { setDefaultEffort(settings, effort) }
}
spawnSession(
settings,
// The id, not the label: labels are editable and the server
@@ -280,6 +322,7 @@ fun SpawnScreen(
if (isLlama) modelKey else model.trim().takeIf { isClaude },
cwd = cwd.trim().takeIf { isClaude },
permissionMode = permissionMode.takeIf { isClaude },
effort = effort.takeIf { isClaude },
// Sent only when set, so blank means "whatever llama.cpp does
// by default" rather than a zero.
params =
@@ -0,0 +1,27 @@
package com.example.aiapp
/**
* Where one transcript lives: a session's own, or one of its subagents'.
*
* The single mechanism [fetchTranscript], [EventStream], [TranscriptSource] and
* [TranscriptCache.session] all take, rather than each growing its own branch between a session and
* a subagent -- see SUBAGENTS.md's "Phone" and "Wire shape". A caller that has only a session id
* builds one with the one-argument constructor; a subagent's screen supplies both ids.
*/
data class TranscriptAddress(val sessionId: String, val subagentId: String? = null) {
/** The URL segment naming this transcript, before `/transcript` or `/events`. */
val urlPath: String
get() =
if (subagentId == null) "sessions/$sessionId"
else "sessions/$sessionId/subagents/$subagentId"
/**
* Where this transcript's cache lives on the phone, relative to the cache root.
*
* A subagent's nests under its session's directory rather than sitting beside it, so deleting a
* session's cache directory takes its subagents' with it -- the same one-way door the server's
* own storage describes.
*/
val cachePath: String
get() = if (subagentId == null) sessionId else "$sessionId/subagents/$subagentId"
}
@@ -35,8 +35,15 @@ class TranscriptCache(
private val root: File,
private val warn: (String) -> Unit = { Log.w("ai-app", it) },
) {
/** The cache for one session, whether or not anything has been stored for it yet. */
fun session(id: String): SessionCache = SessionCache(File(root, id), warn)
/**
* The cache for one transcript, whether or not anything has been stored for it yet.
*
* A subagent's [TranscriptAddress.cachePath] nests it under its session's directory, so
* deleting the session (below) takes its subagents' caches with it -- there is no separate
* purge for one.
*/
fun session(address: TranscriptAddress): SessionCache =
SessionCache(File(root, address.cachePath), warn)
/**
* Deletes every session directory not in [ids], called after a successful list fetch. The path
@@ -175,6 +175,18 @@ sealed class TranscriptItem {
val preTokens: Long?,
val postTokens: Long?,
) : TranscriptItem()
/**
* The account ran out of quota, so the turn stopped here.
*
* A divider rather than an error: nothing failed, and what a reader scrolling back needs from
* it is the same thing a clear or a compaction gives them -- why the conversation stops at this
* line.
*
* [resetsAt] is epoch seconds and null where the session was told nothing, which is a state the
* row has words for rather than a time it invents.
*/
data class LimitNote(override val seq: Long, val resetsAt: Double?) : TranscriptItem()
}
/**
@@ -458,6 +470,7 @@ fun foldEvent(items: List<TranscriptItem>, entry: SeqEvent): List<TranscriptItem
} else {
items + TranscriptItem.ImageItem(entry.seq, event.ref)
}
is SessionEvent.LimitReached -> items + TranscriptItem.LimitNote(entry.seq, event.resetsAt)
is SessionEvent.Cleared -> items + TranscriptItem.ClearedNote(entry.seq)
is SessionEvent.Compacted ->
items + TranscriptItem.CompactedNote(entry.seq, event.preTokens, event.postTokens)
@@ -18,7 +18,7 @@ import java.util.concurrent.atomic.AtomicReference
*/
class TranscriptSource(
private val settings: ServerSettings,
private val sessionId: String,
private val address: TranscriptAddress,
val cache: SessionCache,
) {
private val stream = AtomicReference<EventStream?>(null)
@@ -65,7 +65,7 @@ class TranscriptSource(
val tail = cache.tail() ?: return false
// `before = seq + 1` is the newest event with seq <= the cursor, which is the event *at*
// the cursor when the server still has one there.
val answer = fetchTranscript(settings, sessionId, before = tail.seq + 1, limit = 1)
val answer = fetchTranscript(settings, address, before = tail.seq + 1, limit = 1)
val matches =
answer.size == 1 &&
try {
@@ -83,7 +83,7 @@ class TranscriptSource(
*/
suspend fun fetchOpening(): List<SeqEvent> {
DebugStats.count("transcript page from server")
val page = fetchTranscript(settings, sessionId, limit = OPENING_WINDOW)
val page = fetchTranscript(settings, address, limit = OPENING_WINDOW)
page.forEach { (line, entry) -> cache.append(line, entry.seq) }
cache.flush()
return page.map { it.second }
@@ -108,7 +108,7 @@ class TranscriptSource(
val page =
fetchTranscript(
settings,
sessionId,
address,
before = before,
limit = limit,
coalesce = coalesce,
@@ -131,7 +131,7 @@ class TranscriptSource(
* well lose.
*/
fun follow(after: Long, onOpen: () -> Unit, onReset: () -> Unit, onEvent: (SeqEvent) -> Unit) {
val opened = EventStream(settings, sessionId)
val opened = EventStream(settings, address)
stream.getAndSet(opened)?.close()
try {
opened.run(after, onOpen, onReset) { raw, entry ->
@@ -0,0 +1,6 @@
<?xml version="1.0" encoding="utf-8"?>
<resources>
<!-- Overridden by the `bench` build type's resValue (build.gradle.kts) to "AI Sessions bench",
so the two are never mistaken for each other in the launcher or in Settings. -->
<string name="app_name">AI Sessions</string>
</resources>
@@ -0,0 +1,32 @@
package com.example.aiapp
import java.time.ZoneId
import kotlin.test.Test
import kotlin.test.assertEquals
import kotlin.test.assertTrue
/**
* What the transcript says where a session ran out of quota.
*
* The pair worth a test is the one that reads the same when it goes wrong: a reset time that
* arrived and one that never did. The second must not turn into a plausible-looking time, because a
* reader has no way of telling an invented one from a reported one.
*/
class LimitRowTest {
private val utc = ZoneId.of("UTC")
@Test
fun `a reported reset time is shown as a time`() {
// 2026-09-05T12:00:00Z. Asserted as a prefix and the clock reading rather than as the
// whole string: the platform's own short-time format is what this asks for, and it
// differs by JDK and locale down to which space character separates the meridiem.
val summary = limitSummary(1_788_609_600.0, utc)
assertTrue(summary.startsWith("Usage limit reached • resets "), summary)
assertTrue(summary.contains("12:00"), summary)
}
@Test
fun `a limit with no reset time says only what is known`() {
assertEquals("Usage limit reached", limitSummary(null, utc))
}
}
@@ -23,7 +23,7 @@ class TranscriptCacheTest {
private fun cache() = TranscriptCache(File(temp, "v1/host_8443")) { said += it }
private fun session(id: String = "s") = cache().session(id)
private fun session(id: String = "s") = cache().session(TranscriptAddress(id))
private fun line(seq: Long, type: String = "toolStart") =
"""{"seq":$seq,"ts":1.5,"type":"$type","id":"x"}"""
+28
View File
@@ -0,0 +1,28 @@
# The P0 benchmark fixture
`transcript.jsonl` is a synthetic transcript in the app's own event model (the JSON lines
`GET /sessions/{id}/transcript` returns; see `Events.kt`'s `parseSeqEvent` and
`server/src/session/driver.rs`) -- never a real one. It is what both the Compose `bench` build
and iris's bench build open with no server, so the two apps draw exactly the same content and a
frame-time comparison is measuring the renderer rather than the data.
Generated by `./generate.py` (Python stdlib only, seeded -- `SEED = 20260905` -- so re-running it
reproduces the same file byte for byte). It writes into `assets/` -- a separate directory from this
script and README, because the Compose `bench` build type points its own asset source set straight
at `assets/` (`app/androidApp/build.gradle.kts`'s `sourceSets { getByName("bench") }`), and a Python
script and a markdown file have no business inside an APK:
- `transcript.jsonl` -- 3,601 events. The first 3,200 (`BACKLOG_COUNT`) are the scrolled-back
history the benchmark opens with: user turns, tool calls with kilobyte-scale input/output,
assistant replies built from headings, bold/italic/inline code, a link, fenced code blocks that
rotate through rust/kotlin/python/sh/json/toml, a markdown table, two embedded images, and
periodic `usageDelta`/`compacted` events. The remaining 400 (`STREAM_COUNT`) are not part of the
opening window -- both bench harnesses replay them at a fixed rate (20/s) through the same live
fold path a real SSE reply arrives on, which is P0's "streaming phase."
- `bench1.png`, `bench2.png` -- tiny (8x8) flat-colour PNGs, base64-free on disk but served the
same way a real attachment is (`GET /sessions/{id}/files/{name}`), referenced by the two
`"type":"image"` events in the transcript.
Regenerate after changing the shape (a new event type, a different backlog/stream split) with
`./generate.py`, and commit the result -- it is checked in rather than generated at build time so
both apps' bench builds embed the identical bytes without needing this script at build time.
Binary file not shown.

After

Width:  |  Height:  |  Size: 74 B

Binary file not shown.

After

Width:  |  Height:  |  Size: 74 B

File diff suppressed because it is too large. Load diff
+186
View File
@@ -0,0 +1,186 @@
#!/usr/bin/env python3
"""Generates transcript.jsonl -- the synthetic fixture P0's benchmark opens in both apps.
Deterministic (fixed seed), so a Compose bench APK and an iris bench APK draw byte-identical
content: the point of the fixture is a like-for-like comparison, not a realistic one.
Never a real transcript -- see AGENTS.md's ui-sandbox.sh, which this borrows its vocabulary
style from (headings, code fences, a table, a link) rather than reusing its Claude-Code JSONL
shape. This file's shape is the *app's own event model* instead: one JSON object per line,
matching what GET /sessions/{id}/transcript returns and what Events.kt's parseSeqEvent reads
(server/src/session/driver.rs is the source of truth for the field names).
./generate.py writes transcript.jsonl and bench1.png/bench2.png here
BACKLOG_COUNT events (seq 1..BACKLOG_COUNT) are the scrolled-back history the benchmark opens
with. A further STREAM_COUNT events (seq BACKLOG_COUNT+1..) are not part of the opening window;
both bench harnesses replay them at a fixed rate as the "streaming reply" phase, appended through
the same live path a real SSE reply arrives on. Keeping both halves in one file means one
generator and one seed to keep in sync, rather than two fixtures that can drift apart.
"""
import base64
import json
import random
import struct
import zlib
from pathlib import Path
SEED = 20260905
BACKLOG_COUNT = 3200
STREAM_COUNT = 400
HERE = Path(__file__).resolve().parent / "assets"
random.seed(SEED)
LANGUAGES = ["rust", "kotlin", "python", "sh", "json", "toml"]
CODE_SNIPPETS = {
"rust": '''fn fold_event(items: Vec<Item>, seq: u64) -> Vec<Item> {
// a comment worth keeping: this is the fold the app's own screen runs
let mut out = items;
out.push(Item::new(seq));
out
}''',
"kotlin": '''fun foldEvent(items: List<TranscriptItem>, entry: SeqEvent): List<TranscriptItem> {
// mirrors the server's own event model, one item per line
return items + TranscriptItem.from(entry)
}''',
"python": '''def render_report(frames, cpu_ms, rss_kb):
# printed for a human to paste back, so every number carries its unit
return f"{frames} frames, {cpu_ms}ms cpu, {rss_kb}kb peak rss"''',
"sh": '''#!/bin/sh
# scripted scroll loop, the shape transcript-bench.sh drives on a phone
for i in $(seq 1 24); do
ui-trace record --do "swipe 540 700 540 1600 200"
done''',
"json": '{"seq": 1, "type": "status", "state": "running"}',
"toml": '''[package]
name = "bench-fixture"
version = "0.1.0"''',
}
HEADINGS = [
"## Plan",
"## What changed",
"## Why this approach",
"### Open questions",
"## Results",
]
WORDS = (
"session render report frame budget scroll transcript fold event cache "
"cursor probe stream backlog swipe fixture bench compose iris widget layout "
"measure place draw tool call token context window anchor"
).split()
def paragraph(n=24):
words = [random.choice(WORDS) for _ in range(n)]
words[0] = words[0].capitalize()
text = " ".join(words) + "."
# Sprinkle markdown inline spans so the syntax highlighter/markdown parser sees a real mix.
text = text.replace(" fold ", " **fold** ", 1)
text = text.replace(" cursor ", " *cursor* ", 1)
text = text.replace(" cache ", " `cache` ", 1)
if "bench" in text:
text = text.replace(
" bench ", " [bench](https://example.com/bench) ", 1
)
return text
def make_png(rgb, size=8):
"""A tiny, valid PNG -- flat colour, no external dependency."""
def chunk(tag, data):
c = tag + data
return struct.pack(">I", len(data)) + c + struct.pack(">I", zlib.crc32(c))
sig = b"\x89PNG\r\n\x1a\n"
ihdr = struct.pack(">IIBBBBB", size, size, 8, 2, 0, 0, 0)
raw = b""
for _ in range(size):
raw += b"\x00" + bytes(rgb) * size
idat = zlib.compress(raw)
return sig + chunk(b"IHDR", ihdr) + chunk(b"IDAT", idat) + chunk(b"IEND", b"")
def main():
HERE.mkdir(exist_ok=True)
lines = []
seq = 1
ts = 1_788_000_000.0
def emit(type_, **fields):
nonlocal seq, ts
obj = {"seq": seq, "ts": round(ts, 3), "type": type_}
obj.update(fields)
lines.append(json.dumps(obj, separators=(",", ":")))
seq += 1
ts += random.uniform(0.05, 2.0)
emit("status", state="running")
emit("settings", model="bench-model", permissionMode="auto")
image_refs = []
turn = 0
while seq <= BACKLOG_COUNT:
turn += 1
emit("userMessage", text=f"Turn {turn}: {paragraph(12)}", id=None, attachments=[])
# A tool call with kilobyte-scale input/output every few turns.
if turn % 3 == 0:
tool_id = f"tool-{turn}"
big_input = json.dumps({"path": f"/repo/file_{turn}.rs", "content": paragraph(400)})
emit("toolStart", id=tool_id, tool="Edit", input=big_input)
big_output = "\n".join(paragraph(60) for _ in range(20))
emit("toolUpdate", id=tool_id, output=big_output[: len(big_output) // 2])
emit("toolEnd", id=tool_id, output=big_output)
# A reply: a heading, prose, a fenced block in a rotating language, a table, then deltas.
emit("assistantText", delta=f"{random.choice(HEADINGS)}\n\n")
emit("assistantText", delta=paragraph(30) + "\n\n")
lang = LANGUAGES[turn % len(LANGUAGES)]
emit("assistantText", delta=f"```{lang}\n{CODE_SNIPPETS[lang]}\n```\n\n")
if turn % 5 == 0:
emit(
"assistantText",
delta="| column | value |\n|---|---|\n| a | " + paragraph(3) + " |\n\n",
)
# A run of small deltas -- the shape a live reply actually streams in.
for _ in range(random.randint(3, 8)):
emit("assistantText", delta=paragraph(6) + " ")
# A couple of images, base64 PNGs, the way a real transcript embeds a screenshot.
if turn in (10, 40):
ref = f"bench{len(image_refs) + 1}.png"
image_refs.append(ref)
emit("image", ref=ref, about=None)
emit("usageDelta", tokens=random.randint(200, 4000), context=random.randint(2000, 180000))
if turn % 15 == 0:
emit(
"compacted",
preTokens=180000,
postTokens=20000,
trigger="auto",
)
# The streaming-phase tail: one long reply, built entirely from text deltas, the shape a
# bench harness replays at a fixed events/sec through the live fold path.
emit("userMessage", text="One more, streamed live for the benchmark's timing phase.", id=None, attachments=[])
while seq <= BACKLOG_COUNT + STREAM_COUNT:
emit("assistantText", delta=paragraph(5) + " ")
emit("status", state="idle")
(HERE / "transcript.jsonl").write_text("\n".join(lines) + "\n")
(HERE / "bench1.png").write_bytes(make_png((220, 90, 90)))
(HERE / "bench2.png").write_bytes(make_png((90, 150, 220)))
print(f"wrote {len(lines)} events ({BACKLOG_COUNT} backlog + {STREAM_COUNT} stream) to transcript.jsonl")
if __name__ == "__main__":
main()
+9 -3
View File
@@ -4,6 +4,11 @@
# ./build-apk.sh the release build, signed (what the phone runs)
# ./build-apk.sh debug the debug build, for reproducing something the
# emulator scripts would build anyway
# ./build-apk.sh bench P0's benchmark build (own app id, "AI Sessions
# bench" label, opens straight onto the fixture
# session -- see docs/RUST.md's P0 box and
# app/bench-fixture/README.md). Signed the same
# as release; never touches the CA it pins.
#
# Dev Updater's `.dev-updater.ron` at the checkout root spells these out as
# build modes, one command line each; it passes nothing else, so the word
@@ -25,8 +30,9 @@ VARIANT=${1:-release}
case "$VARIANT" in
release) TASK=assembleRelease ;;
debug) TASK=assembleDebug ;;
bench) TASK=assembleBench ;;
*)
echo "build-apk.sh: unknown variant '$VARIANT' (release, debug)" >&2
echo "build-apk.sh: unknown variant '$VARIANT' (release, debug, bench)" >&2
exit 2
;;
esac
@@ -81,7 +87,7 @@ fi
# uninstalling it first: the signatures differ, and Android refuses to
# update across them.
KEYSTORE="${AI_APP_KEYSTORE:-${XDG_CONFIG_HOME:-$HOME/.config}/ai-app/release.jks}"
if [ "$VARIANT" = release ] && [ ! -f "$KEYSTORE" ]; then
if { [ "$VARIANT" = release ] || [ "$VARIANT" = bench ]; } && [ ! -f "$KEYSTORE" ]; then
KEYTOOL="${JAVA_HOME:+$JAVA_HOME/bin/keytool}"
KEYTOOL="${KEYTOOL:-keytool}"
if ! command -v "$KEYTOOL" >/dev/null 2>&1; then
@@ -97,7 +103,7 @@ if [ "$VARIANT" = release ] && [ ! -f "$KEYSTORE" ]; then
-keyalg RSA -keysize 2048 -validity 10000 \
-storepass "$PASSWORD" -keypass "$PASSWORD" -dname "CN=ai-app" >/dev/null 2>&1)
fi
if [ "$VARIANT" = release ]; then
if [ "$VARIANT" = release ] || [ "$VARIANT" = bench ]; then
AI_APP_KEYSTORE="$KEYSTORE"
AI_APP_KEYSTORE_PASSWORD=$(cat "$KEYSTORE.password")
export AI_APP_KEYSTORE AI_APP_KEYSTORE_PASSWORD
+34
View File
@@ -0,0 +1,34 @@
#!/bin/sh
# RUST.md's I5 "Where iris's frame time goes" pass (2026-09-05). The same
# 24-swipe/6-cycle loop as transcript-bench.sh's, extracted for iris's own
# demo app -- transcript-bench.sh itself is Compose-specific (opens by
# session title through the Compose app's own UI) and cannot be called
# directly against dev.iris.android.demo.
#
# MUST be run from inside this checkout (not /tmp): ui-trace/adb pick which
# emulator to target from the current directory's basename (the
# per-checkout-AVD rule), and a previous pass lost two attempts to a `cd`
# into /tmp that made this resolve to a nonexistent "tmp" checkout.
set -eu
cd "$(dirname "$0")"
. ./android-env.sh >/dev/null 2>&1
cycles=${1:-6}
ui-trace record -d 3000 --do "tap 'Reset frame report'" -o /tmp/iris-bench-reset.txt >/dev/null
adb logcat -c
DO=""
i=0
while [ "$i" -lt "$cycles" ]; do
DO="$DO --do 'swipe 540 700 540 1600 200' --do 'wait 500'"
DO="$DO --do 'swipe 540 700 540 1600 200' --do 'wait 500'"
DO="$DO --do 'swipe 540 1600 540 700 200' --do 'wait 500'"
DO="$DO --do 'swipe 540 1600 540 700 200' --do 'wait 500'"
i=$((i + 1))
done
eval ui-trace record -d $((cycles * 16000 + 20000)) $DO -o /tmp/iris-bench-scroll.txt >/dev/null
ui-trace record -d 3000 --do "tap 'Frame report'" -o /tmp/iris-bench-report.txt >/dev/null
sleep 1
adb logcat -d -s iris-android-app:I | grep "iris frame report:"
+1 -1
View File
@@ -7,7 +7,7 @@ edition = "2024"
# with `server/` via `event-model`), the REST + SSE clients for its HTTP
# surface (see `server/src/routes.rs`'s module doc for the table), the
# transcript fold and cache, the markdown block model, the syntax
# highlighter and the ANSI parser. See CLIENT_CORE.md at the repo root for
# highlighter and the ANSI parser. See `docs/CLIENT_CORE.md` for
# what this holds today, what it does not yet, and how it corresponds to
# the Kotlin it replaces.
#
+152
View File
@@ -0,0 +1,152 @@
//! What a Rust client needs to reach one enrolled server: host, port and
//! bearer token. Mirrors the shape `ServerConfig.kt`/`Api.kt`'s
//! `handleEnrollment` parses out of an `aiapp://enroll?host=H&port=P&token=T`
//! deep link -- the exact link `wg-app-link`'s `enroll` module mints and
//! `app/ui-sandbox.sh`'s banner prints, so any Rust client can enrol from
//! the same text a phone would scan as a QR, with no second format
//! invented for it (RUST.md's E4).
//!
//! What this type deliberately does not decide: where it is persisted, and
//! under what file permissions. A phone seals its token in the Android
//! Keystore; a desktop client has its own `$XDG_CONFIG_HOME/<app>/`
//! directory and its own file-mode conventions (MACHINE.md: owner-only,
//! never in the repo). Both are caller-specific, so they stay out of this
//! crate per the code rules' "ask for the least you need" -- see
//! `iris/desktop-app/src/config.rs` for the desktop instance.
use serde::{Deserialize, Serialize};
/// One enrolled server: reachable at `https://{host}:{port}`, authenticated
/// with `token` as a bearer header. Does not carry the pinned CA -- that is
/// a public certificate rather than a secret, and where to find it differs
/// by caller (a phone pins the one its APK was built against; a desktop
/// client is told a path).
#[derive(Debug, Clone, PartialEq, Serialize, Deserialize)]
pub struct EnrolledServer {
pub host: String,
pub port: u16,
pub token: String,
}
impl EnrolledServer {
/// Parses `aiapp://enroll?host=H&port=P&token=T` (query order does not
/// matter; unrecognised keys are ignored). `token` is percent-decoded,
/// since `ui-sandbox.sh` encodes it precisely because a raw token can
/// contain `+`, which turns into a space if left to a naive splitter.
pub fn parse_link(link: &str) -> Result<Self, String> {
let query = link.split_once('?').map(|(_, q)| q).ok_or_else(|| {
format!(
"'{link}' has no query string (expected \
aiapp://enroll?host=...&port=...&token=...)"
)
})?;
let mut host = None;
let mut port = None;
let mut token = None;
for pair in query.split('&') {
let Some((key, value)) = pair.split_once('=') else {
continue;
};
let value = percent_decode(value);
match key {
"host" => host = Some(value),
"port" => port = Some(value),
"token" => token = Some(value),
_ => {}
}
}
let host = host.ok_or_else(|| format!("'{link}' is missing 'host'"))?;
let port_str = port.ok_or_else(|| format!("'{link}' is missing 'port'"))?;
let port: u16 = port_str
.parse()
.map_err(|e| format!("'{link}''s port ('{port_str}') is not a number: {e}"))?;
let token = token.ok_or_else(|| format!("'{link}' is missing 'token'"))?;
Ok(Self { host, port, token })
}
/// Where a `client_core::api::UreqTransport` reaches this server.
pub fn base_url(&self) -> String {
format!("https://{}:{}", self.host, self.port)
}
}
fn percent_decode(s: &str) -> String {
let bytes = s.as_bytes();
let mut out = Vec::with_capacity(bytes.len());
let mut i = 0;
while i < bytes.len() {
if bytes[i] == b'%'
&& i + 2 < bytes.len()
&& let Ok(byte) =
u8::from_str_radix(std::str::from_utf8(&bytes[i + 1..i + 3]).unwrap_or(""), 16)
{
out.push(byte);
i += 3;
continue;
}
out.push(bytes[i]);
i += 1;
}
String::from_utf8_lossy(&out).into_owned()
}
#[cfg(test)]
mod tests {
use super::*;
#[test]
fn parses_host_port_and_token() {
let server =
EnrolledServer::parse_link("aiapp://enroll?host=127.0.0.1&port=8547&token=abcDEF123")
.unwrap();
assert_eq!(
server,
EnrolledServer {
host: "127.0.0.1".to_string(),
port: 8547,
token: "abcDEF123".to_string(),
}
);
assert_eq!(server.base_url(), "https://127.0.0.1:8547");
}
#[test]
fn field_order_does_not_matter() {
let server =
EnrolledServer::parse_link("aiapp://enroll?token=tok&port=443&host=example.com")
.unwrap();
assert_eq!(server.host, "example.com");
assert_eq!(server.port, 443);
assert_eq!(server.token, "tok");
}
#[test]
fn a_percent_encoded_token_is_decoded() {
// ui-sandbox.sh's own reason for encoding: a raw '+' would
// otherwise arrive as a space.
let server =
EnrolledServer::parse_link("aiapp://enroll?host=h&port=1&token=a%2Bb%2Fc").unwrap();
assert_eq!(server.token, "a+b/c");
}
#[test]
fn a_missing_field_is_named_in_the_error() {
let err = EnrolledServer::parse_link("aiapp://enroll?host=h&port=1").unwrap_err();
assert!(
err.contains("token"),
"error should name the missing field: {err}"
);
}
#[test]
fn a_non_numeric_port_is_named_in_the_error() {
let err = EnrolledServer::parse_link("aiapp://enroll?host=h&port=x&token=t").unwrap_err();
assert!(
err.contains("port"),
"error should name the offending field: {err}"
);
}
}
+2 -1
View File
@@ -1,9 +1,10 @@
//! The app's pure logic, shared between the server and any Rust client --
//! see `CLIENT_CORE.md` at the repo root for what lives here and what does
//! see `docs/CLIENT_CORE.md` for what lives here and what does
//! not yet.
pub mod ansi;
pub mod api;
pub mod config;
pub mod event_stream;
pub mod highlight;
pub mod notifications;
+2 -2
View File
@@ -1,7 +1,7 @@
//! This phone's copy of the transcripts it has already been sent, so
//! reopening a session does not download it again. Ported from
//! `app/.../TranscriptCache.kt`; see `TRANSCRIPT_CACHE.md` at the repo root
//! for the design and `CLIENT_CORE.md` for how this file corresponds to it.
//! `app/.../TranscriptCache.kt`; see `docs/TRANSCRIPT_CACHE.md`
//! for the design and `docs/CLIENT_CORE.md` for how this file corresponds to it.
//!
//! What is stored is the server's own JSON for one event per line, in
//! transcript order. Reading the cache means running the same [`seq_of`]
+115
View File
@@ -606,6 +606,37 @@ pub fn group_tool_runs(items: &[TranscriptItem]) -> Vec<TranscriptRow> {
rows
}
/// Folds a page of raw transcript lines (`ApiClient::fetch_transcript_page`'s
/// `Vec<Value>`) into the flat item list this module works over. A line
/// this build can't parse fails the whole page rather than being skipped --
/// CODE_RULES's "an enumeration must be able to say 'it broke'" -- since
/// silently dropping one event could hide, say, a user message that then
/// looks like it was never sent. Moved here from `desktop-app`'s `app.rs`
/// (RUST.md's E4) when the Android transcript client (I5) needed the same
/// fold: "write the logic once" applies to any caller embedding
/// `transcript-ui` against a live server, not just the first one.
pub fn fold_page(values: &[serde_json::Value]) -> Result<Vec<TranscriptItem>, String> {
let mut items = Vec::new();
for value in values {
let event: SeqEvent = serde_json::from_value(value.clone()).map_err(|e| {
format!("the server sent a transcript line this build couldn't parse: {e}")
})?;
items = fold_event(&items, &event);
}
Ok(items)
}
/// The wire `seq` a raw transcript line carries -- the live-stream resume
/// cursor after loading a page must be this, not a folded item's `seq()`.
/// A folded `AssistantMsg` keeps the seq of the *first* delta it
/// accumulated (`fold_event`'s own doc), so resuming from that seq would
/// re-deliver every delta already folded into it, duplicating the tail of
/// a reply that was mid-stream when the page was fetched -- found via a
/// real screenshot in E4 (RUST.md), where the assistant's line doubled.
pub fn raw_seq(value: &serde_json::Value) -> Option<u64> {
value.get("seq")?.as_u64()
}
#[cfg(test)]
mod tests {
use super::*;
@@ -824,4 +855,88 @@ mod tests {
other => panic!("expected a QuestionCard, got {other:?}"),
}
}
fn line(seq: u64, json: serde_json::Value) -> serde_json::Value {
let mut obj = json;
obj["seq"] = serde_json::json!(seq);
obj["ts"] = serde_json::json!(1.0);
obj
}
/// The regression for a bug a real `run-headless.sh` screenshot found
/// in `desktop-app` (E4, RUST.md): resuming the live stream from the
/// last *item's* seq re-delivers the deltas already folded into a
/// still-open assistant message, doubling its tail. `raw_seq` of the
/// last wire line must be the true high-water mark instead, which for a
/// run of deltas is higher than every item's own `seq()`.
#[test]
fn the_resume_cursor_is_the_last_wire_seq_not_the_last_items_seq() {
let values = vec![
line(1, serde_json::json!({"type": "userMessage", "text": "hi"})),
line(
2,
serde_json::json!({"type": "assistantText", "delta": "a"}),
),
line(
3,
serde_json::json!({"type": "assistantText", "delta": "b"}),
),
line(
4,
serde_json::json!({"type": "assistantText", "delta": "c"}),
),
];
let after = raw_seq(values.last().unwrap()).unwrap();
assert_eq!(after, 4);
let items = fold_page(&values).unwrap();
let assistant_seq = items
.iter()
.find(|i| matches!(i, TranscriptItem::AssistantMsg { .. }))
.unwrap()
.seq();
assert_eq!(assistant_seq, 2);
assert_ne!(
after, assistant_seq,
"the fixed bug: these must differ here"
);
}
#[test]
fn a_page_folds_into_one_settled_assistant_message() {
let values = vec![
line(1, serde_json::json!({"type": "userMessage", "text": "hi"})),
line(
2,
serde_json::json!({"type": "assistantText", "delta": "hel"}),
),
line(
3,
serde_json::json!({"type": "assistantText", "delta": "lo"}),
),
];
let items = fold_page(&values).unwrap();
assert_eq!(
items,
vec![
TranscriptItem::UserMsg {
seq: 1,
text: "hi".to_string(),
attachments: Vec::new(),
},
TranscriptItem::AssistantMsg {
seq: 2,
text: "hello".to_string(),
settled: false,
},
]
);
}
#[test]
fn an_unparseable_line_fails_the_whole_page() {
let values = vec![serde_json::json!({"seq": 1, "ts": 1.0, "type": "not-a-real-type"})];
let err = fold_page(&values).unwrap_err();
assert!(err.contains("couldn't parse"));
}
}
+16
View File
@@ -26,6 +26,7 @@ next (a Masonry or iris transcript screen, most likely).
| `api.rs` | `Api.kt` | Partial -- see below |
| `event_stream.rs` | `EventStream.kt` | Done |
| `transcript_fold.rs` | `TranscriptItems.kt`, `ToolRows.kt` | Partial -- see below |
| `config.rs` | `ServerConfig.kt`'s `handleEnrollment` | New, desktop-only so far -- see below |
| *(not started)* | `TranscriptSource.kt` | Not started |
| *(not ported, and may never be)* | `TranscriptUnits.kt` | Out of scope -- see below |
@@ -103,6 +104,21 @@ deciding how `event_model` itself represents "a shape I don't recognise"
-- a shared-model decision affecting `server/` too, not a `client-core`-only
fix, so it is recorded here rather than silently worked around.
## `config.rs`: `EnrolledServer`
`EnrolledServer` (host, port, bearer token) plus `parse_link`, which reads
the exact `aiapp://enroll?host=H&port=P&token=T` deep link
`wg-app-link`'s `enroll` mints and `ServerConfig.kt`'s `handleEnrollment`
parses on the phone -- so any Rust client enrols from the same text a
phone would scan as a QR, with no second format invented for it (RUST.md's
E4, DECISIONS.md 2026-09-05). Deliberately does not decide where it is
persisted or under what file permissions -- a phone seals its token in the
Android Keystore, `iris/desktop-app/src/config.rs` writes it to
`$XDG_CONFIG_HOME/ai-app-desktop/enrollment.json` at 0600 -- since that is
caller-specific (the code rules' "ask for the least you need"). Its only
caller today is `desktop-app`; a future Android build of this crate would
be a second one, not a reason to move the type.
## What is not started at all
- **`TranscriptSource.kt`** -- the layer that decides whether a page comes
+287
View File
@@ -0,0 +1,287 @@
# Decisions taken for Iris to review
Short list of design choices made by the design agent without asking, so
they can be judged and reversed later. Detail lives in RUST.md (and IRIS.md
for iris API changes); this file is only the summary. Newest first. Items
marked **DEFERRED** are ones the agent chose not to decide alone.
## 2026-09-05
- **iris no longer asks every device for compute-shader limits it never
uses.** `adapter.request_device` (both `iris/src/android/render.rs` and
`iris/src/default/render.rs`) used `Limits::default()` plus an override
for `max_buffer_size`, and `Limits::default()` unconditionally requests
desktop-tier compute limits (`max_compute_workgroups_per_dimension:
65535`, per `wgpu_types`) even though nothing in `iris`/`iris-core`
creates a `ComputePipeline` or writes a `@compute` shader stage —
confirmed by grepping the whole tree, not assumed. That crashed
`request_device` outright on the Android emulator's software GL path
(`EMU_GPU=software`, `--features force-gles`): SwiftShader's GL reports
itself as OpenGL ES 3.0, which has no compute shaders at all, so the
adapter's real limit is 0 against the unconditional request for 65535 —
`RUST.md`'s "Software mode ... crashes for a third, different reason,"
2026-09-05, earlier today. The same would happen on any real
GLES-3.0-only Android device, not just the emulator. Fixed by a new
`iris_core::device_limits()` (`iris/core/src/render/mod.rs`), shared by
both platform backends so the two requests cannot drift, that zeros the
six `max_compute_*` fields explicitly rather than switching to a
downlevel `Limits` preset — `Limits::downlevel_webgl2_defaults()` was
considered and rejected: it also zeros
`max_storage_buffers_per_shader_stage`, and `shader.wgsl`'s vertex stage
reads four `var<storage>` buffers (rects, glyphs, masks, move_offsets),
so that preset would trade the compute crash for a bind-group-layout
one on the same downlevel hardware this is meant to support. No
capability check or fallback path was needed since nothing is being
disabled — the request is simply narrowed to what the pipeline actually
uses. `rigs/gpu-probe`'s own mirrored limits (it is deliberately its own
crate, not a workspace member, so it cannot call `device_limits()`
directly) were updated to match, and confirm `IRIS DEVICE: ok` against
this VM's own Vulkan and GL adapters. **Not verified this pass**: the
specific SwiftShader-ES-3.0 crash this fixes, on-device — the
`EMU_GPU=software` cold boot this needs would have force-restarted this
checkout's emulator while another session was actively running its own
app on it (`com.example.aiapp` had window focus at the time), so it was
left for a pass when the emulator is free rather than disrupting that
session. Everything reachable without the emulator is clean: `cargo
fmt`/`clippy --workspace --all-targets`/`test --workspace`, `cargo ndk
build`/`clippy` for `iris-android-app` with `force-gles`, and
`gpu-probe` against this VM's own Vulkan and GL(ES 3.2, which still has
compute and so would not have reproduced the crash even before this
fix — not a substitute for the real ES-3.0 test).
- **P0's Compose half is built and smoke-tested on the emulator** — the
`bench` build type, the shared `app/bench-fixture/` transcript, and an
in-process fake backend (`BenchFixture.kt`/`BenchNetwork.kt`) that
answers `TranscriptSource`/`EventStream` from an in-memory event log
instead of a real server, so the fold and paging under test are the real
ones. Full account, the smoke run's report, and what is deliberately
left (the iris half, the real on-phone runs) are in RUST.md's P0 box.
Not a decision to review so much as the gate itself now being runnable —
flagged here because it is the first half of something Iris explicitly
asked to see before P1.
- **P0's iris half is also built and smoke-tested on the emulator,
2026-09-05.** A new `bench` Cargo feature on `iris-android-app`, on top
of `transcript-screen`: the same checked-in fixture (`include_str!`, no
asset pipeline needed), the same 24-swipe scroll loop animated through
`List::scroll` and the same 400-event/20s streaming phase through
`fold_event`, "Run benchmark"/"Copy report" as named accessible
controls, and the same three added report fields (process CPU time,
peak RSS, battery current) via direct JNI calls
(`bench_jni.rs::PlatformHandle`) since `android_view` has no
`BatteryManager`/`ClipboardManager` wrapper of its own. One small public
API addition to get there: `AndroidAppState::platform_ready` (`IRIS.md`),
a default-no-op lifecycle hook handing an implementor a `JavaVM` +
`GlobalRef` it can call Java through from any thread. Packaged with a
new `release` build type on `iris-android-app`'s own Gradle project
(there was previously only `debug`), signed with the same key
`app/build-apk.sh` generates. Smoke run and the full report are in
RUST.md's P0 box; not attempted this pass: the real on-phone runs and
Iris's pass/fail call, which is the actual gate.
- **The intermittent touch-scroll dropout is root-caused and fixed: a
missed `ACTION_DOWN` hit-test, not the previously-suspected coalesced
first `ACTION_MOVE`.** Diagnosed by temporary logcat tracing of every
touch event, `DragArbiter` state transition and `Selection::drag`
dispatch (removed once confirmed), reproduced on this checkout's own
emulator against a real sandbox session. The trace showed the actual
mechanism: a gesture's `ACTION_DOWN` lands wherever the finger actually
is, which is not guaranteed to fall inside the same row-local sensor
region a later `ACTION_MOVE` in the same gesture lands in (a row's own
padding/gap, or its non-selectable sender-name header, is
pointer-transparent to `iris::sense::CursorSense`). When that happens,
the widget that ends up handling the gesture never saw `PressStart`, so
`DragArbiter` sits in `Idle` — which answers every subsequent frame with
`Undecided` and has no way to tell "no press is happening" from "a press
is happening but I missed its start," so it never recovers on its own
for the rest of that gesture. One real trace showed exactly this: touch
`Down`/`Move`/`Up` all delivered correctly, but zero `PressStart`
reaching the arbiter, `state=Idle` unchanged from first frame to last.
Fixed at the call site that has the context to recover
(`iris::transcript_ui::selection::Selection::drag`,
`iris/transcript-ui/src/selection.rs`): a new `DragArbiter::is_idle()`
(`iris/src/sense.rs`) lets it notice a `Pressing` frame arriving with the
arbiter still `Idle` — which can only mean a missed `PressStart`, since a
`Pressing` sense requires the button to genuinely be down — and start the
press there instead of where it was missed. Three new unit tests in
`sense.rs`'s `drag_arbiter_tests` and one in `transcript-ui`'s
`selection::tests` (the latter fails on the code before this fix).
Commit follows. Not the same failure the earlier pass's `DECISIONS.md`
DEFERRED item speculated about (a coalesced first `ACTION_MOVE` skipping
slop detection) — that hypothesis is now ruled out; the arbiter's own
slop/long-press logic was never wrong. RUST.md's I5 box,
"Touch-scroll dropout root-caused, 2026-09-05" has the full trace.
- **P0, a phone benchmark gate before any porting, asked for by Iris
2026-09-05**: "before P1 I'd like to see benchmarks & also maybe stress
test on my own phone ... If it doesn't match compose reasonably well then
I don't think I'd wanna continue." Design (RUST.md's P0 box has the
detail): the same embedded synthetic fixture in both apps with no server
needed; the same scripted scroll loop then a streaming phase, run
programmatically since the phone has no usable system tracing and no
agent can drive it; the same report from both (frames, janky %, p50/p90/
p99, process CPU time, peak RSS, battery current where readable) with a
copy button; the iris app under its own id and the Compose one as a new
`bench` build type with an id suffix, so neither replaces her production
install; two arm64 APKs plus instructions delivered under `~/host/bench/`.
The gate is hers: iris within a reasonable margin of Compose release on
p50, p99 and CPU time, no crashes, no visible stutter. If it fails, the
port stops.
- **The rest of the port is one UI crate, `iris/app-ui`, grown out of
`iris/transcript-ui` rather than started beside it.** It holds a
`Screen` enum plus a back stack — the Rust equivalent of `AppRoot.kt`'s
`when` — and `iris/desktop-app`/`iris/android-app` become thin entry
points over it. Chosen over a fresh crate because `transcript-ui`
already has the right generic shape (`Rsc: HasEvents` +
`Rsc::State: FocusHost`) and the `client-core`/`event-model` path
dependencies every later screen needs, so growing it in place is the
smaller diff. Platform-only code (notification service, share target,
QR scanner, Keystore token, deep-link enrolment) stays in the E3/E5
Java shell (`android-shell/` + `app/shellApp`) rather than moving into
this crate, since none of it is a screen. The Android APK is built by
`cargo xtask apk` (E5), merging the app-ui cdylib into the E3 shell so
there is one app rather than a demo shell plus a service shell.
`app/androidApp` (the Compose app) stays untouched and is the baseline
every step is measured against, until parity is reached (P7 decides
the switch, and is itself a load-bearing decision left to Iris). Order
is by risk to the daily-use path: session screen first (P1, where
every hard behaviour already lives), then the shell merge and a real
phone install (P2), then root tabs (P3), the explorer (P4),
settings/enrolment (P5), desktop parity (P6), and the cutover itself
(P7). Full plan: RUST.md's "The port, in order (decided 2026-09-05)".
- **iris gets its own measured frame report, rather than waiting on a
`dumpsys`/`gfxinfo` answer that cannot see a `SurfaceView`'s GPU-drawn
frames.** `iris_core::FrameReport` (`iris/core/src/render/frame_report.rs`)
times each frame's wall clock from the same point `render()`'s redraw
starts to just after `queue.submit` + `present()` — the span Compose's
own render report and `gfxinfo` both count — into a fixed 4096-entry
ring (no allocation per frame; `report()` is the only place that
allocates, and only on a button tap). The report gives total frames,
janky % over the same 16.7ms budget `gfxinfo` uses, P50/P90/P99 and the
worst, plus a reset. Exposed the way the Compose app's copy-button
report already is: two named controls ("Frame report", "Reset frame
report") on the transcript screen, tappable by accessibility name via
`ui-trace`, logging under this crate's fixed `android_logger` tag
(`iris-android-app`) so a script can grep `"iris frame report"` the way
`transcript-bench.sh` greps `"ai-app render report"`. The report's own
`Display` line says plainly that it measures up to the `present()` call
returning, not GPU/compositor completion — wgpu's `present()` is not
fenced against either, so presenting that span as "time to reach the
screen" would be a measured-looking number that is actually inferred,
which the standing UI rule forbids.
- **`ui-trace` gains a hold-then-drag gesture, additive, in
`emulator-tools`.** Neither of its two existing actions can produce
"hold stationary for `LONG_PRESS`, then move without lifting" — `tap`
has no hold and `swipe X1 Y1 X2 Y2 MS` interpolates motion across its
whole duration from t=0. A new action presses, waits, then moves to a
second point and releases as one continuous touch (raw
`sendevent`/`MotionEvent` injection, extending whatever mechanism the
existing `swipe` already uses), so `DragArbiter`'s pan-vs-select rule
(`iris/src/sense.rs`, already covered by 8 unit tests against a
synthetic clock) can finally be driven on a real device instead of only
in a test harness.
- **Touch drag on a transcript row follows Android's own rule**: a vertical
drag pans the list immediately; a stationary press held 500 ms starts a
text selection which further dragging extends; a horizontal drag while
something is already selected extends that selection without the wait.
One `DragArbiter` per list decides it (`iris/src/sense.rs`). Chosen over a
"text layer always wins" or "list always wins" rule because either loses
one of the two gestures a reader expects.
- **E4's desktop shape is a new `iris/desktop-app` crate**: a winit window
holding `transcript-ui`'s screen beside a session list, talking to a real
`ai-server` through `client-core`. It enrols by pasting the same
`aiapp://enroll?…` link a phone scans (`client-core::config::EnrolledServer`)
and keeps it owner-only under `$XDG_CONFIG_HOME/ai-app-desktop/`. The
pinned CA is a path given on the command line, not baked in. Chosen so
the phone and desktop share one enrolment format and no second one is
invented.
- **I5's Android integration extends `iris-android-app` (I2's shell)
behind a Cargo feature (`transcript-screen`), rather than a third
shell crate.** That project already has the Gradle module, the
`IrisView`/`MainActivity` Java, and the JNI registration; the only
thing a second screen needs on top is a different `AndroidAppState`,
the same axis `tabs_ui::build`/`transcript_ui::build` already vary
along on the winit side. `tabs-screen`/`transcript-screen` are
mutually exclusive and each pulls in only its own deps, so the plain
tabs build (I2/I4) is untouched.
- **Order of remaining work, updated 2026-09-05**: the two in-flight
pieces and I5's Android integration are all done; next is giving iris
its own frame-timing report so item 3 below can be decided by a number.
- **DECIDED by Iris, 2026-09-05: iris is the app's framework; Masonry was
the calibration.** Her words: "I think iris definitely makes more sense
based on the limitations we've found." The limitations: Masonry has no
touch scroll on Android (E2), no per-span rich text and no cross-row
selection on the pinned commit (E2), and its keyboard bridge is a TODO
(E1); iris carries the same screen under the Compose baseline on the
host GPU (p50 15.0 ms against Compose's 20.0 ms, RUST.md's I5 box). What
follows: the E-steps are closed as calibration, and the port proceeds
on iris — screens, the shell (E3/E5), and `client-core` underneath.
The item below is kept as the record of what she decided from.
- **Was DEFERRED — whether to commit to iris over Masonry for `ai-app`.**
Updated 2026-09-05 with the clean comparison the recommendation wanted:
same sandbox session content, same emulator, `EMU_GPU=software`, one
session. Headline numbers (RUST.md's I5 box, "Clean scroll comparison,
2026-09-05," has the full table and every caveat):
| app | build | frames | janky % | p50 | p90 | p99 | worst |
|---|---|---|---|---|---|---|---|
| Compose (in-app report) | debug | 1102 | 99.0% late | 33.8ms | 50.6ms | 79.5ms | -- |
| Compose (`dumpsys gfxinfo`) | debug | 1499 | 21.15% (95.66% legacy) | 32ms | 48ms | 150ms (p99) | -- |
| iris (`FrameReport`) | **release** | 299 | 94.65% | 79.1ms | 98.6ms | 117.8ms | 212.6ms |
| iris (`FrameReport`, repeat) | **release** | 233 | 94.42% | 109.3ms | 130.8ms | 147.1ms | 150.5ms |
**Not a clean apples-to-apples reading, stated plainly rather than
smoothed over**: iris had to be built **release** (debug `SIGSEGV`s on
this emulator's Vulkan loader, I4's finding) against Compose's mandated
**debug** build, so this asymmetry likely *understates* iris's gap
rather than the reverse; the three frame-time sources measure different
things (Compose's own phase accounting vs. Android's HWUI deadline-miss
definition vs. iris's redraw-start-to-present window, the last of which
`dumpsys gfxinfo` cannot see at all for iris's `SurfaceView`); and both
figures are emulator numbers under software rasterisation, which
Compose's *own* in-app report shows already costs 20-34ms/frame in
`swap`+`gpu` alone under this GPU mode, so a same-mode iris number well
above 16.7ms was expected going in for either app. A second pair under
`-gpu host` was not taken this pass. The earlier session's suspected
intermittent touch-delivery dropout was **not reproduced** this pass —
the zero-frame results this time traced to this pass's own script bug
(a `cd` that changed which emulator `ui-trace` targeted), not the
emulator; a CPU-load rise during the gesture was observed by a sampler
running throughout, but did not correlate with any failure, so the
original candidate is neither confirmed nor ruled out.
The choice in front of Iris, updated: decide now on the
structural-plus-functional case already made (iris works end-to-end
where Masonry's scroll gesture doesn't exist at all on Android) plus
this table — reading the two build profiles and three jank definitions
with the caveats above rather than as a single number — or ask for a
same-profile, same-GPU-mode rerun first. RUST.md's I5 box has the full
account.
**Updated 2026-09-05, the `-gpu host` pair taken.** Real GPU rendering
(`force-gles` -- the default Vulkan backend has no adapter at all under
plain host-GPU boot, confirmed by the exact `wgpu` error) reverses the
software-mode shape:
| app | build | GPU mode | frames | janky % | p50 | p90 | p99 | worst | cpu p50 | gpu-wait p50 |
|---|---|---|---|---|---|---|---|---|---|---|
| Compose (in-app report) | debug | host (virgl) | 1268 | 96.4% late | 20.0ms | 28.4ms | 37.7ms | -- | -- | -- |
| iris (`FrameReport`) | **release**, `force-gles` | host (virgl) | 62 | 41.94% | 15.0ms | 21.8ms | 37.1ms | 37.1ms | 0.2ms | 12.9ms |
Under real GPU rendering iris's median frame is *faster* than
Compose's, not the 2-3x-slower shape the software-mode table shows. A
new split inside `FrameReport` (redraw-to-submit vs. submit-to-present,
commit `e2a1fad`) says why: iris's own CPU work per frame is a median
0.2ms -- almost the entire frame is time spent handing the frame to the
driver, not in iris's layout/text/primitive code. This is consistent
with the earlier software-mode gap being mostly SwiftShader's CPU
rasterisation cost rather than an iris-specific slowness, but is not
proof of it: a same-mode software `force-gles` run to isolate the
backend crashed for an unrelated reason (SwiftShader's GL path reports
itself as OpenGL ES 3.0, which has no compute shaders, and iris's device
request assumes them unconditionally) — real scope to fix, not done
here — and the two apps' frame populations still differ in kind the same
way the software-mode caveats describe. A real intermittent touch-
scroll dropout was also reproduced this pass (six consecutive swipes
produced zero redraws while taps kept working; an identical retry then
succeeded) and is not explained. RUST.md's I5 box, "Where iris's frame
time goes, 2026-09-05, the `-gpu host` pass," has the full account. The
iris-vs-Masonry choice itself is still Iris's to make.
View File
File renamed without changes.
+439
View File
@@ -0,0 +1,439 @@
# iris: notable public API changes
For Iris to read on her own time. Each entry is a change to iris's public
surface that a widget author or app author would notice: a trait method
added, removed or re-shaped; a type that callers construct differently; a
capability that moved. Small and trivial changes do not go here.
An entry gives the date, what changed, why, and a short before/after where
it helps judge the change without the session that made it. Newest first.
## 2026-09-05: `AndroidAppState::platform_ready` (RUST.md's P0 box, iris half)
Added a second, optional lifecycle method to `iris::android::AndroidAppState`
(`iris/src/android/view.rs`), called once from `new_peer` right after `new`:
```rust
fn platform_ready(&mut self, rsc: &mut AndroidRsc<Self>, vm: JavaVM, view: GlobalRef) {}
```
Default does nothing, so every existing implementor (`Client`,
`TranscriptClient`) is unaffected. It exists for a caller that needs to call
into Java itself beyond what a `RequestRedraw` handle already covers --
P0's bench build (`iris-android-app`'s new `bench` feature,
`bench_client.rs`/`bench_jni.rs`) uses it to hold a `JavaVM` + `GlobalRef`
to the view so its "Copy report" control and once-a-second battery sampler
can call `BatteryManager`/`ClipboardManager` through the view's own
`Context`, from a background tokio task as well as the UI thread. `new`
itself was not extended with these two parameters: most implementors need
nothing here, and `new`'s job is building the widget tree, not holding a
platform handle. `vm`/`view` are independent handles from the ones
`new_peer` keeps for its own `RequestRedraw` (a fresh `get_java_vm`/
`new_global_ref` each), so storing them has no effect on that mechanism.
## 2026-09-05 (later still): `iris_core::device_limits()`, and iris no longer requests compute-shader limits
New public function, `iris_core::device_limits() -> wgpu::Limits`. Why:
`adapter.request_device`'s `required_limits` was `Limits::default()` plus
a `max_buffer_size` override in both platform backends, and
`Limits::default()` requests desktop-tier compute-shader limits
unconditionally (`max_compute_workgroups_per_dimension: 65535`) even
though nothing in `iris`/`iris-core` uses a `ComputePipeline` — that
crashed device creation outright on a downlevel GL adapter reporting
OpenGL ES 3.0 (no compute shaders at all: the Android emulator's
`EMU_GPU=software` path, and any real GLES-3.0-only Android device).
`device_limits()` is what both `android::render::AndroidRenderer::new`
and `default::render::UiRenderer::new` now build their `required_limits`
from, so the request cannot drift between the two backends.
Before: `Limits { max_buffer_size: 1 << 30, ..Default::default() }`
inlined in each backend. After: `iris_core::device_limits()`, which is
the same thing with the six `max_compute_*` fields additionally zeroed.
A caller building its own `DeviceDescriptor` outside these two backends
(there are none today, but a third platform backend would want this)
should call `device_limits()` rather than reaching for
`Limits::default()` directly, unless it genuinely adds a compute pass —
in which case it wants the specific compute limits that pass needs, not
the desktop-tier default for everything.
## 2026-09-05 (later the same day): `iris_core::FrameReport` (RUST.md's I5 box)
New public type, `iris_core::FrameReport` (re-exported from `iris_core`'s
`render` module alongside `FrameStats` and `JANK_THRESHOLD`). Why: `dumpsys
gfxinfo` cannot see a `SurfaceView`'s own GPU-drawn frames at all, so a
`wgpu`-rendered iris screen had no way to ask "was this smooth" the way
Compose's own in-app render report already can -- item 3 of RUST.md's
recommendation was stuck on a one-sided number for exactly this reason.
`FrameReport::record(elapsed: Duration)` is called once per frame (wired
into `android/view.rs`'s `render()`, wrapping the same span from redraw
start to after `queue.submit`+`present()` that Compose's report and
`gfxinfo` both count) and writes into a fixed 4096-entry ring -- no
allocation on the hot path. `FrameReport::report() -> Option<FrameStats>`
gives total frames, janky % (over `JANK_THRESHOLD`, the same 16.7ms 60Hz
budget `gfxinfo` uses), P50/P90/P99 and the worst; `None` if nothing has
been recorded since the last `reset()`, not a zeroed report that would
read as a real measurement. `FrameStats`'s `Display` line says plainly
that it measures up to `present()` being called, not GPU/compositor
completion, since wgpu's `present()` isn't fenced against either.
`AndroidUiState` gained a `pub frame_report: FrameReport` field --
anything with `HasAndroidUiState` can now read or reset it. Before this,
there was no way to ask iris's own render path how long a frame took at
all, on any backend.
Before/after, for a caller that already has `ui_state: &AndroidUiState`:
```rust
// before: no such question could be asked
// after:
match ui_state.frame_report.report() {
Some(stats) => log::info!("iris frame report: {stats}"),
None => log::info!("iris frame report: no frames recorded yet"),
}
ui_state.frame_report.reset(); // via android_state_mut()
```
`iris-android-app`'s transcript screen exposes this as two named,
tappable controls ("Frame report", "Reset frame report") rather than
requiring a caller to wire its own UI -- see `transcript_client.rs`'s
`frame_report_controls`.
## 2026-09-05: `Tasks::redraw_handle` (RUST.md's I5 Android integration)
New public method on `iris::task::Tasks`, `redraw_handle(&self) ->
Arc<dyn RequestRedraw>`. Why: a caller running its own long-lived loop
*inside* one spawned task (a live SSE follow, the Android transcript
client's `select_session`) has no other way to ask for a frame after each
`TaskCtx::update` -- `Tasks::spawn`'s own wrapper only requests one, after
the whole async closure finishes, which fits a single request-then-update
but not a stream that needs to be seen redrawing after *each* event. This
is the same gap `iris/desktop-app`'s module doc names for why it uses
winit's `Proxy<AppEvent>` instead of `Tasks` -- android-view has no
`Proxy`, so this is what closes it there.
**A real bug this uncovered, not a hypothetical**: calling the returned
handle's `request_redraw()` from the background thread crashed the process
(`SIGABRT`, `Result::unwrap() on an Err value: JavaException`) the first
time an Android transcript fetch called it a second time. `android/render.rs`'s
`AndroidRedrawHandle` was already attaching the calling thread to the JVM
correctly, but its `request_redraw` called `View::post_frame_callback`,
whose Java side calls `Choreographer.getInstance()` -- which throws unless
the *calling* thread already has a `Looper`, and a tokio worker thread,
even freshly JNI-attached, has none. Fixed by routing through
`View::post_delayed(0)` instead (Android's own thread-safe "queue work onto
this View's UI thread" primitive, needing no caller-side `Looper`), landing
on a new `IrisViewPeer::delayed_callback` override that drains tasks and
renders -- same body as `do_frame`, on the UI thread where
`post_frame_callback` is safe again. Any future caller of `redraw_handle()`
from a background thread gets this for free; nothing about the fix is
specific to the transcript screen.
## 2026-09-05: `transcript_ui::build_tree` (RUST.md's E4)
`transcript_ui::build` claimed the whole window (`ui_state.set_root(tree)`)
as its last step, which is right for a window that *is* the transcript
screen (the winit example, an eventual Android cdylib) and wrong for the
desktop app, which puts a session list beside it. `build_tree` is `build`
minus that last step: it returns `(TranscriptScreen, StrongWidget)` instead
of just `TranscriptScreen`, and the caller decides where the tree goes —
into `ui_state.set_root`, or into a `WidgetPtr` alongside something else
(`iris/desktop-app`'s `rebuild_transcript`). `build` is now one line calling
`build_tree` and doing the `set_root` itself, so existing callers are
unaffected.
```rust
// before, and still available, for a caller that wants to *be* the window:
let screen = transcript_ui::build(rsc, &mut ui_state, rows);
// new, for a caller embedding the screen beside something else:
let (screen, tree) = transcript_ui::build_tree(rsc, rows);
some_widget_ptr(rsc).set(tree);
```
## 2026-09-05: `DragArbiter`, pan-vs-select for one shared touch gesture (RUST.md's I5)
New public type, `iris::sense::DragArbiter`. Why: a widget author who
registers both a list-level pan and a row-level drag-to-select on the same
touch gesture has no way to arbitrate between them — `core/src/sense.rs`'s
`run_sensors` always gives the innermost layer first refusal, so the inner
one wins every frame it is pressed, not just the frame the press started
(this is exactly what left transcript-ui's touch-drag panning unreachable
until now). `DragArbiter` is one small state machine, one instance per
gesture surface (a whole list, not per row), that a caller drives with its
own `press_start`/`update`/`release` calls and a caller-supplied `Instant`
(so it is unit-testable without a real clock or a render harness). It
decides the way Android itself does: an ordinary vertical drag pans
immediately; a stationary press held `LONG_PRESS` (500ms) starts a
selection, which any further drag then extends; a horizontal drag while
something is already selected extends it immediately, skipping the wait.
```rust
// One per list, held alongside whatever state coordinates the rows:
let mut arbiter = DragArbiter::new();
// On press-down:
arbiter.press_start(pos, Instant::now(), already_selected);
// Every frame the button/finger stays down:
match arbiter.update(pos, Instant::now()) {
DragOutcome::Pan(dy) => list.scroll(-dy),
DragOutcome::SelectStart => selection.begin(...),
DragOutcome::SelectExtend => selection.extend(...),
DragOutcome::Undecided => {}
}
// On release:
arbiter.release();
```
`transcript-ui`'s `Selection::drag` (`transcript-ui/src/selection.rs`) is
the reference caller: every row's `CursorSense::click_or_drag() |
CursorSense::unclick()` handler routes through one `Selection`-owned
arbiter instead of calling `begin`/`extend` directly, so a drag that starts
on a row's own rendered text now pans the list correctly instead of
always starting a selection. 8 new unit tests in `iris/src/sense.rs`'s
`drag_arbiter_tests` module.
### 2026-09-05, later: `DragArbiter::is_idle()`, recovering a missed `press_start`
Follow-up to the above, from a real touch-scroll dropout: a gesture's
`ACTION_DOWN` can land on a caller's own dead space (a row's padding, a
gap, a header with no handler) that never calls `press_start`, so the
first frame the arbiter actually sees is a `Pressing`-shaped `update`
with no matching start. Before this, `update`'s `Idle` arm had no way to
tell that apart from "nothing is happening" and answered `Undecided`
forever for the rest of that gesture. `is_idle(&self) -> bool` lets a
caller notice the gap and recover: if `is_idle()` is true on a frame the
caller knows a press is genuinely down (its own `Pressing`/equivalent
sense fired), call `press_start` right there instead of assuming one
already happened. `transcript-ui`'s `Selection::drag` is the reference
caller — one new match arm, checked before the ordinary `update`-only
case. Any other `DragArbiter` caller with the same "one sensor per
sub-region, no fallback for dead space" shape has the same gap and wants
the same recovery.
## 2026-09-05: `SpanStyle`, per-range text styling (RUST.md's I5)
A `TextBuffer` used to have exactly one style (`TextAttrs`: colour, size,
family, ...) for its whole string, applied via `push_default` into parley's
ranged builder. `SpanStyle` is a second, optional layer: a byte range plus
whichever of colour/family/font size/bold/italic/underline it overrides,
pushed with parley's own `push(property, range)` instead. Why: a transcript
row's markdown (a heading, **bold**, `inline code`, a link) all inside one
wrapped paragraph needs each to carry its own look while the paragraph
still wraps and selects as a single buffer — the thing `masonry`'s
`TextArea` cannot do (`StyleSet` is one style for the whole editor,
`text_area.rs:43-44`'s `// TODO: RichTextInput`), and the reason this
existed at all.
```rust
let (text, spans) = transcript_ui::markdown::render_markdown(src, 16.0);
wtext(text)
.spans(spans) // new: TextBuilder::spans, on both Text and TextEdit
.editable(EditMode::MultiLine)
.add(rsc);
```
Two things a widget author should know before reaching for it:
- **Call `.spans()` before or after `.editable()`, both work** — the field
lives on `TextBuilder` itself, not either output type, and both
`TextOutput::run` and `TextEditOutput::run` apply it to the buffer via
`TextBuffer::set_spans`. **These two call sites are a pair**: adding a
third `TextBuilderOutput` impl without also calling `set_spans` there
reproduces the exact bug this box shipped once already (spans silently
dropped for `TextEdit`, found only by screenshotting, not by any test —
`markdown.rs`'s own unit tests check string/range logic, which is
correct in isolation and proves nothing about whether the render path
ever sees it).
- **Colour is now per-glyph, not per-buffer.** `PlacedGlyph` gained a
`color: UiColor` field (from parley's own per-run `Style::brush`), and
`Painter::glyphs` draws each glyph in its own colour instead of
`RenderedText::color` uniformly. `RenderedText::color` still exists (the
buffer's *base* colour, for a caller that wants it as a whole, e.g. to
tint a cursor) but no longer drives what a glyph actually renders as.
## 2026-09-05: accessibility names via AccessKit (RUST.md's I4)
`.label()` (already in `trait_fns.rs`, previously unused anywhere in-tree)
is now load-bearing: it's the one thing that puts a widget in the AccessKit
tree `iris_core::ui::access::AccessTree` builds and both backends push
out. A widget author who wants a control to be findable by name (and
tappable by name, through `ui-trace`/a real screen reader) calls `.label()`
on it; nothing else is required, and a widget nobody labels is invisible
to this system at zero cost, not just zero UI.
```rust
let button = rect(Color::LIME)
.on(CursorSense::click(), move |_, rsc| { ... })
.label("Add task"); // now findable by uiautomator/AccessKit as "Add task"
```
Two new things a widget author might touch directly:
- **`Widget::access_role(&self) -> accesskit::Role`**, default `Unknown`.
Override it if your widget has a real platform equivalent —
`TextEdit` now returns `TextInput`/`MultilineTextInput` by `EditMode`.
Only consulted for a widget that also has a `.label()`; an unlabelled
widget's `access_role` is never called.
- **`Widgets::named() -> impl Iterator<Item = WidgetId>`** — every widget
with an explicit label, for anything else that wants to walk the same
set `AccessTree` does.
Nothing about `Painter`, `draw`, or the layout/move machinery changed —
this sits entirely beside them, reading `resolved_region`'s output rather
than participating in producing it.
## 2026-09-05: `List`, a virtualised bottom-anchored list (RUST.md's I3)
A new widget, `iris::widget::List` (`iris/src/widget/list.rs` -- read its
module doc first), for the transcript's kind of screen: variable-height
rows, keyed by a `u64`, composed only while visible, moved rather than
re-laid-out on scroll, a scroll anchor that survives a row inserted above
it, "more" sentinels at each end, and "hold the edge nearest the tap" when
a row's height changes (`note_tap`, resolved in the layout pass).
```rust
let mut list = List::new(Axis::Y);
list.push_back(ListRow::new(key, row_widget)); // O(1)
list.push_front(ListRow::new(older_key, row)); // O(1), anchor unaffected
list.set_more_before(Some(spinner_widget)); // sentinel, drawn at the edge
list.note_tap(viewport_y); // before mutating a row's height
let (top, bottom) = list.extent(key).unwrap(); // last frame's on-screen box, if visible
```
Built entirely out of existing primitives (`Painter::widget`/`widget_within`/
`reposition`/`draw_twice`, and `draw_inner`'s own old-children diffing) --
no new mechanism was added to the render core for it. One correctness
lesson worth reading even for other widgets: a row that fills whatever
region it is offered (`Rect`, `is_size_independent`) cannot be measured at
a throwaway oversized region and then merely `reposition`ed into place --
`reposition` only ever writes an offset, never a size, so the oversized
primitive stays oversized. `List` fixes this by caching each row's real
height once measured and placing an already-known row directly at its
exact box; see `list.rs`'s `place` for the full reasoning and
`a_fill_shaped_background_is_not_left_oversized` for the regression test.
## 2026-09-05: a second backend (android-view), and what moved to make room for it
RUST.md's I2. Three changes a widget or app author would notice, all in
service of the same thing: `default` (winit) and the new `android`
(android-view) backends sharing what does not depend on windowing.
- **`Selector`/`Selectable`'s bound changed from `Rsc::State:
HasDefaultUiState` to `Rsc::State: FocusHost`** (new trait, `attr.rs`).
`HasDefaultUiState` still exists and still works — `default/attr.rs` now
implements `FocusHost` for anything that has it — so a winit app's
existing code is unaffected. An Android app implements `FocusHost` via
`HasAndroidUiState` instead. Affects only an app that referenced
`HasDefaultUiState` directly at a `Selectable`/`Selector` call site
rather than through `.attr::<Selectable>(())`, which nothing in-tree
does.
- **`Tasks::init` takes `Arc<dyn RequestRedraw>` instead of
`Arc<winit::window::Window>`.** `RequestRedraw` (`task.rs`) is one method,
`fn request_redraw(&self)`; `winit::window::Window` implements it
(`default/render.rs`), so `Tasks::init(window)` at a call site is
unchanged by inference. Only matters if something constructed a `Tasks`
directly rather than through `DefaultRsc`/`AndroidRsc`.
- **`TextEdit::apply_event`/`TextInputResult` are `#[cfg(not(target_os =
"android"))]`** — they take a `winit::event::KeyEvent`, which does not
exist on Android; `android/input.rs` drives the same primitives
(`backspace`/`delete`/`motion`/`insert`, all still unconditional) from
`ndk::event::Keycode` directly instead. New unconditional getters on the
way: `TextEdit::text()`/`selection_range()`/`caret()`, and
`TextEditCtx::delete_byte_range`/`set_cursor_byte` — the primitives
`android/ime.rs`'s `InputConnection` bridge needed and that were not
previously exposed publicly.
## 2026-09-04: `Widget::draw` reports the size it used; `desired_width`/`desired_height` are gone
A widget used to implement three methods (`draw`, `desired_width`,
`desired_height`); it now implements one, `fn draw(&mut self, painter: &mut
Painter) -> Size`, which draws into `painter.region()` and returns how much
of it was used. Why: the two extra methods routinely re-simulated what
`draw` was about to do anyway (`Span::desired_ortho` copied its own draw
loop to get cross-axis sizing right) — one visit per widget per frame
instead of up to three. A container that needs a child's size before
placing it (alignment, centering) draws the child once at a provisional
region, reads the returned `Size`, and calls the new `Painter::reposition`
to move it into its final spot — an O(1) offset write, not a second draw. A
widget whose drawn output never depends on the size it's given (a
fixed-size `Rect`, a decoded `Image`) overrides the new `fn
is_size_independent(&self) -> bool { false }` to `true`, which skips
redrawing it when only its offered region changes shape.
```rust
// before
fn draw(&mut self, painter: &mut Painter) { /* ... */ }
fn desired_width(&mut self, ctx: &mut SizeCtx) -> Len { /* ... */ }
fn desired_height(&mut self, ctx: &mut SizeCtx) -> Len { /* ... */ }
// after
fn draw(&mut self, painter: &mut Painter) -> Size { /* ... */ }
```
`SizeCtx` and `Cache` are gone with it — see `LAYOUT.md` for the full
design, the move-offset mechanism this shipped alongside, and the file
list.
## 2026-09-04: texture pipeline rebuilt off the binding array
`Textures`/`TextureHandle`, `GlyphPrimitive`, and `UiRenderNode::new` all
changed shape. Why: the old pipeline bound every texture ever drawn in one
`binding_array<texture_2d<f32>>` and asked every device, unconditionally,
for `VK_EXT_descriptor_indexing` — a real share of Android GPUs lack it,
and it failed outright on the Android emulator's software Vulkan. See
TEXTURES.md's "Recommended shape" and "Implemented, 2026-09-04".
- **`UiRenderNode::new` drops its `limits: UiLimits` parameter, and
`UiLimits` is gone.** Before: `UiRenderNode::new(&device, &queue,
&config, UiLimits::default())`. After: `UiRenderNode::new(&device,
&queue, &config)`. Nothing replaces it — there are no more
binding-array limits to size.
- **`src/default/render.rs`'s device request asks for no features and no
binding-array limits.** Before: `required_features:
Features::TEXTURE_BINDING_ARRAY | Features::PARTIALLY_BOUND_BINDING_ARRAY
| Features::SAMPLED_TEXTURE_AND_STORAGE_BUFFER_ARRAY_NON_UNIFORM_INDEXING`
plus two `max_binding_array_*` limits. After: `Features::empty()` (the
`DeviceDescriptor` default) and only `max_buffer_size` set, which was
never about the binding array.
- **`TextureHandle` has no `primitive()` method any more**; a caller
outside `iris` shouldn't have been calling it (it fed the old renderer's
internals), but if something did: use `image_index()` for a standalone
image's bind-group index. There is no equivalent for a page — a page has
no bind group of its own now, see below.
- **`GlyphPrimitive` has no public constructor from a struct literal.**
Before: `GlyphPrimitive { uv_min, uv_max, view_idx, sampler_idx, color,
flags }`. After: `GlyphPrimitive::new(uv_min, uv_max, layer, color,
flags)` — one `layer` (the shared atlas array's layer) instead of a
`view_idx`/`sampler_idx` pair, since a page is now a layer of one array
texture rather than its own bound texture.
- **A widget author drawing images is unaffected**: `Painter::texture`/
`texture_at`/`texture_within` and `Textures::add` keep their signatures.
What changed underneath is that each standalone image now gets its own
`wgpu::BindGroup` and draw call instead of a slot in the shared array —
invisible from the widget API, visible only in `UiRenderNode`'s internals
and in `iris`'s device requirements.
## 2026-09-05: `FrameReport` splits each frame at `queue.submit`
`FrameStats` gains two fields, and `FrameReport` gains a second recording
method, to answer "is a slow frame iris's own CPU work or the driver/GPU"
with a number instead of a guess (RUST.md's I5 box).
- **`FrameReport::record_split(total, submit_to_present)`** is a second way
to record a frame, alongside the existing `record(total)` (unchanged,
and still what a caller with no split should use — it now reads as
`cpu_p50 == total`, `gpu_wait_p50 == 0`, rather than fabricating a
number for a half it never measured).
- **`FrameStats` gains `cpu_p50` and `gpu_wait_p50`**: medians of
redraw-start-to-submit and submit-to-after-`present()` respectively,
independent of each other and of the existing `p50`/`p90`/`p99`/`worst`
(which are unchanged, and still over the whole frame). The Android
renderer's `draw()` now returns the `submit_to_present` `Duration` it
measured, which `android::view::render()` passes to `record_split`.
- **Caveat carried in both doc comments**: `submit_to_present` is not
fenced against the GPU actually finishing — it is "how long the CPU was
blocked handing the frame to the driver," not a confirmed GPU-completion
time. Enough to separate "iris is slow building the frame" from "iris is
slow handing it off," not enough to claim an exact GPU budget.
+387
View File
@@ -0,0 +1,387 @@
# iris: known problems and things still to build
Iris's own list for the library, recorded 2026-09-04 in her words where it
matters, so the agents working through RUST.md pick these up in a sensible
order rather than rediscovering them. Each item says where it sits in the
order and what "done" looks like. Tick and date them in place.
## Fix
- [x] **`request_device` asked for compute-shader limits it never uses
(2026-09-05).** `Limits::default()` (both `iris/src/android/render.rs`
and `iris/src/default/render.rs`) requests desktop-tier compute limits
unconditionally, even though nothing in `iris`/`iris-core` creates a
`ComputePipeline` or writes a `@compute` shader stage — confirmed by
grepping the whole tree, not assumed. That crashed device creation
outright on the Android emulator's software GL path (`EMU_GPU=software`,
`--features force-gles`): SwiftShader's GL reports itself as OpenGL ES
3.0, which has no compute shaders, so the adapter's real limit is 0
against the unconditional request for 65535 — the same would happen on
any real GLES-3.0-only Android device. Fixed by a new, shared
`iris_core::device_limits()` (`iris/core/src/render/mod.rs`) that zeros
exactly the six `max_compute_*` fields rather than switching to a
downlevel `Limits` preset — `downlevel_webgl2_defaults()` also zeros
`max_storage_buffers_per_shader_stage`, which `shader.wgsl`'s vertex
stage needs (four `var<storage>` buffers), so that preset would trade
this crash for a bind-group-layout one on the same hardware.
`rigs/gpu-probe`'s own hand-mirrored `Limits` (it is deliberately its
own crate, not able to call `device_limits()` directly) was updated to
match. See `DECISIONS.md` and RUST.md's I5 box for the account,
including what could not be re-verified on-device this pass (the
emulator was in concurrent use by another session).
- [x] **Input does not fall through by input type (2026-09-04).**
`SensorUi::run_sensors` (`src/default/sense.rs`) used to set "consumed,
stop checking lower layers" from mere hover — a widget registered for
nothing but `click()` blocked a `Scroll` meant for whatever was behind
it, since "the cursor is over this widget" and "this widget handled the
event" were the same check. Fixed by judging consumption per input
kind: with no button transition and no scroll happening this frame
("momentary" activity), the topmost hovered widget still wins, same as
before; when something momentary *is* happening, only a widget whose
registered senses actually include a matching non-hover one (checked
via a new `TypeEventManager::registered`, which lists what a widget
registered without running anything) consumes it, so a widget with only
`Hovering`/click handlers can no longer block a scroll from reaching a
list underneath. `iris/src/sense_tests.rs` builds a button-over-a-list
`Stack` with a plain `HasEvents` impl (no GPU or window) and checks both
directions: a scroll over the button reaches the list, and a real click
still reaches the button — confirmed to fail on the pre-fix code and
pass after.
- [x] **Appending one image to an already-loaded list rebuilds every other
image's bind group (2026-09-05, fixed 2026-09-05).** Found by the
benchmark below: `GpuTextures::update` (`core/src/render/texture.rs`)
triggered `rebuild_image_bind_groups` — a loop over *every live
standalone image*, rebuilding its `BindGroup` — whenever the shared
`masks` or `move_offsets` GPU buffer was resized (`masks_resized ||
moves_resized` in `UiRenderNode::update`, `core/src/render/mod.rs`), and
a widget getting its *first* move-offset slot (LAYOUT.md section 2 —
every widget gets one on first draw) could be exactly what grows that
buffer. So one new message with one new image, appended to a transcript
that already has N images loaded, did not cost O(1): it cost one
`create_image` for the new image plus one `make_image_bind_group` per
*existing* image, because the new widget's own move slot pushed the
arena past its capacity. Measured directly in
`iris/examples/bench_images.rs`: appending a 1,001st image to 1,000
already-settled ones reported **1,001** bind-group creates for that one
frame, not 1 (`./run-bench.sh images`, frame 5 in the transcript below).
**Fix**: `masks`/`move_offsets` never belonged in a standalone image's own
bind group (group 2) in the first place — the group also holds that
image's own texture view, which is the only thing that is genuinely
per-image, so a buffer shared by *everything* forced a rebuild of
*every* group the moment it moved. Gave masks/move_offsets their own
bind group (group 3 in `shader.wgsl` and `UiRenderNode`: `masks_layout`/
`masks_group`), bound once per frame in `UiRenderNode::draw` rather than
once per draw call, instead of duplicating them into every per-image
group. `GpuTextures` and its image bind groups now know nothing about
either buffer — `rebuild_image_bind_groups` is called only from
`grow_array` (the atlas array texture growing, which genuinely does
change what every image's own bind group must reference) — so a
masks/move_offsets resize now touches exactly one bind group, ever,
regardless of how many images are live. Numbers after the fix, same
benchmark and command:
./run-bench.sh images
frame=1 bind_group_creates=1000 (cold load, unchanged)
frame=2 bind_group_creates=0 (was 1000 -- see the item below)
frame=3 bind_group_creates=0
frame=4 bind_group_creates=0
(append one image here)
frame=5 bind_group_creates=1 (was 1001)
frame=6 bind_group_creates=0
`run-headless.sh tabs --shot` still 27266 bytes, byte-for-byte unchanged,
confirming the bind-group restructuring changed nothing about what is
drawn.
- [x] **Bind-group creation takes two frames to reach the steady state, not
one (2026-09-05, closed by the fix above, 2026-09-05).** Same benchmark:
loading 1,000 images cold used to report 1,000 creates on frame 1
(expected — `create_image`, one per new image) *and again* 1,000 on
frame 2, before settling to 0 from frame 3. This was `rebuild_image_bind_groups`
firing a second time for the same masks/move-offsets buffer-growth
reason as the item above, confirming the guess recorded here — the two
were exactly the same root cause measured two different ways. Frame 2
now reports 0 (see the numbers above); not a separate fix.
- [ ] **A read-only text display has no widget of its own — P0's bench
report area is a `TextEdit` standing in for one (2026-09-05).** The only
way to get selectable text on screen today is `.editable(...)` plus
`.attr::<Selectable>(())` (`Selectable` is only implemented for
`TextEdit`, `iris/src/attr.rs`), which also makes the field focusable —
tapping the bench report opens the soft keyboard over text nothing lets
you type into. Harmless for a bench-only debug screen (not fixed this
pass), but a real "selectable, not editable" text primitive would
remove the keyboard side effect and is worth having before another
screen wants the same thing (P1's own transcript rows already read
their content from a `TextEdit` for the same reason).
## Build
- [x] **Benchmarks**, not unit tests, run on demand (2026-09-05; a
`benches/` or a script under `iris/`, never in `cargo test`). The
scenario that matters most is a **message list** — chat apps and this
app's transcript alike — stressed with many messages and many images.
One case in particular: **resizing an input box** (typing enough text to
grow it) that pushes a long list of messages above it must stay very
fast and recalculate almost nothing — a move of everything above, not a
re-layout. That is exactly the O(1) move chain in LAYOUT.md; the
benchmark is what proves it. Done when the numbers are in this file with
the command, and the input-box case reports draws re-run, not just frame
time.
**Built as two rigs**, chosen per scenario by whether a real `wgpu`
device is needed (`UiRenderState`/`Widgets` touch no GPU or window, so
most of this runs as an ordinary binary — the same property
`layout_tests.rs` relies on):
- `iris/benches/message_list.rs` — a plain `Instant`-timed binary
(`[[bench]] harness = false` in `iris/Cargo.toml`), not criterion: see
the file's own header for why (short version — every scenario here
reduces to a *count* `UiRenderState::take_counters` already produces,
which criterion's statistical machinery adds nothing to and which a
new dependency is not worth pulling in for). Covers (a) first-frame
cost of a message list of N wrapped-text rows (one in 20 also carrying
a small in-memory image) for N = 100/1,000/10,000; (b) per-frame cost
of scrolling that list, 200 ticks; (c) the input-box case — a
fixed-height field at the bottom of the screen growing by a line 40
times, with the message list above it filling the rest of the screen.
Run: `cd iris && cargo bench --bench message_list` (always release —
`cargo bench` builds the `bench` profile, which is optimized).
- `iris/examples/bench_images.rs` — needs a real device, so it runs
through `iris/run-headless.sh bench_images`, printing
`UiRenderNode::take_image_bind_group_creates()` (a new counter, added
in `core/src/render/texture.rs` and `core/src/render/mod.rs`,
mirroring `UiRenderState::take_counters`) each frame. Covers (d): 1,000
image rows, checked both cold (does bind-group creation reach zero
once loaded) and after appending one more image once settled (does
*that* stay cheap) — the second question is what actually matters for
a live transcript and is what turned up the two Fix items above.
- `iris/run-bench.sh [list|images]` runs either or both and is what to
run before/after touching `Scroll`, `Span`, `Sized`, the move-offset
chain, or `GpuTextures`.
**Numbers (2026-09-05, release, `cargo bench`/`run-headless.sh`, this
VM: AMD Ryzen 7 3800X, 8 cores, rustc 1.98.0 nightly-2026-09-03):**
cd iris && cargo bench --bench message_list
(a) first frame, N=100: 30.30ms draws=227 rewrites=15 moves=0
(a) first frame, N=1000: 186.04ms draws=2252 rewrites=150 moves=0
(a) first frame, N=10000:1770.36ms draws=22502 rewrites=1500 moves=0
(b) scroll, N=100/1000/10000, 200 ticks each:
draws=200 rewrites=0 moves=200 (identical at every N)
per-tick average: 0.0002ms (identical at every N)
(c) input grows 40 lines, N=100/1000/10000 rows above it:
draws=320 rewrites=40 moves=160 (identical at every N)
per-line average: 0.0012-0.0013ms (identical at every N)
cd iris && ./run-bench.sh images (2026-09-05, before the fix)
frame=1 bind_group_creates=1000 (cold load)
frame=2 bind_group_creates=1000 (see Fix item above)
frame=3 bind_group_creates=0
frame=4 bind_group_creates=0
(append one image here)
frame=5 bind_group_creates=1001 (see Fix item above)
frame=6 bind_group_creates=0
cd iris && ./run-bench.sh images (2026-09-05, after the fix)
frame=1 bind_group_creates=1000 (cold load, unchanged -- genuine work)
frame=2 bind_group_creates=0
frame=3 bind_group_creates=0
frame=4 bind_group_creates=0
(append one image here)
frame=5 bind_group_creates=1 (one image's own create_image, O(1))
frame=6 bind_group_creates=0
**Reading it**: (a) is real, necessary work — shaping and laying out N
never-before-seen text rows — and scales with N as it must, ~10x cost
per 10x N. (b) and (c) are the pass conditions that matter: both are
**exactly flat across N = 100 to 10,000**, confirming LAYOUT.md's O(1)
move chain holds for both scrolling and for a growing input box pushing
the message list — draws/moves per tick or per line do not grow with
list size, and the per-operation cost (a fraction of a microsecond) is
nowhere near a frame budget. (d)'s cold-load and steady-state halves
behave as designed; its *append* half did not, until the fix above moved
masks/move_offsets out of the per-image bind group — now flat at O(1)
the same way (b) and (c) are.
- **I5's transcript screen (`iris/transcript-ui/`, 2026-09-05) — what it
left, each recorded at the point in the code it would go rather than
silently dropped. See RUST.md's I5 box for the full account of what
*was* built (the screen, `SpanStyle`, cross-row selection, the growing
composer).**
- [x] **Android integration for this screen — done, 2026-09-05.**
`iris-android-app`'s `transcript-screen` Cargo feature
(`transcript_client.rs`) runs this screen against a real `ai-server`
through `client-core`, confirmed on-device: real scrolling, real
touch-drag panning, tap-by-name on the composer. Two real bugs found
and fixed along the way (a missing `INTERNET` permission; a
background-thread redraw request that crashed via a `Looper`
requirement, fixed by routing through `View::post_delayed` — see
`IRIS.md`'s `Tasks::redraw_handle` entry). See RUST.md's I5 box,
"The Android integration, done 2026-09-05" for the full account.
- [x] **A render-time number for iris, comparable to Compose's
`transcript-bench.sh` report — instrumentation done and a real number
obtained, 2026-09-05 (later the same day); the clean comparable loop
is not.** `iris_core::FrameReport` (`iris/core/src/render/
frame_report.rs`, `IRIS.md`'s new entry) times every frame from
`render()`'s redraw start to after `queue.submit`+`present()`, exposed
as two named on-screen controls ("Frame report", "Reset frame
report"). Driven against a real on-device touch-drag it read
`frames=34 janky%=61.76 p50=26.5ms p90=48.0ms p99=98.1ms
worst=98.1ms` — real, not inferred, but accumulated across several
gestures rather than one clean 24-swipe loop, because of the new
finding below. See RUST.md's I5 box, "Update, 2026-09-05, later the
same day" for the full account.
- [ ] **New, 2026-09-05: intermittent touch delivery to iris's
`SurfaceView` under this checkout's `EMU_GPU=software` emulator.**
The same swipe coordinates, confirmed (by scanning a screenshot
column for the first non-black pixel) to sit over real row text,
sometimes produced 30+ real frames and a screenshot diff and
sometimes produced zero of either, across otherwise-identical
`ui-trace` invocations. Not the already-understood "already at that
scroll edge" case (reproduced with content confirmed taller than the
viewport, in both directions). Leading candidate, not yet confirmed:
this checkout's emulator was independently observed at ~78% of one
CPU core, continuously, while idle on-screen — SwiftShader's software
rasterisation is CPU-bound by design, and a synthetic touch competing
with that load for delivery is plausible but unmeasured *during* a
failing gesture (the standing rule against diagnosing from
after-the-fact measurements applies here). Needs a sampler (load,
`dumpsys input`, a `-i 0` `ui-trace` capture) running while a failing
gesture is driven, and ideally a comparison under `-gpu host` (real
Vulkan) to see whether it is specific to software rendering. This is
what blocks the clean, comparable 24-swipe loop above.
- [x] **Long-press-then-drag-to-select — confirmed on-device, 2026-09-05
(later the same day).** `ui-trace` gained a `holddrag X1 Y1 X2 Y2
HOLD_MS MOVE_MS` action (`emulator-tools`, additive, extends the same
`MotionEvent`/`injectInputEvent` mechanism `swipe` already used):
press, hold past `LONG_PRESS`, move, release, as one continuous touch.
Driven against a real row (`holddrag 300 1850 300 2050 600 300`) it
produced `iris selection: begin at row ...` then a sequence of
`iris selection: extend to row ...` log lines
(`transcript-ui/src/selection.rs`, a new small `log` dependency since
selection has no accessibility label of its own yet — see the next
item), and a screenshot taken right after shows the expected
highlighted selection spanning multiple rows. `DragArbiter`'s own
unit tests already covered this sequence against a synthetic clock;
this is the first time it has been driven by a real device touch.
- [x] **Touch-drag panning over a row's own rendered text — done,
2026-09-05.** `row.rs` used to register `CursorSense::click_or_drag()`
on each row's `TextEdit` for cross-row selection; `TextEdit::draw`'s
`painter.child_layer()` (`iris/src/widget/text/edit.rs:87`) meant that
registration won `core/src/sense.rs::run_sensors`'s per-layer
arbitration on every frame it was pressed, not just the frame the
press started, so a list pan gesture registered on `List` itself never
got a turn while a row was under the finger. Fixed with
`iris::sense::DragArbiter` (recorded in `IRIS.md`), one small state
machine per list deciding pan vs. select the way Android does (a
vertical drag pans immediately; a stationary press held `LONG_PRESS`
(500ms) starts a selection which further drag extends; a horizontal
drag while something is already selected extends immediately) —
`transcript-ui/src/selection.rs`'s `Selection::drag` is the one place
every row's drag now routes through. 8 new unit tests
(`iris/src/sense.rs`'s `drag_arbiter_tests`); `cargo fmt/clippy/test
--workspace` and `cargo ndk` (both `iris` and `transcript-ui`) all
clean; `run-headless.sh` screenshot byte-identical to before the
change (38578 bytes). See RUST.md's I5 box, "Gap closed, 2026-09-05".
- [x] **Intermittent touch-scroll dropout — root-caused and fixed,
2026-09-05.** Not the coalesced-`ACTION_MOVE` hypothesis the earlier
pass suspected (ruled out): a gesture's `ACTION_DOWN` can land on a
row's own padding/gap or its header, which no `CursorSense` covers,
so `DragArbiter` never gets `press_start` and sits in `Idle`
(answers `Undecided` forever) for that whole gesture. Fixed via a new
`DragArbiter::is_idle()` that `Selection::drag`
(`transcript-ui/src/selection.rs`) checks to recover a missed press
on the next `Pressing` frame. Four new unit tests. See RUST.md's I5
box, "Touch-scroll dropout root-caused, 2026-09-05", for the trace and
what a peer session sharing this checkout's emulator mid-pass
prevented from being re-verified end-to-end (the aggregate
`iris-scroll.sh` three-run confirmation and a re-taken FrameReport
row) — a future pass should finish that once the emulator is free.
- [ ] **Row-level accessibility names.** The composer carries
`.label("Message")`; transcript rows do not carry a `.label()` of
their own yet, so `Widgets::named()` (I4) does not include them —
`row.rs`'s `build_text_row` is where one would go, keyed to something
stable per row (its sender + a short excerpt, matching what a screen
reader announcing a chat message would say).
- [ ] **A tappable link and a background chip behind inline code.**
Both need per-range glyph geometry that `TextEditCtx` does not expose
outside `iris::widget::text` (`edit.rs`'s `layout()` helper is
private) — see `markdown.rs`'s module doc for the exact shape the fix
would take (the same primitive `TextEdit::draw`'s own selection
highlight already uses internally,
`iris/src/widget/text/edit.rs:99`).
- [ ] **`Selection`'s anchor-row shortcut.** The row a drag started in
is selected in full (`select_all`) the moment the drag leaves it,
rather than "from the click point to whichever edge points away from
the drag" — needs the same private `layout()` access as the item
above. `selection.rs`'s module doc has the exact reasoning.
- [ ] **No syntax highlighting inside a fenced code block.**
`client_core::highlight` exists (built for the file explorer) and
could feed per-token `SpanStyle`s into a code block's span; wiring it
in was not attempted this pass.
- [ ] **Masks defined relative to each other.** Wanted: mask A multiplies
by something *and also* applies mask B — a mask can reference a parent
mask, the way the move chain references a parent offset. Today masks
are independent regions. Design it beside the move chain (same shape:
a parent index and a bounded walk in the shader); do it when a real
widget needs it, not before.
- [ ] **Positions as a single float per scroll.** Iris raised, and half
rejected, letting a scroll update one float rather than positions:
input handling cares about most elements in a list, so absolute
positions must be computed on the CPU anyway. LAYOUT.md's design
already lands here (GPU walks the chain, CPU resolves on demand for
hit tests). Keep the CPU resolution lazy and per query; do not
materialise every row's absolute position per frame.
- [ ] **Animations, last.** Cosmetic, so after everything above. Must be
**modular — a piece of the library rather than a core part forced into
everything, the same way input is**. Whatever the mechanism, a widget
that does not animate must pay nothing and import nothing for it.
## Build (for the port)
Widgets `RUST.md`'s "The port, in order (decided 2026-09-05)" needs and
iris does not have yet, one entry per gap, named against the P-step that
first needs it. Move an entry up to "Fix" or tick it in place once built;
do not duplicate it there.
- [ ] **A history-paging cushion measured in on-screen viewports, not a
row count.** (**P1**.) `iris::widget::List` has no equivalent of the
Compose app's `HISTORY_SCREENS` — AGENTS.md's "Things that have
bitten" is explicit that a fixed row count under-fills a screen on a
tool-heavy transcript and over-fills one on a text-heavy one, so
whatever loads the next page has to ask the list how many viewports
are actually on screen, not assume a constant.
- [ ] **A scaled thumbnail/image widget for an in-transcript image.**
(**P1**.) `SessionImage.kt`'s bitmap decode-and-downscale has no iris
counterpart; iris's own image widget (used by `bench_images.rs`) draws
a loaded texture but does nothing about sourcing or scaling one from a
server-produced attachment.
- [ ] **A modal/dialog primitive.** (**P1**, reused by **P3** and
**P5**.) Needed for the session settings dialog, `UsageDialog`'s
equivalent, and the delete-with-`deleteForeign` confirmation with its
toggle switch. Build once, wherever it is first needed, rather than
once per screen that wants one.
- [ ] **A horizontal gauge/bar widget.** (**P1**.) For
`SessionUsageBar`'s equivalent — a bounded fill reflecting a fraction,
nothing fancier.
- [ ] **A `BusyItem` equivalent: a dimmed row carrying an operation
label that does not block its list's own scroll/drag.** (**P3**.) The
Compose version tried an overlay first and it swallowed the drag along
with the tap (AGENTS.md's "Shared appearance") — worth not repeating
that attempt in iris before building the row-level version directly.
- [ ] **A toggle switch.** (**P3**.) For the delete dialog's
`deleteForeign` control; iris has no switch/checkbox widget yet as far
as this pass found.
## Reconsider
- [ ] **`WidgetView`.** Iris is unsure of it: what she wants is an easy way
to compose a widget from others (a button is the main case). With
sizing folded into `draw`, composing may be easy enough that `View` is
redundant. Decide after the layout change lands, by writing a button
both ways and keeping the one that is shorter to explain; delete the
other rather than keeping two ways.
View File
File renamed without changes.
+147 -10
View File
@@ -147,12 +147,37 @@ turn.
Spawn: `claude -p --verbose --input-format stream-json --output-format
stream-json --permission-mode <mode>` in the chosen working directory, plus
`--model`. Wire-format notes are pinned against CLI 2.1.237 in
`session/claude.rs`'s module doc: permissions need the hidden
`--permission-prompt-tool stdio` flag, AskUserQuestion answers ride
`updatedInput.answers` keyed by question text, and `set_model`/`interrupt`
`--model` and, where one has been chosen, `--effort`. Wire-format notes are
pinned against CLI 2.1.237 in `session/claude.rs`'s module doc: permissions
need the hidden `--permission-prompt-tool stdio` flag, AskUserQuestion answers
ride `updatedInput.answers` keyed by question text, and `set_model`/`interrupt`
are control requests.
**The thinking level is settled at launch** (added 2026-09-04, because it is
the largest saving available on a long session: output is about an eighth of
what a session costs and thinking is the bulk of output, against the ~1.5% that
is prose). The CLI's only two setting control requests are `set_model` and
`set_permission_mode` -- checked against the 2.1.258 binary -- so there is no
way to ask a running process to think differently. `set_session_effort` is
therefore shaped like `set_session_cwd` rather than like `set_session_model`:
it records the level and **stops the process**, and the next message or Start
launches one that has it. It lives in the session settings dialog beside the
working directory for that reason, not on the session bar beside the model and
the mode, which do take effect mid-turn. `None` is a level in its own right --
the CLI's own default -- so the picker can return to it; a level this app named
as the default instead would be this app choosing one.
**What a new session starts at is `Config::default_effort`**, applied in
`spawn_session` rather than filled in by the spawn screen, so it holds for an
import and a bare API call as well. It is set by the spawn screen's own
picker, whose label says so: one control, where new sessions are made, rather
than a settings page for a single value. It is not on a provider, because
providers are discovered and the next rediscovery would erase it, and not on
the phone, because a second device would then spawn at a level nobody there
chose. `GET`/`POST /defaults` carry it, as a struct rather than a bare value
so the permission mode -- still hardcoded to `auto` on the spawn screen -- can
move there without a second route.
**`--resume` only ever runs when nothing else has that session open.** That
is the rule behind the import refusal, the single `ClaudeDriver::launch`
entry point, and the `Exited` correction below; two CLIs on one session file
@@ -169,10 +194,33 @@ deliberate and easy to undo by accident:
when the process restarts. That leaves the Claude driver as the odd one
out rather than this one — the CLI's memory is a cache in front of the same
transcript. Resolve any inconsistency in this direction.
- **A llama session on an ssh host is refused.** The model is reached over
HTTP and forwarding that port is not built, so refusing beats silently
talking to the wrong machine. A transport is "run this" plus "reach this
port", and only the first half exists.
- **A llama session runs on whatever machine its setup names** (2026-09-04,
the last of phase 5). A transport is "run this" plus "reach this port", and
the second half is `Transport::reserve_port` — the port the server binds
*there* and the port that reaches it *here*, the same number locally —
carried by `Launch::reaching` onto the connection that already runs the
command. `llama-server` binds loopback on the far machine, so nothing is
served to its network. The far port is a guess from a range below the
ephemeral one, because no portable way to ask a machine for a free port
avoids racing the bind anyway; a collision is not silent, since the server
fails to bind and the readiness poll reports what its log said.
- **The model file lives on the machine that serves it** (2026-09-04). Each
setup names its own models directory (`SshConfig::models_dir`, default
`~/.local/share/ai-app/models` expanded *there*), and a spawn resolves the
key on that machine — one round trip answering "at /abs/path" or "missing",
so a model that is not there is refused at the spawn rather than becoming a
server that never becomes ready. The spawn screen offers
`GET /setups/{id}/models`, that machine's list, rather than `GET /models`,
which is this backend's downloads. Downloading *to* another machine is
deliberately not built: a multi-gigabyte transfer with no progress
anywhere, and the file gets there however anything else on that machine
did.
- **The readiness poll watches the process, not only the port.** A model that
will not load, a port already taken, a flag an older build does not know:
all exit within a second and none will ever answer `/health`, so waiting
out the 300s timeout turned the server's own account of the problem into
"gave up". The failure carries the tail of `llama-server.log`, which on a
remote session is the only copy anybody reading the phone can see.
### Models (2026-08-28)
@@ -203,6 +251,13 @@ deliberate and easy to undo by accident:
A driver says what to run; something above it turns that into a process.
Otherwise transport knowledge sits inside a translator whose job is a wire
format, and every future driver has to remember to do the same.
- **A forwarded launch gets a pty and every other one does not** (measured
2026-09-04). Killing the ssh client ends a CLI because it closes the stdin
that CLI is reading; `llama-server` never reads its stdin, so the same kill
left it running on the far machine with the model loaded — one orphan per
stopped session. With `-tt` the far side takes SIGHUP when the connection
goes. Its log then arrives through a line discipline, which nothing parses.
`-T` stays everywhere else, where a pty would rewrite the JSONL.
- **`command -v` follows ssh's non-login PATH**, which is narrower than an
interactive shell's, so a binary somewhere unusual is invisible to
discovery. Point `command` at an absolute path.
@@ -499,6 +554,25 @@ rate-limited bucket). Poll at ≥180 s, only while a Claude session exists or
the usage screen is open, and cache the last answer. It is undocumented, so
`usage.rs` treats every field as optional and degrades rather than erroring.
**Per provider, not per machine (2026-09-04).** A machine is not what is
metered; the provider a session runs is. One machine offers echo, the Claude
CLI and a local model side by side, and only the second spends anything — so
pairing a session with a snapshot by machine alone drew the CLI's five-hour
window under every echo session on it, a quota that session cannot spend. A
session now names its meter (`usageProvider`, from
`DriverKind::usage_provider`, which `usage::providers_for` reads too, so the
two lists cannot disagree) and `GET /usage` is matched on machine *and*
provider. `None` is a session that meters nothing, and the phone draws
nothing at all for it — not a zero, and not "unknown".
`DriverKind::Echo` names a meter of its own that exists only when a test has
asked for one: `/usage` in an echo session sets an invented answer
(`usage::Fixture`), and with none set there is no snapshot and no bar. That
is what makes those screens' states reachable — a number near the top, a
window between blocks with no reset time, a machine nobody logged into, one
that could not be reached — without spending real quota to arrange them,
which is why none of them had ever been looked at.
**Per machine, not per backend (2026-08-29).** The credential store that
matters is the one on the machine the session runs on, because that is the
account being billed — and in the layout this aims at, `ai-server` is on the
@@ -522,6 +596,68 @@ always running. So absent means **not running**, and only a timestamp that
arrives and cannot be parsed is unknown. `WindowEnd` in `ResetCountdown.kt`
is the one rule both readers go through.
### Auto-resume (2026-09-05)
**A session may pick itself back up when the account's usage limit lifts.**
Off unless somebody switched that session to it, because it spends quota the
moment quota exists and does so with nobody looking — that is not a thing a
default may decide. It sends one message, `continue` unless another was
typed, and then it is done; there is no retry loop around the conversation
itself.
**Running out of quota is a state, not an error.** `Event::LimitReached`
carries the dialect's reset time where it gave one, and recognising it
belongs to the driver — the Claude CLI ends the turn with `is_error` and
`Claude AI usage limit reached|1788546972`, and nothing above the driver
matches on a string. The transcript draws it as a divider, like a clear or a
compaction: what a reader scrolling back wants from it is why the
conversation stops at that line.
**The schedule is a plan to ask, never a plan to send.** Every reset time
available here is untrustworthy in the direction that matters: the dialect's
is written when the turn fails, and the endpoint's moves when the window
does. So the wait ends in a question to `usage.rs`, and only `ok` with no
window at 100% sends anything. A window still spent reschedules to *its own*
reset time — which is what makes a limit that lifts later than promised wait
longer, and one that lifts sooner resume sooner. A meter that cannot be
asked at all is a longer wait too, never a send: "we could not find out"
must not be able to produce the same action as "there is room".
Bounded, because something has to be: a day after the limit was hit the wait
stops and says so in the session's own transcript. A machine that can never
be asked would otherwise be retried for ever with nothing on screen saying
so.
The schedule is persisted on the session (`resume: Some(ScheduledResume)`),
not held in memory: a five-hour window routinely outlasts a backend restart,
and a wait forgotten across one is a session that silently never comes back.
`resume.rs` is the top layer — it holds the manager and the monitor and
neither holds it — which is what lets the decision be a pure function of a
snapshot and a clock. The pump reports limits downward on a broadcast, for
the reason `Shared` exists: the pump runs underneath the manager.
**Exercised with echo, never with a real account.** `/limit [minutes]` in an
echo session reports the same event a real driver does, and `/usage` sets
what the meter answers — deliberately two commands, because the two
disagreeing is the state the whole design is about. The loop was driven end
to end that way on 2026-09-05: the wait moved from the dialect's two minutes
to the meter's seven when the meter changed its mind, and the message went
out on the first check after the meter came back under the limit.
### Subagents (2026-09-05)
**A subagent is a second transcript owned by a session, in the same event
model, with no process and no controls of its own.** Full design and wire
shape in `SUBAGENTS.md`, kept separate because the app half is being built
against it in parallel and it is the shared contract between the two. The
one-paragraph reason: a session's Task-tool helpers already speak the common
event model on the parent's own stdout (each line carrying
`parent_tool_use_id`), so giving each one its own small transcript — same
file format, same paging routes, same SSE stream, reused by addressing rather
than by copying — costs a routing step in the translator and a registry
(`session/subagent.rs`) rather than a second session type with a driver, a
process and a config entry it does not need.
### HTTP surface
**`routes.rs`'s module doc comment is the table.** REST for actions, one SSE
@@ -800,8 +936,9 @@ Noticed and deliberately not fixed, so they are not re-found from scratch.
Phases 13 (the skeleton pipe, the full Claude driver, the usage screen) done
2026-08-24. Phase 4 (llama.cpp: model browsing, downloads, and `llama-server`
through its OpenAI-compatible endpoint) and phase 5 (ssh) done 2026-08-28.
The file explorer and the transcript cache followed in September. What is
through its OpenAI-compatible endpoint) and phase 5 (ssh) done 2026-08-28,
except for the remote `llama-server` and its port forward, which landed
2026-09-04. The file explorer and the transcript cache followed in September. What is
left is real-phone/WireGuard bring-up, which is operational rather than code.
Each phase ended runnable and verified against the real thing. The backend
+1832 -32
View File
File diff suppressed because it is too large. Load diff
View File
File renamed without changes.
View File
File renamed without changes.
File renamed without changes.
+377 -4
View File
@@ -151,7 +151,7 @@ source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "5a15f179cd60c4584b8a8c596927aadc462e27f2ca70c04e0071964a73ba7a75"
dependencies = [
"cfg-if",
"getrandom",
"getrandom 0.3.4",
"once_cell",
"version_check",
"zerocopy",
@@ -534,6 +534,12 @@ dependencies = [
"arrayvec",
]
[[package]]
name = "base64"
version = "0.23.1"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "ac07cdecf99051d9a5238b80f35af32cdeba5b336e55d957b318b50137e18da5"
[[package]]
name = "bit-set"
version = "0.8.0"
@@ -710,6 +716,16 @@ version = "0.2.1"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "613afe47fcd5fac7ccf1db93babcb082c5994d996f20b8b159f2ad1658eb5724"
[[package]]
name = "client-core"
version = "0.1.0"
dependencies = [
"event-model",
"serde",
"serde_json",
"ureq",
]
[[package]]
name = "clipboard-win"
version = "5.4.1"
@@ -755,6 +771,35 @@ dependencies = [
"crossbeam-utils",
]
[[package]]
name = "cookie"
version = "0.18.2"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "1a373e3602691c3cdea496d2f0ee5935151e6168fe87739483c463db1b2f2f87"
dependencies = [
"percent-encoding",
"time",
"version_check",
]
[[package]]
name = "cookie_store"
version = "0.22.1"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "15b2c103cf610ec6cae3da84a766285b42fd16aad564758459e6ecf128c75206"
dependencies = [
"cookie",
"document-features",
"idna",
"indexmap",
"log",
"serde",
"serde_derive",
"serde_json",
"time",
"url",
]
[[package]]
name = "core-foundation"
version = "0.9.4"
@@ -871,6 +916,25 @@ version = "1.2.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "f27ae1dd37df86211c42e150270f82743308803d90a6f6e6651cd730d5e1732f"
[[package]]
name = "deranged"
version = "0.5.8"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "7cd812cc2bc1d69d4764bd80df88b4317eaef9e773c75226407d9bc0876b211c"
[[package]]
name = "desktop-app"
version = "0.1.0"
dependencies = [
"client-core",
"event-model",
"iris",
"serde_json",
"tempfile",
"transcript-ui",
"winit",
]
[[package]]
name = "dispatch"
version = "0.2.0"
@@ -1023,6 +1087,14 @@ dependencies = [
"pin-project-lite",
]
[[package]]
name = "event-model"
version = "0.1.0"
dependencies = [
"serde",
"serde_json",
]
[[package]]
name = "exr"
version = "1.74.0"
@@ -1165,6 +1237,15 @@ version = "0.3.1"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "aa9a19cbb55df58761df49b23516a86d432839add4af60fc256da840f66ed35b"
[[package]]
name = "form_urlencoded"
version = "1.2.2"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "cb4cb245038516f5f85277875cdaa4f7d2c9a0fa0468de06ed190163b1581fcf"
dependencies = [
"percent-encoding",
]
[[package]]
name = "futures-core"
version = "0.3.34"
@@ -1239,6 +1320,26 @@ dependencies = [
"windows-link",
]
[[package]]
name = "getopts"
version = "0.2.24"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "cfe4fbac503b8d1f88e6676011885f34b7174f46e59956bba534ba83abded4df"
dependencies = [
"unicode-width",
]
[[package]]
name = "getrandom"
version = "0.2.17"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "ff2abc00be7fca6ebc474524697ae276ad847ad0a6b3faa4bcb027e9a4614ad0"
dependencies = [
"cfg-if",
"libc",
"wasi",
]
[[package]]
name = "getrandom"
version = "0.3.4"
@@ -1398,6 +1499,22 @@ version = "0.2.1"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "dfa686283ad6dd069f105e5ab091b04c62850d3e4cf5d67debad1933f55023df"
[[package]]
name = "http"
version = "1.5.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "918d3568bebf352712bc2ef3d46a8bcf1a75b373be6539de198e9105cbbf9ce0"
dependencies = [
"bytes",
"itoa",
]
[[package]]
name = "httparse"
version = "1.10.1"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "6dbf3de79e51f3d586ab4cb9d5c3e2c14aa28ed23d180cf89b4df0454a69cc87"
[[package]]
name = "icu_collections"
version = "2.3.0"
@@ -1526,6 +1643,27 @@ version = "2.3.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "ae293c039020f9ec10710af98d29ce6aa2051486638b49c9a6409f3b4a9e98ad"
[[package]]
name = "idna"
version = "1.1.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "3b0875f23caa03898994f6ddc501886a45c7d3d62d04d2d90788d47be1b1e4de"
dependencies = [
"idna_adapter",
"smallvec",
"utf8_iter",
]
[[package]]
name = "idna_adapter"
version = "1.2.2"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "cb68373c0d6620ef8105e855e7745e18b0d00d3bdb07fb532e434244cdb9a714"
dependencies = [
"icu_normalizer",
"icu_properties",
]
[[package]]
name = "image"
version = "0.25.9"
@@ -1641,6 +1779,12 @@ dependencies = [
"either",
]
[[package]]
name = "itoa"
version = "1.0.18"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "8f42a60cbdf9a97f5d2305f08a87dc4e09308d1276d28c869c684d7777685682"
[[package]]
name = "jni"
version = "0.21.1"
@@ -1669,7 +1813,7 @@ version = "0.1.34"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "9afb3de4395d6b3e67a780b6de64b51c978ecf11cb9a462c66be7d4ca9039d33"
dependencies = [
"getrandom",
"getrandom 0.3.4",
"libc",
]
@@ -1978,6 +2122,12 @@ dependencies = [
"num-traits",
]
[[package]]
name = "num-conv"
version = "0.2.2"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "521739c6d2bac4aa25192232afe6841231376b2b26d4d9fae5ecf8ca5772e441"
[[package]]
name = "num-derive"
version = "0.4.2"
@@ -2620,6 +2770,12 @@ dependencies = [
"zerovec",
]
[[package]]
name = "powerfmt"
version = "0.2.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "439ee305def115ba05938db6eb1644ff94165c5ab5e9420d1c1bcedbba909391"
[[package]]
name = "ppv-lite86"
version = "0.2.21"
@@ -2672,6 +2828,25 @@ dependencies = [
"syn 2.0.113",
]
[[package]]
name = "pulldown-cmark"
version = "0.13.4"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "e9f068eba8e7071c5f9511831b44f32c740d5adf574e990f946ddb53db2f314e"
dependencies = [
"bitflags 2.10.0",
"getopts",
"memchr",
"pulldown-cmark-escape",
"unicase",
]
[[package]]
name = "pulldown-cmark-escape"
version = "0.11.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "007d8adb5ddab6f8e3f491ac63566a7d5002cc7ed73901f72057943fa71ae1ae"
[[package]]
name = "pxfm"
version = "0.1.27"
@@ -2746,7 +2921,7 @@ version = "0.9.3"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "99d9a13982dcf210057a8a78572b2217b667c3beacbf3a0d8b454f6f82837d38"
dependencies = [
"getrandom",
"getrandom 0.3.4",
]
[[package]]
@@ -2881,6 +3056,20 @@ version = "0.8.52"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "0c6a884d2998352bb4daf0183589aec883f16a6da1f4dde84d8e2e9a5409a1ce"
[[package]]
name = "ring"
version = "0.17.14"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "a4689e6c2294d81e88dc6261c768b63bc4fcdb852be6d1352498b114f61383b7"
dependencies = [
"cc",
"cfg-if",
"getrandom 0.2.17",
"libc",
"untrusted",
"windows-sys 0.52.0",
]
[[package]]
name = "roxmltree"
version = "0.21.1"
@@ -2922,6 +3111,41 @@ dependencies = [
"windows-sys 0.61.2",
]
[[package]]
name = "rustls"
version = "0.23.43"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "0283386ce02abc0151e1761d08802dfe86c173b0b494af5cbc086574e453da06"
dependencies = [
"log",
"once_cell",
"ring",
"rustls-pki-types",
"rustls-webpki",
"subtle",
"zeroize",
]
[[package]]
name = "rustls-pki-types"
version = "1.15.1"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "2f4925028c7eb5d1fcdaf196971378ed9d2c1c4efc7dc5d011256f76c99c0a96"
dependencies = [
"zeroize",
]
[[package]]
name = "rustls-webpki"
version = "0.103.15"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "f3c3cf1d8b1e7d4927e2d154c3fcb02979afb9939629c62cd9048d4f07b60ac2"
dependencies = [
"ring",
"rustls-pki-types",
"untrusted",
]
[[package]]
name = "rustversion"
version = "1.0.22"
@@ -2998,6 +3222,19 @@ dependencies = [
"syn 2.0.113",
]
[[package]]
name = "serde_json"
version = "1.0.151"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "c841b55ecdae098c80dcae9cf767f6f8a0c2cdb3416bbef72181df4d0fe73f14"
dependencies = [
"itoa",
"memchr",
"serde",
"serde_core",
"zmij",
]
[[package]]
name = "serde_repr"
version = "0.1.21"
@@ -3138,6 +3375,12 @@ version = "0.1.1"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "6637bab7722d379c8b41ba849228d680cc12d0a45ba1fa2b48f2a30577a06731"
[[package]]
name = "subtle"
version = "2.6.1"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "13c2bddecc57b384dee18652358fb23172facb8a2c51ccc10d74c157bdea3292"
[[package]]
name = "swash"
version = "0.2.10"
@@ -3196,7 +3439,7 @@ source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "0136791f7c95b1f6dd99f9cc786b91bb81c3800b639b3478e561ddb7be95e5f1"
dependencies = [
"fastrand",
"getrandom",
"getrandom 0.3.4",
"once_cell",
"rustix 1.1.3",
"windows-sys 0.61.2",
@@ -3265,6 +3508,36 @@ dependencies = [
"zune-jpeg 0.4.21",
]
[[package]]
name = "time"
version = "0.3.55"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "cdb87b95ec50ddfa440816d227a17b2ccbdda963a316a727fda0fc4334f7d134"
dependencies = [
"deranged",
"num-conv",
"powerfmt",
"serde_core",
"time-core",
"time-macros",
]
[[package]]
name = "time-core"
version = "0.1.9"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "9e1c906769ad99c88eaa54e728060edef082f8e358ff32030cb7c7d315e81109"
[[package]]
name = "time-macros"
version = "0.2.32"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "7e689342a48d2ea927c87ea50cabf8594854bf940e9310208848d680d668ed85"
dependencies = [
"num-conv",
"time-core",
]
[[package]]
name = "tiny-skia"
version = "0.11.4"
@@ -3371,6 +3644,17 @@ dependencies = [
"once_cell",
]
[[package]]
name = "transcript-ui"
version = "0.1.0"
dependencies = [
"client-core",
"event-model",
"iris",
"log",
"pulldown-cmark",
]
[[package]]
name = "tree_magic_mini"
version = "3.2.2"
@@ -3409,6 +3693,12 @@ dependencies = [
"keyboard-types",
]
[[package]]
name = "unicase"
version = "2.9.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "dbc4bc3a9f746d862c45cb89d705aa10f187bb96c76001afab07a0d35ce60142"
[[package]]
name = "unicode-ident"
version = "1.0.22"
@@ -3427,6 +3717,62 @@ version = "0.2.2"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "b4ac048d71ede7ee76d585517add45da530660ef4390e49b098733c6e897f254"
[[package]]
name = "untrusted"
version = "0.9.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "8ecb6da28b8a351d773b68d5825ac39017e680750f980f3a1a85cd8dd28a47c1"
[[package]]
name = "ureq"
version = "3.4.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "972d7902c8735f2695410b8aed7df6ed12a47394aa1c8d7af49f0497b731a94d"
dependencies = [
"base64",
"cookie_store",
"flate2",
"log",
"percent-encoding",
"rustls",
"rustls-pki-types",
"serde",
"serde_json",
"ureq-proto",
"utf8-zero",
"webpki-roots",
]
[[package]]
name = "ureq-proto"
version = "0.6.1"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "da5f78b09e6941e1a0f2e30e695e4b120377b54d5e0aec11b594bb57b3971613"
dependencies = [
"base64",
"http",
"httparse",
"log",
]
[[package]]
name = "url"
version = "2.5.8"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "ff67a8a4397373c3ef660812acab3268222035010ab8680ec4215f38ba3d0eed"
dependencies = [
"form_urlencoded",
"idna",
"percent-encoding",
"serde",
]
[[package]]
name = "utf8-zero"
version = "0.8.1"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "b8c0a043c9540bae7c578c88f91dda8bd82e59ae27c21baca69c8b191aaf5a6e"
[[package]]
name = "utf8_iter"
version = "1.0.4"
@@ -3471,6 +3817,12 @@ dependencies = [
"winapi-util",
]
[[package]]
name = "wasi"
version = "0.11.1+wasi-snapshot-preview1"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "ccf3ec651a847eb01de73ccad15eb7d99f80485de043efb2f370cd654f4ea44b"
[[package]]
name = "wasip2"
version = "1.0.1+wasi-0.2.4"
@@ -3667,6 +4019,15 @@ dependencies = [
"wasm-bindgen",
]
[[package]]
name = "webpki-roots"
version = "1.0.9"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "7dcd9d09a39985f5344844e66b0c530a33843579125f23e21e9f0f220850f22a"
dependencies = [
"rustls-pki-types",
]
[[package]]
name = "weezl"
version = "0.1.12"
@@ -4535,6 +4896,12 @@ dependencies = [
"synstructure",
]
[[package]]
name = "zeroize"
version = "1.9.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "e13c156562582aa81c60cb29407084cdb54c4164760106ab78e6c5b0858cf64e"
[[package]]
name = "zerotrie"
version = "0.2.5"
@@ -4570,6 +4937,12 @@ dependencies = [
"syn 3.0.5",
]
[[package]]
name = "zmij"
version = "1.0.23"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "29666d0abbfad1e3dc4dcf6144730dd3a3ab225bbbdac83319345b1b44ccfc1b"
[[package]]
name = "zune-core"
version = "0.4.12"
+16 -1
View File
@@ -57,6 +57,16 @@ send_wrapper = "0.6.0"
# installs it -- this crate never installs a logger itself.
log = "0.4.28"
[features]
# RUST.md's I5 "Where iris's frame time goes" diagnosis: forces the Android
# `wgpu::Instance` to `Backends::GL` instead of `Backends::PRIMARY`, so the
# same build can be measured against SwiftShader's software Vulkan ICD (the
# default) or virgl's GLES path, without a second env-var plumbing path that
# nothing on this machine can hand to an already-launched Android process
# (there is no `am start` environment and no system-property reader here to
# add one). Android-only; `android/render.rs` is the only reader.
force-gles = []
[dev-dependencies]
tokio = { workspace = true, features = ["sync", "rt", "rt-multi-thread", "time"] }
# The tabs example's widget tree. A dev-dependency cycle back to this
@@ -73,7 +83,7 @@ name = "message_list"
harness = false
[workspace]
members = ["core", "macro", "tabs-ui"]
members = ["core", "macro", "tabs-ui", "transcript-ui", "desktop-app"]
# android-app pulls in android-view, which needs the NDK sysroot to link
# -- excluded so `cargo build --workspace --all-targets` on the host stays
# buildable. Cross-compile it from its own directory (its own single-crate
@@ -100,3 +110,8 @@ accesskit = "0.25.0"
iris-core = { path = "core" }
iris-macro = { path = "macro" }
tokio = "1.49.0"
# Current stable as of 2026-09-05 (`cargo search`) -- I5's markdown block
# model, the same crate E2's uncommitted `e2-transcript` experiment used for
# the identical job (RUST.md), rather than reimplementing a CommonMark
# parser.
pulldown-cmark = "0.13.4"
+366
View File
@@ -558,6 +558,12 @@ dependencies = [
"arrayvec",
]
[[package]]
name = "base64"
version = "0.23.1"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "ac07cdecf99051d9a5238b80f35af32cdeba5b336e55d957b318b50137e18da5"
[[package]]
name = "bit-set"
version = "0.8.0"
@@ -734,6 +740,16 @@ version = "0.2.2"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "f079e83a288787bcd14a6aea84cee5c87a67c5a3e660c30f557a3d24761b3527"
[[package]]
name = "client-core"
version = "0.1.0"
dependencies = [
"event-model",
"serde",
"serde_json",
"ureq",
]
[[package]]
name = "clipboard-win"
version = "5.4.1"
@@ -779,6 +795,35 @@ dependencies = [
"crossbeam-utils",
]
[[package]]
name = "cookie"
version = "0.18.2"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "1a373e3602691c3cdea496d2f0ee5935151e6168fe87739483c463db1b2f2f87"
dependencies = [
"percent-encoding",
"time",
"version_check",
]
[[package]]
name = "cookie_store"
version = "0.22.1"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "15b2c103cf610ec6cae3da84a766285b42fd16aad564758459e6ecf128c75206"
dependencies = [
"cookie",
"document-features",
"idna",
"indexmap",
"log",
"serde",
"serde_derive",
"serde_json",
"time",
"url",
]
[[package]]
name = "core-foundation"
version = "0.9.4"
@@ -886,6 +931,12 @@ version = "1.2.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "f27ae1dd37df86211c42e150270f82743308803d90a6f6e6651cd730d5e1732f"
[[package]]
name = "deranged"
version = "0.5.8"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "7cd812cc2bc1d69d4764bd80df88b4317eaef9e773c75226407d9bc0876b211c"
[[package]]
name = "dispatch"
version = "0.2.0"
@@ -1048,6 +1099,14 @@ dependencies = [
"pin-project-lite",
]
[[package]]
name = "event-model"
version = "0.1.0"
dependencies = [
"serde",
"serde_json",
]
[[package]]
name = "exr"
version = "1.74.2"
@@ -1179,6 +1238,15 @@ version = "0.3.1"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "aa9a19cbb55df58761df49b23516a86d432839add4af60fc256da840f66ed35b"
[[package]]
name = "form_urlencoded"
version = "1.2.2"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "cb4cb245038516f5f85277875cdaa4f7d2c9a0fa0468de06ed190163b1581fcf"
dependencies = [
"percent-encoding",
]
[[package]]
name = "futures-core"
version = "0.3.34"
@@ -1253,6 +1321,26 @@ dependencies = [
"windows-link",
]
[[package]]
name = "getopts"
version = "0.2.24"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "cfe4fbac503b8d1f88e6676011885f34b7174f46e59956bba534ba83abded4df"
dependencies = [
"unicode-width",
]
[[package]]
name = "getrandom"
version = "0.2.17"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "ff2abc00be7fca6ebc474524697ae276ad847ad0a6b3faa4bcb027e9a4614ad0"
dependencies = [
"cfg-if",
"libc",
"wasi",
]
[[package]]
name = "getrandom"
version = "0.3.4"
@@ -1423,6 +1511,22 @@ version = "0.2.1"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "dfa686283ad6dd069f105e5ab091b04c62850d3e4cf5d67debad1933f55023df"
[[package]]
name = "http"
version = "1.5.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "918d3568bebf352712bc2ef3d46a8bcf1a75b373be6539de198e9105cbbf9ce0"
dependencies = [
"bytes",
"itoa",
]
[[package]]
name = "httparse"
version = "1.10.1"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "6dbf3de79e51f3d586ab4cb9d5c3e2c14aa28ed23d180cf89b4df0454a69cc87"
[[package]]
name = "icu_collections"
version = "2.3.0"
@@ -1551,6 +1655,27 @@ version = "2.3.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "ae293c039020f9ec10710af98d29ce6aa2051486638b49c9a6409f3b4a9e98ad"
[[package]]
name = "idna"
version = "1.1.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "3b0875f23caa03898994f6ddc501886a45c7d3d62d04d2d90788d47be1b1e4de"
dependencies = [
"idna_adapter",
"smallvec",
"utf8_iter",
]
[[package]]
name = "idna_adapter"
version = "1.2.2"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "cb68373c0d6620ef8105e855e7745e18b0d00d3bdb07fb532e434244cdb9a714"
dependencies = [
"icu_normalizer",
"icu_properties",
]
[[package]]
name = "image"
version = "0.25.10"
@@ -1640,9 +1765,15 @@ version = "0.1.0"
dependencies = [
"android-view",
"android_logger",
"client-core",
"event-model",
"iris",
"libc",
"log",
"serde_json",
"tabs-ui",
"tokio",
"transcript-ui",
]
[[package]]
@@ -1676,6 +1807,12 @@ dependencies = [
"either",
]
[[package]]
name = "itoa"
version = "1.0.18"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "8f42a60cbdf9a97f5d2305f08a87dc4e09308d1276d28c869c684d7777685682"
[[package]]
name = "jni"
version = "0.21.1"
@@ -2096,6 +2233,12 @@ dependencies = [
"num-traits",
]
[[package]]
name = "num-conv"
version = "0.2.2"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "521739c6d2bac4aa25192232afe6841231376b2b26d4d9fae5ecf8ca5772e441"
[[package]]
name = "num-derive"
version = "0.4.2"
@@ -2744,6 +2887,12 @@ dependencies = [
"zerovec",
]
[[package]]
name = "powerfmt"
version = "0.2.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "439ee305def115ba05938db6eb1644ff94165c5ab5e9420d1c1bcedbba909391"
[[package]]
name = "ppv-lite86"
version = "0.2.21"
@@ -2796,6 +2945,25 @@ dependencies = [
"syn 2.0.119",
]
[[package]]
name = "pulldown-cmark"
version = "0.13.4"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "e9f068eba8e7071c5f9511831b44f32c740d5adf574e990f946ddb53db2f314e"
dependencies = [
"bitflags 2.13.1",
"getopts",
"memchr",
"pulldown-cmark-escape",
"unicase",
]
[[package]]
name = "pulldown-cmark-escape"
version = "0.11.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "007d8adb5ddab6f8e3f491ac63566a7d5002cc7ed73901f72057943fa71ae1ae"
[[package]]
name = "pulp"
version = "0.22.3"
@@ -3075,6 +3243,20 @@ version = "0.8.53"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "47b34b781b31e5d73e9fbc8689c70551fd1ade9a19e3e28cfec8580a79290cc4"
[[package]]
name = "ring"
version = "0.17.14"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "a4689e6c2294d81e88dc6261c768b63bc4fcdb852be6d1352498b114f61383b7"
dependencies = [
"cc",
"cfg-if",
"getrandom 0.2.17",
"libc",
"untrusted",
"windows-sys 0.52.0",
]
[[package]]
name = "roxmltree"
version = "0.21.1"
@@ -3125,6 +3307,41 @@ dependencies = [
"windows-sys 0.61.2",
]
[[package]]
name = "rustls"
version = "0.23.43"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "0283386ce02abc0151e1761d08802dfe86c173b0b494af5cbc086574e453da06"
dependencies = [
"log",
"once_cell",
"ring",
"rustls-pki-types",
"rustls-webpki",
"subtle",
"zeroize",
]
[[package]]
name = "rustls-pki-types"
version = "1.15.1"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "2f4925028c7eb5d1fcdaf196971378ed9d2c1c4efc7dc5d011256f76c99c0a96"
dependencies = [
"zeroize",
]
[[package]]
name = "rustls-webpki"
version = "0.103.15"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "f3c3cf1d8b1e7d4927e2d154c3fcb02979afb9939629c62cd9048d4f07b60ac2"
dependencies = [
"ring",
"rustls-pki-types",
"untrusted",
]
[[package]]
name = "rustversion"
version = "1.0.23"
@@ -3207,6 +3424,19 @@ dependencies = [
"syn 3.0.5",
]
[[package]]
name = "serde_json"
version = "1.0.151"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "c841b55ecdae098c80dcae9cf767f6f8a0c2cdb3416bbef72181df4d0fe73f14"
dependencies = [
"itoa",
"memchr",
"serde",
"serde_core",
"zmij",
]
[[package]]
name = "serde_repr"
version = "0.1.21"
@@ -3363,6 +3593,12 @@ version = "0.1.1"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "6637bab7722d379c8b41ba849228d680cc12d0a45ba1fa2b48f2a30577a06731"
[[package]]
name = "subtle"
version = "2.6.1"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "13c2bddecc57b384dee18652358fb23172facb8a2c51ccc10d74c157bdea3292"
[[package]]
name = "swash"
version = "0.2.10"
@@ -3490,6 +3726,36 @@ dependencies = [
"zune-jpeg",
]
[[package]]
name = "time"
version = "0.3.55"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "cdb87b95ec50ddfa440816d227a17b2ccbdda963a316a727fda0fc4334f7d134"
dependencies = [
"deranged",
"num-conv",
"powerfmt",
"serde_core",
"time-core",
"time-macros",
]
[[package]]
name = "time-core"
version = "0.1.9"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "9e1c906769ad99c88eaa54e728060edef082f8e358ff32030cb7c7d315e81109"
[[package]]
name = "time-macros"
version = "0.2.32"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "7e689342a48d2ea927c87ea50cabf8594854bf940e9310208848d680d668ed85"
dependencies = [
"num-conv",
"time-core",
]
[[package]]
name = "tiny-skia"
version = "0.11.4"
@@ -3596,6 +3862,17 @@ dependencies = [
"once_cell",
]
[[package]]
name = "transcript-ui"
version = "0.1.0"
dependencies = [
"client-core",
"event-model",
"iris",
"log",
"pulldown-cmark",
]
[[package]]
name = "tree_magic_mini"
version = "3.2.2"
@@ -3634,6 +3911,12 @@ dependencies = [
"keyboard-types",
]
[[package]]
name = "unicase"
version = "2.9.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "dbc4bc3a9f746d862c45cb89d705aa10f187bb96c76001afab07a0d35ce60142"
[[package]]
name = "unicode-ident"
version = "1.0.24"
@@ -3652,6 +3935,62 @@ version = "0.2.2"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "b4ac048d71ede7ee76d585517add45da530660ef4390e49b098733c6e897f254"
[[package]]
name = "untrusted"
version = "0.9.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "8ecb6da28b8a351d773b68d5825ac39017e680750f980f3a1a85cd8dd28a47c1"
[[package]]
name = "ureq"
version = "3.4.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "972d7902c8735f2695410b8aed7df6ed12a47394aa1c8d7af49f0497b731a94d"
dependencies = [
"base64",
"cookie_store",
"flate2",
"log",
"percent-encoding",
"rustls",
"rustls-pki-types",
"serde",
"serde_json",
"ureq-proto",
"utf8-zero",
"webpki-roots",
]
[[package]]
name = "ureq-proto"
version = "0.6.1"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "da5f78b09e6941e1a0f2e30e695e4b120377b54d5e0aec11b594bb57b3971613"
dependencies = [
"base64",
"http",
"httparse",
"log",
]
[[package]]
name = "url"
version = "2.5.8"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "ff67a8a4397373c3ef660812acab3268222035010ab8680ec4215f38ba3d0eed"
dependencies = [
"form_urlencoded",
"idna",
"percent-encoding",
"serde",
]
[[package]]
name = "utf8-zero"
version = "0.8.1"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "b8c0a043c9540bae7c578c88f91dda8bd82e59ae27c21baca69c8b191aaf5a6e"
[[package]]
name = "utf8_iter"
version = "1.0.4"
@@ -3696,6 +4035,12 @@ dependencies = [
"winapi-util",
]
[[package]]
name = "wasi"
version = "0.11.1+wasi-snapshot-preview1"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "ccf3ec651a847eb01de73ccad15eb7d99f80485de043efb2f370cd654f4ea44b"
[[package]]
name = "wasip2"
version = "1.0.4+wasi-0.2.12"
@@ -3889,6 +4234,15 @@ dependencies = [
"wasm-bindgen",
]
[[package]]
name = "webpki-roots"
version = "1.0.9"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "7dcd9d09a39985f5344844e66b0c530a33843579125f23e21e9f0f220850f22a"
dependencies = [
"rustls-pki-types",
]
[[package]]
name = "weezl"
version = "0.1.12"
@@ -4748,6 +5102,12 @@ dependencies = [
"synstructure",
]
[[package]]
name = "zeroize"
version = "1.9.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "e13c156562582aa81c60cb29407084cdb54c4164760106ab78e6c5b0858cf64e"
[[package]]
name = "zerotrie"
version = "0.2.5"
@@ -4789,6 +5149,12 @@ version = "0.6.7"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "34b31d188d9d685a4f9c7b46d6e36631b07058d2cfe190267adce54dc230bf12"
[[package]]
name = "zmij"
version = "1.0.23"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "29666d0abbfad1e3dc4dcf6144730dd3a3ab225bbbdac83319345b1b44ccfc1b"
[[package]]
name = "zune-core"
version = "0.5.3"
+46 -1
View File
@@ -17,10 +17,55 @@ crate-type = ["cdylib"]
[dependencies]
iris = { path = "../" }
tabs-ui = { path = "../tabs-ui" }
android-view = { git = "https://github.com/rust-mobile/android-view.git", rev = "bec6c62a96cef8239b0fd7fedeef9b184d02e3a1" }
android_logger = "0.15.0"
log = "0.4.28"
# `tabs-screen` (default, I2/I4's demo) and `transcript-screen` (I5's
# Android integration) are mutually exclusive -- one `ActiveClient` type is
# compiled in, never both (`lib.rs`'s doc comment) -- so both sets of deps
# are optional and each screen's feature pulls in only its own. Without
# this, building `--features transcript-screen` alone (default features
# still on) left `tabs-ui` linked but never referenced under that cfg,
# which Cargo's `unused_dependencies` lint (on by default) correctly flags.
tabs-ui = { path = "../tabs-ui", optional = true }
transcript-ui = { path = "../transcript-ui", optional = true }
client-core = { path = "../../client-core", optional = true }
event-model = { path = "../../event-model", optional = true }
serde_json = { version = "1", features = ["float_roundtrip"], optional = true }
# P0's bench build only (docs/RUST.md): `getrusage(RUSAGE_SELF)` for
# process CPU time, matching `libc::getrusage`'s mention in that box over
# parsing `/proc/self/stat` by hand and assuming `USER_HZ`. Already in the
# workspace's own dependency tree transitively (`iris/Cargo.lock`, pinned
# at 0.2.179) -- this makes it a direct dependency at the same version
# rather than a second, possibly-drifting resolution.
libc = { version = "0.2.179", optional = true }
# P0's bench build only: the scroll animation and the streaming phase are
# both a sequence of `sleep`s inside the async task `rsc.spawn_task` already
# runs on iris's own tokio runtime (`iris/src/task.rs`'s `Tasks::init`), and
# the battery sampler is a second, concurrent task on that same runtime
# (`tokio::spawn`) -- so this crate needs `tokio` directly rather than only
# through `iris`. `rt`+`time` only: no I/O, no macros, nothing this crate
# doesn't call. Version matches the one `iris`'s own dependency tree already
# resolves to (`iris/Cargo.lock`), so there is one copy of the runtime, not
# two.
tokio = { version = "1.53.1", features = ["rt", "time"], optional = true }
[features]
default = ["tabs-screen"]
tabs-screen = ["dep:tabs-ui"]
transcript-screen = ["dep:transcript-ui", "dep:client-core", "dep:event-model", "dep:serde_json"]
# RUST.md's I5 "Where iris's frame time goes": forces the GLES backend
# instead of SwiftShader's software Vulkan. See `iris/Cargo.toml`'s own doc
# on the feature this forwards to.
force-gles = ["iris/force-gles"]
# P0's iris half (docs/RUST.md, docs/AGENTS.md's "The rigs"): the same
# checked-in fixture, scroll loop and streaming phase the Compose `bench`
# build type drives, run here against `transcript-ui`'s real screen with no
# server. Depends on `transcript-screen` for `transcript-ui`/`client-core`/
# `event-model` -- `lib.rs`'s `ActiveClient` selection gives this feature
# priority over `transcript-screen`'s own `TranscriptClient` when both are
# listed, which is how this crate's build command names both explicitly.
bench = ["transcript-screen", "dep:libc", "dep:tokio"]
[profile.release]
panic = "abort"
+30
View File
@@ -19,9 +19,39 @@ android {
versionName = "1.0"
}
// A release build must be signed, and the key is per machine rather than per repo -- same
// reasoning and the same key as `app/build-apk.sh` (the Compose app): it is what a phone
// recognises the app by, and a secret never lives in a checkout (the mount is shared with an
// untrusted VM). `build-apk.sh` generates this key once and points at it through the
// environment; without it a release build here is unsigned, which is fine for everything
// except installing.
def keystore = System.getenv("AI_APP_KEYSTORE")
signingConfigs {
if (keystore != null) {
release {
storeFile = file(keystore)
storePassword = System.getenv("AI_APP_KEYSTORE_PASSWORD")
keyAlias = "ai-app"
keyPassword = storePassword
}
}
}
buildTypes {
debug {
}
// P0's iris half (docs/RUST.md's P0 box): the build a phone actually runs. The `.so`
// itself is built separately with `cargo ndk --release --features "transcript-screen
// force-gles bench"` straight into src/main/jniLibs/ (this crate's own Cargo.toml) --
// Gradle here only packages and signs whatever is already there, the same division as the
// debug/tabs-screen build this project started with. `applicationIdSuffix` keeps it
// installable beside a debug build of the tabs demo rather than replacing it.
release {
applicationIdSuffix ".bench"
if (keystore != null) {
signingConfig = signingConfigs.release
}
}
}
compileOptions {
@@ -1,6 +1,15 @@
<?xml version="1.0" encoding="utf-8"?>
<manifest xmlns:android="http://schemas.android.com/apk/res/android">
<!-- Only needed by the transcript-screen feature (RUST.md's I5),
which talks to a real ai-server; the plain tabs demo (I2/I4) makes
no network call and never noticed this was missing. Absent,
UreqTransport::new's connect failed with EPERM (Operation not
permitted), not the ECONNREFUSED/ENETUNREACH a firewall or a dead
server would give: a seccomp-level socket denial reads nothing
like a network problem, which is what made it worth a comment. -->
<uses-permission android:name="android.permission.INTERNET" />
<application
android:allowBackup="true"
android:label="iris android-view demo"
+100
View File
@@ -0,0 +1,100 @@
// Only does anything under the `transcript-screen` feature (RUST.md's I5
// Android integration) -- the plain tabs build (I2/I4) needs none of this
// and stays untouched, same reasoning as the feature gate in Cargo.toml.
//
// Bakes the sandbox server's host, port, token and pinned CA in at build
// time, the same way `app/androidApp/build.gradle.kts`'s
// `GeneratePinnedCert` task bakes the CA for the Compose app -- see that
// file's comment for why reading the machine's own certificate at build
// time is the right trust boundary. This build additionally bakes the
// host/port/token, which the Compose app does not: that app enrolls at
// runtime from a scanned QR/deep link, and a from-scratch enrollment UI
// (Keystore-sealed token storage, a QR/link scanner) is real, separate
// scope this integration does not need to build to answer RUST.md's
// question -- there is nothing here yet resembling `ServerConfig.kt`. So
// this is a **deliberate simplification for this rig only**: an APK built
// this way is good for exactly the emulator/server pair that built it, and
// must never be treated as a template for a real enrollment flow. Recorded
// in RUST.md's I5 box rather than left to be rediscovered.
use std::path::PathBuf;
fn main() {
if std::env::var_os("CARGO_FEATURE_TRANSCRIPT_SCREEN").is_none() {
return;
}
// P0's bench build (docs/RUST.md) opens the checked-in fixture with no
// server at all -- `bench_client.rs` never references the `pinned`
// module this generates, so requiring a live server's host/port/token/
// CA to build it (as plain `transcript-screen` does, below) would be a
// pointless requirement for a build that talks to nothing.
if std::env::var_os("CARGO_FEATURE_BENCH").is_some() {
return;
}
println!("cargo:rerun-if-env-changed=AI_APP_TRANSCRIPT_HOST");
println!("cargo:rerun-if-env-changed=AI_APP_TRANSCRIPT_PORT");
println!("cargo:rerun-if-env-changed=AI_APP_TRANSCRIPT_TOKEN");
println!("cargo:rerun-if-env-changed=AI_APP_CA");
println!("cargo:rerun-if-env-changed=XDG_CONFIG_HOME");
let host = require_env(
"AI_APP_TRANSCRIPT_HOST",
"the sandbox server's host as the emulator reaches it, e.g. 10.0.2.2",
);
let port = require_env(
"AI_APP_TRANSCRIPT_PORT",
"the sandbox server's port -- app/ui-sandbox.sh's start banner prints it",
);
let token = require_env(
"AI_APP_TRANSCRIPT_TOKEN",
"the bearer token -- ~/.config/ai-app/sandbox-token, or the start banner's enrollment link",
);
let ca_path = std::env::var_os("AI_APP_CA")
.map(PathBuf::from)
.unwrap_or_else(|| {
let base = std::env::var_os("XDG_CONFIG_HOME")
.map(PathBuf::from)
.unwrap_or_else(|| {
let home = std::env::var_os("HOME").expect("HOME must be set");
PathBuf::from(home).join(".config")
});
base.join("ai-app").join("certs").join("ca.pem")
});
let ca_pem = std::fs::read_to_string(&ca_path).unwrap_or_else(|e| {
panic!(
"no CA certificate at {} ({e}).\n\
Start ai-server (or app/ui-sandbox.sh) once on this machine first -- it \
generates the CA this build pins. Set AI_APP_CA=/path/to/ca.pem to build \
against a different one.",
ca_path.display()
)
});
let ca_pem = ca_pem.trim();
if !ca_pem.starts_with("-----BEGIN CERTIFICATE-----") {
panic!("{} is not a PEM certificate.", ca_path.display());
}
let out_dir = PathBuf::from(std::env::var_os("OUT_DIR").unwrap());
let generated = format!(
"// Generated by build.rs from {host}:{port} and {ca}. Do not edit.\n\
pub const HOST: &str = {host_lit:?};\n\
pub const PORT: u16 = {port};\n\
pub const TOKEN: &str = {token_lit:?};\n\
pub const CA_PEM: &str = {ca_lit:?};\n",
host = host,
port = port
.parse::<u16>()
.unwrap_or_else(|e| panic!("AI_APP_TRANSCRIPT_PORT={port:?} is not a u16: {e}")),
ca = ca_path.display(),
host_lit = host,
token_lit = token,
ca_lit = ca_pem,
);
std::fs::write(out_dir.join("pinned_config.rs"), generated).unwrap();
}
fn require_env(name: &str, what: &str) -> String {
std::env::var(name).unwrap_or_else(|_| {
panic!("{name} must be set to build the transcript-screen feature -- {what}")
})
}
+418
View File
@@ -0,0 +1,418 @@
//! P0's iris half (docs/RUST.md's P0 box, docs/AGENTS.md's "The rigs"):
//! the same fixture, scroll loop and streaming phase the Compose `bench`
//! build type's `BenchRun.kt`/`BenchFixture.kt` drive, run here against
//! `transcript-ui`'s real screen with no server -- a frame-time comparison
//! that measures the renderer rather than the data or the network.
//!
//! **Reuses `transcript_client.rs`'s shape** (folded items, a full
//! `transcript_ui::build_tree` rebuild per event) with the network half
//! replaced by the checked-in fixture, embedded with `include_str!` --
//! `app/bench-fixture/assets/transcript.jsonl`, 1,915,760 bytes, generated
//! by `app/bench-fixture/generate.py` and never a real transcript (that
//! file's own README). The first 3,200 lines are the opening backlog,
//! folded once through `client_core::transcript_fold::fold_page` exactly
//! as a real `/transcript` page would be; the remaining ~400 are the
//! streaming tail, replayed one at a time through `fold_event` -- the same
//! fold path a live SSE reply arrives on -- by the "Run benchmark"
//! control below.
use crate::bench_jni::PlatformHandle;
use android_view::jni::{JavaVM, objects::GlobalRef};
use client_core::transcript_fold::{TranscriptItem, fold_event, fold_page, group_tool_runs};
use event_model::SeqEvent;
use iris::android::{AndroidAppState, AndroidRsc, AndroidUiState, HasAndroidUiState};
use iris::prelude::*;
use std::sync::Arc;
use std::sync::atomic::{AtomicBool, Ordering};
use std::time::Duration;
/// bench-fixture/README.md: the first `BACKLOG_COUNT` non-blank lines are
/// the opening window; the rest are the streaming tail. Kept in sync with
/// `BenchFixture.kt`'s identical constant by hand -- both read the same
/// checked-in file, so a mismatch would only mean the two apps' bench
/// builds open a different split of it, not a wrong-vs-right answer.
const BACKLOG_COUNT: usize = 3200;
/// `BenchRun.kt`'s own constants -- kept identical so the two apps' bench
/// runs are the same gesture and the same load, which is the entire point
/// of a shared fixture and a shared scripted loop (P0's pass condition).
const CYCLES: usize = 6;
const SWIPE_PX: f32 = 900.0;
const SWIPE_MS: u64 = 200;
const SWIPE_PAUSE_MS: u64 = 500;
const STREAM_EVENTS_PER_SEC: u64 = 20;
const STREAM_SECONDS: u64 = 20;
/// One animation step's target cadence -- close enough to 60Hz that a
/// `List::scroll` swipe is many small moves rather than one jump, so
/// frames are actually rendered along the way (the point of animating it
/// at all rather than calling `scroll` once per swipe).
const ANIM_STEP_MS: u64 = 16;
const FIXTURE_JSONL: &str = include_str!("../../../app/bench-fixture/assets/transcript.jsonl");
pub struct BenchClient {
ui_state: AndroidUiState,
content: WeakWidget<WidgetPtr>,
report_display: WeakWidget<TextEdit>,
screen: Option<transcript_ui::TranscriptScreen>,
items: Vec<TranscriptItem>,
/// The events not yet streamed -- consumed by `start_benchmark`'s own
/// clone, kept here only as the source a second run would need (the
/// button can be pressed more than once; `running` just stops overlap,
/// not repeat).
stream_tail: Vec<SeqEvent>,
platform: Option<Arc<PlatformHandle>>,
last_report: Option<String>,
running: bool,
}
impl HasAndroidUiState for BenchClient {
fn android_state(&self) -> &AndroidUiState {
&self.ui_state
}
fn android_state_mut(&mut self) -> &mut AndroidUiState {
&mut self.ui_state
}
}
/// Parses the fixture once: `serde_json::Value`s for the backlog
/// (`fold_page` takes a page of raw wire JSON, same as a real
/// `/transcript` response) and folded `SeqEvent`s for the tail (`fold_event`
/// takes one live wire event at a time, same as a real SSE frame).
fn parse_fixture() -> (Vec<serde_json::Value>, Vec<SeqEvent>) {
let lines: Vec<&str> = FIXTURE_JSONL
.lines()
.filter(|line| !line.trim().is_empty())
.collect();
let mut backlog = Vec::with_capacity(BACKLOG_COUNT.min(lines.len()));
let mut stream_tail = Vec::new();
for (i, line) in lines.iter().enumerate() {
let value: serde_json::Value =
serde_json::from_str(line).expect("bench fixture is generated JSON, always valid");
if i < BACKLOG_COUNT {
backlog.push(value);
} else {
let event: SeqEvent = serde_json::from_value(value)
.expect("bench fixture event matches event-model's SeqEvent");
stream_tail.push(event);
}
}
(backlog, stream_tail)
}
fn placeholder<Rsc: HasEvents>(rsc: &mut Rsc, message: &str) -> StrongWidget {
wtext(message.to_string())
.color(Color::WHITE)
.wrap(true)
.pad(16)
.add_strong(rsc)
.any()
}
/// `getrusage(RUSAGE_SELF)`'s user+system time, in ms -- `None` only if
/// the syscall itself fails, which UI_RULES.md's "never present an
/// inferred value as a measured one" says to keep apart from a real (and
/// here, impossible) zero.
fn process_cpu_ms() -> Option<u64> {
// SAFETY: `rusage` is a plain-old-data struct `getrusage` fully
// initialises on success; on failure it is never read.
unsafe {
let mut usage: libc::rusage = std::mem::zeroed();
if libc::getrusage(libc::RUSAGE_SELF, &mut usage) != 0 {
return None;
}
let user_ms = usage.ru_utime.tv_sec as u64 * 1000 + usage.ru_utime.tv_usec as u64 / 1000;
let sys_ms = usage.ru_stime.tv_sec as u64 * 1000 + usage.ru_stime.tv_usec as u64 / 1000;
Some(user_ms + sys_ms)
}
}
/// `VmHWM` from `/proc/self/status` -- the process's peak RSS since it
/// started, in kB. Same source `BenchRun.kt`'s `peakRssLine` reads, so the
/// two reports' numbers mean the same thing.
fn peak_rss_kb() -> Option<u64> {
std::fs::read_to_string("/proc/self/status")
.ok()?
.lines()
.find_map(|line| line.strip_prefix("VmHWM:"))
.and_then(|rest| rest.trim().strip_suffix("kB"))
.and_then(|n| n.trim().parse().ok())
}
fn battery_line(samples: &[i32]) -> String {
if samples.is_empty() {
return " battery current: unavailable on this device".to_string();
}
let mean = samples.iter().map(|&v| v as i64).sum::<i64>() / samples.len() as i64;
let min = samples.iter().min().unwrap();
let max = samples.iter().max().unwrap();
format!(
" battery current: mean {mean}\u{b5}A over {} samples (min {min}, max {max})",
samples.len()
)
}
impl AndroidAppState for BenchClient {
fn new(mut ui_state: AndroidUiState, rsc: &mut AndroidRsc<Self>) -> Self {
let content = WidgetPtr::new().add(rsc);
let loading = placeholder(rsc, "Loading fixture...");
content(rsc).set(loading);
let report_display = wtext("")
.editable(EditMode::MultiLine)
.text_align(Align::LEFT)
.wrap(true)
.size(14)
.color(Color::WHITE)
.attr::<Selectable>(())
.label("Benchmark report")
.add(rsc);
let controls = bench_controls(rsc);
let tree = (
controls,
content.height(rest(2)),
report_display.height(rest(1)).pad(8),
)
.span(Dir::DOWN)
.add_strong(rsc)
.any();
ui_state.set_root(tree);
let mut client = Self {
ui_state,
content,
report_display,
screen: None,
items: Vec::new(),
stream_tail: Vec::new(),
platform: None,
last_report: None,
running: false,
};
let (backlog, stream_tail) = parse_fixture();
client.stream_tail = stream_tail;
match fold_page(&backlog) {
Ok(items) => {
client.items = items;
client.rebuild_transcript(rsc);
}
Err(message) => {
client.show_message(rsc, &format!("Couldn't fold the bench fixture: {message}"))
}
}
client
}
fn platform_ready(&mut self, _rsc: &mut AndroidRsc<Self>, vm: JavaVM, view: GlobalRef) {
self.platform = Some(Arc::new(PlatformHandle::new(vm, view)));
}
fn back_pressed(&mut self, _rsc: &mut AndroidRsc<Self>, _render: &mut UiRenderState) -> bool {
false
}
}
type Rsc = AndroidRsc<BenchClient>;
fn bench_controls(rsc: &mut Rsc) -> WeakWidget {
let run_rect = rect(Color::rgb(40, 70, 40))
.on(
CursorSense::click(),
|ctx: EventIdCtx<'_, Rsc, _, _>, rsc: &mut Rsc| {
ctx.state.start_benchmark(rsc);
},
)
.label("Run benchmark");
let run = (
run_rect,
wtext("Run benchmark").size(18).text_align(Align::CENTER),
)
.stack()
.pad(8)
.add(rsc);
let copy_rect = rect(Color::rgb(50, 50, 60))
.on(
CursorSense::click(),
|ctx: EventIdCtx<'_, Rsc, _, _>, _rsc: &mut Rsc| {
ctx.state.copy_report();
},
)
.label("Copy report");
let copy = (
copy_rect,
wtext("Copy report").size(18).text_align(Align::CENTER),
)
.stack()
.pad(8)
.add(rsc);
(run, copy).span(Dir::RIGHT).height(56).add(rsc)
}
impl BenchClient {
fn show_message(&mut self, rsc: &mut Rsc, message: &str) {
let widget = placeholder(rsc, message);
(self.content)(rsc).set(widget);
self.screen = None;
}
fn rebuild_transcript(&mut self, rsc: &mut Rsc) {
let rows = group_tool_runs(&self.items);
let (screen, tree) = transcript_ui::build_tree(rsc, rows);
(self.content)(rsc).set(tree);
self.screen = Some(screen);
}
fn copy_report(&mut self) {
let Some(report) = &self.last_report else {
log::info!("iris bench report: nothing to copy -- run the benchmark first");
return;
};
let Some(platform) = &self.platform else {
log::info!("iris bench report: no platform handle, can't reach the clipboard");
return;
};
if platform.copy_to_clipboard("iris bench report", report) {
log::info!("iris bench report: copied to clipboard");
} else {
log::info!("iris bench report: clipboard copy failed");
}
}
/// P0's scripted run: `BenchRun.kt`'s scroll loop, then its streaming
/// phase, then the report -- run in-process for the same reason that
/// file's own doc gives (no usable system tracing on a real phone, no
/// agent that can drive one).
fn start_benchmark(&mut self, rsc: &mut Rsc) {
if self.running {
log::info!("iris bench report: already running");
return;
}
self.running = true;
self.android_state_mut().frame_report.reset();
self.report_display.edit(rsc).set("Running benchmark...");
let redraw = rsc.tasks.redraw_handle();
let platform = self.platform.clone();
let stream_tail = self.stream_tail.clone();
let cpu_start = process_cpu_ms();
rsc.spawn_task(async move |mut ctx| {
// The swipe loop: two drags toward newer content, two back --
// a cycle returns to where it started, so the whole loop
// measures steady-state scrolling. `BenchRun.kt`'s own
// comment on this shape.
for _ in 0..CYCLES {
for delta in [SWIPE_PX, SWIPE_PX, -SWIPE_PX, -SWIPE_PX] {
animate_scroll(&mut ctx, &redraw, delta, SWIPE_MS).await;
tokio::time::sleep(Duration::from_millis(SWIPE_PAUSE_MS)).await;
}
}
// Pinned to the newest end before streaming starts, matching
// `stream-bench.sh`'s "Jump to latest" tap.
ctx.update(|state: &mut BenchClient, rsc| {
if let Some(screen) = &state.screen {
(screen.list)(rsc).jump_to_end();
}
});
redraw.request_redraw();
// The battery sampler runs concurrently with the streaming
// phase, once a second, the same cadence `BatterySampler` uses
// on the Compose side -- via its own JNI-attached thread, not
// `ctx.update`, since a sample needs no widget-tree access.
let sampler_done = Arc::new(AtomicBool::new(false));
let samples = Arc::new(std::sync::Mutex::new(Vec::<i32>::new()));
let sampler = platform.clone().map(|platform| {
let done = sampler_done.clone();
let samples = samples.clone();
tokio::spawn(async move {
while !done.load(Ordering::Relaxed) {
if let Some(value) = platform.battery_current_ua() {
samples.lock().unwrap().push(value);
}
tokio::time::sleep(Duration::from_secs(1)).await;
}
})
});
let total = (STREAM_EVENTS_PER_SEC * STREAM_SECONDS) as usize;
let mut sent = 0usize;
for event in stream_tail.into_iter().take(total) {
ctx.update(move |state: &mut BenchClient, rsc| {
state.items = fold_event(&state.items, &event);
state.rebuild_transcript(rsc);
});
redraw.request_redraw();
sent += 1;
tokio::time::sleep(Duration::from_millis(1000 / STREAM_EVENTS_PER_SEC)).await;
}
// Lets the last few deltas land and draw before the report is
// read -- `BenchRun.kt`'s own closing delay.
tokio::time::sleep(Duration::from_millis(300)).await;
sampler_done.store(true, Ordering::Relaxed);
if let Some(sampler) = sampler {
let _ = sampler.await;
}
let battery = battery_line(&samples.lock().unwrap());
let cpu_line = match (cpu_start, process_cpu_ms()) {
(Some(start), Some(end)) => {
format!(" process CPU time over this run: {}ms", end.saturating_sub(start))
}
_ => " process CPU time over this run: unavailable".to_string(),
};
let rss_line = match peak_rss_kb() {
Some(kb) => format!(" peak RSS: {kb}kB"),
None => " peak RSS: unavailable (/proc/self/status unreadable)".to_string(),
};
ctx.update(move |state: &mut BenchClient, rsc| {
state.running = false;
let scroll_line = format!(
" scroll: {CYCLES} cycles ({} swipes), streamed {sent}/{total} fixture events",
CYCLES * 4
);
let frames_line = match state.android_state().frame_report.report() {
Some(stats) => format!("{stats}"),
None => "no frames recorded".to_string(),
};
let report = format!(
"iris bench report\n{frames_line}\n{scroll_line}\n{cpu_line}\n{rss_line}\n{battery}"
);
log::info!("iris bench report: {report}");
state.report_display.edit(rsc).set(&report);
state.last_report = Some(report);
});
redraw.request_redraw();
});
}
}
/// Moves `List::scroll` by `total_px` over `duration_ms`, in ~60Hz steps,
/// so the swipe is many rendered frames rather than one jump -- the same
/// shape `animateScrollBy(SWIPE_PX, tween(SWIPE_MS))` gives on the Compose
/// side, in the one place the two backends have to differ (iris's `List`
/// has no built-in tween, so this drives it by hand).
async fn animate_scroll(
ctx: &mut iris::task::TaskCtx<Rsc>,
redraw: &Arc<dyn iris::task::RequestRedraw>,
total_px: f32,
duration_ms: u64,
) {
let steps = (duration_ms / ANIM_STEP_MS).max(1);
let step_px = total_px / steps as f32;
for _ in 0..steps {
ctx.update(move |state: &mut BenchClient, rsc| {
if let Some(screen) = &state.screen {
(screen.list)(rsc).scroll(step_px);
}
});
redraw.request_redraw();
tokio::time::sleep(Duration::from_millis(ANIM_STEP_MS)).await;
}
}
+134
View File
@@ -0,0 +1,134 @@
//! JNI calls the `bench` feature needs that go through the shell's own
//! Java side rather than anything `iris`/`android-view` already wraps:
//! `BatteryManager.getIntProperty(BATTERY_PROPERTY_CURRENT_NOW)` for the
//! per-second battery sample, and `ClipboardManager.setPrimaryClip` for
//! the "Copy report" control (P0's iris half, docs/RUST.md). Neither is
//! part of `android_view::context`'s own `Context`/`Resources` wrappers
//! (that file's own `// TODO: more methods?`), so this calls them
//! directly rather than growing that crate's wrapper for two one-off
//! calls this crate alone needs.
//!
//! Holds its own `JavaVM` + `GlobalRef` to the view (handed in through
//! [`iris::android::AndroidAppState::platform_ready`]) so it can attach
//! whichever thread calls it -- the battery sampler runs on a background
//! tokio task, not the UI thread the rest of `IrisViewPeer`'s JNI calls
//! run on. `JavaVM::attach_current_thread` is safe to call from a thread
//! already attached (the `jni` crate detects it and does not double
//! attach), so no caller here needs to know or care which thread it is.
use android_view::jni::{
JNIEnv, JavaVM,
objects::{GlobalRef, JObject, JValue},
};
/// `android.os.BatteryManager.BATTERY_PROPERTY_CURRENT_NOW` -- not exposed
/// as a constant anywhere reachable without the Android SDK jar, so named
/// here with its source rather than left as a bare `2`.
const BATTERY_PROPERTY_CURRENT_NOW: i32 = 2;
pub struct PlatformHandle {
vm: JavaVM,
view: GlobalRef,
}
impl PlatformHandle {
pub fn new(vm: JavaVM, view: GlobalRef) -> Self {
Self { vm, view }
}
fn context<'e>(&self, env: &mut JNIEnv<'e>) -> Option<JObject<'e>> {
env.call_method(
self.view.as_obj(),
"getContext",
"()Landroid/content/Context;",
&[],
)
.ok()?
.l()
.ok()
}
fn system_service<'e>(
&self,
env: &mut JNIEnv<'e>,
context: &JObject<'e>,
name: &str,
) -> Option<JObject<'e>> {
let jname = env.new_string(name).ok()?;
env.call_method(
context,
"getSystemService",
"(Ljava/lang/String;)Ljava/lang/Object;",
&[JValue::Object(jname.as_ref())],
)
.ok()?
.l()
.ok()
}
/// One sample of `BATTERY_PROPERTY_CURRENT_NOW`, in microamps. `None`
/// on any JNI failure, on a device with no `BatteryManager` service,
/// or when the platform itself answers "not supported" -- `0` or
/// `Integer.MIN_VALUE` are both documented SDK answers for that, and
/// both would read as a real (and wrong) measurement if folded into an
/// average rather than named apart. UI_RULES.md: never present an
/// inferred value as a measured one.
pub fn battery_current_ua(&self) -> Option<i32> {
let mut guard = self.vm.attach_current_thread().ok()?;
let env: &mut JNIEnv = &mut guard;
let context = self.context(env)?;
let battery_manager = self.system_service(env, &context, "batterymanager")?;
let value = env
.call_method(
&battery_manager,
"getIntProperty",
"(I)I",
&[JValue::Int(BATTERY_PROPERTY_CURRENT_NOW)],
)
.ok()?
.i()
.ok()?;
if value == 0 || value == i32::MIN {
None
} else {
Some(value)
}
}
/// Puts `text` on the system clipboard through `ClipboardManager` --
/// `true` only if the whole JNI chain (service lookup, `ClipData`,
/// `setPrimaryClip`) succeeded.
pub fn copy_to_clipboard(&self, label: &str, text: &str) -> bool {
self.try_copy_to_clipboard(label, text).is_some()
}
fn try_copy_to_clipboard(&self, label: &str, text: &str) -> Option<()> {
let mut guard = self.vm.attach_current_thread().ok()?;
let env: &mut JNIEnv = &mut guard;
let context = self.context(env)?;
let clipboard = self.system_service(env, &context, "clipboard")?;
let jlabel = env.new_string(label).ok()?;
let jtext = env.new_string(text).ok()?;
let clip = env
.call_static_method(
"android/content/ClipData",
"newPlainText",
"(Ljava/lang/CharSequence;Ljava/lang/CharSequence;)Landroid/content/ClipData;",
&[
JValue::Object(jlabel.as_ref()),
JValue::Object(jtext.as_ref()),
],
)
.ok()?
.l()
.ok()?;
env.call_method(
&clipboard,
"setPrimaryClip",
"(Landroid/content/ClipData;)V",
&[JValue::Object(&clip)],
)
.ok()?;
Some(())
}
}
+48 -3
View File
@@ -1,5 +1,5 @@
//! The android-view demo app: iris's `tabs` widget tree (`tabs_ui::build`,
//! shared with the winit example) running through
//! The android-view demo app: by default, iris's `tabs` widget tree
//! (`tabs_ui::build`, shared with the winit example) running through
//! `iris::android`'s `ViewPeer`. This is RUST.md's I2 pass condition made
//! concrete -- there is no UI here beyond what `tabs-ui` already draws.
//!
@@ -9,6 +9,32 @@
//! wrapping `iris::android::new_peer`'s generic function in a concrete
//! `extern "system" fn`, since `register_view_class` wants a plain
//! function pointer.
//!
//! **`transcript-screen` feature (RUST.md's I5 Android integration):** with
//! `--features transcript-screen`, `new_view_peer` instantiates
//! `transcript_client::TranscriptClient` instead of the tabs `Client`
//! below, against a real `ai-server` (see that module's doc). Chosen over a
//! third shell crate: this one already has the Gradle project, the
//! `IrisView`/`MainActivity` Java, and the JNI registration I2 built and
//! measured against, and the only thing a transcript screen needs on top
//! is a different `AndroidAppState` -- the same axis `tabs_ui::build` vs.
//! `transcript_ui::build` already varies along on the winit side (compare
//! `iris/examples/tabs.rs` and `iris/transcript-ui/examples/transcript.rs`).
//! A build picks one screen or the other, never both, so `Client` and
//! `TranscriptClient` are cfg-gated apart rather than switched at runtime --
//! there is no in-app navigation to switch *to* on either side yet.
//!
//! **`bench` feature (P0's iris half, docs/RUST.md):** a third
//! `AndroidAppState`, `bench_client::BenchClient`, on the same axis --
//! `transcript_ui::build_tree` again, this time against the checked-in
//! fixture (`app/bench-fixture/assets/transcript.jsonl`) instead of a real
//! server, with a "Run benchmark" control that drives the same scroll loop
//! and streaming phase the Compose `bench` build type's `BenchRun.kt`
//! does. `bench` depends on `transcript-screen` (Cargo.toml) for
//! `transcript-ui`/`client-core`/`event-model`, so both features end up
//! enabled together -- `ActiveClient` below gives `bench` priority in that
//! case, the same way `transcript-screen` already takes priority over the
//! default `tabs-screen`.
use android_view::{
Context, View,
@@ -18,19 +44,30 @@ use android_view::{
},
register_view_class,
};
#[cfg(not(feature = "transcript-screen"))]
use iris::android::{AndroidAppState, AndroidRsc, AndroidUiState, HasAndroidUiState};
#[cfg(not(feature = "transcript-screen"))]
use iris::prelude::*;
use log::LevelFilter;
use std::ffi::c_void;
#[cfg(feature = "bench")]
mod bench_client;
#[cfg(feature = "bench")]
mod bench_jni;
#[cfg(all(feature = "transcript-screen", not(feature = "bench")))]
mod transcript_client;
/// The app's `View` subclass, matching the Java side's package --
/// `app/src/main/java/dev/iris/android/demo/IrisView.java`.
const VIEW_CLASS: &str = "dev/iris/android/demo/IrisView";
#[cfg(not(feature = "transcript-screen"))]
pub struct Client {
ui_state: AndroidUiState,
}
#[cfg(not(feature = "transcript-screen"))]
impl HasAndroidUiState for Client {
fn android_state(&self) -> &AndroidUiState {
&self.ui_state
@@ -40,6 +77,7 @@ impl HasAndroidUiState for Client {
}
}
#[cfg(not(feature = "transcript-screen"))]
impl AndroidAppState for Client {
fn new(mut ui_state: AndroidUiState, rsc: &mut AndroidRsc<Self>) -> Self {
// `widgets.info` is the winit example's frame-debug readout, kept
@@ -61,12 +99,19 @@ impl AndroidAppState for Client {
}
}
#[cfg(not(feature = "transcript-screen"))]
type ActiveClient = Client;
#[cfg(all(feature = "transcript-screen", not(feature = "bench")))]
type ActiveClient = transcript_client::TranscriptClient;
#[cfg(feature = "bench")]
type ActiveClient = bench_client::BenchClient;
extern "system" fn new_view_peer<'local>(
env: JNIEnv<'local>,
view: View<'local>,
context: Context<'local>,
) -> jlong {
iris::android::new_peer::<Client>(env, view, context)
iris::android::new_peer::<ActiveClient>(env, view, context)
}
/// # Safety
+375
View File
@@ -0,0 +1,375 @@
//! RUST.md's I5 Android integration: `transcript-ui`'s screen filling the
//! whole window on android-view, against a real `ai-server` through
//! `client-core` -- the missing half `iris-android-app` (I2) only had for
//! `tabs-ui` until now. Behind the `transcript-screen` Cargo feature so the
//! plain build (`cargo ndk build`, no `--features`) stays exactly the tabs
//! demo I2/I4 already measured against.
//!
//! **Deliberate simplification, recorded rather than left to be
//! rediscovered (RUST.md's I5 box has the full account)**: there is no
//! session list and no enrollment UI here. The server, port, token and
//! pinned CA are baked in at build time (`build.rs`'s
//! `AI_APP_TRANSCRIPT_HOST`/`_PORT`/`_TOKEN`/`AI_APP_CA`), and the first
//! session `ApiClient::fetch_sessions` returns is opened automatically --
//! there is nothing to tap to get there, which is what `transcript-bench.sh`
//! and `ui-trace` need to land straight on the screen under test. A real
//! app needs `desktop-app`'s `EnrolledServer`/QR-link flow or E3's
//! Keystore-sealed `ServerConfig.kt`; building a second one of those was
//! not this pass's job.
//!
//! **Reuses `iris/desktop-app`'s `app.rs` shape almost exactly** --
//! `fold_event`/`group_tool_runs`/`fold_page`/`raw_seq` from
//! `client_core::transcript_fold`, a `generation` counter guarding against
//! a stale background response, and a full rebuild of the widget tree on
//! every event (same tradeoff, same reason: `push_row` cannot update a row
//! already on screen, and this rig's conversations are small). What
//! differs is only the redraw mechanism: android-view has no
//! `winit::EventLoopProxy`, so this uses `iris::task::Tasks::redraw_handle`
//! (new, added alongside this box) to request a frame after each
//! `TaskCtx::update` instead of relying on `Tasks::spawn`'s single
//! end-of-future redraw -- see that method's own doc for why.
use client_core::api::{ApiClient, UreqTransport};
use client_core::event_stream::{StreamItem, follow_session_events};
use client_core::transcript_fold::{TranscriptItem, fold_event, fold_page, group_tool_runs};
use event_model::SeqEvent;
use iris::android::{AndroidAppState, AndroidRsc, AndroidUiState, HasAndroidUiState};
use iris::prelude::*;
use std::sync::Arc;
use std::sync::atomic::{AtomicU64, Ordering};
mod pinned {
include!(concat!(env!("OUT_DIR"), "/pinned_config.rs"));
}
pub struct TranscriptClient {
ui_state: AndroidUiState,
/// The screen's own content -- everything under the fixed
/// [`frame_report_controls`] bar, which is built once (`new`, below)
/// and never touched by `show_message`/`rebuild_transcript`'s own
/// `set` calls the way `desktop-app`'s `transcript_ptr` isn't touched
/// by rebuilding the session list beside it.
content: WeakWidget<WidgetPtr>,
screen: Option<transcript_ui::TranscriptScreen>,
/// The folded transcript as of the last rebuild -- kept here (not
/// re-derived) for the same reason `desktop-app`'s `Client::items`
/// exists: a live `StreamEvent` only carries one new wire event, and
/// `fold_event` needs everything folded so far to fold it in.
items: Vec<TranscriptItem>,
/// The session currently open -- `None` only before the first fetch
/// resolves. Read back by `apply_event`'s rebuild, which has no session
/// id of its own (a live `SeqEvent` doesn't carry one).
session_id: Option<String>,
/// Bumped every time a new session load starts; a background response
/// checks it before touching state, so a slow reply for a session this
/// screen has moved on from can't overwrite what replaced it. There is
/// only ever one session here (no list to switch away to), but the
/// guard still matters for the *first* fetch racing a `stop`/`start`.
generation: Arc<AtomicU64>,
}
impl HasAndroidUiState for TranscriptClient {
fn android_state(&self) -> &AndroidUiState {
&self.ui_state
}
fn android_state_mut(&mut self) -> &mut AndroidUiState {
&mut self.ui_state
}
}
/// Builds one `UreqTransport` from the config `build.rs` baked in. Called
/// twice per session load, same as `desktop-app`'s `build_transport`
/// closure -- `ApiClient` and the live-stream follow each need their own,
/// since `UreqTransport` holds its own `ureq::Agent`.
fn build_transport() -> Result<UreqTransport, String> {
let base_url = format!("https://{}:{}", pinned::HOST, pinned::PORT);
UreqTransport::new(
base_url,
pinned::TOKEN.to_string(),
pinned::CA_PEM.as_bytes(),
)
.map_err(|e| e.to_string())
}
fn placeholder<Rsc: HasEvents>(rsc: &mut Rsc, message: &str) -> StrongWidget {
wtext(message.to_string())
.color(Color::WHITE)
.wrap(true)
.pad(16)
.add_strong(rsc)
.any()
}
/// The two named controls RUST.md's I5 box ("Measurements taken" (b))
/// drives by name over `ui-trace`, e.g. `ui-trace record --do "tap 'Frame
/// report'"`. `dumpsys gfxinfo` cannot see this screen's own GPU-drawn
/// frames at all -- this is the screen's own equivalent of the Compose
/// app's "Copy render timings" control, logged rather than clipboarded
/// (no clipboard wiring exists here) under this crate's own fixed
/// `android_logger` tag (`iris-android-app`, `lib.rs`'s `JNI_OnLoad`),
/// grep-able on the fixed string `"iris frame report"` the way
/// `transcript-bench.sh` greps `"ai-app render report"`.
fn frame_report_controls(rsc: &mut AndroidRsc<TranscriptClient>) -> WeakWidget {
type Rsc = AndroidRsc<TranscriptClient>;
let report_rect = rect(Color::rgb(50, 50, 60))
.on(
CursorSense::click(),
|ctx: EventIdCtx<'_, Rsc, _, _>, _rsc: &mut Rsc| match ctx
.state
.android_state()
.frame_report
.report()
{
Some(stats) => log::info!("iris frame report: {stats}"),
None => log::info!(
"iris frame report: no frames recorded -- scroll first, then press this"
),
},
)
.label("Frame report");
let report = (
report_rect,
wtext("Frame report").size(18).text_align(Align::CENTER),
)
.stack()
.pad(8)
.add(rsc);
let reset_rect = rect(Color::rgb(70, 40, 40))
.on(
CursorSense::click(),
|ctx: EventIdCtx<'_, Rsc, _, _>, _rsc: &mut Rsc| {
ctx.state.android_state_mut().frame_report.reset();
log::info!("iris frame report: reset");
},
)
.label("Reset frame report");
let reset = (
reset_rect,
wtext("Reset").size(18).text_align(Align::CENTER),
)
.stack()
.pad(8)
.add(rsc);
(report, reset).span(Dir::RIGHT).height(56).add(rsc)
}
impl AndroidAppState for TranscriptClient {
fn new(mut ui_state: AndroidUiState, rsc: &mut AndroidRsc<Self>) -> Self {
let content = WidgetPtr::new().add(rsc);
let loading = placeholder(rsc, "Loading sessions...");
content(rsc).set(loading);
let tree = (frame_report_controls(rsc), content.height(rest(1)))
.span(Dir::DOWN)
.add_strong(rsc)
.any();
ui_state.set_root(tree);
let mut client = Self {
ui_state,
content,
screen: None,
items: Vec::new(),
session_id: None,
generation: Arc::new(AtomicU64::new(0)),
};
client.spawn_fetch_sessions(rsc);
client
}
fn back_pressed(&mut self, _rsc: &mut AndroidRsc<Self>, _render: &mut UiRenderState) -> bool {
// No screen stack of its own -- same "let the activity finish"
// answer `iris-android-app`'s tabs `Client` already gives.
false
}
}
impl TranscriptClient {
fn show_message(&mut self, rsc: &mut AndroidRsc<Self>, message: &str) {
let widget = placeholder(rsc, message);
(self.content)(rsc).set(widget);
self.screen = None;
}
fn spawn_fetch_sessions(&mut self, rsc: &mut AndroidRsc<Self>) {
let redraw = rsc.tasks.redraw_handle();
let my_generation = self.generation.load(Ordering::SeqCst);
let generation = self.generation.clone();
rsc.spawn_task(async move |mut ctx| {
let outcome = match build_transport() {
Ok(transport) => ApiClient::new(transport)
.fetch_sessions()
.map_err(|e| e.to_string()),
Err(e) => Err(format!("couldn't set up TLS: {e}")),
};
ctx.update(move |state: &mut TranscriptClient, rsc| {
if generation.load(Ordering::SeqCst) != my_generation {
return;
}
match outcome {
Ok(sessions) => match sessions.into_iter().next() {
Some(session) => state.select_session(rsc, session.id),
None => state.show_message(rsc, "No sessions on the sandbox server."),
},
Err(message) => {
state.show_message(rsc, &format!("Couldn't list sessions: {message}"))
}
}
});
redraw.request_redraw();
});
}
/// Loads the opening page, then follows the live SSE stream for the
/// rest of this session's life -- `desktop-app`'s `select_session`
/// almost verbatim, with `Proxy::send_event` replaced by `ctx.update` +
/// `redraw.request_redraw()` (see this module's doc).
fn select_session(&mut self, rsc: &mut AndroidRsc<Self>, session_id: String) {
let my_generation = self.generation.fetch_add(1, Ordering::SeqCst) + 1;
self.items.clear();
self.session_id = Some(session_id.clone());
self.show_message(rsc, "Loading transcript...");
let redraw = rsc.tasks.redraw_handle();
let live_generation = self.generation.clone();
rsc.spawn_task(async move |mut ctx| {
let transports =
build_transport().and_then(|rest| build_transport().map(|stream| (rest, stream)));
let (rest, stream_transport) = match transports {
Ok(pair) => pair,
Err(e) => {
let message = format!("couldn't set up TLS: {e}");
ctx.update(move |state: &mut TranscriptClient, rsc| {
if live_generation.load(Ordering::SeqCst) == my_generation {
state.show_message(rsc, &message);
}
});
redraw.request_redraw();
return;
}
};
let api = ApiClient::new(rest);
// The most recent 200 events, coalesced -- the same page size
// `desktop-app` uses; RUST.md's I3/history-paging work is what
// a real scrollback would reuse (out of scope here, same as
// E4).
let page: Result<Vec<serde_json::Value>, String> = api
.fetch_transcript_page(&session_id, None, 200, true)
.map_err(|e| e.to_string());
// The wire `seq` of the last line, not a folded item's `seq()`
// -- see `client_core::transcript_fold::raw_seq`'s doc for why
// resuming from the latter re-delivers deltas already folded
// into an in-progress reply.
let after = page
.as_ref()
.ok()
.and_then(|values| values.last())
.and_then(client_core::transcript_fold::raw_seq)
.unwrap_or(0);
let result = page.and_then(|values| fold_page(&values));
{
let live_generation = live_generation.clone();
ctx.update(move |state: &mut TranscriptClient, rsc| {
if live_generation.load(Ordering::SeqCst) != my_generation {
return;
}
match result {
Ok(items) => {
state.items = items;
state.rebuild_transcript(rsc);
}
Err(message) => {
state.show_message(rsc, &format!("Couldn't load transcript: {message}"))
}
}
});
}
redraw.request_redraw();
if live_generation.load(Ordering::SeqCst) != my_generation {
return;
}
// The outer closure here is an `FnMut` -- `follow_session_events`
// calls it once per line -- so it captures `live_generation` by
// move and re-clones it for each inner `ctx.update` closure
// rather than moving a shared `stop`-style helper into itself:
// a value moved out of an `FnMut`'s captures on one call leaves
// nothing there for the next.
let _ =
follow_session_events(
&stream_transport,
&session_id,
after,
move |item| match item {
StreamItem::Open | StreamItem::Reset => {
live_generation.load(Ordering::SeqCst) == my_generation
}
StreamItem::Event { event, .. } => {
if live_generation.load(Ordering::SeqCst) != my_generation {
return false;
}
let live_generation = live_generation.clone();
ctx.update(move |state: &mut TranscriptClient, rsc| {
if live_generation.load(Ordering::SeqCst) != my_generation {
return;
}
state.apply_event(rsc, &event);
});
redraw.request_redraw();
true
}
},
);
});
}
/// Rebuilds the whole widget tree from `self.items` -- same tradeoff as
/// `desktop-app`'s `rebuild_transcript` (this module's doc comment).
/// Reads `self.session_id` rather than taking one, since every caller
/// (the opening page, and every live event) already has it set there.
fn rebuild_transcript(&mut self, rsc: &mut AndroidRsc<Self>) {
let in_progress = self
.screen
.as_ref()
.map(|screen| screen.composer.field.edit(rsc).text.text().to_string())
.filter(|t| !t.is_empty());
let rows = group_tool_runs(&self.items);
let (screen, tree) = transcript_ui::build_tree(rsc, rows);
if let Some(text) = in_progress {
screen.composer.field.edit(rsc).set(&text);
}
if let Some(session_id) = self.session_id.clone() {
let field = screen.composer.field;
rsc.register_event(field, Submit, move |ctx, rsc| {
let text = field.edit(rsc).take();
let text = text.trim().to_string();
if !text.is_empty() {
ctx.state.send_message(session_id.clone(), text);
}
});
}
(self.content)(rsc).set(tree);
self.screen = Some(screen);
}
fn apply_event(&mut self, rsc: &mut AndroidRsc<Self>, event: &SeqEvent) {
self.items = fold_event(&self.items, event);
self.rebuild_transcript(rsc);
}
fn send_message(&mut self, session_id: String, text: String) {
std::thread::spawn(move || {
if let Ok(transport) = build_transport() {
let api = ApiClient::new(transport);
let _ = api.send_message(&session_id, &text, &[]);
}
});
}
}
+113 -6
View File
@@ -1,8 +1,9 @@
use crate::{Align, GlyphAtlas, GlyphKey, PlacedGlyph, RegionAlign, Textures, UiColor, util::Vec2};
use parley::{
Alignment, AlignmentOptions, FontContext, FontFamily, FontFamilyName, GenericFamily, Layout,
LayoutContext, LineHeight, PositionedLayoutItem, StyleProperty,
Alignment, AlignmentOptions, FontContext, FontFamily, FontFamilyName, FontStyle, FontWeight,
GenericFamily, Layout, LayoutContext, LineHeight, PositionedLayoutItem, StyleProperty,
};
use std::ops::Range;
use swash::{
FontRef,
scale::{Render, ScaleContext, Source, StrikeWith},
@@ -51,6 +52,72 @@ impl Family {
}
}
/// One styled run inside a `TextBuffer`, overriding `TextAttrs`' base style
/// over `range` (a byte range into the buffer's text). Every field is
/// optional so a span only says what it changes -- e.g. a link span sets
/// `color` and `underline` and leaves weight/family at the paragraph's own
/// default. This is I5's answer to RUST.md's inline-rich-text ceiling
/// (`masonry/src/widgets/text_area.rs`'s `StyleSet` is one style for the
/// whole editor, with `// TODO: RichTextInput` beside it): parley's own
/// `RangedBuilder::push` already takes a style and a range, so per-span
/// bold/italic/monospace/colour/underline only needed plumbing this struct
/// through to it and giving each glyph its own colour at draw time (see
/// `PlacedGlyph::color` and `TextData::place` below) instead of the one
/// `RenderedText::color` every glyph used to share.
#[derive(Clone, PartialEq)]
pub struct SpanStyle {
pub range: Range<usize>,
pub color: Option<UiColor>,
pub family: Option<Family>,
/// Overrides `TextAttrs::font_size` for just this range -- what lets a
/// heading inside a transcript row's single `TextEdit` be bigger than
/// the paragraph text around it, so a whole markdown-folded row (block
/// and inline styling both) can stay one selectable text buffer instead
/// of one widget per block.
pub font_size: Option<f32>,
pub bold: bool,
pub italic: bool,
pub underline: bool,
}
impl SpanStyle {
pub fn new(range: Range<usize>) -> Self {
Self {
range,
color: None,
family: None,
font_size: None,
bold: false,
italic: false,
underline: false,
}
}
pub fn color(mut self, color: UiColor) -> Self {
self.color = Some(color);
self
}
pub fn family(mut self, family: Family) -> Self {
self.family = Some(family);
self
}
pub fn font_size(mut self, size: f32) -> Self {
self.font_size = Some(size);
self
}
pub fn bold(mut self) -> Self {
self.bold = true;
self
}
pub fn italic(mut self) -> Self {
self.italic = true;
self
}
pub fn underline(mut self) -> Self {
self.underline = true;
self
}
}
#[derive(Clone, PartialEq)]
pub struct TextAttrs {
pub color: UiColor,
@@ -86,8 +153,12 @@ impl Default for TextAttrs {
pub struct TextBuffer {
text: String,
layout: Layout<UiColor>,
spans: Vec<SpanStyle>,
/// What the current layout was built for, so `shape` can decline to redo
/// work that would come out the same.
/// work that would come out the same. Spans are not part of this key --
/// `set_spans` forces `shaped` to `None` directly, the same way `edit`
/// does, since spans change far less often than a naive equality check
/// on the whole `Vec` would cost to compute every frame.
shaped: Option<(TextAttrs, Option<f32>)>,
}
@@ -96,10 +167,19 @@ impl TextBuffer {
Self {
text: text.into(),
layout: Layout::new(),
spans: Vec::new(),
shaped: None,
}
}
/// Replace this buffer's per-range style overrides (I5's rich text --
/// see `SpanStyle`). Invalidates the layout unconditionally, mirroring
/// `set_text`.
pub fn set_spans(&mut self, spans: Vec<SpanStyle>) {
self.spans = spans;
self.shaped = None;
}
pub fn new_empty() -> Self {
Self::new("")
}
@@ -150,6 +230,27 @@ impl TextBuffer {
attrs.line_height,
)));
builder.push_default(StyleProperty::Brush(attrs.color));
for span in &self.spans {
let range = span.range.clone();
if let Some(color) = span.color {
builder.push(StyleProperty::Brush(color), range.clone());
}
if let Some(family) = &span.family {
builder.push(StyleProperty::FontFamily(family.family()), range.clone());
}
if let Some(size) = span.font_size {
builder.push(StyleProperty::FontSize(size), range.clone());
}
if span.bold {
builder.push(StyleProperty::FontWeight(FontWeight::BOLD), range.clone());
}
if span.italic {
builder.push(StyleProperty::FontStyle(FontStyle::Italic), range.clone());
}
if span.underline {
builder.push(StyleProperty::Underline(true), range.clone());
}
}
builder.build_into(&mut self.layout, &self.text);
self.layout.break_all_lines(width);
self.layout
@@ -175,6 +276,7 @@ impl TextData {
let font = run.run().font();
let font_size = run.run().font_size();
let coords = run.run().normalized_coords();
let run_color = run.style().brush;
let Some(font_ref) = FontRef::from_index(font.data.as_ref(), font.index as usize)
else {
continue;
@@ -227,6 +329,7 @@ impl TextData {
glyph.x.floor() + entry.left as f32,
glyph.y.floor() - entry.top as f32,
),
color: run_color,
});
}
}
@@ -245,11 +348,15 @@ fn hash_coords(coords: &[i16]) -> u64 {
h
}
/// A laid-out string, ready to draw: where each glyph goes, how big the whole
/// thing is, and what colour to tint the atlas with.
/// A laid-out string, ready to draw: where each glyph goes and how big the
/// whole thing is.
///
/// Cheap to clone and to keep, which is the point -- a widget holds one across
/// frames and re-emits its quads without going near the rasteriser.
/// frames and re-emits its quads without going near the rasteriser. `color`
/// is the buffer's *base* colour (`TextAttrs::color`) for a caller that wants
/// it as a whole (e.g. tinting a cursor to match); the colour each glyph is
/// actually drawn in is `PlacedGlyph::color`, which a `SpanStyle` can
/// override per range.
#[derive(Clone)]
pub struct RenderedText {
pub glyphs: std::sync::Arc<Vec<PlacedGlyph>>,
+9 -1
View File
@@ -10,7 +10,7 @@
//! it, and a resize re-emits quads without touching the GPU's copy at all.
use crate::{
PatchRect, TextureHandle, Textures,
PatchRect, TextureHandle, Textures, UiColor,
util::{HashMap, Vec2},
};
use image::RgbaImage;
@@ -228,8 +228,16 @@ fn write_glyph(page: &mut RgbaImage, image: &Image, x: u32, y: u32) {
}
/// Where a glyph goes on screen, in pixels relative to the text's origin.
///
/// `color` is per-glyph (read from the parley run's own `Brush`, since
/// `UiColor` is parley's brush type here) rather than a single colour for
/// the whole `RenderedText`, so that a span pushed with its own
/// `StyleProperty::Brush` (I5's inline rich text: a link, a diff of colour
/// inside one wrapped paragraph) actually renders in that colour instead of
/// the buffer's base one.
#[derive(Clone, Copy)]
pub struct PlacedGlyph {
pub entry: GlyphEntry,
pub offset: Vec2,
pub color: UiColor,
}
+299
View File
@@ -0,0 +1,299 @@
use std::time::Duration;
/// The frame budget `dumpsys gfxinfo` also uses to call a frame "janky": the
/// 60Hz vsync period. Kept as the same threshold so a percentage from this
/// report and a percentage from `gfxinfo` mean the same thing.
pub const JANK_THRESHOLD: Duration = Duration::from_nanos(16_666_667);
/// Enough frames for several minutes of scrolling before the oldest ones
/// start being overwritten -- the same "diagnostic, not a log" sizing
/// `FrameStats.kt`'s `CAP` uses on the Compose side, chosen independently
/// here since a `Duration` is smaller than the six `Long` arrays it keeps.
const RING_CAPACITY: usize = 4096;
/// A per-frame wall-time report iris keeps of itself, because `dumpsys
/// gfxinfo` cannot see a `SurfaceView`'s own GPU-drawn frames at all
/// (RUST.md's I5 box, "Measurements taken" (b)): it instruments Android's
/// ordinary Skia/HWUI View-drawing pipeline, which a `wgpu`-rendered
/// `SurfaceView` bypasses entirely. `record` is meant to be called once per
/// frame, wrapping the same span Compose's own render report and `gfxinfo`
/// count -- from the frame's redraw/update start to after the frame is
/// handed to the platform to present.
///
/// **What this does not measure**: wgpu's `present()` call queues the frame
/// with the compositor and returns; it is not fenced against the GPU
/// actually finishing the frame or the compositor actually showing it, the
/// way `gfxinfo`'s own `GPU_DURATION`/vsync accounting is. So a sample here
/// is "how long the CPU took to build and submit this frame", not
/// "how long the frame took to reach the screen" -- named in
/// [`FrameStats`]'s own `Display` line rather than presented as the latter,
/// per the standing rule against showing an inferred number as a measured
/// one where the two differ.
///
/// Fixed-size ring, no allocation on the hot path -- `report()` is the only
/// place that allocates (a sort over the current ring), and it is only
/// ever called from a button tap, not once per frame.
pub struct FrameReport {
ring: Box<[Duration; RING_CAPACITY]>,
/// The `submit_to_present` half of each sample in `ring`, same index,
/// same lifetime -- kept as a second ring rather than a ring of pairs so
/// the existing `ring`/percentile code above is untouched (RUST.md's I5
/// "Where iris's frame time goes" CPU/GPU split, added 2026-09-05).
/// `ring[i] - submit_ring[i]` is that frame's `redraw_to_submit` half.
submit_ring: Box<[Duration; RING_CAPACITY]>,
/// How many of `ring`'s slots hold a real sample -- saturates at
/// `RING_CAPACITY`, unlike `total_frames` below which keeps counting.
len: usize,
pos: usize,
/// All frames recorded since the last `reset`, even past `RING_CAPACITY`
/// -- what `janky_percent` divides by, so a long run's percentage stays
/// correct even once the ring itself only holds the most recent frames.
total_frames: u64,
janky_frames: u64,
}
/// One resolved reading. `Display` is the log line both the "Frame report"
/// button and `transcript-bench.sh`-style scripts read, grep-able on
/// `"iris frame report"`.
pub struct FrameStats {
pub total_frames: u64,
pub janky_percent: f64,
pub p50: Duration,
pub p90: Duration,
pub p99: Duration,
pub worst: Duration,
/// Median of `redraw_to_submit` -- iris's own CPU work (layout, text,
/// primitive building) up to and including building the `queue.submit`
/// call, per frame. RUST.md's I5 "Where iris's frame time goes" split,
/// added 2026-09-05 to answer "CPU or GPU?" with a number rather than a
/// guess.
pub cpu_p50: Duration,
/// Median of `submit_to_present` -- the `queue.submit` call itself plus
/// `present()`, i.e. wherever the driver/GPU/compositor wait actually
/// happens. Same caveat as the type's own doc: `present()` is not
/// fenced against the GPU actually finishing, so this is "how long the
/// CPU was blocked handing the frame off", not the frame's true GPU
/// time -- still enough to separate "iris is slow building the frame"
/// from "iris is slow handing it to the driver".
pub gpu_wait_p50: Duration,
}
impl std::fmt::Display for FrameStats {
fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result {
write!(
f,
"frames={} janky%={:.2} p50={:.1}ms p90={:.1}ms p99={:.1}ms worst={:.1}ms \
(measures redraw-start to after present() is called, not GPU/compositor \
completion)",
self.total_frames,
self.janky_percent,
self.p50.as_secs_f64() * 1000.0,
self.p90.as_secs_f64() * 1000.0,
self.p99.as_secs_f64() * 1000.0,
self.worst.as_secs_f64() * 1000.0,
)?;
write!(
f,
" cpu_p50={:.1}ms gpu_wait_p50={:.1}ms (redraw-start-to-submit vs. \
submit-to-after-present)",
self.cpu_p50.as_secs_f64() * 1000.0,
self.gpu_wait_p50.as_secs_f64() * 1000.0,
)
}
}
impl FrameReport {
pub fn new() -> Self {
Self {
ring: Box::new([Duration::ZERO; RING_CAPACITY]),
submit_ring: Box::new([Duration::ZERO; RING_CAPACITY]),
len: 0,
pos: 0,
total_frames: 0,
janky_frames: 0,
}
}
/// Record one frame's elapsed wall time, with no CPU/GPU split (the
/// `submit_to_present` half is recorded as zero, so `cpu_p50` reads as
/// the whole frame and `gpu_wait_p50` as nothing -- honest for a caller
/// that never measured the split, rather than fabricating one). O(1),
/// no allocation.
pub fn record(&mut self, elapsed: Duration) {
self.record_split(elapsed, Duration::ZERO);
}
/// Record one frame's elapsed wall time, split at `queue.submit`:
/// `submit_to_present` is the `queue.submit()` call plus `present()`;
/// `total - submit_to_present` is everything before it (layout, text,
/// primitive building). RUST.md's I5 "Where iris's frame time goes"
/// CPU/GPU split, added 2026-09-05. O(1), no allocation.
pub fn record_split(&mut self, total: Duration, submit_to_present: Duration) {
self.ring[self.pos] = total;
self.submit_ring[self.pos] = submit_to_present;
self.pos = (self.pos + 1) % RING_CAPACITY;
self.len = (self.len + 1).min(RING_CAPACITY);
self.total_frames += 1;
if total > JANK_THRESHOLD {
self.janky_frames += 1;
}
}
/// Clears every counter and every sample -- what the "Reset frame
/// report" control calls, so a report covers only what was scrolled
/// after the button was pressed (the same reason `FrameStats.kt`'s
/// `reset()` exists on the Compose side).
pub fn reset(&mut self) {
self.len = 0;
self.pos = 0;
self.total_frames = 0;
self.janky_frames = 0;
}
/// `None` if nothing has been recorded since the last reset -- the
/// "no frames recorded, scroll first" case, not a zeroed report that
/// would read as a real (perfect) measurement.
pub fn report(&self) -> Option<FrameStats> {
if self.len == 0 {
return None;
}
let mut samples: Vec<Duration> = self.ring[..self.len].to_vec();
samples.sort_unstable();
let pct = |p: usize| samples[(samples.len() * p / 100).min(samples.len() - 1)];
// Separate arrays rather than subtracting the two medians above:
// medians do not distribute over subtraction, and each needs its
// own sort.
let submit_samples: Vec<Duration> = self.submit_ring[..self.len].to_vec();
let cpu_samples: Vec<Duration> = self.ring[..self.len]
.iter()
.zip(self.submit_ring[..self.len].iter())
.map(|(&total, &submit_to_present)| total.saturating_sub(submit_to_present))
.collect();
let median = |mut v: Vec<Duration>| {
v.sort_unstable();
v[v.len() / 2]
};
Some(FrameStats {
total_frames: self.total_frames,
janky_percent: 100.0 * self.janky_frames as f64 / self.total_frames as f64,
p50: pct(50),
p90: pct(90),
p99: pct(99),
worst: *samples.last().expect("len > 0 checked above"),
cpu_p50: median(cpu_samples),
gpu_wait_p50: median(submit_samples),
})
}
}
impl Default for FrameReport {
fn default() -> Self {
Self::new()
}
}
#[cfg(test)]
mod tests {
use super::*;
#[test]
fn no_frames_reports_none() {
assert!(FrameReport::new().report().is_none());
}
#[test]
fn one_frame_is_every_percentile_and_the_worst() {
let mut r = FrameReport::new();
r.record(Duration::from_millis(10));
let stats = r.report().unwrap();
assert_eq!(stats.total_frames, 1);
assert_eq!(stats.p50, Duration::from_millis(10));
assert_eq!(stats.p99, Duration::from_millis(10));
assert_eq!(stats.worst, Duration::from_millis(10));
assert_eq!(stats.janky_percent, 0.0);
}
#[test]
fn percentiles_and_worst_over_a_known_set() {
let mut r = FrameReport::new();
// 100 samples, 1ms..=100ms, fed out of order so the ring's own
// order is not what gives the right answer -- the sort has to.
for ms in (1..=100).rev() {
r.record(Duration::from_millis(ms));
}
let stats = r.report().unwrap();
assert_eq!(stats.total_frames, 100);
assert_eq!(stats.p50, Duration::from_millis(51));
assert_eq!(stats.p90, Duration::from_millis(91));
assert_eq!(stats.p99, Duration::from_millis(100));
assert_eq!(stats.worst, Duration::from_millis(100));
}
#[test]
fn jank_threshold_matches_gfxinfos_60hz_budget() {
let mut r = FrameReport::new();
r.record(Duration::from_nanos(16_666_667)); // exactly on budget: not janky
r.record(Duration::from_nanos(16_666_668)); // one ns over: janky
let stats = r.report().unwrap();
assert_eq!(stats.janky_percent, 50.0);
}
#[test]
fn janky_percent_is_over_all_time_frames_not_just_the_ring() {
// Fewer than RING_CAPACITY frames, all janky, then a fresh reset --
// the percentage must reset to 0, not divide by a stale count.
let mut r = FrameReport::new();
for _ in 0..10 {
r.record(Duration::from_millis(50));
}
assert_eq!(r.report().unwrap().janky_percent, 100.0);
r.reset();
assert!(r.report().is_none());
r.record(Duration::from_millis(1));
assert_eq!(r.report().unwrap().janky_percent, 0.0);
}
#[test]
fn record_without_a_split_reports_the_whole_frame_as_cpu() {
// A caller that never measured the split (plain `record`) should
// not fabricate a GPU-wait number -- it reads as zero, and the CPU
// half reads as the whole frame.
let mut r = FrameReport::new();
r.record(Duration::from_millis(20));
let stats = r.report().unwrap();
assert_eq!(stats.cpu_p50, Duration::from_millis(20));
assert_eq!(stats.gpu_wait_p50, Duration::ZERO);
}
#[test]
fn record_split_reports_each_halfs_own_median() {
let mut r = FrameReport::new();
// Three frames: total is always 30ms, but the CPU/GPU-wait split
// moves, so the two medians must be independent of each other and
// of `total`'s own median.
r.record_split(Duration::from_millis(30), Duration::from_millis(5));
r.record_split(Duration::from_millis(30), Duration::from_millis(10));
r.record_split(Duration::from_millis(30), Duration::from_millis(20));
let stats = r.report().unwrap();
assert_eq!(stats.p50, Duration::from_millis(30));
assert_eq!(stats.gpu_wait_p50, Duration::from_millis(10));
assert_eq!(stats.cpu_p50, Duration::from_millis(20));
}
#[test]
fn ring_wraps_without_growing_past_capacity() {
let mut r = FrameReport::new();
for i in 0..(RING_CAPACITY * 2) {
r.record(Duration::from_millis(1 + (i % 5) as u64));
}
let stats = r.report().unwrap();
// total_frames keeps the full count even once the ring has wrapped.
assert_eq!(stats.total_frames, (RING_CAPACITY * 2) as u64);
// but every sample the ring can report on is still one of the five
// values fed in, since a wrap can only overwrite with more of the
// same pattern here.
assert!(stats.worst <= Duration::from_millis(5));
}
}
+44
View File
@@ -11,16 +11,60 @@ use wgpu::{
mod atlas;
mod data;
mod frame_report;
mod primitive;
mod texture;
mod util;
pub use atlas::*;
pub use data::{Mask, MaskIdx, MoveIdx, MoveOffset};
pub use frame_report::{FrameReport, FrameStats, JANK_THRESHOLD};
pub use primitive::*;
const SHAPE_SHADER: &str = include_str!("./shader.wgsl");
/// The `wgpu::Limits` both platform backends (`android::render::
/// AndroidRenderer::new`, `default::render::UiRenderer::new`) ask
/// `Adapter::request_device` for -- shared so the two copies cannot drift,
/// per AGENTS.md's "write the logic once."
///
/// Built from `Limits::default()`, **not** a downlevel variant: the shader
/// (`shader.wgsl`) reads four `var<storage>` buffers (rects, glyphs, masks,
/// move_offsets) from the vertex stage, and `Limits::downlevel_webgl2_defaults()`
/// zeroes `max_storage_buffers_per_shader_stage` along with the compute
/// limits below -- switching to it would trade one `request_device` crash
/// for a bind-group-layout one on the same downlevel hardware this is meant
/// to support. `max_buffer_size` is raised for the growing instance/atlas
/// buffers (`ArrBuf`, `GpuTextures`); everything else is `default()`'s
/// desktop-tier value, unchanged.
///
/// The six `max_compute_*` fields are zeroed because nothing in this crate
/// creates a `ComputePipeline` or writes a `@compute` shader stage --
/// grepped for both across `iris`/`iris-core` before writing this, found
/// none. `Limits::default()` requests desktop-tier compute limits
/// unconditionally (`max_compute_workgroups_per_dimension: 65535`) even
/// though nothing asks a device to actually support compute, which is what
/// crashed `request_device` on the Android emulator's software GL path
/// (`EMU_GPU=software`, `force-gles`): SwiftShader's GL reports itself as
/// OpenGL ES 3.0, which has no compute shaders, so the adapter's real limit
/// is 0 and the unconditional request fails outright
/// (`RUST.md`'s "Software mode ... crashes for a third, different reason").
/// The same would happen on a real GLES-3.0-only Android device. If a
/// future change adds a compute pass, request the specific limits it needs
/// here rather than reverting to the desktop-tier default for everything.
pub fn device_limits() -> Limits {
Limits {
max_buffer_size: 1 << 30,
max_compute_workgroup_storage_size: 0,
max_compute_invocations_per_workgroup: 0,
max_compute_workgroup_size_x: 0,
max_compute_workgroup_size_y: 0,
max_compute_workgroup_size_z: 0,
max_compute_workgroups_per_dimension: 0,
..Default::default()
}
}
pub struct UiRenderNode {
uniform_group: BindGroup,
primitive_layout: BindGroupLayout,
+1 -1
View File
@@ -194,7 +194,7 @@ impl<'a> Painter<'a> {
glyph.entry.uv_min,
glyph.entry.uv_max,
glyph.entry.layer,
text.color,
glyph.color,
flags_for(glyph.entry.is_color),
),
region,
+27
View File
@@ -0,0 +1,27 @@
[package]
name = "desktop-app"
version.workspace = true
edition.workspace = true
# RUST.md's E4: the same transcript-ui screen (I5) in a winit window on the
# desktop, beside a session list, talking to a real `ai-server` through
# `client-core`'s REST + SSE clients. Enrolment reuses the phone's own
# `aiapp://enroll?...` link (`client-core::config`) rather than inventing a
# second format -- see DECISIONS.md's 2026-09-05 entry. An ordinary
# workspace member (unlike `android-app`): nothing here needs the NDK, so
# `cargo build --workspace --all-targets` at the host stays clean with it
# included.
[dependencies]
iris = { path = ".." }
transcript-ui = { path = "../transcript-ui" }
client-core = { path = "../../client-core" }
event-model = { path = "../../event-model" }
# Already pulled in transitively through client-core; used directly here
# only to persist `EnrolledServer` as the app's own tiny config file (see
# `config.rs`) -- no new dependency.
serde_json = { version = "1", features = ["float_roundtrip"] }
winit = { workspace = true }
[dev-dependencies]
tempfile = "3"
+425
View File
@@ -0,0 +1,425 @@
//! RUST.md's E4: a session list on the left, `transcript-ui`'s screen (I5)
//! filling the rest, both against a real `ai-server` reached through
//! `client-core`. The layout is the simplest thing that shows both at
//! once -- a fixed-width column and `rest(1)` for everything else, using
//! `iris::widget::{Span, WidgetPtr}` the way `tabs-ui` already switches
//! panes, rather than anything desktop-specific:
//!
//! ```text
//! +-----------+--------------------------------------+
//! | session | transcript_ui::TranscriptScreen |
//! | list | (List of folded rows + composer) |
//! | (WidgetPtr| |
//! | swapped | (WidgetPtr swapped whole on session |
//! | on data) | switch or a new transcript event) |
//! +-----------+--------------------------------------+
//! ```
//!
//! **Deliberately left simple, and why**: every incoming SSE event refolds
//! the *entire* transcript (`client_core::transcript_fold::fold_event` is
//! already `O(items)` and a desktop session's conversation is small) and
//! rebuilds the whole right-hand widget tree from scratch, rather than
//! reaching for `TranscriptScreen::push_row`'s incremental append.
//! `push_row` cannot update a row already on screen -- only append a new
//! one -- and a streaming assistant reply is exactly a row whose *text*
//! keeps changing after it first appears (see `transcript-ui`'s own doc on
//! `fold_event` folding deltas into one growing item). A full rebuild
//! shows that growth correctly at the cost of redrawing everything each
//! time; fine for this proof, wrong for a long, fast-streaming transcript
//! -- the incremental path that fixes it needs `transcript-ui` to expose
//! updating a row in place, which it does not yet. The composer's
//! in-progress text survives a rebuild (`rebuild_transcript`'s
//! `in_progress` local) since the user typing a followup while a reply
//! streams in is the one case a naive rebuild would otherwise lose data
//! on.
//!
//! Background network I/O (`client_core::api`/`event_stream`, both
//! blocking by design -- see `client-core`'s `Cargo.toml`) runs on plain
//! `std::thread`s that report back through `winit`'s `EventLoopProxy`
//! (`Proxy<AppEvent>`), rather than through iris's own `Tasks`/`task_on`:
//! `Tasks` only requests a redraw once, after its whole async closure
//! finishes, which fits a single request-then-update but not a live SSE
//! loop that needs to be seen redrawing after *each* event it relays.
//! `Proxy::send_event` wakes the window's event loop immediately, once per
//! event, which is what a stream wants.
use client_core::api::{ApiClient, SessionSummary, UreqTransport};
use client_core::event_stream::{StreamItem, follow_session_events};
use client_core::transcript_fold::{
TranscriptItem, fold_event, fold_page, group_tool_runs, raw_seq,
};
use event_model::SeqEvent;
use iris::prelude::*;
use std::sync::Arc;
use std::sync::atomic::{AtomicU64, Ordering};
/// The session list column's width -- a fixed size for the simplest
/// layout that shows both panels at once (UI_RULES's text-truncation and
/// no-shrink rules apply to what's drawn inside it, not to this choice of
/// column width itself).
const LIST_WIDTH: f32 = 260.0;
/// Everything a background thread hands back to the window's event loop.
/// `generation` on the session-scoped variants is the generation
/// `select_session` was on when the thread started (`Client::generation`)
/// -- compared back against the current one before being applied, so a
/// slow response from a session the reader has since clicked away from
/// can't overwrite what replaced it.
enum AppEvent {
Sessions(Result<Vec<SessionSummary>, String>),
TranscriptLoaded {
session_id: String,
generation: u64,
result: Result<Vec<TranscriptItem>, String>,
},
StreamEvent {
session_id: String,
generation: u64,
event: SeqEvent,
},
StreamEnded {
session_id: String,
generation: u64,
message: Option<String>,
},
SendFailed(String),
}
pub fn run() {
DefaultApp::<Client>::run();
}
#[derive(DefaultUiState)]
struct Client {
ui_state: DefaultUiState,
api: Arc<ApiClient<UreqTransport>>,
/// A second, independent `UreqTransport` to the same server, used only
/// by `select_session`'s live-follow loop. `ApiClient` keeps its
/// transport private (rightly -- nothing outside it should reach past
/// the typed calls), so a caller that also needs the raw
/// `Transport::stream` for SSE, as this one does, builds its own
/// rather than the crate growing a getter whose only purpose would be
/// letting one caller reach around its own abstraction.
stream_transport: Arc<UreqTransport>,
proxy: Proxy<AppEvent>,
sessions: Vec<SessionSummary>,
selected: Option<String>,
items: Vec<TranscriptItem>,
list_ptr: WeakWidget<WidgetPtr>,
transcript_ptr: WeakWidget<WidgetPtr>,
screen: Option<transcript_ui::TranscriptScreen>,
/// Bumped every time the selected session changes; see `AppEvent`'s
/// doc for what it guards against.
generation: Arc<AtomicU64>,
}
impl DefaultAppState for Client {
type Event = AppEvent;
fn new(
mut ui_state: DefaultUiState,
rsc: &mut DefaultRsc<Self>,
proxy: Proxy<AppEvent>,
) -> Self {
// Re-validated here rather than threaded through from `main` --
// `DefaultApp::run()` takes no payload, so there is no other way
// to get `main`'s parsed CLI/config into this constructor. `main`
// already called this once to fail fast before a window opens;
// this call only fails if the filesystem changed underneath the
// process in between, which is not a case worth a nicer message.
let (server, ca_pem) = crate::load_startup_config().unwrap_or_else(|e| {
eprintln!("desktop-app: {e}");
std::process::exit(2);
});
let build_transport =
|| UreqTransport::new(server.base_url(), server.token.clone(), &ca_pem);
let (rest_transport, stream_transport) = build_transport()
.and_then(|rest| build_transport().map(|stream| (rest, stream)))
.unwrap_or_else(|e| {
eprintln!(
"desktop-app: couldn't set up TLS to {}: {e}",
server.base_url()
);
std::process::exit(1);
});
let api = Arc::new(ApiClient::new(rest_transport));
let stream_transport = Arc::new(stream_transport);
let list_ptr = WidgetPtr::new().add(rsc);
let transcript_ptr = WidgetPtr::new().add(rsc);
let loading = placeholder(rsc, "Loading sessions...");
transcript_ptr(rsc).set(loading);
(list_ptr.width(LIST_WIDTH), transcript_ptr.width(rest(1)))
.span(Dir::RIGHT)
.set_root(rsc, &mut ui_state);
let client = Self {
ui_state,
api,
stream_transport,
proxy,
sessions: Vec::new(),
selected: None,
items: Vec::new(),
list_ptr,
transcript_ptr,
screen: None,
generation: Arc::new(AtomicU64::new(0)),
};
client.spawn_fetch_sessions();
client
}
fn event(&mut self, event: AppEvent, rsc: &mut DefaultRsc<Self>, _render: &mut UiRenderState) {
match event {
AppEvent::Sessions(Ok(sessions)) => {
self.sessions = sessions;
self.rebuild_list(rsc);
if self.selected.is_none() {
self.show_message(rsc, "Select a session.");
}
}
AppEvent::Sessions(Err(message)) => {
self.show_message(rsc, &format!("Couldn't list sessions: {message}"));
}
AppEvent::TranscriptLoaded {
session_id,
generation,
result,
} => {
if self.current(&session_id, generation) {
match result {
Ok(items) => {
self.items = items;
self.rebuild_transcript(rsc);
}
Err(message) => {
self.show_message(
rsc,
&format!("Couldn't load {session_id}: {message}"),
);
}
}
}
}
AppEvent::StreamEvent {
session_id,
generation,
event,
} => {
if self.current(&session_id, generation) {
self.items = fold_event(&self.items, &event);
self.rebuild_transcript(rsc);
}
}
AppEvent::StreamEnded {
session_id,
generation,
message: Some(message),
} => {
if self.current(&session_id, generation) {
eprintln!("desktop-app: {session_id}'s live connection ended: {message}");
}
}
AppEvent::StreamEnded { .. } => {}
AppEvent::SendFailed(message) => {
eprintln!("desktop-app: couldn't send: {message}");
}
}
self.ui_state.window.request_redraw();
}
}
impl Client {
fn current(&self, session_id: &str, generation: u64) -> bool {
self.selected.as_deref() == Some(session_id)
&& self.generation.load(Ordering::SeqCst) == generation
}
/// Replaces the right-hand panel with a line of text -- built before
/// `transcript_ptr` is reached for, since building the message and
/// swapping it in both need `rsc` and can't overlap as one borrow.
fn show_message(&mut self, rsc: &mut DefaultRsc<Self>, message: &str) {
let widget = placeholder(rsc, message);
(self.transcript_ptr)(rsc).set(widget);
}
fn spawn_fetch_sessions(&self) {
let api = self.api.clone();
let proxy = self.proxy.clone();
std::thread::spawn(move || {
let result = api.fetch_sessions().map_err(|e| e.to_string());
let _ = proxy.send_event(AppEvent::Sessions(result));
});
}
fn rebuild_list(&mut self, rsc: &mut DefaultRsc<Self>) {
let list = Span::empty(Dir::DOWN).gap(2).add(rsc);
for session in &self.sessions {
let selected = self.selected.as_deref() == Some(session.id.as_str());
let row = session_row(rsc, session, selected);
list(rsc).push(row);
}
let tree = list
.background(rect(Color::rgb(24, 24, 28)))
.add_strong(rsc)
.any();
(self.list_ptr)(rsc).set(tree);
}
/// Selecting a session starts a fresh generation: any thread still
/// working for the previous one checks `Client::current` before
/// touching state, so a slow response for a session the reader has
/// clicked away from is silently dropped rather than overwriting what
/// replaced it.
fn select_session(&mut self, rsc: &mut DefaultRsc<Self>, session_id: String) {
let generation = self.generation.fetch_add(1, Ordering::SeqCst) + 1;
self.selected = Some(session_id.clone());
self.items.clear();
self.screen = None;
self.rebuild_list(rsc);
self.show_message(rsc, "Loading transcript...");
let api = self.api.clone();
let stream_transport = self.stream_transport.clone();
let proxy = self.proxy.clone();
let live_generation = self.generation.clone();
std::thread::spawn(move || {
// The most recent 200 events, coalesced -- plenty for a
// desktop proof; RUST.md's I3/history-paging work is what a
// real scrollback would reuse, out of scope here (E4 is only
// "the same screen runs in a window").
let page: Result<Vec<serde_json::Value>, String> = api
.fetch_transcript_page(&session_id, None, 200, true)
.map_err(|e| e.to_string());
// The raw wire `seq` of the last line fetched -- not the seq of
// the last *folded item*. A `TranscriptItem::AssistantMsg` keeps
// the seq of the first delta it accumulated (`fold_event`'s own
// doc: "a row whose identity changed with every delta would be
// a new row every frame"), so resuming the live stream from
// that seq re-delivers every delta already folded into it,
// duplicating the tail of whatever reply was mid-stream when
// the page was fetched. Found by screenshotting a real reply
// through `run-headless.sh`: the assistant's line read "You
// said: ... testsaid: ... test", the back half being deltas 2
// through N replayed onto an already-complete message.
let after = page
.as_ref()
.ok()
.and_then(|values| raw_seq(values.last()?))
.unwrap_or(0);
let result = page.and_then(|values| fold_page(&values));
let _ = proxy.send_event(AppEvent::TranscriptLoaded {
session_id: session_id.clone(),
generation,
result,
});
// Follows live from here in the same thread -- sequential
// rather than a second thread, since there is nothing to do
// with the stream until the page above has been sent anyway.
let stop = || live_generation.load(Ordering::SeqCst) != generation;
if stop() {
return;
}
let outcome =
follow_session_events(&*stream_transport, &session_id, after, |item| match item {
StreamItem::Open | StreamItem::Reset => !stop(),
StreamItem::Event { event, .. } => {
if stop() {
return false;
}
let _ = proxy.send_event(AppEvent::StreamEvent {
session_id: session_id.clone(),
generation,
event,
});
true
}
});
let _ = proxy.send_event(AppEvent::StreamEnded {
session_id,
generation,
message: outcome.err().map(|e| e.to_string()),
});
});
}
fn send_message(&mut self, session_id: String, text: String) {
let api = self.api.clone();
let proxy = self.proxy.clone();
std::thread::spawn(move || {
if let Err(e) = api.send_message(&session_id, &text, &[]) {
let _ = proxy.send_event(AppEvent::SendFailed(e.to_string()));
}
});
}
fn rebuild_transcript(&mut self, rsc: &mut DefaultRsc<Self>) {
let in_progress = self
.screen
.as_ref()
.map(|screen| screen.composer.field.edit(rsc).text.text().to_string())
.filter(|t| !t.is_empty());
let rows = group_tool_runs(&self.items);
let (screen, tree) = transcript_ui::build_tree(rsc, rows);
if let Some(text) = in_progress {
screen.composer.field.edit(rsc).set(&text);
}
if let Some(session_id) = self.selected.clone() {
let field = screen.composer.field;
rsc.register_event(field, Submit, move |ctx, rsc| {
let text = field.edit(rsc).take();
let text = text.trim().to_string();
if !text.is_empty() {
ctx.state.send_message(session_id.clone(), text);
}
});
}
(self.transcript_ptr)(rsc).set(tree);
self.screen = Some(screen);
}
}
/// One row in the session list: title on top, status below, highlighted
/// when it's the one currently shown.
fn session_row(
rsc: &mut DefaultRsc<Client>,
session: &SessionSummary,
selected: bool,
) -> StrongWidget {
let bg = if selected {
Color::rgb(58, 90, 138)
} else {
Color::rgb(38, 38, 44)
};
let id = session.id.clone();
let label = format!("{}\n{}", session.title, session.status);
wtext(label)
.color(Color::WHITE)
.wrap(true)
.pad(10)
.width(rest(1))
.background(rect(bg))
.on(
CursorSense::click(),
move |ctx, rsc: &mut DefaultRsc<Client>| {
ctx.state.select_session(rsc, id.clone());
},
)
.add_strong(rsc)
.any()
}
fn placeholder(rsc: &mut DefaultRsc<Client>, message: &str) -> StrongWidget {
wtext(message.to_string())
.color(Color::WHITE)
.wrap(true)
.pad(16)
.add_strong(rsc)
.any()
}
+134
View File
@@ -0,0 +1,134 @@
//! Where the desktop app keeps the enrollment it should not have to be
//! told about a second time: `client_core::config::EnrolledServer`,
//! persisted at `$XDG_CONFIG_HOME/ai-app-desktop/enrollment.json`,
//! owner-only (0600) -- MACHINE.md's rule for anything holding a bearer
//! token, and the reason `client_core::config`'s own doc comment leaves
//! persistence and file mode to the caller.
//!
//! JSON rather than the project's usual RON: `wg-app-link`'s RON house
//! rules (`format`) are for configs a person hand-edits, and this file
//! never is one -- only this program ever writes or reads it, and
//! `serde_json` is already in the dependency graph through `client-core`,
//! so nothing new is added to reach for it.
use client_core::config::EnrolledServer;
use std::io;
use std::path::{Path, PathBuf};
/// `$XDG_CONFIG_HOME/ai-app-desktop`, falling back to `~/.config` the way
/// the XDG basedir spec says to when the variable is unset -- the same
/// fallback `wg_app_link::xdg::config_home` uses, reimplemented here
/// rather than depended on: that helper lives in the `wg-app-link`
/// submodule, which `server/` needs but this desktop-only crate does not,
/// and pulling in a git submodule for one path join would cost more than
/// it saves.
pub fn config_dir() -> PathBuf {
let base = std::env::var_os("XDG_CONFIG_HOME")
.map(PathBuf::from)
.unwrap_or_else(|| {
let home = std::env::var_os("HOME").expect("HOME must be set");
PathBuf::from(home).join(".config")
});
base.join("ai-app-desktop")
}
fn enrollment_file(dir: &Path) -> PathBuf {
dir.join("enrollment.json")
}
/// Persists `server` under `dir` (`config_dir()` for real use; a tempdir in
/// the tests below), creating it if needed, and sets the file owner-only --
/// it carries a bearer token, the same reason `server/`'s own token store
/// is 0600.
pub fn save_enrollment_in(dir: &Path, server: &EnrolledServer) -> io::Result<()> {
std::fs::create_dir_all(dir)?;
let path = enrollment_file(dir);
let json = serde_json::to_vec_pretty(server)
.expect("EnrolledServer holds nothing that fails to serialise");
std::fs::write(&path, json)?;
#[cfg(unix)]
{
use std::os::unix::fs::PermissionsExt;
std::fs::set_permissions(&path, std::fs::Permissions::from_mode(0o600))?;
}
Ok(())
}
/// `Ok(None)` when nothing has been enrolled yet, rather than an error --
/// "not enrolled" is an ordinary first-run state, not a failure (UI_RULES'
/// "a deliberate choice is not a problem to report" applies just as well
/// to a file that simply hasn't been written yet).
pub fn load_enrollment_in(dir: &Path) -> io::Result<Option<EnrolledServer>> {
let path = enrollment_file(dir);
match std::fs::read(&path) {
Ok(bytes) => {
let server = serde_json::from_slice(&bytes).map_err(|e| {
io::Error::new(
io::ErrorKind::InvalidData,
format!("{} is not a valid enrollment ({e})", path.display()),
)
})?;
Ok(Some(server))
}
Err(e) if e.kind() == io::ErrorKind::NotFound => Ok(None),
Err(e) => Err(e),
}
}
pub fn save_enrollment(server: &EnrolledServer) -> io::Result<()> {
save_enrollment_in(&config_dir(), server)
}
pub fn load_enrollment() -> io::Result<Option<EnrolledServer>> {
load_enrollment_in(&config_dir())
}
#[cfg(test)]
mod tests {
use super::*;
#[test]
fn a_saved_enrollment_reads_back_the_same() {
let dir = tempfile::tempdir().unwrap();
let server = EnrolledServer {
host: "127.0.0.1".to_string(),
port: 8547,
token: "tok".to_string(),
};
save_enrollment_in(dir.path(), &server).unwrap();
let read_back = load_enrollment_in(dir.path()).unwrap();
assert_eq!(read_back, Some(server));
}
#[test]
fn nothing_saved_yet_is_none_not_an_error() {
let dir = tempfile::tempdir().unwrap();
assert_eq!(load_enrollment_in(dir.path()).unwrap(), None);
}
#[test]
#[cfg(unix)]
fn the_saved_file_is_owner_only() {
use std::os::unix::fs::PermissionsExt;
let dir = tempfile::tempdir().unwrap();
let server = EnrolledServer {
host: "h".to_string(),
port: 1,
token: "t".to_string(),
};
save_enrollment_in(dir.path(), &server).unwrap();
let mode = std::fs::metadata(enrollment_file(dir.path()))
.unwrap()
.permissions()
.mode();
assert_eq!(mode & 0o777, 0o600);
}
#[test]
fn a_corrupt_file_is_named_in_the_error() {
let dir = tempfile::tempdir().unwrap();
std::fs::write(enrollment_file(dir.path()), b"not json").unwrap();
let err = load_enrollment_in(dir.path()).unwrap_err();
assert!(err.to_string().contains("enrollment.json"));
}
}
+95
View File
@@ -0,0 +1,95 @@
//! RUST.md's E4: the transcript screen (`transcript-ui`, I5) in a real
//! winit window on the desktop, with a session list beside it, talking to
//! a real `ai-server` over `client-core`'s REST + SSE clients. See
//! `app.rs`'s module doc for the widget tree and the event flow.
//!
//! Usage:
//!
//! desktop-app --ca /path/to/ca.pem --link 'aiapp://enroll?host=H&port=P&token=T'
//! desktop-app --ca /path/to/ca.pem # after the first run above
//!
//! `--link` is the same text `app/ui-sandbox.sh`'s banner prints and a
//! phone would scan as a QR (DECISIONS.md, 2026-09-05) -- pasted rather
//! than scanned, since a desktop has no camera to assume. It is parsed and
//! saved to `config::save_enrollment` once; later runs read it back and
//! `--link` is only needed again to enrol against a different server. The
//! CA is never persisted -- it is a public certificate whose path a
//! caller is expected to already know (`AGENTS.md`'s "prefer exercising
//! the server directly": the same `certs/ca.pem` a `curl --cacert` call
//! uses).
mod app;
mod config;
use client_core::config::EnrolledServer;
struct Args {
ca_path: std::path::PathBuf,
link: Option<String>,
}
fn parse_args() -> Result<Args, String> {
let mut ca_path = None;
let mut link = None;
let mut args = std::env::args().skip(1);
while let Some(arg) = args.next() {
match arg.as_str() {
"--ca" => {
ca_path = Some(std::path::PathBuf::from(
args.next().ok_or("--ca needs a path")?,
))
}
"--link" => link = Some(args.next().ok_or("--link needs a value")?),
other => return Err(format!("unrecognised argument '{other}'")),
}
}
Ok(Args {
ca_path: ca_path.ok_or(
"--ca PATH is required (the pinned CA's certificate, e.g. \
~/.config/ai-app/certs/ca.pem)",
)?,
link,
})
}
/// What `app.rs`'s `Client::new` needs to talk to the server: the enrolled
/// server (freshly parsed from `--link`, or read back from last time) and
/// the CA's PEM bytes. Loading is a pure function of the process's own
/// argv and config file, so it is safe to call again from `Client::new` --
/// see that call site's comment for why it is not threaded through some
/// other way (`DefaultApp::run()` takes no payload).
fn load_startup_config() -> Result<(EnrolledServer, Vec<u8>), String> {
let args = parse_args()?;
let server = match args.link {
Some(link) => {
let server = EnrolledServer::parse_link(&link)?;
config::save_enrollment(&server)
.map_err(|e| format!("couldn't save the enrollment: {e}"))?;
server
}
None => config::load_enrollment()
.map_err(|e| format!("couldn't read the saved enrollment: {e}"))?
.ok_or_else(|| {
format!(
"no server enrolled yet under {} -- pass --link 'aiapp://enroll?...' \
once (app/ui-sandbox.sh's start banner prints one)",
config::config_dir().display()
)
})?,
};
let ca_pem = std::fs::read(&args.ca_path)
.map_err(|e| format!("couldn't read the CA at {}: {e}", args.ca_path.display()))?;
Ok((server, ca_pem))
}
fn main() {
// Validated once here so a bad `--ca`/`--link` is reported on stderr
// before any window opens; `Client::new` calls this same function
// again once the window exists, so this first call is a fast-fail
// rather than the only place the values come from.
if let Err(e) = load_startup_config() {
eprintln!("desktop-app: {e}");
std::process::exit(2);
}
app::run();
}
+23 -4
View File
@@ -4,6 +4,16 @@
# ./run-headless.sh tabs [-- cargo args]
# ./run-headless.sh tabs --shot /tmp/tabs.png --seconds 4
#
# `--bin` runs a real crate binary instead of an example (E4's
# `desktop-app`, which is a window a person runs, not a demo) --
# `cargo build --bin NAME` instead of `--example NAME`, and
# `target/debug/NAME` instead of `target/debug/examples/NAME`. Its own
# argv (the CLI flags a real binary takes, as opposed to `cargo build`'s
# own flags after `--`) comes through `$RUN_HEADLESS_ARGS`, word-split on
# purpose -- an example never needed one, so there was nowhere to plumb it
# through positionally without disturbing the existing `-- cargo args`
# convention above.
#
# The VM has a virtio-gpu render node (Vulkan 1.4 through Venus, GL 4.6
# through virgl), so wgpu runs on the host's real GPU -- what is missing is
# only a compositor to give winit a surface. So: a headless sway, the same
@@ -20,16 +30,18 @@ run="${XDG_RUNTIME_DIR:-/tmp}/iris-headless"
seconds=3
shot=""
example=""
kind=example
while [ $# -gt 0 ]; do
case "$1" in
--shot) shot=$2; shift 2 ;;
--seconds) seconds=$2; shift 2 ;;
--bin) kind=bin; shift ;;
--) shift; break ;;
*) example=$1; shift ;;
esac
done
[ -n "$example" ] || { echo "usage: $0 EXAMPLE [--shot PNG] [--seconds N] [-- cargo args]" >&2; exit 2; }
[ -n "$example" ] || { echo "usage: $0 NAME [--bin] [--shot PNG] [--seconds N] [-- cargo args]" >&2; exit 2; }
mkdir -p "$run"
export SWAYSOCK="$run/sway.sock"
@@ -67,10 +79,17 @@ export WAYLAND_DISPLAY
echo "run-headless: $WAYLAND_DISPLAY (sway $(swaymsg -t get_version --raw | sed -n 's/.*"human_readable":"\([^"]*\)".*/\1/p'))" >&2
cd "$here"
cargo build --example "$example" "$@" >&2
bin="$here/target/debug/examples/$example"
if [ "$kind" = bin ]; then
cargo build --bin "$example" "$@" >&2
bin="$here/target/debug/$example"
else
cargo build --example "$example" "$@" >&2
bin="$here/target/debug/examples/$example"
fi
"$bin" >"$run/$example.log" 2>&1 &
# shellcheck disable=SC2086 -- deliberately word-split: this is the
# binary's own argv, not a single path.
"$bin" ${RUN_HEADLESS_ARGS:-} >"$run/$example.log" 2>&1 &
pid=$!
trap 'kill "$pid" 2>/dev/null || true' EXIT INT TERM
+47 -10
View File
@@ -6,6 +6,7 @@ use android_view::{
};
use iris_core::{UiData, UiRenderNode, UiRenderState};
use pollster::FutureExt;
use std::time::{Duration, Instant};
use wgpu::{
rwh::{DisplayHandle, HandleError, HasDisplayHandle, HasWindowHandle, WindowHandle},
*,
@@ -51,8 +52,19 @@ pub struct AndroidRenderer {
impl AndroidRenderer {
pub fn new(window: NativeWindow, width: u32, height: u32) -> Self {
// `force-gles` (RUST.md's I5 "Where iris's frame time goes") swaps
// the software-Vulkan (SwiftShader) path for GLES/virgl on the same
// build, to isolate whether the backend itself explains the frame
// time gap against Compose. `cfg!` rather than a runtime switch:
// there is no way to hand an env var to an already-launched Android
// process on this machine (see the feature's doc in Cargo.toml).
let backends = if cfg!(feature = "force-gles") {
Backends::GL
} else {
Backends::PRIMARY
};
let instance = Instance::new(&InstanceDescriptor {
backends: Backends::PRIMARY,
backends,
..Default::default()
});
@@ -74,12 +86,11 @@ impl AndroidRenderer {
// Same request as the winit backend's `UiRenderer::new` -- no
// binding-array features, see TEXTURES.md's "Recommended shape".
// `iris_core::device_limits()` is shared between the two backends;
// see its own doc for why it is not simply `Limits::default()`.
let (device, queue) = adapter
.request_device(&DeviceDescriptor {
required_limits: Limits {
max_buffer_size: 1 << 30,
..Default::default()
},
required_limits: iris_core::device_limits(),
..Default::default()
})
.block_on()
@@ -128,7 +139,18 @@ impl AndroidRenderer {
self.ui.update(&self.device, &self.queue, ui, render);
}
pub fn draw(&mut self) {
/// Draws and presents one frame, returning the time spent in
/// `queue.submit` plus `present()` -- wherever a driver/GPU/compositor
/// wait would actually show up. The caller (`android::view::render`)
/// already times the whole frame from its own `redraw_to_submit` start;
/// subtracting this from that total is `redraw_to_submit` itself
/// (layout, text, primitive building, and this method's own render-pass
/// recording). RUST.md's I5 "Where iris's frame time goes" diagnosis,
/// added 2026-09-05 -- see `iris_core::FrameReport::record_split`'s own
/// doc for the caveat this shares: `present()` is not fenced against
/// the GPU actually finishing, so this is "how long the CPU was blocked
/// handing the frame off", not confirmed GPU time.
pub fn draw(&mut self) -> Duration {
let output = self.surface.get_current_texture().unwrap();
let view = output
.texture
@@ -151,8 +173,10 @@ impl AndroidRenderer {
self.ui.draw(render_pass);
}
let submit_start = Instant::now();
self.queue.submit(std::iter::once(encoder.finish()));
output.present();
submit_start.elapsed()
}
pub fn size(&self) -> iris_core::util::Vec2 {
@@ -169,9 +193,22 @@ impl AndroidRenderer {
/// `Tasks`' redraw handle on Android: a background task finishes on the
/// tokio thread `Tasks::init` spawned, which is not attached to the JVM, so
/// asking for a frame means attaching first. `post_frame_callback` needs a
/// live `View` reference; the global ref is what survives past the JNI call
/// that handed it to us.
/// asking for a frame means attaching first. The global ref is what
/// survives past the JNI call that handed the `View` to us.
///
/// **Goes through `View::post_delayed`, not `post_frame_callback`
/// directly** -- found the hard way (RUST.md's I5 Android integration):
/// `post_frame_callback`'s Java side calls `Choreographer.getInstance()`,
/// which throws `IllegalStateException` unless the *calling* thread already
/// has a `Looper` (`Choreographer.getInstance()`'s own contract). A tokio
/// worker thread, even freshly attached to the JVM, has none -- the crash
/// was a `JavaException` inside `View::post_frame_callback`'s `.unwrap()`,
/// aborting the process on the second `redraw.request_redraw()` any
/// android transcript-screen fetch made. `View.postDelayed(Runnable, 0)`
/// is the ordinary Android answer to "queue work onto a View's own UI
/// thread from any thread" and needs no Looper of its own; `delayed_callback`
/// below is what that Runnable resolves to on the UI thread, where a real
/// `post_frame_callback` is safe again.
pub struct AndroidRedrawHandle {
vm: JavaVM,
view: GlobalRef,
@@ -189,6 +226,6 @@ impl RequestRedraw for AndroidRedrawHandle {
return;
};
let local = env.new_local_ref(&self.view).unwrap();
View(local).post_frame_callback(&mut env);
View(local).post_delayed(&mut env, 0);
}
}
+50 -3
View File
@@ -4,7 +4,7 @@ use accesskit_android::Adapter as AccessAdapter;
use android_view::{
AccessibilityNodeInfo, AccessibilityNodeProvider, Bundle, CallbackCtx, Context,
InputConnection, KeyEvent, MotionEvent, Rect, View, ViewPeer,
jni::{JNIEnv, sys::jint},
jni::{JNIEnv, JavaVM, objects::GlobalRef, sys::jint},
ndk::event::{Keycode, MotionAction},
};
// `marker::Sized` explicitly: `crate::prelude::*` below also brings in the
@@ -57,6 +57,11 @@ pub struct AndroidUiState {
/// The AccessKit tree itself -- see `iris_core::AccessTree`'s doc
/// comment.
pub access: AccessTree,
/// iris's own frame-time report (RUST.md's I5 box, "Measurements
/// taken" (b)) -- `render()` below records into it once per frame,
/// because `dumpsys gfxinfo` cannot see a `SurfaceView`'s own
/// GPU-drawn frames at all. See `iris_core::FrameReport`'s own doc.
pub frame_report: FrameReport,
}
impl AndroidUiState {
@@ -72,6 +77,7 @@ impl AndroidUiState {
shared,
access_adapter: Default::default(),
access: AccessTree::new(),
frame_report: FrameReport::new(),
}
}
@@ -101,6 +107,19 @@ pub trait AndroidAppState: HasAndroidUiState {
fn back_pressed(&mut self, rsc: &mut AndroidRsc<Self>, render: &mut UiRenderState) -> bool {
false
}
/// Called once, right after `new`, with a fresh `JavaVM` handle and a
/// global reference to this app's own `View` -- for a caller that
/// needs to call into Java itself beyond what a [`RequestRedraw`]
/// handle already covers (P0's bench build calling
/// `BatteryManager`/`ClipboardManager` through the view's `Context`,
/// docs/RUST.md). Not folded into `new` itself: most implementors need
/// nothing here, and `new`'s job is building the widget tree, not
/// holding a platform handle -- the default does nothing. `vm`/`view`
/// are independent handles from the ones `new_peer` keeps for its own
/// `RequestRedraw` (a fresh `get_java_vm`/`new_global_ref` each), so
/// storing them has no effect on that mechanism.
#[allow(unused_variables)]
fn platform_ready(&mut self, rsc: &mut AndroidRsc<Self>, vm: JavaVM, view: GlobalRef) {}
}
/// The android-view analogue of `default::DefaultRsc` -- identical in
@@ -262,6 +281,13 @@ impl<State: AndroidAppState> IrisViewPeer<State> {
.and_then(|r| self.render.window_region(r, &self.rsc)),
self.window_size(),
);
// iris's own frame-time report (RUST.md's I5 box, "Measurements
// taken" (b)): started here, at the same point a redraw request
// fires, and stopped after `renderer.draw()`'s `queue.submit` +
// `present()` -- the span Compose's render report and `gfxinfo`
// both count. See `iris_core::FrameReport`'s own doc for exactly
// what this does and does not measure.
let frame_start = Instant::now();
let ui_state = self.state.android_state_mut();
self.render.update(&ui_state.root, &mut self.rsc);
let ui_state = self.state.android_state_mut();
@@ -269,7 +295,11 @@ impl<State: AndroidAppState> IrisViewPeer<State> {
return;
};
renderer.update(&mut self.rsc.ui, &mut self.render);
renderer.draw();
let submit_to_present = renderer.draw();
self.state
.android_state_mut()
.frame_report
.record_split(frame_start.elapsed(), submit_to_present);
let ui_state = self.state.android_state();
log::debug!(
"render(): after update active={} root_px={:?}",
@@ -432,6 +462,20 @@ impl<State: AndroidAppState> ViewPeer for IrisViewPeer<State> {
self.render(ctx);
}
/// Where `AndroidRedrawHandle::request_redraw` (`android/render.rs`)
/// actually lands: `View.postDelayed`'s Runnable resolves to this, on
/// the UI thread, which is what makes it safe to call from a background
/// task's own thread when `post_frame_callback`'s `Choreographer`
/// requirement (a `Looper` on the *calling* thread) is not. Same body
/// as `do_frame` -- draining tasks and rendering immediately is a
/// perfectly good answer to "a background fetch has new state," and
/// avoids a second frame-scheduling path to keep in sync with the real
/// one.
fn delayed_callback(&mut self, ctx: &mut CallbackCtx) {
self.drain_tasks();
self.render(ctx);
}
fn as_input_connection(&mut self) -> Option<&mut dyn InputConnection> {
Some(self)
}
@@ -530,7 +574,10 @@ pub fn new_peer<'local, State: AndroidAppState>(
};
let shared = Rc::new(RefCell::new(Shared::default()));
let ui_state = AndroidUiState::new(shared.clone());
let state = State::new(ui_state, &mut rsc);
let mut state = State::new(ui_state, &mut rsc);
let platform_vm = env.get_java_vm().unwrap();
let platform_view = env.new_global_ref(&view.0).unwrap();
state.platform_ready(&mut rsc, platform_vm, platform_view);
let peer = IrisViewPeer {
rsc,
render: UiRenderState::new(),
+4 -5
View File
@@ -102,13 +102,12 @@ impl UiRenderer {
// needs descriptor indexing. See TEXTURES.md's "Recommended shape"
// for why the old binding array asked for
// VK_EXT_descriptor_indexing unconditionally and did not survive a
// real share of Android GPUs.
// real share of Android GPUs. `iris_core::device_limits()` is
// shared with the Android backend; see its own doc for why it is
// not simply `Limits::default()`.
let (device, queue) = adapter
.request_device(&DeviceDescriptor {
required_limits: Limits {
max_buffer_size: 1 << 30,
..Default::default()
},
required_limits: iris_core::device_limits(),
..Default::default()
})
.block_on()
+306
View File
@@ -2,6 +2,7 @@ use crate::prelude::*;
use std::{
ops::{BitOr, Deref, DerefMut},
rc::Rc,
time::{Duration, Instant},
};
#[derive(Clone, Copy, PartialEq)]
@@ -357,3 +358,308 @@ impl BitOr<CursorSense> for CursorSenses {
self
}
}
/// How long a stationary press has to be held before it is treated as a
/// long-press rather than the start of a pan.
pub const LONG_PRESS: Duration = Duration::from_millis(500);
/// How far a press has to move, in pixels, before it counts as a drag
/// rather than jitter -- for both the pan-vs-select axis test and the
/// "did this actually move" long-press guard.
pub const DRAG_SLOP: f32 = 8.0;
/// What a [`DragArbiter`] decided a frame's drag should mean. `Undecided`
/// means neither a pan nor a selection has committed yet, so the caller
/// should do nothing observable this frame.
#[derive(Debug, Clone, Copy, PartialEq)]
pub enum DragOutcome {
Undecided,
/// Scroll the enclosing list by this many window-space pixels along
/// the drag axis (the delta since the arbiter's last decided frame).
Pan(f32),
/// A selection should begin at the arbiter's press origin.
SelectStart,
/// A selection already underway should extend to the current position.
SelectExtend,
}
#[derive(Clone, Copy, PartialEq)]
enum ArbiterState {
Idle,
Undecided { already_selected: bool },
Panning,
Selecting,
}
/// Decides, one shared instance per gesture surface (a transcript's whole
/// row list here), whether a touch drag that starts on a row's own
/// selectable text is panning the list or extending a text selection --
/// RUST.md's I5 finding that both wanted the same `CursorSense::
/// click_or_drag()` gesture, with the inner text layer winning every frame
/// regardless of which one the reader meant. Decided the way Android
/// itself decides it, so a reader's existing muscle memory carries over:
///
/// - An ordinary vertical drag pans -- checked first, and immediately,
/// so a swipe never waits on the long-press timer.
/// - A stationary press held past [`LONG_PRESS`] starts a selection.
/// Every drag frame after that extends it, whichever direction it goes.
/// - A drag that starts **horizontally** while something is already
/// selected extends that selection right away, skipping the long-press
/// wait -- the "drag the selection handle" gesture a reader reaches for
/// once text is already highlighted.
///
/// Pure state, no rendering or widget access, so it is unit-testable
/// exactly like the rest of this module (`sense_tests.rs`'s style) with a
/// caller-supplied `Instant` rather than a real clock.
pub struct DragArbiter {
state: ArbiterState,
origin: Vec2,
origin_at: Instant,
last: Vec2,
}
impl Default for DragArbiter {
fn default() -> Self {
Self {
state: ArbiterState::Idle,
origin: Vec2::ZERO,
origin_at: Instant::now(),
last: Vec2::ZERO,
}
}
}
impl DragArbiter {
pub fn new() -> Self {
Self::default()
}
/// A fresh press-down at `pos`. `already_selected` is whatever the
/// caller's selection state was *before* this press -- it decides
/// whether an early horizontal move extends that selection instead of
/// waiting for a long-press.
pub fn press_start(&mut self, pos: Vec2, now: Instant, already_selected: bool) {
self.origin = pos;
self.origin_at = now;
self.last = pos;
self.state = ArbiterState::Undecided { already_selected };
}
/// Whether this arbiter has no press in flight -- either it has never
/// seen [`Self::press_start`], or the last one it saw was released.
/// What a caller whose own `PressStart` sense can miss (see
/// [`Self::update`]'s doc) uses to notice that a `Pressing` frame has
/// arrived with no matching start, and recover by starting one now.
pub fn is_idle(&self) -> bool {
matches!(self.state, ArbiterState::Idle)
}
/// The press continues (still down) at `pos`. Call once per frame
/// while the button/finger is down; returns what this frame means.
///
/// A caller must not call this while [`Self::is_idle`] is true for a
/// press that is genuinely still down -- `Idle` has no way to tell
/// "no press is happening" from "a press is happening but this
/// arbiter never got its `press_start`," so it always answers
/// `Undecided` and never leaves `Idle` on its own. That second case is
/// real: a touch's `ACTION_DOWN` lands whatever pixel the finger
/// actually hit, which is not guaranteed to be inside the same
/// row-local sensor region a later `ACTION_MOVE` in the same gesture
/// lands in (a row's own padding/gap, or its non-selectable header, is
/// pointer-transparent to `CursorSense`) -- so the widget that
/// receives the gesture's first `Pressing` frame may never have seen
/// its `PressStart`. `iris::transcript_ui::selection::Selection::drag`
/// is the caller that recovers from this, via `is_idle`.
pub fn update(&mut self, pos: Vec2, now: Instant) -> DragOutcome {
match self.state {
ArbiterState::Idle => DragOutcome::Undecided,
ArbiterState::Panning => {
let dy = pos.y - self.last.y;
self.last = pos;
DragOutcome::Pan(dy)
}
ArbiterState::Selecting => {
self.last = pos;
DragOutcome::SelectExtend
}
ArbiterState::Undecided { already_selected } => {
let dx = pos.x - self.origin.x;
let dy = pos.y - self.origin.y;
if already_selected && dx.abs() > DRAG_SLOP && dx.abs() > dy.abs() {
self.state = ArbiterState::Selecting;
self.last = pos;
DragOutcome::SelectExtend
} else if dy.abs() > DRAG_SLOP && dy.abs() >= dx.abs() {
self.state = ArbiterState::Panning;
self.last = pos;
DragOutcome::Pan(dy)
} else if now.duration_since(self.origin_at) >= LONG_PRESS
&& dx.abs() <= DRAG_SLOP
&& dy.abs() <= DRAG_SLOP
{
self.state = ArbiterState::Selecting;
self.last = pos;
DragOutcome::SelectStart
} else {
DragOutcome::Undecided
}
}
}
}
/// The press was released -- back to idle for the next one.
pub fn release(&mut self) {
self.state = ArbiterState::Idle;
}
}
#[cfg(test)]
mod drag_arbiter_tests {
use super::*;
fn t(ms: u64) -> Instant {
// A fixed base plus an offset, rather than `Instant::now()` per
// call -- keeps every test's timing deterministic instead of at
// the mercy of how long the test itself took to run.
Instant::now() - Duration::from_secs(3600) + Duration::from_millis(ms)
}
#[test]
fn small_jitter_stays_undecided() {
let mut a = DragArbiter::new();
a.press_start(Vec2::new(0.0, 0.0), t(0), false);
assert_eq!(a.update(Vec2::new(1.0, 1.0), t(10)), DragOutcome::Undecided);
}
#[test]
fn a_vertical_drag_pans_immediately() {
let mut a = DragArbiter::new();
a.press_start(Vec2::new(0.0, 0.0), t(0), false);
assert_eq!(
a.update(Vec2::new(0.0, 20.0), t(10)),
DragOutcome::Pan(20.0)
);
// Subsequent frames keep panning, by the delta since last frame.
assert_eq!(
a.update(Vec2::new(0.0, 35.0), t(20)),
DragOutcome::Pan(15.0)
);
}
#[test]
fn a_horizontal_drag_with_nothing_selected_does_not_select() {
let mut a = DragArbiter::new();
a.press_start(Vec2::new(0.0, 0.0), t(0), false);
// Horizontal movement alone, with no prior selection, is not any
// of the three named gestures -- it stays undecided rather than
// guessing (it will resolve to a long-press-selection if the
// finger then stops moving, or nothing if it lifts).
assert_eq!(
a.update(Vec2::new(20.0, 0.0), t(10)),
DragOutcome::Undecided
);
}
#[test]
fn a_long_press_without_moving_starts_a_selection() {
let mut a = DragArbiter::new();
a.press_start(Vec2::new(5.0, 5.0), t(0), false);
assert_eq!(a.update(Vec2::new(5.0, 5.0), t(10)), DragOutcome::Undecided);
assert_eq!(
a.update(Vec2::new(6.0, 5.0), t(LONG_PRESS.as_millis() as u64 + 1)),
DragOutcome::SelectStart
);
}
#[test]
fn after_a_long_press_any_further_drag_extends() {
let mut a = DragArbiter::new();
a.press_start(Vec2::new(0.0, 0.0), t(0), false);
assert_eq!(
a.update(Vec2::new(0.0, 0.0), t(LONG_PRESS.as_millis() as u64 + 1)),
DragOutcome::SelectStart
);
// Even a vertical move now extends the selection rather than
// panning -- once a selection has started, it owns the gesture
// until release.
assert_eq!(
a.update(Vec2::new(0.0, 40.0), t(600)),
DragOutcome::SelectExtend
);
}
#[test]
fn a_horizontal_drag_on_already_selected_text_extends_immediately() {
let mut a = DragArbiter::new();
a.press_start(Vec2::new(0.0, 0.0), t(0), true);
assert_eq!(
a.update(Vec2::new(20.0, 2.0), t(10)),
DragOutcome::SelectExtend
);
}
#[test]
fn a_vertical_drag_still_pans_even_with_a_prior_selection() {
let mut a = DragArbiter::new();
a.press_start(Vec2::new(0.0, 0.0), t(0), true);
assert_eq!(
a.update(Vec2::new(0.0, 20.0), t(10)),
DragOutcome::Pan(20.0)
);
}
/// A fresh arbiter that never saw `press_start` -- the state a widget's
/// own arbiter is left in when a gesture's `ACTION_DOWN` landed on a
/// pixel no sensor covered (a row's padding/gap, or its header) and
/// only a later `ACTION_MOVE` reached this widget. `update` must not
/// silently swallow the whole rest of the gesture here; `is_idle` is
/// what a caller checks to notice and recover (RUST.md's I5
/// intermittent-touch-scroll-dropout finding, 2026-09-05) --
/// `transcript_ui::selection::Selection::drag` is the real caller,
/// this is the pure-state half of the fix.
#[test]
fn is_idle_reports_a_press_that_was_never_started() {
let a = DragArbiter::new();
assert!(a.is_idle());
}
#[test]
fn update_on_an_idle_arbiter_stays_undecided_forever_without_recovery() {
// Documents the failure this fix works around: calling `update`
// (as if the arbiter were mid-gesture) without ever having called
// `press_start` leaves it stuck answering `Undecided`, even for a
// movement well past `DRAG_SLOP` that would otherwise pan
// immediately.
let mut a = DragArbiter::new();
assert_eq!(
a.update(Vec2::new(0.0, 100.0), t(10)),
DragOutcome::Undecided
);
assert!(a.is_idle());
}
#[test]
fn a_caller_can_recover_a_missed_press_start_via_is_idle() {
let mut a = DragArbiter::new();
// Simulates the real call site: a `Pressing` frame arrives with no
// matching `PressStart` ever having reached this arbiter.
assert!(a.is_idle());
a.press_start(Vec2::new(0.0, 700.0), t(0), false);
assert_eq!(
a.update(Vec2::new(0.0, 720.0), t(10)),
DragOutcome::Pan(20.0)
);
assert!(!a.is_idle());
}
#[test]
fn release_resets_to_idle() {
let mut a = DragArbiter::new();
a.press_start(Vec2::new(0.0, 0.0), t(0), false);
a.update(Vec2::new(0.0, 20.0), t(10));
a.release();
assert_eq!(
a.update(Vec2::new(0.0, 999.0), t(20)),
DragOutcome::Undecided
);
}
}
+12
View File
@@ -72,6 +72,18 @@ impl<Rsc: HasState> Tasks<Rsc> {
)
}
/// The same redraw handle `spawn`'s wrapper calls once, after a whole
/// task's future completes -- exposed so a caller running its own
/// longer-lived loop *inside* a spawned task (a live SSE follow, here)
/// can ask for a frame after each `TaskCtx::update`, not just at the
/// end. Without this a caller has no way to get a redraw mid-stream,
/// which is exactly the gap `iris/desktop-app`'s `app.rs` module doc
/// names for why it uses winit's `Proxy` instead of `Tasks` -- Android
/// has no `Proxy`, so this is what closes the same gap there.
pub fn redraw_handle(&self) -> Arc<dyn RequestRedraw> {
self.redraw.clone()
}
pub fn spawn<F: AsyncFnOnce(TaskCtx<Rsc>) + 'static + std::marker::Send>(&mut self, task: F)
where
F::CallOnceFuture: Send,
+16 -2
View File
@@ -4,6 +4,7 @@ use std::marker::{PhantomData, Sized};
pub struct TextBuilder<State, O = TextOutput, H: WidgetOption<State> = ()> {
pub content: String,
pub attrs: TextAttrs,
pub spans: Vec<SpanStyle>,
pub hint: H,
pub output: O,
state: PhantomData<State>,
@@ -39,10 +40,19 @@ impl<State, O, H: WidgetOption<State>> TextBuilder<State, O, H> {
self.attrs.wrap = wrap;
self
}
/// Per-range style overrides -- I5's inline rich text (bold, italic,
/// inline-code monospace, link colour/underline) within one wrapped
/// paragraph. See `SpanStyle`'s doc for why this exists and what it
/// replaces.
pub fn spans(mut self, spans: Vec<SpanStyle>) -> Self {
self.spans = spans;
self
}
pub fn editable(self, mode: EditMode) -> TextBuilder<State, TextEditOutput, H> {
TextBuilder {
content: self.content,
attrs: self.attrs,
spans: self.spans,
hint: self.hint,
output: TextEditOutput { mode },
state: PhantomData,
@@ -58,6 +68,7 @@ impl<Rsc: UiRsc, O> TextBuilder<Rsc, O> {
TextBuilder {
content: self.content,
attrs: self.attrs,
spans: self.spans,
hint: move |rsc: &mut Rsc| Some(hint.add_strong(rsc).any()),
output: self.output,
state: PhantomData,
@@ -81,7 +92,8 @@ impl<Rsc: UiRsc> TextBuilderOutput<Rsc> for TextOutput {
state: &mut Rsc,
builder: TextBuilder<Rsc, Self, H>,
) -> Self::Output {
let buf = TextBuffer::new(&builder.content);
let mut buf = TextBuffer::new(&builder.content);
buf.set_spans(builder.spans);
let hint = builder.hint.get(state);
let mut text = Text {
content: builder.content.into(),
@@ -103,7 +115,8 @@ impl<State: UiRsc> TextBuilderOutput<State> for TextEditOutput {
state: &mut State,
builder: TextBuilder<State, Self, H>,
) -> Self::Output {
let buf = TextBuffer::new(&builder.content);
let mut buf = TextBuffer::new(&builder.content);
buf.set_spans(builder.spans);
TextEdit::new(
TextView::new(buf, builder.attrs, builder.hint.get(state)),
builder.output.mode,
@@ -125,6 +138,7 @@ pub fn wtext<State>(content: impl Into<String>) -> TextBuilder<State> {
TextBuilder {
content: content.into(),
attrs: TextAttrs::default(),
spans: Vec::new(),
hint: (),
output: TextOutput,
state: PhantomData,
+26
View File
@@ -0,0 +1,26 @@
[package]
name = "transcript-ui"
version.workspace = true
edition.workspace = true
# I5 (RUST.md): the transcript screen's widget tree, built the same way
# `tabs-ui` is -- its own crate, generic over `Rsc: HasEvents` +
# `Rsc::State: FocusHost`, so the winit example and an eventual
# `iris-android-app`-style cdylib call the same `build`. See its own module
# doc for the design and RUST.md's I5 box for what is and is not proved yet.
#
# `client-core`/`event-model` by path, real code and not a reimplementation
# -- the same dependency shape E2's uncommitted Masonry experiment used for
# the identical job.
[dependencies]
iris = { path = ".." }
client-core = { path = "../../client-core" }
event-model = { path = "../../event-model" }
pulldown-cmark = { workspace = true }
# Selection has no accessibility label of its own yet (IRIS_TODO.md's
# "Row-level accessibility names"), so a logcat line at
# begin/extend is the smallest way to confirm a real long-press-then-drag
# reached `Selection` on-device (RUST.md's I5 box, "Measurements taken"
# (c)) -- version pinned to match `iris-android-app`'s own dependency.
log = "0.4.28"
+122
View File
@@ -0,0 +1,122 @@
//! I5's desktop proof: the transcript screen built from synthetic
//! `client_core::transcript_fold` rows (no network, no server -- see
//! `lib.rs`'s doc for why `transcript-ui` itself never fetches anything),
//! run via `iris/run-headless.sh transcript -- -p transcript-ui` for a
//! screenshot on the winit backend, or `cargo run --example transcript -p
//! transcript-ui` with a real compositor.
//!
//! The rows exercise every one of the seven "hard to get back" behaviours
//! this box's markdown/selection work is meant to show: a heading, bold,
//! italic, an inline code span, a link, a fenced code block (rich inline
//! text), a multi-message conversation (bottom-anchored virtualised list),
//! and a three-call tool run (collapsed by default -- tap it, or drive it
//! with `ui-trace record --do "tap 'Tools'"` on Android, to prove
//! hold-the-edge expand).
use client_core::transcript_fold::{TranscriptItem, TranscriptRow as FoldedRow};
use iris::prelude::*;
fn main() {
DefaultApp::<Client>::run();
}
#[derive(DefaultUiState)]
pub struct Client {
ui_state: DefaultUiState,
#[allow(dead_code)]
screen: transcript_ui::TranscriptScreen,
}
fn msg(seq: u64, from_user: bool, text: &str) -> FoldedRow {
FoldedRow::Single(if from_user {
TranscriptItem::UserMsg {
seq,
text: text.to_string(),
attachments: Vec::new(),
}
} else {
TranscriptItem::AssistantMsg {
seq,
text: text.to_string(),
settled: true,
}
})
}
fn synthetic_rows() -> Vec<FoldedRow> {
vec![
msg(
1,
true,
"Can you show me a **bold** word, some *italic* text, and `inline code`?",
),
msg(
2,
false,
"# Sure\n\nHere's a [link to the repo](https://example.com/ai-app-2) and a fenced block:\n\n```rust\nfn main() {\n println!(\"hi\");\n}\n```",
),
FoldedRow::Tools(vec![
TranscriptItem::ToolRun {
seq: 3,
id: "t1".into(),
run_id: "run1".into(),
tool: "Read".into(),
input: "{\"file\": \"src/main.rs\"}".into(),
output: "fn main() {}\n".into(),
done: true,
asks: Vec::new(),
images: Vec::new(),
},
TranscriptItem::ToolRun {
seq: 4,
id: "t2".into(),
run_id: "run1".into(),
tool: "Edit".into(),
input: "{\"file\": \"src/main.rs\"}".into(),
output: "ok".into(),
done: true,
asks: Vec::new(),
images: Vec::new(),
},
TranscriptItem::ToolRun {
seq: 5,
id: "t3".into(),
run_id: "run1".into(),
tool: "Bash".into(),
input: "cargo build".into(),
output: "Compiling...\nFinished.".into(),
done: true,
asks: Vec::new(),
images: Vec::new(),
},
]),
msg(6, true, "Looks good, thanks!"),
msg(
7,
false,
"You're welcome. Let me know if you'd like anything else.",
),
]
}
impl DefaultAppState for Client {
fn new(
mut ui_state: DefaultUiState,
rsc: &mut DefaultRsc<Self>,
_: Proxy<Self::Event>,
) -> Self {
let screen = transcript_ui::build(rsc, &mut ui_state, synthetic_rows());
// Exercises `push_row`/`ItemKey` beyond construction time, matching
// how a live SSE loop appends -- a row arriving after the screen
// already exists must land at the bottom without disturbing what's
// above it (I3's `push_back`/`snap_end`).
screen.push_row(
rsc,
&FoldedRow::Single(TranscriptItem::CommandRow {
seq: 8,
text: "clear".into(),
}),
);
Self { ui_state, screen }
}
}
+48
View File
@@ -0,0 +1,48 @@
//! The message composer at the bottom of the transcript screen: a
//! multi-line editable field with a natural (not fixed) height, so it
//! grows as typed into -- IRIS_TODO.md's "input box" benchmark case
//! (`iris/benches/message_list.rs` exercises the mechanism in isolation;
//! this wires the same `TextEdit`-with-no-`Sized`-wrapper idiom into the
//! real screen). `lib.rs` gives the transcript `List` `.height(rest(1))`
//! beside this widget in a `Span::down`, so the list's own draw already
//! measures whatever vertical space is left each frame -- nothing here
//! computes a height by hand, and growing this field is exactly the
//! O(1)-move-chain case LAYOUT.md and I3's benchmark already measured.
use iris::prelude::*;
/// `field` is exposed so the caller can read its content on submit
/// (`field.edit(rsc).text()`) and clear it afterward
/// (`field.edit(rsc).set("")`).
pub struct Composer {
pub field: WeakWidget<TextEdit>,
}
/// Returns the composer plus its own bar as a **weak** id -- the caller
/// (`lib.rs::build`) embeds it in the screen's own top-level tuple, whose
/// `set_root` performs the one real strong registration. Calling
/// `.add_strong`/`.upgrade` a second time on an id already strong-owned
/// panics ("was already added", `core/src/widget/like.rs:12`) -- the same
/// mistake this box's `row.rs` first made with its sender-label header, see
/// that file's comment for the fuller account.
pub fn build_composer<Rsc: HasEvents>(rsc: &mut Rsc) -> (Composer, WeakWidget)
where
Rsc::State: FocusHost,
{
let field = wtext("")
.editable(EditMode::MultiLine)
.text_align(Align::LEFT)
.wrap(true)
.size(18)
.color(UiColor::WHITE)
.attr::<Selectable>(())
.label("Message")
.add(rsc);
let bar: WeakWidget = (field.pad(12).width(rest(1)),)
.span(Dir::RIGHT)
.background(rect(UiColor::new(40, 40, 46, 255)))
.add(rsc);
(Composer { field }, bar)
}
+145
View File
@@ -0,0 +1,145 @@
//! The transcript screen, in iris -- RUST.md's I5. Built the same way
//! `tabs-ui` is: its own crate, generic over `Rsc: HasEvents` +
//! `Rsc::State: FocusHost`, so the winit example (`iris/examples/
//! transcript.rs`) and an eventual `iris-android-app`-style cdylib call the
//! same [`build`]. See RUST.md's I5 box for the full account of what is
//! and is not proved yet, and this doc for the shape.
//!
//! ```text
//! +------------------------------------------+
//! | iris::widget::List (transcript_ui::row) | <- .height(rest(1))
//! | row 1: sender label + one TextEdit |
//! | row 2: sender label + one TextEdit |
//! | row 3 (Tools): collapsed/expanded |
//! | ... |
//! +------------------------------------------+
//! | composer bar (transcript_ui::composer) | <- natural height
//! +------------------------------------------+
//! ```
//!
//! **What this crate does not do itself**: fetch anything over the network
//! or read the transcript cache. [`build`] takes an already-folded
//! `Vec<client_core::transcript_fold::TranscriptRow>` and
//! [`TranscriptScreen::push_row`] takes one more as it arrives -- the
//! caller (an app's own `main`, or a future `iris-android-app`-shaped
//! cdylib) owns `client_core::ApiClient`/
//! `event_stream::follow_session_events` and the transcript cache, per the
//! code rules' "ask for the least you need": a widget-tree builder that
//! also knew how to make an HTTPS request would be untestable without a
//! server and unable to be driven by `run-headless.sh` with synthetic rows.
//!
//! **Gap closed, 2026-09-05**: a touch-drag that starts on a row's
//! rendered text used to always begin a cross-row *selection* (`row.rs`'s
//! `CursorSense::click_or_drag()` on each row's `TextEdit`), never a
//! *scroll* of the list, because both wanted the same gesture over the
//! same screen region and `core/src/sense.rs`'s `run_sensors` gave the
//! widget in the *inner* layer (a row's own `TextEdit`) first refusal
//! every frame it was pressed. `row.rs` now routes every row's drag
//! through one shared `iris::sense::DragArbiter`
//! (`Selection::drag`, `selection.rs`), which decides pan vs. select the
//! way Android itself does -- see `DragArbiter`'s own doc and
//! `DECISIONS.md` for the exact rule. `List` scrolls correctly when
//! driven programmatically (I3's benchmark), via the mouse wheel (wired
//! below, `CursorSense::Scroll`), and now via a touch pan starting on a
//! row's own text too.
pub mod composer;
pub mod markdown;
pub mod row;
pub mod selection;
use client_core::transcript_fold::TranscriptRow as FoldedRow;
use iris::prelude::*;
use selection::Selection;
use std::{cell::RefCell, rc::Rc};
pub struct TranscriptScreen {
/// The transcript's own `List` -- exposed so a caller can read
/// `.extent()`/call `.jump_to_end()` etc. directly for anything this
/// crate does not already wrap.
pub list: WeakWidget<List>,
pub composer: composer::Composer,
selection: Rc<RefCell<Selection>>,
}
impl TranscriptScreen {
/// Append one more folded row at the live end of the transcript --
/// what a caller's SSE loop or a sent message calls as new events
/// arrive. `List::push_back` is O(1) and keeps the view pinned to the
/// newest content when it already was (I3).
pub fn push_row<Rsc: HasEvents>(&self, rsc: &mut Rsc, row: &FoldedRow)
where
Rsc::State: FocusHost,
{
let (key, widget) = row::build_row(rsc, self.list, self.selection.clone(), row);
(self.list)(rsc).push_back(ListRow::new(key, widget));
}
/// The concatenated text of whatever is currently selected across one
/// or more rows, `None` if nothing is -- what a copy command reads.
pub fn selected_text(&self, rsc: &mut impl UiRsc) -> Option<String> {
self.selection.borrow().selected_text(rsc)
}
}
pub fn build<Rsc: HasEvents>(
rsc: &mut Rsc,
ui_state: &mut impl HasRoot,
rows: Vec<FoldedRow>,
) -> TranscriptScreen
where
Rsc::State: FocusHost,
{
let (screen, tree) = build_tree(rsc, rows);
ui_state.set_root(tree);
screen
}
/// The same widget tree [`build`] makes, without claiming the window's
/// whole root -- what a caller embedding this screen alongside something
/// else of its own needs (RUST.md's E4: a session list beside the
/// transcript on the desktop). `build` is `build_tree` plus
/// `ui_state.set_root(tree)`; kept as its own function since most callers
/// (the winit example, an eventual Android cdylib) want the screen to *be*
/// the window and don't need the strong handle back.
pub fn build_tree<Rsc: HasEvents>(
rsc: &mut Rsc,
rows: Vec<FoldedRow>,
) -> (TranscriptScreen, StrongWidget)
where
Rsc::State: FocusHost,
{
let selection = Rc::new(RefCell::new(Selection::new()));
let list = List::new(Axis::Y).add(rsc);
for row in &rows {
let (key, widget) = row::build_row(rsc, list, selection.clone(), row);
list(rsc).push_back(ListRow::new(key, widget));
}
// Wheel/trackpad scrolling -- the same idiom `trait_fns.rs`'s
// `scrollable()` uses for `Scroll`, applied directly to `List` since
// `List` already does its own placement and needs no `Scroll` wrapper.
// Real touch-drag panning is the known gap in this module's doc.
list.on(CursorSense::Scroll, |ctx, rsc| {
let delta = ctx.data.scroll_delta.y * 50.0;
ctx.widget(rsc).scroll(delta);
})
.add(rsc);
let (composer, composer_bar) = composer::build_composer(rsc);
let tree = (list.width(rest(1)).height(rest(1)), composer_bar)
.span(Dir::DOWN)
.add_strong(rsc)
.any();
(
TranscriptScreen {
list,
composer,
selection,
},
tree,
)
}
+223
View File
@@ -0,0 +1,223 @@
//! Markdown -> one plain string plus a `Vec<SpanStyle>`, for I5's row
//! builder to hand to a single `TextEdit` (`row.rs`). This is the crate's
//! answer to RUST.md's E2 finding against Masonry ("rich inline text --
//! block-level yes, inline no, and both for the same reason": `TextArea`'s
//! `StyleSet` is one style for the whole editor,
//! `masonry/src/widgets/text_area.rs:43-44`'s `// TODO: RichTextInput`
//! beside it). iris's `SpanStyle` (`core/src/primitive/text.rs`, added for
//! this box) is per-range, so bold/italic/inline-code/links/headings inside
//! one wrapped paragraph render in their own style *and* the paragraph
//! still wraps and selects as one buffer -- there is no second widget per
//! span the way E2's block-level `Prose`-per-heading was.
//!
//! **What this deliberately does not attempt**, each for a reason recorded
//! here rather than silently dropped (see IRIS_TODO.md's dated entries for
//! the same list):
//! - **No background chip behind inline code.** Drawing one needs the
//! glyph run's own geometry (the way `TextEdit::draw`'s selection
//! highlight uses `selection.geometry(layout)`,
//! `iris/src/widget/text/edit.rs:99`), which is `TextEdit`-internal and
//! not exposed to a caller building spans externally. `SpanStyle` gives
//! the code range a monospace family and a dimmer text colour instead --
//! visually distinct, just not chip-shaped.
//! - **A link is styled (colour + underline) but not tappable.** Following
//! it needs the same kind of per-range hit-testing a chip's background
//! would (which byte range did the tap land in, then look up its URL),
//! which is exactly the same missing primitive.
//! - **Tables render as plain paragraphs of their cell text**, no columns.
//! `pulldown_cmark::Tag::Table` is walked but not laid out -- a real grid
//! needs its own widget, out of scope for a row builder.
//! - **A fenced code block's language is not syntax-highlighted.**
//! `client-core::highlight` exists and could feed per-token `SpanStyle`s,
//! but wiring it in is real work belonging to whoever needs it next
//! (IRIS_TODO.md).
//!
//! A heading's `SpanStyle::font_size` override does not also raise its
//! `line_height` (a buffer has one, set from the *base* font size in
//! `TextAttrs`), so a heading's own line looks slightly tighter than a
//! paragraph's -- visible, not incorrect, and not fixed here since it needs
//! `SpanStyle` to carry line-height too, which nothing in this crate needed
//! badly enough yet to justify.
use iris::prelude::*;
use pulldown_cmark::{Event, HeadingLevel, Options, Parser, Tag, TagEnd};
// `UiColor` is `Color<u8>` (`core/src/lib.rs`), not the 0..1 float triples
// its brighter/darker helpers might suggest -- these are plain 0..255 RGB.
pub const CODE_COLOR: UiColor = UiColor::new(140, 217, 242, 255);
pub const LINK_COLOR: UiColor = UiColor::new(140, 190, 255, 255);
const STRIKETHROUGH_COLOR: UiColor = UiColor::new(150, 150, 150, 255);
/// A block-level separator: two blocks never run into each other with no
/// gap, but an empty `out` (the very first block) gets no leading blank.
fn ensure_blank_line(out: &mut String) {
if !out.is_empty() && !out.ends_with("\n\n") {
out.push_str("\n\n");
}
}
fn heading_size(level: HeadingLevel) -> f32 {
match level {
HeadingLevel::H1 => 28.0,
HeadingLevel::H2 => 24.0,
HeadingLevel::H3 => 21.0,
_ => 19.0,
}
}
/// One markdown source string rendered into plain text plus the spans that
/// style it. `base_size` is the row's ordinary paragraph font size, needed
/// only so a heading's override is relative to it rather than a hardcoded
/// absolute the caller cannot retune.
pub fn render_markdown(src: &str, base_size: f32) -> (String, Vec<SpanStyle>) {
let _ = base_size; // headings use fixed sizes today; kept for callers that may want relative sizing later
let mut out = String::new();
let mut spans = Vec::new();
// Stack of start byte offsets for whatever inline/block styling is
// currently open -- pulldown-cmark's `Start`/`End` events are always
// balanced and each `End` already names its own kind (`TagEnd`), so a
// plain offset stack (rather than a tree, or repeating the kind here
// too) is enough.
let mut open: Vec<usize> = Vec::new();
let mut list_depth: u32 = 0;
let parser = Parser::new_ext(src, Options::ENABLE_STRIKETHROUGH | Options::ENABLE_TABLES);
for event in parser {
match event {
Event::Start(tag) => match tag {
Tag::Heading { .. }
| Tag::Emphasis
| Tag::Strong
| Tag::Strikethrough
| Tag::Link { .. } => open.push(out.len()),
Tag::CodeBlock(_) => {
ensure_blank_line(&mut out);
open.push(out.len());
}
Tag::Item => {
out.push_str(&" ".repeat(list_depth.saturating_sub(1) as usize));
out.push_str("\u{2022} ");
}
Tag::List(_) => list_depth += 1,
Tag::Paragraph | Tag::BlockQuote(_) => ensure_blank_line(&mut out),
_ => {}
},
// Only the tag kinds that pushed onto `open` (Start, above) are
// popped here -- `List`/`Item`/`Paragraph`/`BlockQuote`/`Table`
// and friends push nothing, since they need no span, and must
// not touch this stack or they would pop an unrelated styled
// range still open around them.
Event::End(
tag_end @ (TagEnd::Heading(_)
| TagEnd::Emphasis
| TagEnd::Strong
| TagEnd::Strikethrough
| TagEnd::Link
| TagEnd::CodeBlock),
) => {
let Some(start) = open.pop() else {
continue;
};
let range = start..out.len();
if range.is_empty() {
continue;
}
match tag_end {
TagEnd::Heading(level) => {
spans.push(SpanStyle::new(range).font_size(heading_size(level)).bold());
}
TagEnd::Emphasis => spans.push(SpanStyle::new(range).italic()),
TagEnd::Strong => spans.push(SpanStyle::new(range).bold()),
TagEnd::Strikethrough => {
spans.push(SpanStyle::new(range).color(STRIKETHROUGH_COLOR));
}
TagEnd::Link => {
spans.push(SpanStyle::new(range).color(LINK_COLOR).underline());
}
TagEnd::CodeBlock => {
spans.push(
SpanStyle::new(range)
.family(Family::Monospace)
.color(CODE_COLOR),
);
}
_ => unreachable!("filtered by the outer match arm"),
}
}
Event::Text(text) => out.push_str(&text),
// Inline code (single backticks) is one atomic event with no
// `Start`/`End` pair of its own, unlike a fenced block -- so it
// is spanned directly here instead of through the `open` stack.
Event::Code(text) => {
let start = out.len();
out.push_str(&text);
spans.push(
SpanStyle::new(start..out.len())
.family(Family::Monospace)
.color(CODE_COLOR),
);
}
Event::SoftBreak => out.push(' '),
Event::HardBreak => out.push('\n'),
Event::Rule => {
if !out.ends_with('\n') {
out.push('\n');
}
out.push_str("\u{2500}\u{2500}\u{2500}\n");
}
Event::End(TagEnd::List(_)) => list_depth = list_depth.saturating_sub(1),
_ => {}
}
}
(out, spans)
}
#[cfg(test)]
mod tests {
use super::*;
#[test]
fn plain_paragraph_has_no_spans() {
let (text, spans) = render_markdown("just some words", 16.0);
assert_eq!(text, "just some words");
assert!(spans.is_empty());
}
#[test]
fn bold_and_italic_produce_spans_over_the_right_range() {
let (text, spans) = render_markdown("a **bold** and *italic* word", 16.0);
assert_eq!(text, "a bold and italic word");
let bold = spans.iter().find(|s| s.bold && !s.italic).unwrap();
assert_eq!(&text[bold.range.clone()], "bold");
let italic = spans.iter().find(|s| s.italic).unwrap();
assert_eq!(&text[italic.range.clone()], "italic");
}
#[test]
fn heading_gets_a_bigger_font_size_span() {
let (text, spans) = render_markdown("# A Title\n\nbody text", 16.0);
assert!(text.starts_with("A Title"));
let heading = spans.iter().find(|s| s.font_size.is_some()).unwrap();
assert_eq!(&text[heading.range.clone()], "A Title");
assert_eq!(heading.font_size, Some(28.0));
}
#[test]
fn link_is_styled_and_keeps_its_visible_text() {
let (text, spans) = render_markdown("see [the docs](https://example.com) for more", 16.0);
assert!(text.contains("the docs"));
assert!(
!text.contains("example.com"),
"the URL should not leak into the visible text"
);
let link = spans.iter().find(|s| s.underline).unwrap();
assert_eq!(&text[link.range.clone()], "the docs");
}
#[test]
fn fenced_code_block_is_monospaced() {
let (text, spans) = render_markdown("before\n\n```\nlet x = 1;\n```\n\nafter", 16.0);
let code = spans.iter().find(|s| s.family.is_some()).unwrap();
assert!(text[code.range.clone()].contains("let x = 1;"));
}
}
+301
View File
@@ -0,0 +1,301 @@
//! One `iris::widget::list::ListRow` per folded transcript row
//! (`client_core::transcript_fold::TranscriptRow`). Each row's whole text
//! -- headings, paragraphs, inline styling -- goes through `markdown` into
//! **one** `TextEdit`, which is what makes it one thing `Selection`
//! (`selection.rs`) can select and what lets it wrap and scroll as a
//! single buffer, matching RUST.md's "hard to get back" behaviour 2 (rich
//! inline text) and half of behaviour 1 (selectable within a row; across
//! rows is `selection.rs`'s job).
//!
//! A `TranscriptRow::Tools` (a run of adjacent tool calls, grouped by
//! `client_core::transcript_fold::group_tool_runs`) is the row that proves
//! behaviour 3's "hold the edge nearest the tap" on expand: tapping its
//! header calls `List::note_tap` at the row's own on-screen position
//! (read back from `List::extent`, since the tap event only knows its
//! position *within* this row) before toggling a `WidgetPtr` between the
//! collapsed summary and the full detail -- the same two-step contract
//! `list.rs`'s module doc describes for `AGENTS.md`'s `holdTopEdge`.
use crate::markdown::render_markdown;
use crate::selection::Selection;
use client_core::transcript_fold::{QuestionCard, TranscriptItem, TranscriptRow as FoldedRow};
use iris::prelude::*;
use std::{cell::RefCell, rc::Rc, time::Instant};
/// The paragraph size every row's `TextEdit` is built at; markdown headings
/// inside a row scale relative to a fixed set of sizes rather than this one
/// (`markdown::heading_size`), since a heading is meant to look the same
/// regardless of which row's base size surrounds it.
pub const BASE_SIZE: f32 = 16.0;
/// `ItemKey::Seq` already is the `RowKey` (`u64`) this crate's `List` wants.
/// `ItemKey::RunId` is a string (a tool call's own id), so it is hashed into
/// one -- collisions are not a correctness risk worth guarding against here
/// (a `DefaultHasher` collision across the run ids one session produces is
/// astronomically unlikely, and the consequence of one would only be two
/// tool-call rows sharing a list slot, not data loss), and the high bit is
/// forced on so a hashed key can never collide with a real sequence number
/// (this build never produces 2^63 events).
pub fn row_key(key: &client_core::transcript_fold::ItemKey) -> RowKey {
use client_core::transcript_fold::ItemKey;
use std::hash::{Hash, Hasher};
match key {
ItemKey::Seq(seq) => *seq,
ItemKey::RunId(id) => {
let mut h = std::collections::hash_map::DefaultHasher::new();
id.hash(&mut h);
h.finish() | (1 << 63)
}
}
}
/// The sender label shown above a row's text, and the markdown source to
/// render below it. `None` for a system-style note that has no sender.
fn item_content(item: &TranscriptItem) -> (Option<&str>, String) {
match item {
TranscriptItem::UserMsg { text, .. } => (Some("You"), text.clone()),
TranscriptItem::AssistantMsg { text, .. } => (Some("Claude"), text.clone()),
TranscriptItem::ErrorMsg { message, .. } => (Some("Error"), message.clone()),
TranscriptItem::CommandRow { text, .. } => (Some("Command"), format!("`/{text}`")),
TranscriptItem::PeerNote { from, text, .. } => (Some(from.as_str()), text.clone()),
TranscriptItem::Note { text, .. } => (None, text.clone()),
TranscriptItem::ClearedNote { .. } => (None, "_Context cleared._".to_string()),
TranscriptItem::CompactedNote {
pre_tokens,
post_tokens,
..
} => (
None,
match (pre_tokens, post_tokens) {
(Some(pre), Some(post)) => format!("_Compacted: {pre} -> {post} tokens._"),
_ => "_Compacted._".to_string(),
},
),
TranscriptItem::ImageItem { r#ref, .. } => (None, format!("_[image: {ref}]_")),
TranscriptItem::QuestionCard(card) => (Some("Question"), question_markdown(card)),
TranscriptItem::ToolRun {
tool,
input,
output,
..
} => (Some(tool.as_str()), tool_call_markdown(tool, input, output)),
}
}
fn question_markdown(card: &QuestionCard) -> String {
let mut out = card.prompt.clone();
for opt in &card.options {
out.push_str(&format!("\n- {}", opt.label));
}
out
}
fn tool_call_markdown(tool: &str, input: &str, output: &str) -> String {
let mut out = format!("**{tool}**\n\n```\n{input}\n```");
if !output.is_empty() {
out.push_str(&format!("\n\n```\n{output}\n```"));
}
out
}
/// Build one `TextEdit` from a sender label plus markdown source, register
/// it with `selection` under `key`, and wire the pointer handlers that
/// drive `Selection::drag` -- shared by every row variant below, since a
/// selectable row is always "one TextEdit plus this wiring" regardless of
/// what folded it. `list` is threaded through so that same drag can pan
/// the list instead of selecting, per `Selection::drag`'s own doc.
fn build_text_row<Rsc: HasEvents>(
rsc: &mut Rsc,
list: WeakWidget<List>,
selection: Rc<RefCell<Selection>>,
key: RowKey,
sender: Option<&str>,
markdown_src: &str,
) -> StrongWidget
where
Rsc::State: FocusHost,
{
let (text, spans) = render_markdown(markdown_src, BASE_SIZE);
let field = wtext(text)
.spans(spans)
.editable(EditMode::MultiLine)
.text_align(Align::LEFT)
.wrap(true)
.size(BASE_SIZE)
.color(UiColor::WHITE)
.add(rsc);
selection.borrow_mut().register(key, field);
field
// `| CursorSense::unclick()` on top of the usual click-or-drag set
// -- the arbiter inside `Selection::drag` needs the release too,
// to go back to idle for the next press (`DragArbiter::release`).
.on(
CursorSense::click_or_drag() | CursorSense::unclick(),
move |ctx, rsc| {
selection.borrow_mut().drag(
rsc,
list,
key,
ctx.data.pos,
ctx.data.size,
ctx.data.cursor.pos,
ctx.data.sense,
Instant::now(),
);
},
)
.add(rsc);
// `.add` (weak), not `.add_strong` -- `header` is about to be embedded
// as a child of the `.span(Dir::DOWN)` below, whose own composition is
// what performs the *one* real strong registration each child gets.
// Calling `.add_strong`/`.upgrade` here too, then feeding a `.weak()`
// copy into that composition, tried to strong-register the same id
// twice and panicked with "was already added"
// (`core/src/widget/like.rs:12`) -- found running this crate's own
// `run-headless.sh` example, the first real render of a row.
let header: WeakWidget = match sender {
Some(name) => wtext(name.to_string())
.size(13.0)
.color(UiColor::new(150, 150, 160, 255))
.add(rsc),
None => Span::empty(Dir::DOWN).add(rsc),
};
(header, field.width(rest(1)))
.span(Dir::DOWN)
.gap(4)
.pad(10)
.add_strong(rsc)
.any()
}
fn build_single<Rsc: HasEvents>(
rsc: &mut Rsc,
list: WeakWidget<List>,
selection: Rc<RefCell<Selection>>,
key: RowKey,
item: &TranscriptItem,
) -> StrongWidget
where
Rsc::State: FocusHost,
{
let (sender, markdown_src) = item_content(item);
build_text_row(rsc, list, selection, key, sender, &markdown_src)
}
/// A run of adjacent tool calls: collapsed to a one-line summary by
/// default, expanding in place to every call's own tool/input/output on
/// tap -- see the module doc for the hold-the-edge contract this wires
/// against `list`.
fn build_tools<Rsc: HasEvents>(
rsc: &mut Rsc,
list: WeakWidget<List>,
selection: Rc<RefCell<Selection>>,
key: RowKey,
calls: Vec<TranscriptItem>,
) -> StrongWidget
where
Rsc::State: FocusHost,
{
let expanded = Rc::new(RefCell::new(false));
// `.add_strong` (not `.add`) because nothing else in the tree holds a
// strong reference to this `WidgetPtr` the way a container's own
// `add_strong`-on-its-children does for an ordinary child -- this row
// *is* the top of its own subtree, so it has to own itself.
let ptr_strong = WidgetPtr::new().add_strong(rsc);
let ptr = ptr_strong.weak();
let summary_text = format!("\u{25b8} {} tool calls", calls.len());
let full_text = calls
.iter()
.map(|c| match c {
TranscriptItem::ToolRun {
tool,
input,
output,
..
} => tool_call_markdown(tool, input, output),
other => item_content(other).1,
})
.collect::<Vec<_>>()
.join("\n\n");
fn build_content<Rsc: HasEvents>(
rsc: &mut Rsc,
list: WeakWidget<List>,
selection: Rc<RefCell<Selection>>,
key: RowKey,
expanded: bool,
summary: &str,
full: &str,
) -> StrongWidget
where
Rsc::State: FocusHost,
{
let text = if expanded { full } else { summary };
build_text_row(rsc, list, selection, key, Some("Tools"), text)
}
let content = build_content(
rsc,
list,
selection.clone(),
key,
false,
&summary_text,
&full_text,
);
ptr(rsc).set(content);
ptr.on(CursorSense::click(), move |ctx, rsc| {
// `List::note_tap` wants a viewport-relative position, but the
// click event only knows where inside *this row* it landed
// (`ctx.data.pos`) -- `List::extent` (last frame's on-screen box
// for this row's key) is what turns the two into the position
// `list.rs`'s hold-the-edge layout pass resolves against, per the
// module doc's contract.
let (top, _bottom) = list(rsc).extent(key).unwrap_or((0.0, 0.0));
list(rsc).note_tap(top + ctx.data.pos.y);
let was_expanded = *expanded.borrow();
*expanded.borrow_mut() = !was_expanded;
let content = build_content(
rsc,
list,
selection.clone(),
key,
!was_expanded,
&summary_text,
&full_text,
);
// The old content's `StrongWidget` is freed when this drops --
// the removal half of the row this click just replaced.
let _old = ptr(rsc).replace(content);
})
.add(rsc);
ptr_strong.any()
}
pub fn build_row<Rsc: HasEvents>(
rsc: &mut Rsc,
list: WeakWidget<List>,
selection: Rc<RefCell<Selection>>,
row: &FoldedRow,
) -> (RowKey, StrongWidget)
where
Rsc::State: FocusHost,
{
match row {
FoldedRow::Single(item) => {
let key = row_key(&item.key());
(key, build_single(rsc, list, selection, key, item))
}
FoldedRow::Tools(calls) => {
let key = row_key(&calls[0].key());
(key, build_tools(rsc, list, selection, key, calls.clone()))
}
}
}
+368
View File
@@ -0,0 +1,368 @@
//! Selection spanning multiple transcript rows -- RUST.md's "hard to get
//! back" behaviour 1, and the one E2 found flatly impossible on Masonry:
//! `TextArea` wraps exactly one `parley::PlainEditor`, and there is no
//! `SelectionContainer`-shaped type anywhere in `masonry`/`masonry_core`/
//! `xilem` (RUST.md's E2 box, citing
//! `masonry/src/widgets/text_area.rs:414-459`). Each transcript row here is
//! still its own `TextEdit` (one per row, not one per transcript, since a
//! row is what `List` virtualises), so this is not literally "one
//! `PlainEditor`" either -- iris's answer is a coordinator that drives each
//! visible row's *own* selection primitives (`TextEditCtx::select`/
//! `select_all`/`deselect`, already built for a single field) from one
//! pointer drag that crosses row boundaries, giving the same reader-facing
//! result (a selection that runs from a reply into the tool output beneath
//! it, one copy) without needing a single shared text buffer underneath.
//!
//! Rows are keyed by `RowKey` (`iris::widget::list`), which every real row
//! source (a transcript's sequence number) already assigns in the order the
//! reader reads them in -- so "between the anchor and the current row" is
//! answered by ordinary integer comparison via a `BTreeMap`, not a second
//! copy of the list's own ordering.
//!
//! **Scoped shortcut, recorded rather than hidden**: the anchor row (the
//! one the drag started in) is selected in full (`select_all`) the moment
//! the drag leaves it, rather than "from the click point to whichever edge
//! points away from the drag" -- the exact partial selection would need
//! that row's own laid-out size, which `TextEditCtx` does not expose to a
//! caller outside `iris::widget::text` (`edit.rs`'s `layout()` helper is
//! private). Only the row currently *under the pointer* gets a true partial
//! selection (from its own start or end, per direction, to the pointer's
//! exact point) -- see `extend`. Re-entering the anchor row is still exact,
//! since that branch never goes through the approximation.
use iris::prelude::*;
use std::{collections::BTreeMap, time::Instant};
pub struct Selection {
rows: BTreeMap<RowKey, WeakWidget<TextEdit>>,
anchor: Option<(RowKey, Vec2)>,
/// One arbiter shared by every row's drag handler -- RUST.md's I5
/// gesture conflict (a row's own `click_or_drag()` and a list-level
/// pan wanting the same touch gesture). See `drag` below, and
/// `iris::sense::DragArbiter`'s own doc for the decision itself.
arbiter: DragArbiter,
}
impl Default for Selection {
fn default() -> Self {
Self::new()
}
}
impl Selection {
pub fn new() -> Self {
Self {
rows: BTreeMap::new(),
anchor: None,
arbiter: DragArbiter::new(),
}
}
/// A row's selectable text became visible/known. Every addition here
/// needs its removal (`unregister`) -- called when `List` evicts the
/// row (`pop_front`/`pop_back`), so this map never outgrows however
/// many rows are actually loaded.
pub fn register(&mut self, key: RowKey, text: WeakWidget<TextEdit>) {
self.rows.insert(key, text);
}
pub fn unregister(&mut self, key: RowKey) {
self.rows.remove(&key);
if self.anchor.map(|(k, _)| k) == Some(key) {
self.anchor = None;
}
}
/// A fresh press: clears whatever was selected elsewhere (an ordinary
/// click starts a new selection, it does not extend the old one) and
/// gives `key`'s row a collapsed caret at `pos` -- a plain click that
/// never turns into a drag leaves exactly this and nothing else
/// selected.
pub fn begin(&mut self, ui: &mut impl UiRsc, key: RowKey, pos: Vec2, size: Vec2) {
let rows: Vec<RowKey> = self.rows.keys().copied().collect();
for k in rows {
if k != key
&& let Some(w) = self.rows.get(&k)
{
w.edit(ui).deselect();
}
}
if let Some(w) = self.rows.get(&key) {
w.edit(ui).select(pos, size, false, false);
}
self.anchor = Some((key, pos));
}
/// The drag continues, now over `key`'s row at `pos`. See the module
/// doc for the anchor-row shortcut.
pub fn extend(&mut self, ui: &mut impl UiRsc, key: RowKey, pos: Vec2, size: Vec2) {
let Some((anchor_key, _anchor_pos)) = self.anchor else {
return;
};
if key == anchor_key {
if let Some(w) = self.rows.get(&key) {
w.edit(ui).select(pos, size, true, false);
}
return;
}
let (lo, hi) = if anchor_key < key {
(anchor_key, key)
} else {
(key, anchor_key)
};
let in_range: Vec<RowKey> = self.rows.range(lo..=hi).map(|(&k, _)| k).collect();
for k in &in_range {
let Some(w) = self.rows.get(k).copied() else {
continue;
};
if *k == key {
// The row under the pointer: partial selection from
// whichever of its own edges faces the anchor, extended to
// the exact pointer point.
let start = if key > anchor_key { Vec2::ZERO } else { size };
w.edit(ui).select(start, size, false, false);
w.edit(ui).select(pos, size, true, false);
} else {
w.edit(ui).select_all();
}
}
let outside: Vec<RowKey> = self
.rows
.keys()
.copied()
.filter(|k| *k < lo || *k > hi)
.collect();
for k in outside {
if let Some(w) = self.rows.get(&k) {
w.edit(ui).deselect();
}
}
}
/// Whether any row currently has a non-empty selection -- what a fresh
/// press consults so `drag` knows whether an early horizontal move is
/// "start dragging the selection handle" rather than an ordinary tap.
fn has_selection(&self, ui: &mut impl UiRsc) -> bool {
self.rows
.values()
.any(|w| w.edit(ui).text.selected_text().is_some())
}
/// One row's `CursorSense::click_or_drag()` handler, for every row,
/// routes its raw pointer data through here rather than calling
/// `begin`/`extend` directly -- this is the single place that decides
/// whether the gesture pans `list` or extends a selection, so the
/// decision is made once per gesture rather than independently by
/// whichever row happens to be under the finger this frame (see
/// `DragArbiter`'s own doc for why one shared instance, not one per
/// row, is what makes that consistent as a drag crosses row
/// boundaries).
///
/// `pos_row`/`size` are row-local, as `begin`/`extend` want;
/// `pos_window` is in window space, since a pan's delta has to stay
/// meaningful even when this frame's event landed on a different row
/// than the last one.
#[allow(clippy::too_many_arguments)]
pub fn drag(
&mut self,
ui: &mut impl UiRsc,
list: WeakWidget<List>,
key: RowKey,
pos_row: Vec2,
size: Vec2,
pos_window: Vec2,
sense: CursorSense,
now: Instant,
) {
let outcome = match sense {
CursorSense::PressStart(_) => {
let already_selected = self.has_selection(ui);
self.arbiter.press_start(pos_window, now, already_selected);
self.arbiter.update(pos_window, now)
}
CursorSense::PressEnd(_) => {
self.arbiter.release();
return;
}
// A `Pressing` frame with the arbiter still `Idle` means this
// gesture's `ACTION_DOWN` landed somewhere no row's sensor
// covers (a row's own padding/gap, or a header with no
// selection handler) and this row is only now getting the
// touch as it moves across it -- the touch is definitely still
// down (that's what `Pressing` means), so without this the
// arbiter would sit in `Idle` answering `Undecided` for the
// rest of the gesture (`DragArbiter::update`'s own doc).
// Recovered by starting the press here instead of where it
// was missed -- RUST.md's I5 intermittent-touch-scroll-dropout
// finding, 2026-09-05.
_ if self.arbiter.is_idle() => {
let already_selected = self.has_selection(ui);
self.arbiter.press_start(pos_window, now, already_selected);
self.arbiter.update(pos_window, now)
}
_ => self.arbiter.update(pos_window, now),
};
match outcome {
DragOutcome::Undecided => {}
DragOutcome::Pan(dy) => list(ui).scroll(-dy),
DragOutcome::SelectStart => {
// Grep-able on "iris selection" the way the frame report is
// on "iris frame report" -- selection has no accessibility
// label of its own yet, so this is the smallest way to
// confirm a real on-device long-press-then-drag actually
// reached here (RUST.md's I5 box, "Measurements taken" (c)).
log::info!("iris selection: begin at row {key:?}");
self.begin(ui, key, pos_row, size);
}
DragOutcome::SelectExtend => {
log::info!("iris selection: extend to row {key:?}");
self.extend(ui, key, pos_row, size);
}
}
}
/// The concatenated selected text, in row order, `None` if nothing is
/// selected -- what a copy command reads. Joins with a blank line
/// between rows, matching how the transcript itself separates them.
pub fn selected_text(&self, ui: &mut impl UiRsc) -> Option<String> {
let mut parts = Vec::new();
for w in self.rows.values() {
if let Some(text) = w.edit(ui).text.selected_text() {
parts.push(text);
}
}
if parts.is_empty() {
None
} else {
Some(parts.join("\n\n"))
}
}
}
#[cfg(test)]
mod tests {
use super::*;
// Pure range-membership logic, independent of any widget/render
// machinery (the same reasoning `begin`/`extend` apply per-row) --
// exercised directly so the "which rows fall between anchor and
// current" arithmetic has a test that needs no `UiRenderState`.
fn in_range(anchor: RowKey, current: RowKey, keys: &[RowKey]) -> Vec<RowKey> {
let (lo, hi) = if anchor < current {
(anchor, current)
} else {
(current, anchor)
};
keys.iter()
.copied()
.filter(|k| *k >= lo && *k <= hi)
.collect()
}
#[test]
fn selection_spans_forward_across_rows() {
let keys = [1, 2, 3, 4, 5];
assert_eq!(in_range(2, 4, &keys), vec![2, 3, 4]);
}
#[test]
fn selection_spans_backward_across_rows() {
let keys = [1, 2, 3, 4, 5];
assert_eq!(in_range(4, 2, &keys), vec![2, 3, 4]);
}
#[test]
fn selection_within_one_row_is_just_that_row() {
let keys = [1, 2, 3];
assert_eq!(in_range(2, 2, &keys), vec![2]);
}
struct TestRsc {
ui: UiData,
}
impl UiRsc for TestRsc {
fn ui(&self) -> &UiData {
&self.ui
}
fn ui_mut(&mut self) -> &mut UiData {
&mut self.ui
}
}
/// The RUST.md I5 intermittent-touch-scroll-dropout regression: a
/// gesture whose `ACTION_DOWN` landed where no row's sensor covers
/// (padding, a gap, a header with no handler) delivers this row only
/// `Pressing` frames, never `PressStart`. Before the fix, the shared
/// `DragArbiter` stayed `Idle` for the whole gesture (`DragArbiter::
/// update`'s own doc), which is exactly what a real device trace
/// showed for four of twenty-four otherwise-identical swipes in one
/// run -- `ui-trace`-driven touch coordinates land on different row
/// content each time the list actually scrolls, so whether `DOWN`
/// happens to hit a sensor is intermittent by construction. This test
/// fails on the code before `Selection::drag`'s `_ if self.arbiter.
/// is_idle()` branch existed, because the arbiter would still report
/// `is_idle()` after both calls below.
#[test]
fn a_missed_press_start_recovers_on_the_next_pressing_frame() {
let mut rsc = TestRsc {
ui: UiData::default(),
};
let field = rsc
.ui
.widgets
.add_strong(TextEdit::new(
TextView::new(TextBuffer::new_empty(), TextAttrs::default(), None),
EditMode::MultiLine,
))
.weak();
let list = rsc.ui.widgets.add_strong(List::new(Axis::Y)).weak();
let mut sel = Selection::new();
sel.register(1, field);
assert!(sel.arbiter.is_idle());
let now = Instant::now();
let size = Vec2::new(100.0, 20.0);
// No `PressStart` is ever sent -- only the `Pressing` frames a
// widget whose sensor missed the `ACTION_DOWN` would actually see.
sel.drag(
&mut rsc,
list,
1,
Vec2::ZERO,
size,
Vec2::new(540.0, 700.0),
CursorSense::Pressing(CursorButton::Left),
now,
);
assert!(
!sel.arbiter.is_idle(),
"a Pressing frame with the arbiter still Idle must recover \
the press rather than leaving it stuck"
);
}
#[test]
fn unregister_forgets_the_row_and_clears_a_matching_anchor() {
let mut rsc = TestRsc {
ui: UiData::default(),
};
let field = rsc
.ui
.widgets
.add_strong(TextEdit::new(
TextView::new(TextBuffer::new_empty(), TextAttrs::default(), None),
EditMode::MultiLine,
))
.weak();
let mut sel = Selection::new();
sel.register(5, field);
sel.anchor = Some((5, Vec2::ZERO));
assert_eq!(sel.rows.len(), 1);
sel.unregister(5);
assert!(sel.rows.is_empty());
assert!(sel.anchor.is_none());
}
}
+29 -4
View File
@@ -35,6 +35,34 @@ fn iris_features() -> Features {
/// kept for the big storage buffers behind rects/glyphs.
const IRIS_MAX_BUFFER_SIZE: u64 = 1 << 30;
/// Mirrors `iris_core::device_limits()` (`iris/core/src/render/mod.rs`) --
/// cannot call it directly, since this rig is deliberately its own crate,
/// not a workspace member (this file's own Cargo.toml comment). Keep the
/// two in sync by hand when one changes; this rig's whole purpose is "does
/// the device iris actually builds come back," so a stale copy here would
/// silently stop answering that question. Zeroed rather than left at
/// `Limits::default()`'s desktop-tier values because nothing in iris
/// creates a `ComputePipeline` or a `@compute` shader stage -- found by
/// grepping the whole `iris`/`iris-core` tree before this rig's comment was
/// written -- and the unconditional default request is what crashed
/// `request_device` on the Android emulator's software GL path
/// (`EMU_GPU=software`, `force-gles`: SwiftShader's GL reports itself as
/// OpenGL ES 3.0, which has no compute shaders at all, so the adapter's
/// real limit is 0). The same would happen on a real GLES-3.0-only Android
/// device.
fn iris_limits() -> Limits {
Limits {
max_buffer_size: IRIS_MAX_BUFFER_SIZE,
max_compute_workgroup_storage_size: 0,
max_compute_invocations_per_workgroup: 0,
max_compute_workgroup_size_x: 0,
max_compute_workgroup_size_y: 0,
max_compute_workgroup_size_z: 0,
max_compute_workgroups_per_dimension: 0,
..Default::default()
}
}
fn main() {
vk::report();
@@ -104,10 +132,7 @@ fn main() {
// limits requested, this is expected to succeed everywhere -- this rig
// is what turned that from an assumption into a measurement, first on
// this emulator's software Vulkan.
let wanted = Limits {
max_buffer_size: IRIS_MAX_BUFFER_SIZE,
..Default::default()
};
let wanted = iris_limits();
match pollster::block_on(adapter.request_device(&DeviceDescriptor {
required_features: iris_features(),
required_limits: wanted,
+1
View File
@@ -36,6 +36,7 @@ dependencies = [
"sha2",
"tempfile",
"thiserror",
"time",
"tokio",
"tokio-stream",
"tower",
+6
View File
@@ -57,6 +57,12 @@ ureq = { version = "3", features = ["json"] }
# both in the graph rustls refuses to auto-select one.
rustls = "0.23"
libc = "0.2.189"
# One ISO-8601 timestamp: the reset time on the invented rate-limit window
# an echo session's `/usage` puts up. Already in the tree behind the
# certificate machinery, so this is a direct name for what is compiled
# anyway rather than a new crate -- and the alternative was hand-rolling a
# civil-from-days conversion to print one line.
time = { version = "0.3", features = ["formatting"] }
[dev-dependencies]
tempfile = "3"
+14
View File
@@ -207,6 +207,20 @@ mod tests {
/// like any other.
#[tokio::test]
async fn a_spooled_enrollment_is_adopted_on_first_use() {
// Under a subscriber, like every other exercise of this middleware.
// `tracing` caches a callsite's interest process-wide the first time it
// is reached, so the refusal at the end of this test -- reached with no
// subscriber on this thread -- could cache the rejection warning as
// never-enabled and make the tripwire above see an empty log. That
// failed about one full-suite run in ten, in the test that exists to
// notice a credential leak, which is the worst place for a flake.
let _guard = tracing::subscriber::set_default(
tracing_subscriber::fmt()
.with_max_level(tracing::Level::TRACE)
.with_writer(std::io::sink)
.finish(),
);
let dir = tempfile::tempdir().expect("tempdir");
let manager = manager_with_token(dir.path(), "first");
let spooled = generate_token();
+155
View File
@@ -30,6 +30,18 @@ pub struct Config {
pub tokens: Vec<TokenEntry>,
pub setups: Vec<SetupConfig>,
pub sessions: Vec<SessionConfig>,
/// What a new session's thinking level is when nothing chose one.
///
/// Here rather than on a provider because providers are *discovered*: a
/// default written onto one would be erased by the next rediscovery, which
/// is the kind of setting that looks like it stuck until the day it did
/// not. Here rather than on the phone because a second device would then
/// spawn sessions the first one's owner did not expect.
///
/// `None` is the CLI's own default, and stays reachable: this is a level
/// somebody chose, not a level this app picked for them.
#[serde(default, skip_serializing_if = "Option::is_none")]
pub default_effort: Option<String>,
}
/// A machine, and the things it can run.
@@ -78,6 +90,16 @@ pub struct ProviderConfig {
pub models: Vec<String>,
}
impl ProviderConfig {
/// The executable to run for this provider: its override, or its kind's
/// default.
pub fn program(&self) -> &str {
self.command
.as_deref()
.unwrap_or(self.kind.default_program())
}
}
/// How to reach a setup that isn't this machine, with the system `ssh` client
/// -- so `~/.ssh/config`, agents and jump hosts all keep working, and there is
/// one place to configure connections. A remote session is the identical
@@ -94,6 +116,19 @@ pub struct SshConfig {
/// Extra `-o` settings, each written as `Key=value`.
#[serde(default, skip_serializing_if = "Vec::is_empty")]
pub options: Vec<String>,
/// Where this machine keeps the GGUF models it can serve, absent for
/// the same default this backend uses (`~/.local/share/ai-app/models`
/// -- `$XDG_DATA_HOME` is not read on the far side, since it is this
/// machine's environment that would answer). A `~` prefix is the
/// remote home.
///
/// Here rather than on the provider because it is a fact about the
/// machine, and because a machine reached over ssh is where the model
/// has to be: a llama.cpp session serves the file from the machine
/// that runs `llama-server`, and this backend's own downloads are on
/// whichever machine that is only when they are the same one.
#[serde(default, skip_serializing_if = "Option::is_none")]
pub models_dir: Option<PathBuf>,
/// Where a file attached from the phone is put on this machine so the
/// session can read it. Absent means the session's own working directory,
/// or the login home for a session that has none. A `~` prefix is the
@@ -145,6 +180,45 @@ impl DriverKind {
}
}
/// Which paid service meters a session of this kind, and `None` for one
/// that costs nothing.
///
/// What decides which account -- if any -- a rate-limit bar is about is the
/// provider a session runs, not the machine it runs on: an echo session on
/// a machine that also has the Claude CLI was drawn with that CLI's
/// five-hour window, a quota it cannot spend.
///
/// Echo names a meter of its own that exists only when a test has asked for
/// one (`/usage` in `session::echo`), which is how the bar's states are
/// reached without an account. With none set there is no snapshot, and the
/// phone draws nothing.
///
/// The string is a [`crate::usage::UsageProvider::name`], and it is what
/// pairs a session with one of `GET /usage`'s snapshots -- so
/// `usage::providers_for` reads this rather than matching on kinds again.
pub fn usage_provider(self) -> Option<&'static str> {
match self {
Self::ClaudeCli => Some(crate::usage::CLAUDE),
Self::Echo => Some(crate::usage::ECHO),
Self::LlamaCpp => None,
}
}
/// The executable a provider of this kind runs when it names none.
///
/// Here rather than at each spawn site because it is not only the spawn
/// that runs it: `usage` runs the Claude CLI too, to have it refresh its
/// own OAuth token, and a default that disagreed with the driver's would
/// ask the wrong binary on a machine with two installs.
pub fn default_program(self) -> &'static str {
match self {
Self::ClaudeCli => "claude",
Self::LlamaCpp => "llama-server",
// Echo is translated in-process; nothing is spawned for it.
Self::Echo => "echo",
}
}
/// Whether the conversation exists outside this app, so that deleting the
/// session here does not end it.
///
@@ -162,6 +236,24 @@ impl DriverKind {
Self::Echo | Self::LlamaCpp => false,
}
}
/// Whether a thinking level means anything to this kind, so the phone can
/// offer the control only where it does something.
///
/// Reported from here rather than decided on the phone, and asked of the
/// *kind* rather than branched on: the alternative is the session-type
/// `if` this app does not have anywhere else. `--effort` is the Claude
/// CLI's; a llama session's sampling is `params`, and echo does not think.
///
/// It matters more than a control that would simply do nothing, because
/// choosing a level stops the process -- so on a session that cannot use
/// one it is a button whose only effect is the cost.
pub fn takes_effort(self) -> bool {
match self {
Self::ClaudeCli => true,
Self::Echo | Self::LlamaCpp => false,
}
}
}
#[derive(Debug, Clone, Serialize, Deserialize)]
@@ -196,6 +288,17 @@ pub struct SessionConfig {
/// the CLI stays the one authority on which modes exist.
#[serde(skip_serializing_if = "Option::is_none")]
pub permission_mode: Option<String>,
/// How hard the model thinks, passed straight to `--effort`. A string for
/// the same reason `permission_mode` is: the CLI owns which levels exist.
///
/// Unlike the model and the mode, there is no control request that changes
/// one -- checked against 2.1.258, whose only two are `set_model` and
/// `set_permission_mode` -- so this is settled at launch and `None` means
/// whatever the CLI's own default is. That is a state the phone has to be
/// able to *choose*, not just start in, which is why it is an option
/// rather than a level with a default written here.
#[serde(skip_serializing_if = "Option::is_none")]
pub effort: Option<String>,
/// Settings the driver interprets, chosen at spawn.
///
/// Deliberately untyped: what a temperature or a context size means is the
@@ -215,6 +318,30 @@ pub struct SessionConfig {
/// turned off in one tap where one that never arrived is not diagnosable.
#[serde(default = "notify_default")]
pub notify: bool,
/// Whether a session stopped by the account's usage limit sends itself a
/// message once the limit lifts, instead of waiting for a person.
///
/// Off unless somebody asked for it. It spends quota the moment it becomes
/// available and it does so while nobody is looking, which is exactly the
/// kind of thing that must not happen because a default said so.
#[serde(default, skip_serializing_if = "not_set")]
pub auto_resume: bool,
/// What that message says. `None` is [`DEFAULT_RESUME_MESSAGE`], and stays
/// reachable: it is this app's word, not one somebody chose, so clearing
/// the field goes back to it rather than sending an empty message.
#[serde(default, skip_serializing_if = "Option::is_none")]
pub auto_resume_message: Option<String>,
/// The message this session owes itself once the limit lifts, and when to
/// try. Written when a limit is hit, moved when the wait turns out to be
/// wrong, and cleared when the message goes out or auto-resume is turned
/// off -- see [`ScheduledResume`].
///
/// Persisted rather than held in memory because the wait outlives the
/// process doing it: a five-hour window and a weekly one both routinely
/// outlast a backend restart, and a resume forgotten across one is a
/// session that silently never comes back.
#[serde(default, skip_serializing_if = "Option::is_none")]
pub resume: Option<ScheduledResume>,
/// Whether this session's process is stopped when the server exits, instead
/// of being left running for the next start to adopt.
///
@@ -231,6 +358,28 @@ pub struct SessionConfig {
pub created: f64,
}
/// A message owed to a session whose account ran out, and when to try sending
/// it.
///
/// `since` is the whole reason this is a struct: the wait is rescheduled every
/// time the meter is asked and still says no, so `at` alone cannot say how long
/// this has been going on -- and something has to, or a machine that can never
/// be asked is retried until somebody notices. See `crate::resume`.
#[derive(Debug, Clone, Copy, Serialize, Deserialize)]
#[serde(rename_all = "camelCase")]
pub struct ScheduledResume {
/// Epoch seconds: when the limit is next worth checking. Never a promise
/// that the message goes out then -- the meter is asked first.
pub at: f64,
/// Epoch seconds the limit was hit.
pub since: f64,
}
/// What an auto-resume says when nothing else was chosen. One word, because
/// the session already knows what it was doing and this is only the nudge that
/// lets it carry on.
pub const DEFAULT_RESUME_MESSAGE: &str = "continue";
fn notify_default() -> bool {
true
}
@@ -385,6 +534,7 @@ mod tests {
port: Some(2222),
identity_file: None,
options: Vec::new(),
models_dir: None,
attachments_dir: None,
}),
providers: vec![ProviderConfig {
@@ -395,6 +545,7 @@ mod tests {
}],
},
],
default_effort: Some("low".to_string()),
sessions: vec![SessionConfig {
id: "abc123".to_string(),
setup: "vm".to_string(),
@@ -403,8 +554,12 @@ mod tests {
model: None,
cwd: None,
permission_mode: None,
effort: None,
params: BTreeMap::new(),
notify: true,
auto_resume: false,
auto_resume_message: None,
resume: None,
throwaway: false,
created: 1234.5,
}],
+11 -1
View File
@@ -16,6 +16,7 @@ mod config;
mod files;
mod media;
mod models;
mod resume;
mod routes;
mod session;
mod setups;
@@ -264,7 +265,16 @@ async fn main() -> Result<()> {
// No providers listed here any more: which machines can be asked, and about
// what, comes from the setups at the moment the screen is opened -- so a
// machine added from the phone reports its limits without a restart.
let monitor = Arc::new(usage::UsageMonitor::new());
// The fixture is the manager's, because that is where the `/usage` command
// that sets it is typed; the monitor is what serves it.
let monitor = Arc::new(usage::UsageMonitor::new(manager.usage_fixture()));
// The one thing in here that acts without a request behind it: a session
// switched to auto-resume waits out its account's usage limit and picks
// itself back up. Started whether or not any session has it on, because
// the setting is per session and changes from the phone -- see
// `resume::run`.
tokio::spawn(resume::run(Arc::clone(&manager), Arc::clone(&monitor)));
// The bearer-token middleware wraps the entire router -- routes and fallback
// alike -- here and only here, so a new route can't forget auth.
+79
View File
@@ -29,6 +29,8 @@ use serde::Serialize;
use wg_app_link::private;
use crate::session::transport::{Launch, Transport};
/// Identifies this client to HuggingFace. They ask for one, and a request
/// without it is more likely to be rate-limited.
const USER_AGENT: &str = concat!("ai-server/", env!("CARGO_PKG_VERSION"));
@@ -521,6 +523,83 @@ fn collect(root: &Path, dir: &Path, found: &mut Vec<LocalModel>) {
}
}
/// Where a machine reached over ssh keeps its models, when its setup does
/// not say.
///
/// The same place this backend puts its own downloads, written out rather
/// than derived: `$XDG_DATA_HOME` here describes *this* machine's
/// environment, and the far machine's is the far machine's business. A
/// setup whose models are elsewhere says so (`SshConfig::models_dir`).
const FAR_MODELS_DIR: &str = "~/.local/share/ai-app/models";
/// Which directory holds the models on the machine `transport` reaches.
///
/// One answer, because two things ask: the list a spawn screen offers,
/// and the path a session hands `llama-server`. A machine that listed one
/// directory and served from another would offer models that then failed
/// to load, which reads as the model being broken.
pub fn dir_on(transport: &Transport, local: &Path) -> String {
match transport {
Transport::Here => local.to_string_lossy().into_owned(),
Transport::Ssh { ssh, .. } => ssh
.models_dir
.as_ref()
.map_or(FAR_MODELS_DIR.to_string(), |dir| {
dir.to_string_lossy().into_owned()
}),
}
}
/// Every GGUF on the machine a setup names, which is the machine that
/// would have to serve it.
///
/// The local half of this is [`ModelStore::list`], reading the same shape
/// off this machine's disk; a caller picks by transport, since a setup
/// with no ssh *is* this machine and asking a shell about it would be a
/// slower way to the same answer. What must not happen is offering this
/// backend's downloads for a session on another machine: the file has to
/// be where `llama-server` runs, and a list that says otherwise is a
/// claim about the wrong filesystem.
///
/// `dir` is that machine's models directory, `~` included -- expanded on
/// the far side, which is the only place that knows what it is. A
/// directory that is not there is an empty list rather than a failure: a
/// machine that has never had a model put on it is an ordinary state, and
/// the same one as a machine whose directory exists and is empty.
pub async fn on_machine(transport: &Transport, dir: &str) -> Result<Vec<LocalModel>> {
let script = "p=$1; case $p in \"~\") p=$HOME;; \"~/\"*) p=$HOME/${p#\"~/\"};; esac; \
[ -d \"$p\" ] || exit 0; \
find \"$p\" -type f -name '*.gguf' -printf '%s\\t%P\\0'";
let launch = Launch::new(
"sh",
vec![
"-c".to_string(),
script.to_string(),
"sh".to_string(),
dir.to_string(),
],
None,
);
let out = transport.capture(&launch).await?;
let mut found: Vec<LocalModel> = out
.split('\0')
.filter(|record| !record.is_empty())
// Two fields, and the name last, so a `\t` in a filename survives.
.filter_map(|record| record.split_once('\t'))
.filter_map(|(bytes, key)| {
let (repo, file) = key.rsplit_once('/')?;
Some(LocalModel {
key: key.to_string(),
repo: repo.to_string(),
file: file.to_string(),
bytes: bytes.trim().parse().unwrap_or(0),
})
})
.collect();
found.sort_by(|a, b| a.key.cmp(&b.key));
Ok(found)
}
/// A model repository on HuggingFace, as the browse screen shows it.
#[derive(Debug, Clone, Serialize)]
#[serde(rename_all = "camelCase")]
+382
View File
@@ -0,0 +1,382 @@
//! Auto-resume: picking a session back up when its account's usage limit
//! lifts.
//!
//! Off unless a session was switched to it, because this spends quota the
//! moment quota exists and does it while nobody is watching. What it does is
//! narrow on purpose: it sends one message -- "continue" unless something else
//! was typed -- to a session that stopped because the account ran out, and
//! then it is done. There is no retry loop around the conversation itself.
//!
//! **The schedule is a plan to ask, never a plan to send.** A reset time is
//! the one thing here that cannot be trusted: the dialect's is a hint written
//! when the turn failed, the endpoint's moves when the window moves, and both
//! are wrong across the case this exists for -- a limit that lifts later than
//! it said. So the wait ends in a *question* to [`crate::usage`], and only an
//! answer that says the limits no longer apply sends anything. Every other
//! answer, including one that cannot be got at all, becomes a new wait.
//!
//! This is the top layer: it holds the session manager and the usage monitor
//! and neither holds it. That is what lets the decision below be a pure
//! function of a snapshot and a clock, which is the whole of what is worth
//! testing here.
use std::sync::Arc;
use std::time::Duration;
use crate::session::{LimitHit, OwedResume, SessionManager, now};
use crate::usage::{UsageMonitor, UsageSnapshot, UsageState};
/// How often to look at the schedule. Coarse deliberately: a wait measured in
/// hours does not deserve a fine-grained clock, and the meter behind it is
/// cached for three minutes anyway.
const TICK: Duration = Duration::from_secs(60);
/// How close to a scheduled check is close enough to ask the meter. Anything
/// further out is left alone, so a session waiting five hours costs nothing
/// until the last few minutes of it.
const NEARLY: f64 = 300.0;
/// How long to wait after an answer that decided nothing -- the machine could
/// not be asked, or it says the limit is still on with no reset time.
const BACKOFF: f64 = 300.0;
/// The least time to wait before asking again, whatever a reset time says. A
/// window that claims to reset in the past would otherwise be asked about on
/// every tick.
const AT_LEAST: f64 = 60.0;
/// How long after the limit was hit to stop waiting.
///
/// Something has to bound it, or a machine that can never be asked -- an
/// unplugged laptop, a setup somebody edited away -- is retried for ever with
/// nothing on screen saying so. A day is past the longest window Claude
/// reports, so reaching this means the wait was never going to end on its own.
const GIVE_UP: f64 = 24.0 * 60.0 * 60.0;
/// The percentage at which a window is spent. The API counts up to 100, so
/// this is an equality in all but name; written as a threshold because a
/// figure arriving slightly over is a full window, not a corrupt one.
const SPENT: f64 = 100.0;
/// What to do about one owed resume, having asked the meter.
#[derive(Debug, Clone, Copy, PartialEq)]
pub enum Step {
/// The limits no longer apply: send the message.
Send,
/// Ask again at this epoch second.
WaitUntil(f64),
/// This has been waiting longer than anything real would take.
GiveUp,
}
/// Runs the schedule until the server stops.
///
/// Two things wake it: the tick, and a session reporting that it has just run
/// out. The second is not an optimisation -- a limit hit is what *creates* a
/// schedule, and a tick that happened a moment before it would otherwise leave
/// the session unrecorded until the next one.
pub async fn run(manager: Arc<SessionManager>, monitor: Arc<UsageMonitor>) {
let mut limits = manager.subscribe_limits();
loop {
tokio::select! {
_ = tokio::time::sleep(TICK) => {}
hit = limits.recv() => match hit {
Ok(LimitHit { session_id, resets_at }) => note(&manager, &session_id, resets_at),
// Lagged: some reports were dropped, and a session that hit a
// limit while this was busy has no schedule. Nothing is lost
// for good -- the sweep below reads the config, and the
// session will report again the next time it is poked -- but
// it is worth saying, because until then that session waits
// for a person.
Err(tokio::sync::broadcast::error::RecvError::Lagged(missed)) => {
tracing::warn!("auto-resume missed {missed} limit reports");
}
Err(tokio::sync::broadcast::error::RecvError::Closed) => return,
},
}
sweep(&manager, &monitor).await;
}
}
/// Records a limit against the session that hit it, if it is one that resumes.
pub(crate) fn note(manager: &SessionManager, session_id: &str, resets_at: Option<f64>) {
match manager.note_limit(session_id, resets_at) {
Ok(true) => tracing::info!("session {session_id} hit its usage limit; auto-resume is on"),
Ok(false) => {}
Err(err) => tracing::error!("couldn't schedule a resume for {session_id}: {err:#}"),
}
}
/// One pass over everything owed a message.
async fn sweep(manager: &SessionManager, monitor: &Arc<UsageMonitor>) {
let at = now();
for owed in manager.owed_resumes() {
if owed.scheduled.at - at > NEARLY {
continue;
}
// Asked per session rather than once for the whole sweep: the answer
// is cached per machine and per meter, so several sessions on one
// account share one fetch, and a machine nobody is waiting on is not
// dialled at all.
let snapshot = snapshot_for(Arc::clone(monitor), manager, &owed).await;
match decide(snapshot.as_ref(), &owed, now()) {
Step::Send => match manager.resume_now(&owed.session_id) {
Ok(message) => tracing::info!(
"the limit on {} has lifted; sent \"{message}\" to {}",
owed.setup,
owed.session_id
),
Err(err) => {
tracing::error!("couldn't resume {}: {err:#}", owed.session_id)
}
},
Step::WaitUntil(next) => {
if let Err(err) = manager.reschedule_resume(&owed.session_id, next) {
tracing::error!(
"couldn't move {}'s resume to {next}: {err:#}",
owed.session_id
);
}
}
Step::GiveUp => {
// About the machine rather than in the state's own words: the
// detail is in the log, and what lands in the transcript has
// to read on a phone.
let why = match snapshot.as_ref().map(|snapshot| &snapshot.state) {
Some(UsageState::Ok) => "the limit has not lifted in a day".to_string(),
_ => format!("{} could not be asked for a day", owed.setup),
};
if let Err(err) = manager.abandon_resume(&owed.session_id, &why) {
tracing::error!("couldn't clear {}'s resume: {err:#}", owed.session_id);
}
}
}
}
}
/// The numbers for the machine and the meter this session is billed against,
/// and `None` when nothing reports on it.
///
/// Blocking work, so it goes to a blocking thread: the fetch behind it reads a
/// credential file over ssh and then makes an HTTP call.
async fn snapshot_for(
monitor: Arc<UsageMonitor>,
manager: &SessionManager,
owed: &OwedResume,
) -> Option<UsageSnapshot> {
let setups: Vec<_> = manager
.setups()
.into_iter()
.filter(|setup| setup.id == owed.setup)
.collect();
if setups.is_empty() {
return None;
}
let provider = owed.provider;
tokio::task::spawn_blocking(move || {
monitor
.snapshots(&setups)
.into_iter()
.find(|snapshot| snapshot.provider == provider)
})
.await
.unwrap_or_default()
}
/// What one owed resume should do, given what the meter said and the time.
///
/// A pure function of the two, which is what makes the rule inspectable: every
/// answer that is not "the limits no longer apply" is a longer wait, and the
/// only thing that ends the waiting other than success is the clock.
///
/// The reset time comes from the *snapshot* rather than from the schedule, so
/// a window that turns out to reset later than the dialect said pushes the
/// check back, and one that resets sooner pulls it forward. That is the case
/// the whole design is about: the first answer was a guess, this one is a
/// measurement.
pub fn decide(snapshot: Option<&UsageSnapshot>, owed: &OwedResume, at: f64) -> Step {
let step = match snapshot {
// The meter answered with numbers, which is the only answer that can
// send anything.
Some(snapshot) if snapshot.state == UsageState::Ok => {
let spent: Vec<&crate::usage::UsageWindow> = snapshot
.windows
.iter()
.filter(|window| window.percent >= SPENT)
.collect();
if spent.is_empty() {
Step::Send
} else {
// The earliest of the spent windows: it is the first moment
// the situation can have changed, and if the others are still
// full this comes straight back here.
match spent
.iter()
.filter_map(|window| epoch_of(window.resets_at.as_deref()))
.min_by(f64::total_cmp)
{
Some(resets) => Step::WaitUntil(resets),
// Spent with no reset time anybody could read. Not a
// reason to send: what is known is that the limit is on.
None => Step::WaitUntil(at + BACKOFF),
}
}
}
// Logged out, unreachable, or the endpoint refused us -- and nothing
// at all, which is a session whose machine or provider has gone. None
// of them says the limit has lifted, and sending on any of them is
// exactly the "inferred value presented as a measured one" this is
// built to avoid.
_ => Step::WaitUntil(at + BACKOFF),
};
match step {
// Waiting past the point where a real window would have reset means
// whatever is wrong is not going to fix itself.
Step::WaitUntil(_) if at - owed.scheduled.since > GIVE_UP => Step::GiveUp,
Step::WaitUntil(next) => Step::WaitUntil(next.max(at + AT_LEAST)),
other => other,
}
}
/// An RFC-3339 timestamp as epoch seconds, and `None` for one that is absent
/// or unreadable -- the same two answers the phone's countdown makes, kept
/// apart from each other nowhere here because both mean "this cannot decide
/// when to ask".
fn epoch_of(resets_at: Option<&str>) -> Option<f64> {
let text = resets_at?;
time::OffsetDateTime::parse(text, &time::format_description::well_known::Rfc3339)
.ok()
.map(|at| at.unix_timestamp() as f64)
}
#[cfg(test)]
mod tests {
use super::*;
use crate::config::ScheduledResume;
use crate::usage::UsageWindow;
fn owed(since: f64) -> OwedResume {
OwedResume {
session_id: "s1".to_string(),
setup: "local".to_string(),
provider: crate::usage::CLAUDE,
scheduled: ScheduledResume { at: since, since },
}
}
fn snapshot(state: UsageState, windows: Vec<UsageWindow>) -> UsageSnapshot {
UsageSnapshot {
provider: crate::usage::CLAUDE.to_string(),
setup: "local".to_string(),
setup_name: "this machine".to_string(),
state,
windows,
fetched_at: 0.0,
}
}
fn window(percent: f64, resets_at: Option<&str>) -> UsageWindow {
UsageWindow {
kind: "session".to_string(),
label: "5-hour window".to_string(),
percent,
resets_at: resets_at.map(str::to_string),
active: true,
}
}
#[test]
fn a_meter_with_room_in_it_is_the_only_thing_that_sends() {
let clear = snapshot(UsageState::Ok, vec![window(41.0, None)]);
assert_eq!(decide(Some(&clear), &owed(0.0), 100.0), Step::Send);
}
#[test]
fn a_window_still_spent_moves_the_check_to_its_own_reset_time() {
// The case the feature exists for: the wait was scheduled for one
// time, the limit is still on, and the endpoint now names another.
let at = 1_788_546_972.0;
let later = "2026-09-05T12:00:00+00:00";
let spent = snapshot(UsageState::Ok, vec![window(100.0, Some(later))]);
assert_eq!(
decide(Some(&spent), &owed(at - 60.0), at),
Step::WaitUntil(epoch_of(Some(later)).expect("parses"))
);
}
#[test]
fn a_reset_time_already_past_still_waits_a_little() {
let at = 1_788_546_972.0;
let spent = snapshot(
UsageState::Ok,
vec![window(100.0, Some("2020-01-01T00:00:00+00:00"))],
);
assert_eq!(
decide(Some(&spent), &owed(at - 60.0), at),
Step::WaitUntil(at + AT_LEAST)
);
}
#[test]
fn the_earliest_spent_window_is_the_one_worth_waiting_on() {
let at = 1_788_546_972.0;
let soon = "2026-09-05T12:00:00+00:00";
let far = "2026-09-09T12:00:00+00:00";
let mut weekly = window(100.0, Some(far));
weekly.kind = "weekly_all".to_string();
let spent = snapshot(UsageState::Ok, vec![window(100.0, Some(soon)), weekly]);
assert_eq!(
decide(Some(&spent), &owed(at - 60.0), at),
Step::WaitUntil(epoch_of(Some(soon)).expect("parses"))
);
}
#[test]
fn a_meter_that_could_not_be_asked_never_sends() {
let at = 1_788_546_972.0;
for state in [
UsageState::NotLoggedIn,
UsageState::Unreachable {
detail: "no route".to_string(),
},
UsageState::Failed {
detail: "429".to_string(),
},
] {
let broken = snapshot(state.clone(), Vec::new());
assert_eq!(
decide(Some(&broken), &owed(at - 60.0), at),
Step::WaitUntil(at + BACKOFF),
"{state:?}"
);
}
// And no snapshot at all -- a machine or provider edited away under a
// session that was waiting on it.
assert_eq!(
decide(None, &owed(at - 60.0), at),
Step::WaitUntil(at + BACKOFF)
);
}
#[test]
fn waiting_longer_than_any_real_window_gives_up_rather_than_retrying_for_ever() {
let at = 1_788_546_972.0;
let broken = snapshot(
UsageState::Unreachable {
detail: "no route".to_string(),
},
Vec::new(),
);
assert_eq!(
decide(Some(&broken), &owed(at - GIVE_UP - 1.0), at),
Step::GiveUp
);
// A meter that answers is still allowed to send on the same tick: the
// ceiling bounds waiting, not resuming.
let clear = snapshot(UsageState::Ok, vec![window(3.0, None)]);
assert_eq!(
decide(Some(&clear), &owed(at - GIVE_UP - 1.0), at),
Step::Send
);
}
}
+251 -13
View File
@@ -7,6 +7,7 @@
//! POST /setups add {name, ssh?} -- providers are discovered
//! POST /setups/probe dry run {ssh?}: what would be found there
//! GET /setups/{id} one machine, for refetching after a change
//! GET /setups/{id}/models GGUFs on that machine, for a llama session
//! GET /setups/{id}/dir?path=P entries of directory P, and P resolved
//! GET /setups/{id}/file?path=P content of file P, or why not
//! PUT /setups/{id}/file {path, content, ifSha256} -> new size/mtime/sha256
@@ -28,6 +29,12 @@
//! GET /sessions/{id}/transcript a page of history: ?before=N (newest when absent),
//! ?limit=N, ?coalesce=true to count rows not deltas,
//! ?after=N to floor it at what the caller already holds
//! GET /sessions/{id}/subagents [{id, title, status, created, lastActivity}], oldest
//! first -- see SUBAGENTS.md
//! GET /sessions/{id}/subagents/{sub}/transcript exactly the transcript route above,
//! against that subagent's own transcript
//! GET /sessions/{id}/subagents/{sub}/events?after=N exactly the events route above,
//! against that subagent's own stream
//! POST /sessions/{id}/message {text, attachmentIds?}
//! (starts the process first if it has exited)
//! POST /sessions/{id}/unqueue {messageId} -- take back one not read yet
@@ -41,6 +48,8 @@
//! which starts again in the new one
//! POST /sessions/{id}/model {model}
//! POST /sessions/{id}/permission-mode {permissionMode}
//! POST /sessions/{id}/effort {effort} -- null for the CLI's default;
//! settled at launch, so this stops the process
//! POST /sessions/{id}/command {text} -- /compact, /clear, /rename x, or the dialect's own
//! (starts the process first if it has exited)
//! POST /sessions/{id}/compact
@@ -49,8 +58,12 @@
//! DELETE /sessions/{id} kill process, delete transcript + files
//! (?deleteForeign=true removes the machine's own copy too)
//! POST /sessions/{id}/notify {notify} -- announce this one or not
//! POST /sessions/{id}/auto-resume {autoResume, message?} -- carry on by itself
//! once the account's usage limit lifts
//! GET /notifications SSE: every session's attention-wanting
//! moments, live only (see `notifications`)
//! GET /defaults {effort} -- what a new session starts at
//! POST /defaults {effort} -- null for the CLI's own default
//! GET /usage cached usage windows per provider
//! GET /models downloaded GGUFs, and what is being fetched
//! GET /models/search?q=Q HuggingFace repositories matching Q
@@ -92,6 +105,7 @@ use tokio_stream::wrappers::{BroadcastStream, ReceiverStream};
use crate::session::driver::{SessionCommand, Unqueued};
use crate::session::pending::Operation;
use crate::session::subagent::{Subagent, SubagentInfo};
use crate::session::transcript::{CATCH_UP_LIMIT, CatchUp, SeqEvent, catch_up};
use crate::session::{LiveSession, SessionInfo, SessionManager, SpawnSpec};
@@ -109,6 +123,8 @@ pub fn router(manager: Arc<SessionManager>) -> Router {
"/setups/{id}",
get(read_setup).put(update_setup).delete(delete_setup),
)
// The models on the machine a setup names, for a llama session there.
.route("/setups/{id}/models", get(setup_models))
// The filesystem of the machine a setup names. Under the setup
// rather than under a session because a filesystem is a property of
// a machine; a session only says where to start looking.
@@ -121,6 +137,15 @@ pub fn router(manager: Arc<SessionManager>) -> Router {
.route("/sessions/{id}", get(read_session).delete(delete_session))
.route("/sessions/{id}/events", get(events))
.route("/sessions/{id}/transcript", get(transcript))
.route("/sessions/{id}/subagents", get(list_subagents))
.route(
"/sessions/{id}/subagents/{sub}/transcript",
get(subagent_transcript),
)
.route(
"/sessions/{id}/subagents/{sub}/events",
get(subagent_events),
)
.route("/sessions/{id}/message", post(message))
.route("/sessions/{id}/unqueue", post(unqueue))
.route("/sessions/{id}/answer", post(answer))
@@ -131,7 +156,10 @@ pub fn router(manager: Arc<SessionManager>) -> Router {
.route("/sessions/{id}/cwd", post(set_cwd))
.route("/sessions/{id}/model", post(set_model))
.route("/sessions/{id}/permission-mode", post(set_permission_mode))
.route("/sessions/{id}/effort", post(set_effort))
.route("/defaults", get(defaults).post(set_defaults))
.route("/sessions/{id}/notify", post(set_notify))
.route("/sessions/{id}/auto-resume", post(set_auto_resume))
.route("/notifications", get(notifications))
.route("/sessions/{id}/compact", post(compact))
.route("/sessions/{id}/command", post(command))
@@ -197,6 +225,18 @@ fn lookup(manager: &SessionManager, id: &str) -> Result<Arc<LiveSession>, ApiErr
.ok_or_else(|| ApiError::NotFound(format!("no session {id}")))
}
/// A session's subagent by id -- the second half of the lookup every
/// `/sessions/{id}/subagents/{sub}/...` route needs. `Arc` because reopening
/// one from disk (a subagent this process has not touched yet) inserts it
/// into the registry, and a route holding a borrow across that would be
/// holding the registry's lock the whole request.
fn lookup_subagent(session: &LiveSession, sub: &str) -> Result<Arc<Subagent>, ApiError> {
session
.subagents()
.get(sub)
.ok_or_else(|| ApiError::NotFound(format!("no subagent {sub}")))
}
async fn list_sessions(State(manager): State<Arc<SessionManager>>) -> axum::Json<Vec<SessionInfo>> {
axum::Json(manager.sessions())
}
@@ -291,6 +331,9 @@ struct SshRequest {
/// Where attached files land on that machine; see `SshConfig`.
#[serde(default)]
attachments_dir: Option<String>,
/// Where that machine keeps its GGUF models; see `SshConfig`.
#[serde(default)]
models_dir: Option<String>,
}
impl SshRequest {
@@ -320,6 +363,14 @@ impl SshRequest {
.map(str::trim)
.filter(|dir| !dir.is_empty())
.map(std::path::PathBuf::from),
// The same rule, and for the same reason: this directory is
// on the other machine, so a `~` in it is that machine's home.
models_dir: self
.models_dir
.as_deref()
.map(str::trim)
.filter(|dir| !dir.is_empty())
.map(std::path::PathBuf::from),
})
}
}
@@ -495,6 +546,27 @@ struct PathQuery {
path: String,
}
/// The models **that machine** has, which is the list a llama.cpp session
/// on it can choose from.
///
/// Not `GET /models`, which is this backend's own downloads: those are on
/// the machine a session runs on only when they are the same machine. A
/// spawn screen offering this backend's list for a remote setup would be
/// naming files that are not there, and the session would fail at the
/// point of loading rather than at the point of choosing.
async fn setup_models(
State(manager): State<Arc<SessionManager>>,
UrlPath(id): UrlPath<String>,
) -> Result<axum::Json<Vec<crate::models::LocalModel>>, ApiError> {
let setup = setup_by_id(&manager, &id)?;
let transport = crate::session::transport::Transport::for_setup(&setup);
let dir = crate::models::dir_on(&transport, manager.models_dir());
crate::models::on_machine(&transport, &dir)
.await
.map(axum::Json)
.map_err(from_machine)
}
/// What is in a directory, and what that directory resolved to.
async fn list_dir(
State(manager): State<Arc<SessionManager>>,
@@ -609,6 +681,8 @@ struct SpawnRequest {
cwd: Option<PathBuf>,
#[serde(default)]
permission_mode: Option<String>,
#[serde(default)]
effort: Option<String>,
/// Whatever the chosen driver understands -- llama.cpp's context size and
/// sampling. Opaque here on purpose: see `SessionConfig::params`.
#[serde(default)]
@@ -796,6 +870,7 @@ async fn start_import(
model: body.model.clone(),
cwd: None,
permission_mode: body.permission_mode.clone(),
effort: body.effort.clone(),
params: std::collections::BTreeMap::new(),
import: Some(session.clone()),
};
@@ -828,6 +903,8 @@ struct ImportRequest {
model: Option<String>,
#[serde(default)]
permission_mode: Option<String>,
#[serde(default)]
effort: Option<String>,
}
/// Runs `work` on the server, marked as in flight for as long as it takes.
@@ -978,6 +1055,7 @@ async fn spawn(manager: &Arc<SessionManager>, body: SpawnRequest) -> Result<Sess
.filter(|cwd| cwd.as_os_str() != "")
}),
permission_mode: body.permission_mode,
effort: body.effort,
params: body.params,
};
@@ -1311,6 +1389,60 @@ struct PermissionModeRequest {
mode: String,
}
/// What new sessions start at. One field today; a struct rather than a bare
/// value because "the defaults" is the thing a phone asks for, and the next
/// one to move here -- the permission mode, which the spawn screen still
/// hardcodes -- must not need a second route.
#[derive(Serialize, Deserialize)]
#[serde(rename_all = "camelCase")]
#[serde(deny_unknown_fields)]
struct Defaults {
#[serde(default, skip_serializing_if = "Option::is_none")]
effort: Option<String>,
}
async fn defaults(State(manager): State<Arc<SessionManager>>) -> axum::Json<Defaults> {
axum::Json(Defaults {
effort: manager.default_effort(),
})
}
/// Sets what a new session's thinking level is. Applied when a session is
/// spawned, so nothing already running changes underneath anybody.
async fn set_defaults(
State(manager): State<Arc<SessionManager>>,
axum::Json(body): axum::Json<Defaults>,
) -> Result<StatusCode, ApiError> {
manager
.set_default_effort(body.effort.as_deref())
.map_err(bad_request)?;
Ok(StatusCode::NO_CONTENT)
}
#[derive(Deserialize)]
#[serde(rename_all = "camelCase")]
#[serde(deny_unknown_fields)]
struct EffortRequest {
/// Absent or null is the CLI's own default, which is a choice somebody can
/// make rather than only a state to start in.
#[serde(default)]
effort: Option<String>,
}
/// Records how hard this session thinks, and stops the process so the next one
/// is launched with it -- `--effort` has no control request behind it. See
/// [`SessionManager::set_session_effort`].
async fn set_effort(
State(manager): State<Arc<SessionManager>>,
UrlPath(id): UrlPath<String>,
axum::Json(body): axum::Json<EffortRequest>,
) -> Result<StatusCode, ApiError> {
manager
.set_session_effort(&id, body.effort.as_deref())
.map_err(bad_request)?;
Ok(StatusCode::NO_CONTENT)
}
async fn set_permission_mode(
State(manager): State<Arc<SessionManager>>,
UrlPath(id): UrlPath<String>,
@@ -1339,6 +1471,33 @@ async fn set_notify(
Ok(StatusCode::NO_CONTENT)
}
#[derive(Deserialize)]
#[serde(rename_all = "camelCase", deny_unknown_fields)]
struct AutoResumeRequest {
auto_resume: bool,
/// What to send when the limit lifts. Absent -- and empty, which is what a
/// cleared field sends -- means this app's own default word, which is a
/// choice a caller has to be able to make rather than only start in.
#[serde(default)]
message: Option<String>,
}
/// Turns auto-resume on or off, and sets what it would say.
///
/// One request for both, because they are one decision: switching it on
/// without saying what to send is the ordinary case, and changing the words
/// while it is off is how somebody sets it up before it is needed.
async fn set_auto_resume(
State(manager): State<Arc<SessionManager>>,
UrlPath(id): UrlPath<String>,
axum::Json(body): axum::Json<AutoResumeRequest>,
) -> Result<StatusCode, ApiError> {
manager
.set_session_auto_resume(&id, body.auto_resume, body.message.as_deref())
.map_err(bad_request)?;
Ok(StatusCode::NO_CONTENT)
}
#[derive(Deserialize)]
#[serde(deny_unknown_fields)]
struct CommandRequest {
@@ -1616,8 +1775,36 @@ async fn transcript(
Query(query): Query<TranscriptQuery>,
) -> Result<axum::Json<Vec<crate::session::transcript::SeqEvent>>, ApiError> {
let session = lookup(&manager, &id)?;
transcript_page(session.transcript_path(), &id, query)
}
/// Exactly [`transcript`]'s route and answer, against one subagent's own
/// transcript instead of its session's -- see `SUBAGENTS.md`'s wire shape.
async fn subagent_transcript(
State(manager): State<Arc<SessionManager>>,
UrlPath((id, sub)): UrlPath<(String, String)>,
Query(query): Query<TranscriptQuery>,
) -> Result<axum::Json<Vec<crate::session::transcript::SeqEvent>>, ApiError> {
let session = lookup(&manager, &id)?;
let subagent = lookup_subagent(&session, &sub)?;
transcript_page(
&subagent.transcript_path(),
&format!("{id}/subagents/{sub}"),
query,
)
}
/// A page of history at `path`, newest first to open with -- the one
/// implementation [`transcript`] and [`subagent_transcript`] share, since a
/// subagent's transcript is read exactly the way a session's is. `label` is
/// only for the debug line below.
fn transcript_page(
path: &Path,
label: &str,
query: TranscriptQuery,
) -> Result<axum::Json<Vec<crate::session::transcript::SeqEvent>>, ApiError> {
let events = crate::session::transcript::read_window(
session.transcript_path(),
path,
query.before,
query.after,
query.limit,
@@ -1629,7 +1816,7 @@ async fn transcript(
// for events and draws rows, and the ratio between them is a property of
// the conversation. `RUST_LOG=ai_server=debug`.
tracing::debug!(
session = %id,
session = %label,
before = ?query.before,
after = ?query.after,
limit = query.limit,
@@ -1647,23 +1834,74 @@ async fn events(
headers: HeaderMap,
) -> Result<Sse<impl tokio_stream::Stream<Item = Result<SseEvent, Infallible>>>, ApiError> {
let session = lookup(&manager, &id)?;
let cursor = headers
.get("last-event-id")
.and_then(|value| value.to_str().ok())
.and_then(|value| value.parse().ok())
.unwrap_or(query.after);
let cursor = cursor_of(&headers, query.after);
// Subscribe before reading the file so nothing can land in the gap
// between replay and live; overlap is deduplicated by seq.
let live = session.subscribe();
let (tx, stream) = mpsc::channel(64);
tokio::spawn(stream_session(
Ok(sse_stream(
session.transcript_path().to_path_buf(),
cursor,
live,
tx,
));
Ok(Sse::new(ReceiverStream::new(stream).map(Ok)).keep_alive(KeepAlive::default()))
))
}
/// Exactly [`events`]'s route and answer, against one subagent's own stream
/// instead of its session's -- see `SUBAGENTS.md`'s wire shape.
async fn subagent_events(
State(manager): State<Arc<SessionManager>>,
UrlPath((id, sub)): UrlPath<(String, String)>,
Query(query): Query<EventsQuery>,
headers: HeaderMap,
) -> Result<Sse<impl tokio_stream::Stream<Item = Result<SseEvent, Infallible>>>, ApiError> {
let session = lookup(&manager, &id)?;
let subagent = lookup_subagent(&session, &sub)?;
let cursor = cursor_of(&headers, query.after);
let live = subagent.subscribe();
Ok(sse_stream(subagent.transcript_path(), cursor, live))
}
/// The cursor an SSE reconnect resumes from: the native `Last-Event-ID`
/// takes precedence over the query parameter, same cursor either way.
fn cursor_of(headers: &HeaderMap, query_after: u64) -> u64 {
headers
.get("last-event-id")
.and_then(|value| value.to_str().ok())
.and_then(|value| value.parse().ok())
.unwrap_or(query_after)
}
/// Spawns the backlog-then-live task and wraps it as the response, the one
/// piece [`events`] and [`subagent_events`] share.
fn sse_stream(
transcript: PathBuf,
cursor: u64,
live: broadcast::Receiver<SeqEvent>,
) -> Sse<impl tokio_stream::Stream<Item = Result<SseEvent, Infallible>>> {
let (tx, stream) = mpsc::channel(64);
tokio::spawn(stream_session(transcript, cursor, live, tx));
Sse::new(ReceiverStream::new(stream).map(Ok)).keep_alive(KeepAlive::default())
}
/// `GET /sessions/{id}/subagents`: every subagent this session has started,
/// oldest first, with a status read from its own transcript -- see
/// `SUBAGENTS.md`'s wire shape. A subagent whose last status is `Running` is
/// reported `Unknown` instead when the session itself is not running: its
/// process was the session's, and a session with none has nothing left to
/// ask.
async fn list_subagents(
State(manager): State<Arc<SessionManager>>,
UrlPath(id): UrlPath<String>,
) -> Result<axum::Json<Vec<SubagentInfo>>, ApiError> {
let session = lookup(&manager, &id)?;
// Anything but `Exited` or `Unknown` has a process behind it, which is
// what decides whether a subagent still reading `Running` from its own
// transcript can be believed -- see `SUBAGENTS.md`'s wire shape.
let running = !matches!(
session.status(),
crate::session::driver::SessionStatus::Exited
| crate::session::driver::SessionStatus::Unknown
);
Ok(axum::Json(session.subagents().list(running)))
}
/// Every session's attention-wanting moments, on one stream.
+29 -6
View File
@@ -59,6 +59,7 @@ use tokio::sync::mpsc;
use super::driver::{AttachmentRef, Driver, Event, EventSink, SessionStatus, Unqueued};
use super::process;
use super::subagent::Subagents;
use super::transport::{Launch, Streams, Transport};
use crate::config::{ProviderConfig, SessionConfig};
use translate::{AnswerOutcome, Setting, Translator, starts_a_model_call};
@@ -213,8 +214,12 @@ impl ClaudeDriver {
transport: &Transport,
session_dir: &Path,
sink: EventSink,
subagents: Arc<Subagents>,
) -> Result<Self> {
let state = Arc::new(Mutex::new(Translator::new(session_dir.to_path_buf())));
let state = Arc::new(Mutex::new(Translator::new(
session_dir.to_path_buf(),
subagents,
)));
let queue = Arc::new(Mutex::new(Queue::default()));
let reading = Arc::new(AtomicBool::new(true));
@@ -351,6 +356,12 @@ impl ClaudeDriver {
if let Some(mode) = &meta.permission_mode {
push("--permission-mode", mode);
}
// Launch-only: see `SessionConfig::effort`. Omitted entirely when
// unset, so the CLI's own default is what an unchosen session gets
// rather than a level this app decided to call the default.
if let Some(effort) = &meta.effort {
push("--effort", effort);
}
// Named at birth, so this session is the same session in the CLI's own
// picker and in what other agents see.
//
@@ -386,7 +397,7 @@ impl ClaudeDriver {
let stdout = create_log(&session_dir.join(STDOUT_LOG))?;
let stderr = create_log(&session_dir.join(STDERR_LOG))?;
let program = provider.command.as_deref().unwrap_or("claude");
let program = provider.program();
let launch = Launch::new(program, args, meta.cwd.as_deref());
let child = transport.spawn(
&launch,
@@ -1163,7 +1174,10 @@ mod tests {
/// this" and "the transcript records that".
fn events_from_lines(lines: &[&str]) -> Vec<Event> {
let dir = tempfile::tempdir().expect("temp dir");
let state = Arc::new(Mutex::new(Translator::new(dir.path().to_path_buf())));
let state = Arc::new(Mutex::new(Translator::new(
dir.path().to_path_buf(),
Arc::new(Subagents::new(dir.path().to_path_buf())),
)));
let queue = Arc::new(Mutex::new(Queue::default()));
let (sink, mut out) = mpsc::unbounded_channel::<Event>();
for line in lines {
@@ -1187,7 +1201,10 @@ mod tests {
interject: impl FnOnce(&Arc<Mutex<Queue>>),
) -> Vec<Event> {
let dir = tempfile::tempdir().expect("temp dir");
let state = Arc::new(Mutex::new(Translator::new(dir.path().to_path_buf())));
let state = Arc::new(Mutex::new(Translator::new(
dir.path().to_path_buf(),
Arc::new(Subagents::new(dir.path().to_path_buf())),
)));
let queue = Arc::new(Mutex::new(Queue::default()));
let (sink, mut out) = mpsc::unbounded_channel::<Event>();
let mut interject = Some(interject);
@@ -1481,7 +1498,10 @@ mod tests {
// doing.
let dir = tempfile::tempdir().expect("tempdir");
let (sink, mut received) = mpsc::unbounded_channel();
let state = Arc::new(Mutex::new(Translator::new(dir.path().to_path_buf())));
let state = Arc::new(Mutex::new(Translator::new(
dir.path().to_path_buf(),
Arc::new(Subagents::new(dir.path().to_path_buf())),
)));
let queue = Arc::new(Mutex::new(Queue::default()));
let text = r#"{"type":"stream_event","event":{"type":"content_block_delta","index":0,"delta":{"type":"text_delta","text":"working"}},"parent_tool_use_id":null}"#;
@@ -1526,7 +1546,10 @@ mod tests {
// session back to work.
let dir = tempfile::tempdir().expect("tempdir");
let (sink, mut received) = mpsc::unbounded_channel();
let state = Arc::new(Mutex::new(Translator::new(dir.path().to_path_buf())));
let state = Arc::new(Mutex::new(Translator::new(
dir.path().to_path_buf(),
Arc::new(Subagents::new(dir.path().to_path_buf())),
)));
let queue = Arc::new(Mutex::new(Queue::default()));
queue.lock().unwrap().close(&sink, "the session ended");
+512 -40
View File
@@ -11,10 +11,12 @@
use std::collections::HashMap;
use std::path::{Path, PathBuf};
use std::sync::{Arc, Mutex};
use serde_json::{Value, json};
use super::super::driver::{Event, QuestionOption, SessionStatus, context_tokens};
use super::super::subagent::Subagents;
/// Whether this line is the CLI opening a fresh model call.
///
@@ -92,10 +94,20 @@ pub(super) struct Translator {
/// reports none rather than repeating the previous turn's.
context: Option<u64>,
session_dir: PathBuf,
/// This session's subagents, shared with every child translator below --
/// see `SUBAGENTS.md`. One registry per session, so a subagent started
/// through this translator or any of its children lands in the same
/// place a route reads it back from.
subagents: Arc<Subagents>,
/// One translator per subagent id, holding *its* streaming and
/// tool-tracking state -- separate from the parent's because tool ids
/// are unique but a `stream_event`'s content-block index is not, and
/// parallel subagents interleave their deltas on one stdout.
children: HashMap<String, Arc<Mutex<Translator>>>,
}
impl Translator {
pub(super) fn new(session_dir: PathBuf) -> Self {
pub(super) fn new(session_dir: PathBuf, subagents: Arc<Subagents>) -> Self {
Self {
session_id: None,
pending: HashMap::new(),
@@ -103,6 +115,8 @@ impl Translator {
interrupting: false,
context: None,
session_dir,
subagents,
children: HashMap::new(),
}
}
@@ -122,13 +136,77 @@ impl Translator {
pub(super) fn translate(&mut self, message: &Value) -> Vec<Event> {
// Events from subagents (Task tool internals) carry a
// parent_tool_use_id; the transcript shows the Task tool's own
// start/end instead of every nested step.
if message
.get("parent_tool_use_id")
.is_some_and(|id| !id.is_null())
{
return Vec::new();
// start/end instead of every nested step. Routed into that
// subagent's own transcript rather than dropped -- see
// `SUBAGENTS.md`.
if let Some(parent_id) = message.get("parent_tool_use_id").and_then(Value::as_str) {
return self.translate_child(parent_id, message);
}
self.dispatch(message)
}
/// A line belonging to a subagent rather than to this translator's own
/// session. Always returns nothing to the *caller*: everything it
/// produces goes into the subagent's own transcript instead.
fn translate_child(&mut self, id: &str, message: &Value) -> Vec<Event> {
match self.subagents.get(id) {
Some(subagent) if !subagent.is_open() => {
// Not stale: the Task tool runs in the background by
// default, so a finished subagent can still be sent another
// message later (SendMessage) and start working again. A
// line arriving after `finish` means exactly that, not a
// conversation that is over -- see `SUBAGENTS.md`.
self.subagents.reopen(id);
}
Some(_) => {}
None => {
// Nobody has heard of this id yet: the Task call itself
// either has not been seen or never will be. Started here
// with the best title available -- the tool name of this
// first line -- since SUBAGENTS.md's real title only
// arrives with the Task call.
self.subagents.start(id, &fallback_title(message), None);
}
}
let child = self
.children
.entry(id.to_string())
.or_insert_with(|| {
Arc::new(Mutex::new(Translator::new(
self.session_dir.clone(),
Arc::clone(&self.subagents),
)))
})
.clone();
let events = child.lock().unwrap().dispatch(message);
for event in events {
// The subagent's own vocabulary is Running/Exited/Unknown, never
// Idle -- a background Task is either working or it has ended,
// never merely "between turns" the way a session is. Dropped
// here rather than never produced, so a `result` line's own
// `Idle` (dispatch's ordinary end-of-turn event, for a subagent
// dialect that ever sends one) is caught the same way a
// `message_delta` would be.
if !matches!(
event,
Event::Status {
state: SessionStatus::Idle
}
) {
self.subagents.record(id, event);
}
}
// What actually ends a subagent's turn: not the parent's
// `tool_result`, which for a background Task arrives at launch
// ("Async agent launched...") long before the work is done -- see
// `SUBAGENTS.md`.
if ends_a_turn(message) {
self.subagents.finish(id);
}
Vec::new()
}
fn dispatch(&mut self, message: &Value) -> Vec<Event> {
match message.get("type").and_then(Value::as_str) {
Some("system") => self.translate_system(message),
// The CLI's own announcement that `/clear` took effect, sent just
@@ -224,12 +302,12 @@ impl Translator {
.and_then(Value::as_bool)
.unwrap_or(false)
{
events.push(Event::Error {
message: message
.get("result")
.and_then(Value::as_str)
.unwrap_or("the turn ended with an error")
.to_string(),
let said = message.get("result").and_then(Value::as_str);
events.push(match said.and_then(usage_limit) {
Some(resets_at) => Event::LimitReached { resets_at },
None => Event::Error {
message: said.unwrap_or("the turn ended with an error").to_string(),
},
});
}
let context = self.context.take();
@@ -379,22 +457,47 @@ impl Translator {
content
.iter()
.filter(|block| block.get("type").and_then(Value::as_str) == Some("tool_use"))
.map(|block| Event::ToolStart {
id: block
.map(|block| {
let id = block
.get("id")
.and_then(Value::as_str)
.unwrap_or_default()
.to_string(),
tool: block
.to_string();
let tool = block
.get("name")
.and_then(Value::as_str)
.unwrap_or_default()
.to_string(),
input: block.get("input").cloned().unwrap_or(Value::Null),
.to_string();
let input = block.get("input").cloned().unwrap_or(Value::Null);
// A subagent this call is about to start -- see
// `SUBAGENTS.md`'s lifecycle #1. The parent's own transcript
// still shows only the Task call itself, below.
if tool == "Task" || tool == "Agent" {
self.start_subagent_from_task(&id, &input);
}
Event::ToolStart { id, tool, input }
})
.collect()
}
/// Starts the subagent a Task call names, with the title and prompt
/// SUBAGENTS.md describes: the call's `description`, then
/// `(<subagent_type>)` when one is given, falling back to the tool's own
/// name when there is no description to build one from.
fn start_subagent_from_task(&self, id: &str, input: &Value) {
let description = text_field(input, "description");
let subagent_type = text_field(input, "subagent_type");
let prompt = input.get("prompt").and_then(Value::as_str);
let title = match (description, subagent_type) {
(Some(description), Some(subagent_type)) => {
format!("{description} ({subagent_type})")
}
(Some(description), None) => description,
(None, _) => "Task".to_string(),
};
self.subagents.start(id, &title, prompt);
}
fn translate_control_request(&mut self, message: &Value) -> Vec<Event> {
let request = &message["request"];
if request.get("subtype").and_then(Value::as_str) != Some("can_use_tool") {
@@ -598,11 +701,90 @@ impl Translator {
id: about.clone(),
output: texts.join("\n"),
});
// Deliberately does *not* finish a subagent `about` might name:
// the Task tool runs in the background by default, so this
// `tool_result` -- "Async agent launched..." -- arrives at
// launch, long before the subagent's own work is done. What
// ends it is its own turn ending, handled in `translate_child`.
}
events
}
}
/// The title to start a subagent under when its own first line arrives
/// before (or without) its Task call ever being seen: the tool name of that
/// first line, which is the only thing known about it yet. `"subagent"` for
/// a line this cannot even find a tool name in, such as one that opens with
/// something other than a tool call.
fn fallback_title(message: &Value) -> String {
message["message"]["content"]
.as_array()
.into_iter()
.flatten()
.find(|block| block.get("type").and_then(Value::as_str) == Some("tool_use"))
.and_then(|block| block.get("name"))
.and_then(Value::as_str)
.unwrap_or("subagent")
.to_string()
}
/// Whether this line is a subagent's *own* turn ending -- the only thing
/// that does, per `SUBAGENTS.md`: not the parent's `tool_result`, which for
/// a background Task arrives at launch rather than at completion.
///
/// Checked on the raw line rather than on what `dispatch` returns, so this
/// never has to touch the shared `translate_stream_event`/`dispatch` code a
/// top-level session's own turn-ending also goes through -- a subagent's
/// idea of "ended" must not change when a real session's does.
///
/// `message_delta` is the raw API's own signal, carrying the stop reason:
/// `end_turn` is genuinely done, `tool_use` means the model is about to call
/// one and there is more coming. A `result` line is the CLI's own shape for
/// a top-level turn; a subagent has not been observed to send one, but
/// SUBAGENTS.md counts it too in case a future CLI version does.
fn ends_a_turn(message: &Value) -> bool {
match message.get("type").and_then(Value::as_str) {
Some("stream_event") => {
let event = &message["event"];
event.get("type").and_then(Value::as_str) == Some("message_delta")
&& event["delta"].get("stop_reason").and_then(Value::as_str) == Some("end_turn")
}
Some("result") => true,
_ => false,
}
}
/// Whether a failed turn failed because the account is out of quota, and when
/// the CLI said the limit lifts.
///
/// The wording is the CLI's: a turn stopped by the limit ends with `is_error`
/// and a result of `Claude AI usage limit reached|1788546972`, the reset being
/// epoch seconds after a pipe. Matched on the sentence rather than on a code
/// because the CLI sends none, so this is deliberately loose about everything
/// but the four words.
///
/// The two `None`s mean different things and both are real. The outer one is
/// "some other failure". The inner one is "the limit is reached and the CLI did
/// not say until when" -- which is not a reason to invent a time: `crate::resume`
/// asks the usage endpoint before sending anything, and that answer is the one
/// that decides.
///
/// Milliseconds are accepted as well as seconds and told apart by magnitude,
/// since a wrong guess would schedule a resume tens of thousands of years out
/// and look exactly like auto-resume being broken.
fn usage_limit(result: &str) -> Option<Option<f64>> {
if !result.to_ascii_lowercase().contains("usage limit reached") {
return None;
}
let stamp = result
.rsplit('|')
.next()
.and_then(|tail| tail.trim().parse::<f64>().ok())
.filter(|stamp| *stamp > 0.0)
.map(|stamp| if stamp > 1e11 { stamp / 1000.0 } else { stamp });
Some(stamp)
}
/// A string field that is there and not empty, or `None`. The CLI omits these
/// rather than sending them empty, but a caller that sends `""` means the same
/// thing and should not produce a description that draws as a blank line.
@@ -665,10 +847,18 @@ mod tests {
.collect()
}
/// A fresh, empty subagent registry over the same temp dir a test's
/// translator writes into -- every test here is about the parent's own
/// events, so what a registry does with a subagent is `subagent.rs`'s
/// tests to make, not these.
fn test_subagents(dir: &tempfile::TempDir) -> Arc<Subagents> {
Arc::new(Subagents::new(dir.path().to_path_buf()))
}
#[test]
fn captures_the_resume_token_and_the_settings_from_init() {
let dir = tempfile::tempdir().expect("tempdir");
let mut translator = Translator::new(dir.path().to_path_buf());
let mut translator = Translator::new(dir.path().to_path_buf(), test_subagents(&dir));
let events = translate_lines(
&mut translator,
&[
@@ -690,7 +880,7 @@ mod tests {
#[test]
fn a_setting_is_reported_when_the_cli_accepts_it_and_not_before() {
let dir = tempfile::tempdir().expect("tempdir");
let mut translator = Translator::new(dir.path().to_path_buf());
let mut translator = Translator::new(dir.path().to_path_buf(), test_subagents(&dir));
// What `set_model` does: remember, send, and say nothing yet.
translator.expect_setting("req-a".to_string(), Setting::Model("sonnet".to_string()));
@@ -767,7 +957,7 @@ mod tests {
// The line it sends just after answering `set_permission_mode`, which is
// also how a mode changed from the terminal arrives.
let dir = tempfile::tempdir().expect("tempdir");
let mut translator = Translator::new(dir.path().to_path_buf());
let mut translator = Translator::new(dir.path().to_path_buf(), test_subagents(&dir));
let events = translate_lines(
&mut translator,
&[
@@ -787,7 +977,7 @@ mod tests {
fn streams_text_deltas_and_skips_the_consolidated_copy() {
// Real lines (trimmed) from the 2.1.237 probe.
let dir = tempfile::tempdir().expect("tempdir");
let mut translator = Translator::new(dir.path().to_path_buf());
let mut translator = Translator::new(dir.path().to_path_buf(), test_subagents(&dir));
let events = translate_lines(
&mut translator,
&[
@@ -807,7 +997,7 @@ mod tests {
#[test]
fn tool_use_and_result_become_tool_events() {
let dir = tempfile::tempdir().expect("tempdir");
let mut translator = Translator::new(dir.path().to_path_buf());
let mut translator = Translator::new(dir.path().to_path_buf(), test_subagents(&dir));
let events = translate_lines(
&mut translator,
&[
@@ -834,7 +1024,7 @@ mod tests {
#[test]
fn subagent_events_are_not_duplicated_into_the_transcript() {
let dir = tempfile::tempdir().expect("tempdir");
let mut translator = Translator::new(dir.path().to_path_buf());
let mut translator = Translator::new(dir.path().to_path_buf(), test_subagents(&dir));
let events = translate_lines(
&mut translator,
&[
@@ -844,10 +1034,243 @@ mod tests {
assert!(events.is_empty());
}
/// A child line does not just vanish from the parent -- it lands in its
/// own subagent's transcript, with that transcript's own sequence
/// numbers, starting at 1 like any other.
#[test]
fn a_child_line_lands_in_its_own_subagents_transcript() {
let dir = tempfile::tempdir().expect("tempdir");
let subagents = test_subagents(&dir);
let mut translator = Translator::new(dir.path().to_path_buf(), Arc::clone(&subagents));
translate_lines(
&mut translator,
&[
r#"{"type":"assistant","message":{"content":[{"type":"tool_use","id":"toolu_c1","name":"Bash","input":{"command":"echo hi"}}]},"parent_tool_use_id":"toolu_parent"}"#,
],
);
let subagent = subagents.get("toolu_parent").expect("subagent started");
let lines = crate::session::transcript::read_after(&subagent.transcript_path(), 0)
.expect("read subagent transcript");
assert_eq!(lines[0].seq, 1);
assert_eq!(
lines[0].event,
Event::Status {
state: SessionStatus::Running
}
);
assert!(
lines.iter().any(
|entry| matches!(&entry.event, Event::ToolStart { tool, .. } if tool == "Bash")
)
);
}
/// The title and prompt shown for a subagent come from the Task call
/// that started it, not from anything guessed at its first line.
#[test]
fn the_subagent_takes_its_title_and_prompt_from_the_task_call() {
let dir = tempfile::tempdir().expect("tempdir");
let subagents = test_subagents(&dir);
let mut translator = Translator::new(dir.path().to_path_buf(), Arc::clone(&subagents));
translate_lines(
&mut translator,
&[
r#"{"type":"assistant","message":{"content":[{"type":"tool_use","id":"toolu_task","name":"Task","input":{"description":"Investigate the bug","prompt":"Find why X fails","subagent_type":"general-purpose"}}]},"parent_tool_use_id":null}"#,
],
);
let rows = subagents.list(true);
assert_eq!(rows.len(), 1);
assert_eq!(rows[0].title, "Investigate the bug (general-purpose)");
let subagent = subagents.get(&rows[0].id).expect("subagent");
let lines = crate::session::transcript::read_after(&subagent.transcript_path(), 0)
.expect("read subagent transcript");
assert!(lines.iter().any(
|entry| matches!(&entry.event, Event::UserMessage { text, .. } if text == "Find why X fails")
));
}
/// The parent's `tool_result` for the Task id is what ends the
/// subagent -- SUBAGENTS.md's lifecycle #3 -- and nothing else does.
#[test]
fn the_parents_tool_result_does_not_finish_the_subagent() {
// The Task tool runs in the background by default: this
// `tool_result` is "Async agent launched...", arriving the moment
// the subagent *starts*, while it goes on working for however long
// its own turn takes. Finishing it here was the bug -- a running
// background agent read as "finished" with its transcript truncated
// at launch.
let dir = tempfile::tempdir().expect("tempdir");
let subagents = test_subagents(&dir);
let mut translator = Translator::new(dir.path().to_path_buf(), Arc::clone(&subagents));
translate_lines(
&mut translator,
&[
r#"{"type":"assistant","message":{"content":[{"type":"tool_use","id":"toolu_task2","name":"Task","input":{"description":"helper"}}]},"parent_tool_use_id":null}"#,
],
);
let subagent = subagents.get("toolu_task2").expect("subagent started");
assert!(subagent.is_open());
translate_lines(
&mut translator,
&[
r#"{"type":"user","message":{"role":"user","content":[{"type":"tool_result","tool_use_id":"toolu_task2","content":"Async agent launched","is_error":false}]},"parent_tool_use_id":null}"#,
],
);
assert!(subagent.is_open());
}
/// What actually ends a subagent: the raw API's own `message_delta`
/// saying its turn stopped with `end_turn`. Never written into the
/// subagent's own transcript as `Idle` -- its vocabulary has no such
/// state.
#[test]
fn the_subagents_own_end_turn_finishes_it() {
let dir = tempfile::tempdir().expect("tempdir");
let subagents = test_subagents(&dir);
let mut translator = Translator::new(dir.path().to_path_buf(), Arc::clone(&subagents));
translate_lines(
&mut translator,
&[
r#"{"type":"assistant","message":{"content":[{"type":"tool_use","id":"toolu_task3","name":"Task","input":{"description":"helper"}}]},"parent_tool_use_id":null}"#,
r#"{"type":"stream_event","event":{"type":"message_delta","delta":{"stop_reason":"end_turn"}},"parent_tool_use_id":"toolu_task3"}"#,
],
);
let subagent = subagents.get("toolu_task3").expect("subagent started");
assert!(!subagent.is_open());
let lines = crate::session::transcript::read_after(&subagent.transcript_path(), 0)
.expect("read subagent transcript");
assert!(
!lines
.iter()
.any(|entry| matches!(&entry.event, Event::Status { state } if *state == SessionStatus::Idle)),
"a subagent's transcript must never carry Idle: {lines:?}"
);
assert_eq!(
lines.last().unwrap().event,
Event::Status {
state: SessionStatus::Exited
}
);
}
/// `stop_reason: "tool_use"` is the model about to call a tool, with
/// more of the turn still coming -- not an end.
#[test]
fn a_stop_reason_of_tool_use_does_not_finish_the_subagent() {
let dir = tempfile::tempdir().expect("tempdir");
let subagents = test_subagents(&dir);
let mut translator = Translator::new(dir.path().to_path_buf(), Arc::clone(&subagents));
translate_lines(
&mut translator,
&[
r#"{"type":"assistant","message":{"content":[{"type":"tool_use","id":"toolu_task4","name":"Task","input":{"description":"helper"}}]},"parent_tool_use_id":null}"#,
r#"{"type":"stream_event","event":{"type":"message_delta","delta":{"stop_reason":"tool_use"}},"parent_tool_use_id":"toolu_task4"}"#,
],
);
assert!(
subagents
.get("toolu_task4")
.expect("subagent started")
.is_open()
);
}
/// A background Task can be sent another message long after its first
/// turn ended -- a further child line for it reopens rather than being
/// dropped, and the same transcript and child translator carry on.
#[test]
fn a_line_after_finish_reopens_the_subagent_rather_than_being_dropped() {
let dir = tempfile::tempdir().expect("tempdir");
let subagents = test_subagents(&dir);
let mut translator = Translator::new(dir.path().to_path_buf(), Arc::clone(&subagents));
translate_lines(
&mut translator,
&[
r#"{"type":"assistant","message":{"content":[{"type":"tool_use","id":"toolu_task5","name":"Task","input":{"description":"helper"}}]},"parent_tool_use_id":null}"#,
r#"{"type":"stream_event","event":{"type":"message_delta","delta":{"stop_reason":"end_turn"}},"parent_tool_use_id":"toolu_task5"}"#,
],
);
let subagent = subagents.get("toolu_task5").expect("subagent started");
assert!(!subagent.is_open());
translate_lines(
&mut translator,
&[
r#"{"type":"assistant","message":{"content":[{"type":"tool_use","id":"toolu_more","name":"Bash","input":{}}]},"parent_tool_use_id":"toolu_task5"}"#,
],
);
assert!(subagent.is_open());
let lines = crate::session::transcript::read_after(&subagent.transcript_path(), 0)
.expect("read subagent transcript");
// Running, [prompt], Exited, Running (reopened), then the new line's
// own ToolStart -- the same transcript throughout, not a new one.
assert!(
lines.iter().any(
|entry| matches!(&entry.event, Event::ToolStart { tool, .. } if tool == "Bash")
)
);
assert_eq!(
lines
.iter()
.filter(|entry| matches!(
&entry.event,
Event::Status {
state: SessionStatus::Running
}
))
.count(),
2,
"expected one Running at creation and one at the reopen: {lines:?}"
);
}
/// Two subagents running at once keep two separate transcripts: tool ids
/// are unique but a `stream_event`'s content-block index is not, so
/// sharing translation state between them would cross their streams.
#[test]
fn two_parallel_subagents_keep_separate_transcripts() {
let dir = tempfile::tempdir().expect("tempdir");
let subagents = test_subagents(&dir);
let mut translator = Translator::new(dir.path().to_path_buf(), Arc::clone(&subagents));
translate_lines(
&mut translator,
&[
r#"{"type":"assistant","message":{"content":[{"type":"tool_use","id":"toolu_a","name":"Bash","input":{}}]},"parent_tool_use_id":"toolu_task_a"}"#,
r#"{"type":"assistant","message":{"content":[{"type":"tool_use","id":"toolu_b","name":"Read","input":{}}]},"parent_tool_use_id":"toolu_task_b"}"#,
],
);
let a = subagents.get("toolu_task_a").expect("subagent a");
let b = subagents.get("toolu_task_b").expect("subagent b");
let a_events = crate::session::transcript::read_after(&a.transcript_path(), 0)
.expect("read a's transcript");
let b_events = crate::session::transcript::read_after(&b.transcript_path(), 0)
.expect("read b's transcript");
assert!(
a_events.iter().any(
|entry| matches!(&entry.event, Event::ToolStart { tool, .. } if tool == "Bash")
)
);
assert!(
b_events.iter().any(
|entry| matches!(&entry.event, Event::ToolStart { tool, .. } if tool == "Read")
)
);
assert!(
!a_events.iter().any(
|entry| matches!(&entry.event, Event::ToolStart { tool, .. } if tool == "Read")
)
);
assert!(
!b_events.iter().any(
|entry| matches!(&entry.event, Event::ToolStart { tool, .. } if tool == "Bash")
)
);
}
#[test]
fn a_permission_request_becomes_an_allow_deny_question() {
let dir = tempfile::tempdir().expect("tempdir");
let mut translator = Translator::new(dir.path().to_path_buf());
let mut translator = Translator::new(dir.path().to_path_buf(), test_subagents(&dir));
let events = translate_lines(
&mut translator,
&[
@@ -896,7 +1319,7 @@ mod tests {
#[test]
fn denying_a_permission_sends_deny() {
let dir = tempfile::tempdir().expect("tempdir");
let mut translator = Translator::new(dir.path().to_path_buf());
let mut translator = Translator::new(dir.path().to_path_buf(), test_subagents(&dir));
translate_lines(
&mut translator,
&[
@@ -914,7 +1337,7 @@ mod tests {
// The real 2.1.237 shape, verified live: answers go back inside
// updatedInput, keyed by the question text.
let dir = tempfile::tempdir().expect("tempdir");
let mut translator = Translator::new(dir.path().to_path_buf());
let mut translator = Translator::new(dir.path().to_path_buf(), test_subagents(&dir));
let events = translate_lines(
&mut translator,
&[
@@ -966,7 +1389,7 @@ mod tests {
// in the event: a phone that had to read this dialect's tool input to
// find them would be the only place that knew how.
let dir = tempfile::tempdir().expect("tempdir");
let mut translator = Translator::new(dir.path().to_path_buf());
let mut translator = Translator::new(dir.path().to_path_buf(), test_subagents(&dir));
let events = translate_lines(
&mut translator,
&[
@@ -1014,7 +1437,7 @@ mod tests {
#[test]
fn images_in_tool_results_are_saved_and_referenced() {
let dir = tempfile::tempdir().expect("tempdir");
let mut translator = Translator::new(dir.path().to_path_buf());
let mut translator = Translator::new(dir.path().to_path_buf(), test_subagents(&dir));
// A 1x1 PNG, the smallest real payload worth round-tripping.
let png = "iVBORw0KGgoAAAANSUhEUgAAAAEAAAABCAYAAAAfFcSJAAAADUlEQVR42mNk+M9QDwADhgGAWjR9awAAAABJRU5ErkJggg==";
let line = format!(
@@ -1043,7 +1466,7 @@ mod tests {
#[test]
fn a_turn_result_reports_usage_and_returns_to_idle() {
let dir = tempfile::tempdir().expect("tempdir");
let mut translator = Translator::new(dir.path().to_path_buf());
let mut translator = Translator::new(dir.path().to_path_buf(), test_subagents(&dir));
let events = translate_lines(
&mut translator,
&[
@@ -1078,7 +1501,7 @@ mod tests {
#[test]
fn a_turn_started_by_another_agent_records_who_and_what() {
let dir = tempfile::tempdir().expect("tempdir");
let mut translator = Translator::new(dir.path().to_path_buf());
let mut translator = Translator::new(dir.path().to_path_buf(), test_subagents(&dir));
let events = translate_lines(
&mut translator,
&[
@@ -1112,7 +1535,7 @@ mod tests {
#[test]
fn an_ordinary_turn_carries_no_peer_note() {
let dir = tempfile::tempdir().expect("tempdir");
let mut translator = Translator::new(dir.path().to_path_buf());
let mut translator = Translator::new(dir.path().to_path_buf(), test_subagents(&dir));
let events = translate_lines(
&mut translator,
&[
@@ -1137,7 +1560,7 @@ mod tests {
#[test]
fn the_context_is_what_the_last_message_held_not_the_turn_added_up() {
let dir = tempfile::tempdir().expect("tempdir");
let mut translator = Translator::new(dir.path().to_path_buf());
let mut translator = Translator::new(dir.path().to_path_buf(), test_subagents(&dir));
let events = translate_lines(
&mut translator,
&[
@@ -1177,7 +1600,7 @@ mod tests {
// Note the snake_case keys -- the CLI's transcript file writes the same
// records in camelCase.
let dir = tempfile::tempdir().expect("tempdir");
let mut translator = Translator::new(dir.path().to_path_buf());
let mut translator = Translator::new(dir.path().to_path_buf(), test_subagents(&dir));
let events = translate_lines(
&mut translator,
&[
@@ -1207,7 +1630,7 @@ mod tests {
#[test]
fn a_failed_compaction_says_why_and_leaves_the_turn_running() {
let dir = tempfile::tempdir().expect("tempdir");
let mut translator = Translator::new(dir.path().to_path_buf());
let mut translator = Translator::new(dir.path().to_path_buf(), test_subagents(&dir));
let events = translate_lines(
&mut translator,
&[
@@ -1234,7 +1657,7 @@ mod tests {
#[test]
fn a_boundary_without_counts_says_so_rather_than_inventing_them() {
let dir = tempfile::tempdir().expect("tempdir");
let mut translator = Translator::new(dir.path().to_path_buf());
let mut translator = Translator::new(dir.path().to_path_buf(), test_subagents(&dir));
let events = translate_lines(
&mut translator,
&[
@@ -1254,7 +1677,7 @@ mod tests {
#[test]
fn an_error_result_surfaces_the_message() {
let dir = tempfile::tempdir().expect("tempdir");
let mut translator = Translator::new(dir.path().to_path_buf());
let mut translator = Translator::new(dir.path().to_path_buf(), test_subagents(&dir));
let events = translate_lines(
&mut translator,
&[
@@ -1275,6 +1698,55 @@ mod tests {
);
}
/// Running out of quota is a state, not a failure of the work.
///
/// The naive reading -- an error result like any other -- is what shipped
/// before this: the transcript said "Claude AI usage limit reached|…" in
/// red, which is neither readable nor actionable, and nothing above the
/// driver could tell it apart from a broken tool call.
#[test]
fn a_turn_stopped_by_the_usage_limit_says_so_and_carries_the_reset() {
let dir = tempfile::tempdir().expect("tempdir");
let mut translator = Translator::new(dir.path().to_path_buf(), test_subagents(&dir));
let events = translate_lines(
&mut translator,
&[
r#"{"type":"result","subtype":"error_during_execution","is_error":true,"result":"Claude AI usage limit reached|1788546972","usage":{}}"#,
],
);
assert_eq!(
events[0],
Event::LimitReached {
resets_at: Some(1_788_546_972.0)
}
);
}
#[test]
fn a_limit_the_cli_gave_no_reset_for_is_reported_without_one() {
let dir = tempfile::tempdir().expect("tempdir");
let mut translator = Translator::new(dir.path().to_path_buf(), test_subagents(&dir));
let events = translate_lines(
&mut translator,
&[
r#"{"type":"result","subtype":"error_during_execution","is_error":true,"result":"Claude AI usage limit reached","usage":{}}"#,
],
);
// Not a time this side invented: the meter is asked before anything is
// sent, and a made-up reset would only decide when to ask.
assert_eq!(events[0], Event::LimitReached { resets_at: None });
}
#[test]
fn a_reset_in_milliseconds_is_not_read_as_the_year_58000() {
assert_eq!(
usage_limit("Claude AI usage limit reached|1788546972000"),
Some(Some(1_788_546_972.0))
);
// And anything that is not the limit stays an ordinary failure.
assert_eq!(usage_limit("something broke"), None);
}
/// Pressing Stop is not a failure, and the CLI cannot tell you which it was.
///
/// An interrupted turn arrives as exactly the same shape a broken one does,
@@ -1286,7 +1758,7 @@ mod tests {
#[test]
fn a_turn_stopped_on_purpose_is_not_an_error() {
let dir = tempfile::tempdir().expect("tempdir");
let mut translator = Translator::new(dir.path().to_path_buf());
let mut translator = Translator::new(dir.path().to_path_buf(), test_subagents(&dir));
let stopped_result = r#"{"type":"result","subtype":"error_during_execution","is_error":true,"result":"Interrupted by user","usage":{}}"#;
translator.expect_interrupt();
@@ -1318,7 +1790,7 @@ mod tests {
#[test]
fn replayed_and_synthetic_user_text_is_skipped() {
let dir = tempfile::tempdir().expect("tempdir");
let mut translator = Translator::new(dir.path().to_path_buf());
let mut translator = Translator::new(dir.path().to_path_buf(), test_subagents(&dir));
let events = translate_lines(
&mut translator,
&[
Loaded 100 of 108 files, more files were not shown because too many files have changed in this diff. Show more