Author SHA1 Message Date
irisandClaude Opus 5 76fcbdccb9 iris: a List gives back its overscroll in the frame that found it
The last place in iris that corrected itself on a later frame, and the
item docs/IRIS_TODO.md carried from the Scroll change. Iris's rule:
"nothing in the framework should ever self heal because it should not be
drawn incorrectly in the first place. If you need 2 draws to get
something into the correct position then that should happen within the
same frame."

`clamp_to_content` measured the gap past the end of the content from the
edges the walk had just placed, wrote it to the anchor and asked for
another frame -- so one frame was drawn with the list past its own end,
and a fling that had already stopped was not going to ask for the frame
that fixed it. Now the walk outward from the anchor is `List::lay_out`,
`overscroll_gap` is a pure measurement of the same gap (no painter, no
redraw handle), and `draw` moves the anchor and walks a second time
inside the same frame.

One further pass always settles it: the gap comes from the edges the
first walk placed, so moving the anchor by it puts that edge exactly on
the viewport's, and the opposite end can only open a new gap when the
content is shorter than the viewport, which `overscroll_gap` declines to
touch. The second walk is paid only on an overscrolled frame and re-offers
every row the same cached-height box at a new offset, which `draw_inner`
dispatches as an O(1) move.

`Painter::draw_again` had no other caller and is removed with it, so the
framework no longer offers a way to ask for a corrective frame at all.

Simplification in the same change: a placement is one pinned edge plus a
height, so `Placement::edges(height)` gives the box and `place`'s
top-known and bottom-known cases stop being two copies of the same
arithmetic -- four match arms down to two.

Four tests draw no settling frame on purpose and fail without the change:
`fling_toward_the_start_stops_at_the_first_row` and the new
`scrolling_past_the_start_is_given_back_in_the_same_frame` (list.rs), and
`scrolling_past_the_first_row_settles_on_it` /
`scrolling_past_the_last_row_settles_on_it` (layer 1, top_edge.rs).

Verified: cargo fmt --check, clippy --workspace --all-targets clean,
cargo test --workspace and ./run-tests.sh green, the phone-shaped
headless window replaying flick-120hz.touch draws the transcript
correctly, and the arm64 release APK builds.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-08 17:26:51 -04:00
irisandClaude Opus 5 a00376994e iris: a Scroll measures and places its content in the same frame
Follows Iris on the previous commit: "nothing in the framework should
ever self heal because it should not be drawn incorrectly in the first
place. If you need 2 draws to get something into the correct position
then that should happen within the same frame. Layout should never be
frame dependent, it should be a pure function of the state."

So `Scroll::draw` no longer places its child against last frame's
content length and asks for a corrective frame. It draws the child once
at that length purely to measure it, then places it at the length just
measured, with the end-pin and the clamp applied only to the second
placement -- the measure-then-place idiom `Span::draw` and `List::place`
already use. Last frame's length survives as a hint that keeps the
common case cheap: when the content's length did not change the two
regions are identical, so the first call is `draw_inner`'s O(1) `mov`
and the second returns at its first line. Nothing drawn depends on the
hint.

Reverts the frame-loop change from the previous commit (a frame that
left anything dirty asked for another), which existed only to deliver
that corrective frame and would have made any widget marking itself
dirty spin at full rate.

Knock-on: an end-anchored Scroll now sits at its end on its first drawn
frame rather than its second, since the end-pin no longer waits for a
length. Two layout tests that scroll down from what they assumed was the
top now build their area with `at_end: false`, which is what they meant.

`List::clamp_to_content` is the only next-frame correction left. Its
comment cited Scroll's lag as precedent, which no longer exists; it now
says it is a deviation from the rule, and docs/IRIS_TODO.md carries it.

Verified: the layer-1 test draws no settling frame and still passes; on
the emulator the caret's bottom is 1509 against a bar edge of 1535, 26px
inside a 31px padding.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-08 17:12:26 -04:00
irisandClaude Opus 5 ba57086361 iris: a scroll area whose content grew asks to be drawn again
Iris's phone: typing newlines into the composer with the keyboard up
dropped the caret flush against the bar's bottom edge, eating the 12dp
padding, and closing the keyboard fixed it.

`Scroll::draw` offers its child last frame's content length on purpose,
so an ordinary scroll tick is an O(1) move rather than a redraw. The
comment claimed the lag self-corrects on the next frame; nothing asked
for that frame. A keystroke dirties the field, that frame draws it in a
box one line short of its text, and the tree is clean afterwards -- so
the stale placement is the last one drawn. The composer's text is
centred in its box, so one line short hung half a line past each end and
put the caret's line box a whole padding low. Closing the keyboard
rewrote the bar's inset, dirtied it, and forced the missing redraw.

`Scroll::draw` now calls `Painter::draw_again` when what it measured
differs from what it offered, and a frame that leaves anything dirty asks
for another frame on both backends -- `draw_again` sets its mark during
the update, after the input path's own check has run, so nothing asked
before this (which applied to `List::clamp_to_content` too).

Verified at layer 1 (the new test fails on the old code with the caret
exactly on the bar's edge) and on the emulator: the caret's bottom moved
from 1535 -- the bar's own bottom edge -- to 1509, 26px inside a 31px
padding, the remainder being parley's line box overhanging its line
height. `phone.rs` grew `--typed TEXT`, which enters text over frames
rather than preloading it; only that reproduces this.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-08 16:59:06 -04:00
irisandClaude Opus 5 1a9655414e docs: Iris's idea for retiring masked_by -- a Stack that names its mask
She asked whether `masked_by` earns its place, since
`.background(x).masked()` looks like the same thing. Measured: for a
square-cornered surface it is (identical to the pixel on the composer at
the phone's size and density), and what the pair cannot express is a
clip that is not a box, which is why the method stands for now.

Her suggestion, in IRIS_TODO.md's "Reconsider": let `Stack` name where
its mask comes from the way `StackSize::Child(n)` already names where
its size comes from, at which point `masked_by` and `Masked::shape` both
go away.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-08 16:38:05 -04:00
irisandClaude Opus 5 5e34dba2fd iris: a press only reaches the widget the pointer is on
Iris's 2026-09-08 report, both halves, and her own diagnosis of the
second: "tapping outside of something that a fling is currently active
for should have no code in common with the fling that could influence
it."

`run_sensors` runs a widget one frame after the pointer leaves it
(`ActivationState::End`, which is not `Off`) so `HoverEnd` can fire, and
`should_run` derived the press and wheel senses from raw button state
without consulting `hover`. That farewell frame carried a `PressStart` to
a widget the finger was nowhere near -- and a press on already-coasting
content is a catch, which commits to a pan with no `DRAG_SLOP`, so the
widget captured the pointer and swallowed the whole gesture. Its hover
was stale because a gesture that ends while captured returns from the
capture branch, which never reaches the loop that updates it.

Measured before the fix on the real screen: a fence flicked sideways,
then a finger down on a row 500px above it dragged 160px down the screen
-- the list moved by zero, the fence moved by zero, and the fence held
the pointer throughout. After: the list follows the finger and the
fence's fling carries on coasting, which is what she asked for and falls
out of the fix rather than being arranged.

`should_run` now requires `hover.is_on()` for every non-hover sense.
`Drop`/`Cancel` are unaffected -- they are delivered deliberately to a
widget that is not under the pointer, with an explicit `On`.

Also: the composer is clipped to its own bar rather than inside its
padding (`.masked_by(rect(BAR_FILL))` in place of a `.masked()` +
`.background()` pair) -- "the box should be clipped rather than the inset
text". A long message was being sliced mid-glyph 12dp in from the bar's
edge, leaving a band of bare surface above the cut.

New: `Scroll::is_scrolling`, the name `List` already uses; the phone
rig's `--message TEXT` and `--ime PX`, since the composer's overflowing
and keyboard-open states cannot otherwise be looked at headlessly.

Tests fail on the old code, one per layer:
`a_press_does_not_reach_a_widget_the_pointer_has_just_left` (sensors, no
screen) and `a_drag_away_from_a_coasting_fence_scrolls_the_list_and_
leaves_it_coasting` (the report itself, layer 1).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-08 16:23:32 -04:00
irisandClaude Opus 5 cc8148cbec deps: every crate to its latest version, wgpu 28 -> 30
`cargo upgrade --incompatible` in each of the nine workspaces here, then
`cargo update`. Most of it is version numbers only -- log, winit,
bytemuck, image, tokio, libc, android_logger, proc-macro2/quote, and syn
2 -> 3 with no source change. The wg-app-link submodule's twelve
dependencies were already at their latest majors, so that shared
repository needs no commit.

wgpu 28 -> 30 (and pollster 0.4 -> 1.0) is the part with API in it:
bind-group and vertex-buffer slots are optional now, `Instance::new`
takes an owned `InstanceDescriptor` carrying the platform's display
handle (the desktop passes winit's, since wgpu wants it for a GLES
surface presented on Wayland -- which is what this machine's fallback
produces; Android passes none), `RequestAdapterOptions` and
`SurfaceConfiguration` each gained a field kept at its historical value,
`get_current_texture` answers with an enum instead of a Result, and
`present` moved onto the queue.

The one that would not have failed at compile time: naga now requires
`@interpolate(flat)` on integer varyings, so `shader.wgsl`'s three u32
outputs were rejected at `create_shader_module` -- an abort on the device
rather than a build error. Flat is the only interpolation an integer can
have, so this states what the hardware already did.

Checked: build, clippy, fmt and tests in all nine workspaces (iris 196,
server 160); layer 2 screenshots on Vulkan and on force-gles, identical;
the arm64 release APK builds and the x86_64 bench ran a full
fling/stream/type/keyboard cycle on the emulator's GLES adapter.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-08 15:47:59 -04:00
irisandClaude Opus 5 a2e5e5881c docs: the two warnings the bench Android build still prints, and why they stand
Both are pre-existing and both are decisions rather than cleanups.
`show_diagnostics_overlay` and the Java overlay behind it are an escape
hatch that draws a report even when iris itself has stopped drawing --
the one case the in-iris diagnostics pane cannot cover -- so deleting
them to clear the warning would remove a fallback, and Iris has no
logcat on her phone.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-08 15:28:21 -04:00
irisandClaude Opus 5 c121bc0725 iris-android-app: FIELDS_PER_LINE is gated with the reader that uses it
`cargo ndk check` on the default features warned that it was never used:
its only reader is `line_fields`, which is `#[cfg(feature =
"transcript-screen")]` because the tabs demo links no `client-core` and
so has no ring to lay out.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-08 15:27:31 -04:00
irisandClaude Opus 5 fe7dc9c728 docs: Iris's second 2026-09-08 phone report, and the workaround list closed
RUST.md gets the report verbatim with what each of the four defects
actually was, the tests that pin them, and two traps worth not
re-finding (a fixed-coordinate tap that "failed" by 544px because it had
toggled a tool group, and a layer-1 repro that only reproduces inside a
`List`). IRIS.md and DECISIONS.md get the design half: one `Flinger`
whose seam puts the sign convention and the content's end with the
caller, a cancel as a first-class end to a gesture, and why a row is
drawn twice on the frame its height changes.

LAYOUT.md gains the two rules those turned on, since both govern the
layout rather than this pass: padding works in any container and is an
inset or an outset depending on how tight the parent's region is (Iris's
own words), and a widget offered a box it does not fit is drawn again at
its true box in the same frame rather than the next one.

IRIS_TODO.md's "worked around in tool.rs rather than fixed here" is gone
-- Iris, 2026-09-08: "There should never be workaround code." Two of the
four entries are ticked; the two that remain are missing capabilities
rather than defects being dodged, and each now carries a diagnosis of
what building it costs instead of a workaround: an overflow ellipsis
needs `TextBuffer` to have a displayed string distinct from its source
(parley has none of its own, and every byte-offset consumer -- spans,
`byte_at`, `Selection`, `apply_delta` -- moves if the buffer is
truncated), and selectable tool-card text needs a register/unregister
lifecycle across the three routes that rebuild a card, which is where a
stale `Selection` handle panics.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-08 15:26:12 -04:00
irisandClaude Opus 5 02b277e7ad iris: every scroll area flings, on either axis, through one Flinger
Iris, 2026-09-08: "Flinging doesn't work in horizontal scroll areas.
Flinging should be enabled by default in all scroll areas on android to
match composes behavior." Compose's `scrollable` attaches
`ScrollableDefaults.flingBehavior()` on every axis it is given and it is
not something a caller opts into, so neither is this.

`iris::sense::Flinger` is the fling `List` already had, taken out of it:
the `FlingCalculator` curve, the clock (started at the first tick, not
the release, so a caller on an explicit clock is not handed a fling that
has already expired), the incremental delta, Compose's two release
thresholds and the trace line. What it deliberately does *not* know is
which way a positive delta moves the content or whether there is content
left to move into -- a `List` scrolls its anchor one way and a `Scroll`
moves its `amt` the other, so the caller applies `tick`'s delta in its
own convention and calls `stop` at its own wall. `List` keeps
`fling`/`tick_fling`/`is_scrolling`/`cancel_fling` unchanged as a
surface, now three lines each over the shared type.

`Scroll` gains it, plus the two things a coasting widget needs and it
had no reason to have before: the display density (read from the painter
in `draw`, since the deceleration is physical -- a hardcoded 1.0 made a
one-second coast run for 45 on a list), and `PressState::scrolling`, so
a finger put down on a coasting fence stops it there from the first
sample rather than after `DRAG_SLOP`. `Scroll::drag` now answers whether
it started a fling, which is what `WidgetLike::scroll_area` needs to
call `UiData::animate` -- the same split `List::fling`'s doc describes,
and for the same reason: only the caller can reach the frame loop.

`Scroll::axis()` is public for a caller that found the widget rather
than built it.

Tests: `scroll.rs`'s three (a released pan coasts and decelerates on both
axes; both walls stop it; a press on coasting content catches it with no
slop), and `transcript-fixture/tests/fence_fling.rs`, which flicks a real
markdown fence in the real transcript screen and reads the fence's own
`Scroll` back out of what was drawn. Confirmed to fail with the release
arm removed ("the fence stopped dead at the release: 272 -> 272").

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-08 15:22:03 -04:00
irisandClaude Opus 5 fc82d9d7e8 iris: a cancelled gesture is not a release, and a row height is not last frame's
Three of the four defects in Iris's 2026-09-08 report, each with a
layer-1 repro that fails without the change.

**A gesture the platform takes away is now a cancel, not a release**
(`CursorState::cancelled`, `SensorUi::run_sensors`). Android mapped
`ACTION_CANCEL` onto the same arm as `ACTION_UP`, so the system's own
swipe up from the bottom edge to leave the app reached iris as a flick
released at speed and the transcript flung while the app was in the
background -- "leaving and reopening the app also randomly moved the
vertical scroll". A cancelled sample now delivers `CursorSense::Cancel`
to the capture holder *and* every widget still tracking the press,
clears both, and derives nothing else: no tap, no selection, no fling.
The harness's `TouchAction::Cancel` says the same thing, so it is
testable from a `.touch` file.

**A `DragGesture` ignores a `Cancel` when it is the one holding the
capture.** A cancel goes to every pressed widget that did not capture,
and one gesture is routinely driven by several of those -- a transcript
row's text block feeds `Selection`'s shared gesture, which captures
under the *list's* id, so the block is a "loser" on the very frame its
own pan committed. Acting on that released the pan the frame it started
(`catch_a_fling.rs` fails without the guard). With it, a row's block can
register the whole `drag_senses()` set, `Cancel` included, which is what
the doc on that set has always said a widget driving a gesture must do.

**A row whose measurement disagrees with the box it was offered is drawn
again at its true box, this frame** (`List::place`, both placements).
A row is offered its *cached* height and a `.background(rect(..))` fills
whatever box it is handed, so on the frame a row changed height its text
laid out at the new height and its background painted at the old one --
"collapsing and opening an edit card draws the card background a frame
late, so it looks closed even when there's text". The bottom-anchored
half used a `reposition`, which writes an offset and never a size, so it
could not fix it either.

**The nested-`Span` workaround in `tool.rs` is gone**, restoring the 4dp
inset a tool group holds its cards off its edge by. "A `Span` of
`Pad`ded children inside another `Span` places those children a slot out
of step" is **not reproducible on 2026-09-08** -- verified both with
`IRIS_TOOLS_EXPANDED=1 run-headless.sh transcript --shot` and with a new
layer-1 test.

Tests: `transcript-fixture/tests/gesture_cancel.rs` (three, including a
real code fence pushed into the screen so the pan has something to
capture it), `list.rs`'s
`a_row_that_changes_height_draws_its_background_at_the_new_height_immediately`,
`layout_tests.rs`'s
`a_span_of_padded_children_inside_a_span_draws_each_where_its_box_is`.
Each was confirmed to fail with the change backed out.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-08 15:15:29 -04:00
irisandClaude Opus 5 9e301f30c6 iris: ship an icon font subset, and delete the drawn mark
Iris asked why `mark` existed at all -- "the font should be working if
it's working for compose and nerd fonts are bundled". It was not: the
Compose app draws its icons from its own committed Nerd Fonts subset,
while iris, which bundles no font since 2026-09-07, was setting the
disclosure mark with bare geometric codepoints (U+25B8/25BE/25B4) out of
whatever face the platform resolved -- an empty box on her phone, a dot
on this VM. The 2026-09-07 note that "iris had no equivalent icon font to
keep" is the gap: it had none because it had never had one.

So iris ships the same kind of subset. iris/core/build-icon-font.sh is
the Compose script with its own GLYPHS list, writing a 992-byte
nerd_icons.ttf with three Material Design glyphs from the Mono face;
iris::icon names the codepoints; Family::Icons is how text asks for them.
The variant names an intention rather than a font name -- only TextData
knows what the file registered as, and it resolves it during shaping --
and it is a named family, never a generic one, so nothing falls back into
it for text and an icon cannot fall back out of it onto a system face
that happens to have the codepoint.

every_icon_is_in_the_bundled_font maps each constant through the shipped
font's charmap, which is the guard the script's "the two lists have to
agree" comment asks for. FontDiagnostics gains icon_family, so a build
whose font failed to register says so instead of drawing tofu; the
emulator reports icons=Some("Symbols Nerd Font Mono").

widget/mark.rs is deleted. It drew one correct triangle, but every
further icon would have been another rasteriser, and an icon as text
takes the size, colour and baseline of the line it sits in for free.

Looked at rather than only compiled: closed and open marks in
run-headless.sh phone --phone either side of a tap, and the collapse
bar's up mark under IRIS_TOOLS_EXPANDED=1.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-08 14:42:16 -04:00
irisandClaude Opus 5 341b7a5922 iris: a device change re-uploads its textures, and a mark is one texture per shape
The bench APK panicked on frame 1 on the emulator:

    iris panic at iris/core/src/render/texture.rs:461:22:
    texture slot 89 is not a live standalone image: None

widget::mark called Textures::add per widget, so a folded card per tool
call meant a standalone image, a bind group and a draw call each --
hundreds of copies of three pictures. Textures::reset, which the Android
surface-rebuild path calls for a genuinely new renderer, then threw the
slot numbering away with the pixels, leaving every one of those live
handles naming a slot nothing recognised. Its doc had said the only
standalone image in the workspace was tabs-ui's, "confirmed by grep" --
true when written, false the moment mark existed.

Textures::reupload replaces reset: queue every slot for upload again in
slot order, empty slots included, so the new device gets the same slot
numbering and a handle a widget has been holding still names its own
texture. The glyph atlas is no longer cleared on that path either, so an
app switch stops re-rasterising every glyph on screen.

Textures::shared(key, make) is one texture per description, keyed by a
SharedTextureKey the caller packs exactly rather than hashes. mark keys on
direction and colour: three mark textures for the screen, not one a card.

And the devlog can finally show a panic. After a crash, Dev Updater's
query starts the app process for the provider alone, so no activity ran,
so set_crash_dir never replayed the panic hook's file -- the Runtime tab
held one line, the provider announcing itself. DevLogProvider.nativeReady
takes the files directory and does the replay from onCreate; the hook also
saves the dying run's last 80 lines beside the panic, read through a new
non-blocking LogRing::try_tail_text so a panic holding the ring's lock
cannot deadlock the hook.

Verified on this checkout's emulator: opens clean, survives 33 full-screen
scrolls back through the fixture, image_bind_group_creates_prev=1; a real
panic replays into the next launch, and a hand-written last-panic.txt
replays in a process started by a provider query with no activity.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-08 14:27:09 -04:00
irisandClaude Opus 5 c8785b6091 docs/IRIS.md: it is the log of how iris is being built, not an API changelog
Iris, 2026-09-08: 'any major additions or design things should be added
there, not just public API stuff. You may as well remove the public API
bit at this point.' Widened the header, pointed AGENTS.md at the new
scope, and added the design point behind the scroll bug -- a cached
measurement needs its own value for 'not measured yet'.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-08 14:03:29 -04:00
irisandClaude Opus 5 9c560e3492 docs: tick the drawn chevron
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-08 14:01:42 -04:00
irisandClaude Opus 5 e5a90c6135 iris: mark() -- a drawn disclosure triangle, instead of a codepoint the phone lacks
The tool cards' open/closed marks were U+25B8/25BE/25B4 in whatever face
resolved. That worked while iris bundled its own fonts; since the move to
the platform collection on 2026-09-07 Iris's phone draws an empty box and
this machine draws a dot -- UI_RULES' 'don't rely on characters the
platform might not have'.

iris::widget::mark rasterises one oversampled, antialiased triangle into
the ordinary texture path and scales it into the box the caller asks for,
so it needs no new primitive and is correct at any density. Its two tests
check the shape points where it was asked to and leaves its corners
clear, which is the half nobody would look at on a device that renders it
wrong.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-08 14:01:31 -04:00
irisandClaude Opus 5 af1b0c5ab2 docs/RUST.md: what the folded-card sanity check found
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-08 13:57:35 -04:00
irisandClaude Opus 5 cce4b28324 iris: a scroll area no longer opens at the end of content it has not measured
Scroll::content_len was 0.0 both for 'nothing here' and for 'not drawn
yet', so the first frame's clamp found a scroll range of zero, read
amt == len as 'sitting at the end' and set snap_end -- and the next
frame, now knowing the real length, jumped to it. On a phone that put a
code fence at the end of its longest line, mid-word, before anybody
touched it. It is an Option now, and the clamp does not answer a question
it cannot yet answer.

Which edge an area opens at is also a caller's decision rather than a
default: scrollable_on starts at the beginning (what is read),
scrollable_to_end pins to the end while content grows (what is typed --
the composer), both through one Scroll::new(inner, axis, at_end).

And tool.rs's raw_block pans sideways again: the 2026-09-06 'a
scrollable_on(Axis::X) around a non-editable Text draws nothing' defect
does not reproduce, most likely fixed by the shaped-mask work, so a long
command is readable rather than clipped.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-08 13:57:16 -04:00
irisandClaude Opus 5 94d8373289 iris-android-app: the bench observes the fling instead of driving it at 60Hz
The fling phase called List::tick_fling itself every 16ms, so on a 120Hz
phone every second frame redrew a position already drawn -- Iris saw the
benchmark scroll visibly less smoothly than her own finger, and it was
the rig rather than the renderer. A real fling is advanced once per frame
by UiData::tick_animations from the frame callback, so the phase now
starts one the way a gesture does (fling + animate) and polls
is_scrolling to know when it settled. ANIM_STEP_MS becomes POLL_MS,
which is what it always was here.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-08 13:48:17 -04:00
irisandClaude Opus 5 8310431497 iris: pin the nested-scroll axis rule, and record the capture fix in RUST.md
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-08 13:46:39 -04:00
irisandClaude Opus 5 b863f9f3df iris: a capture cancels every other gesture, and the pointer leaves UiRenderState
Two defects Iris reported from her phone on 2026-09-08, one root cause
each, both in how a gesture ends.

A widget that takes pointer capture cuts every other widget off from the
press completely -- no PressEnd, no Drop -- so anything else tracking it
was left with an open gesture at a stale origin, and the *next* touch
anywhere was measured from that origin. That is the transcript jumping on
a tap after a code fence was panned sideways. CursorSense::Cancel is the
missing state: delivered once to each loser of a capture race, the way
Android sends ACTION_CANCEL and the web sends pointercancel.

And  registered click_or_drag|unclick, which never matches
a Drop, so a Scroll that had captured never saw its own gesture end and
stayed panning from where the finger left. That is the horizontal snap
back. CursorSense::drag_senses() states the rule once for every widget
driving a DragGesture instead of per call site.

The pointer's own state (who holds capture, who is tracking the press) no
longer lives in a Mutex on UiRenderState. It is Event::Global for the
cursor senses -- owned by the event manager that runs the dispatch,
reached by &mut, with a per-dispatch PointerRequests slot for handlers --
per Iris: never reach for locks first, and input-wide state belongs to
the general input handler. What had forced the lock was a Data: Send
bound on task_on that nothing needed; the spawned future never sees the
event's data.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-08 13:44:51 -04:00
irisandClaude Opus 5 cdeb7b0857 docs/RUST.md: the work is done inline, not handed to subagents (Iris, 2026-09-08)
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-08 13:31:02 -04:00
irisandClaude Opus 5 476609e1d3 docs/RUST.md: Iris's 2026-09-08 phone report -- tap-jump, nested scroll, folded cards, and the bench's 60Hz gesture
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-08 13:27:46 -04:00
irisandClaude Opus 5 2756087e1c emulator: settle on GLES, and make every run say which adapter drew it
Iris's call, after the guest measurement: the emulator is a GLES rig and
nothing chases hardware Vulkan in it; the Vulkan path is covered by the
desktop build and by her phone.

Nothing had to be forced. The emulator has no hardware Vulkan at all --
its only Vulkan is SwiftShader in software -- and its GLES is the host's
real RX 7900 XT through virgl at ES 3.1, so iris's existing runtime
fallback lands there by itself. Verified end to end with an ordinary
debug APK: "no Backends(VULKAN|...) adapter on this device, falling back
to GLES", then "Android Emulator OpenGL ES Translator (virgl (AMD Radeon
RX 7900 XT ...)) (Gl, OpenGL ES 3.1 ...) on Backends(GL)". So the
emulator and the phone run the same binary, differing only in what it
finds -- which is the point, and `force-gles` must not be reintroduced to
arrange the emulator's backend.

What changed:
- The Android renderer logs the full adapter line at startup, as the
  desktop already did. Only the backend enum was logged, which cannot
  tell `Gl` on the host's GPU from `Gl` on SwiftShader; the same rule was
  written on one member of the pair and not the other.
- `run-bench.sh` prints that line before any number.
- build-apk.sh, Cargo.toml and RUST.md's "Vulkan in the emulator" carried
  the stale premise that the emulator defaults to software Vulkan and has
  to be steered off it. The recipes are marked superseded rather than
  deleted, since the record of why host Vulkan is unavailable is still
  worth having.
- Drive-by: an `#[allow]`-free clippy warning in android/platform.rs
  (useless JObject conversion) that only appears on the android target.

No Vulkan requirement was found in iris itself to remove: neither backend
asks for a feature, `device_limits()` stays at wgpu's defaults with the
compute fields zeroed, and both probe rather than expect an adapter.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-08 13:12:21 -04:00
irisandClaude Opus 5 af7d5f3782 docs: what the emulator actually gives a GPU app, measured in the guest
`gpu-probe` cross-compiled with cargo-ndk and run inside a default
`emu up`: the guest's GL adapter is the host's real RX 7900 XT through
virgl, reporting OpenGL ES 3.1 with compute shaders, 1024 invocations per
workgroup and 64 KB of workgroup storage -- the same numbers the desktop
gets. `Backends::PRIMARY` still finds nothing, because the guest's only
Vulkan is SwiftShader.

So GPU acceleration in the emulator is not a thing to get working; it is
the default, and it is GLES. What is missing is GPU-accelerated Vulkan,
and the Venus retry on mesa 26.2.2 fails exactly as it did on 26.1.7 with
no newer emulator package to try.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-08 13:04:15 -04:00
iris db0a41a7cd gpu-probe: say whether each adapter has compute, not just the preferred one
Compute is a downlevel capability rather than a feature -- unconditional
on any Vulkan 1.0 device, GLES 3.1 and up -- so the question shadows and
blur raise is what the *weakest* adapter iris can fall back to offers.
Both here answer yes: Venus and virgl each report COMPUTE_SHADERS, 1024
invocations per workgroup and 64 KB of workgroup storage, virgl because
it is ES 3.2. The only no-compute machine in this project is the
emulator's SwiftShader software GL at ES 3.0.

RUST.md gains that table, why DRM native context is unrelated to it, and
what each of shadows/blur/paths actually needs -- only vello proper turns
the compute question on.
2026-09-08 12:55:47 -04:00
irisandClaude Opus 5 b9924e7617 iris: the GPU test's crash was the Vulkan loader unloading Mesa, not wgpu
`mask_sdf` SIGSEGVd after printing `test result: ok`, and the workaround
was to hand the device to the process with `mem::forget` on the reading
that "dropping a wgpu device on Venus segfaults". Every part of that
except the symptom was wrong.

`rigs/gpu-probe`'s new `teardown` bin is the experiment, one variable per
mode: the same open-and-close exits 0 on the main thread and SIGSEGVs on
a spawned one; it needs no GPU work and no device, only an instance; raw
`ash` does it with no wgpu involved at all; and keeping the instance
alive fixes it. Destroying the last VkInstance makes the loader dlclose
the ICD, and Mesa's ICD here registers a pthread_key_create destructor
into its own text without `-z nodelete`, so glibc calls it through
unmapped memory when the thread exits. libtest runs every #[test] on a
spawned thread, which is the whole reason this looked like a drop bug.
`VK_LOADER_DISABLE_DYNAMIC_LIBRARY_UNLOADING=1` confirms the mechanism.

So the fix is one `wgpu::Instance` for the process -- what wgpu asks for
anyway -- and the device, queue and everything else drop normally again.
The escape and its paragraph of reasons are gone.

Also: the machine-level graphics notes duplicated in docs/RUST.md,
run-headless.sh and two source comments now point at the
`this-machine-graphics` skill, which is the only copy.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-08 12:22:47 -04:00
irisandClaude Opus 5 f014e8d9cf docs/RUST.md: point at the this-machine-graphics skill
The GPU findings from 2026-09-08 would bite any project on this machine,
not just this one, so they are now a skill (AGENTS.md's own rule about
where a machine-wide lesson belongs). This section keeps the iris- and
port-specific half and names the skill for the rest.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-08 12:09:45 -04:00
irisandClaude Opus 5 0ccc444246 iris: cut test debug info, say which adapter drew, and log on the desktop
Three findings from one morning, all of them things that were invisible
rather than wrong. docs/RUST.md's two new sections have the full account.

**`cargo test --workspace` was taking half an hour, and it was debug
info.** rustc's default `debug = true`, times eight test binaries each
statically linking the whole wgpu + naga + winit + parley graph, means
every one of them gets a private copy of that graph's DWARF written into
it: the linkers for one run had written ~54 GB between them and were
still going at thirty minutes -- the worst single one 16.9 GB for one
test binary -- leaving an 88 GB target/. It was not CPU: the machine was
87% idle, and rust-lld's threads were in D state in btrfs
`handle_reserve_ticket`, blocked on space reservation at 83% full. So
`debug = "line-tables-only"` on both `profile.dev` and `profile.test` --
both, because `cargo test` builds dependencies under one and the test
targets under the other. Cold, with all 19 suites run: 69 s and a 3.7 GB
target. Backtraces keep file and line; `RUSTFLAGS="-C debuginfo=2"` per
run buys back variable inspection when a debugger actually needs it.

**The desktop had no logger at all**, so every `log::` call on that side
went to `log`'s no-op default -- including the GLES fallback warning
added hours earlier. `DefaultApp::run` installs a stderr logger
(`src/default/logging.rs`, no new dependency: a level and a line is a
page of code against env_logger plus its filter dialect), and the
renderer now says which adapter won at `info`. That line is the point:
with a silent fallback, a layer-2 screenshot rendered by llvmpipe and one
rendered by the host's GPU are the same PNG, and which one it was is
exactly what the screenshot is being taken to judge.

**`tests/mask_sdf.rs` is a render pass now, not a compute pass.** It
asked for `adapter.limits()` because `iris_core::device_limits()`
deliberately zeroes the six `max_compute_*` fields -- a decision on
record since 2026-09-05, which this quietly worked around instead of
following. It now asks for what iris asks for and calls the function from
the fragment stage, where the renderer calls it. The compute pass was
*not* why it crashed, and the record should not say it was: the rewrite
crashes identically. What the crash is: dropping a wgpu device on this
VM's Venus adapter segfaults, after the test has produced its answer
(worst CPU/shader disagreement 5.8e-6). Narrowed -- plain Vulkan creating
and destroying five VkDevices on the same adapter is clean, and the same
binary with Vulkan hidden falls back to GL and exits clean. Worked around
at `Gpu::leak`, with the reason and the delete-me condition written
there.

`rigs/virtgpu-probe` is the new rig behind the Venus half: which capsets
the host offers (0x16 -- VIRGL, VIRGL2, VENUS; no capset 6, so no DRM
native context without host-side work), whether the device has compute
(it does: 1024 invocations/workgroup -- the "no compute" finding on
record is about the Android emulator's SwiftShader, a different machine),
and whether plain Vulkan teardown is clean.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-08 12:06:02 -04:00
irisandClaude Opus 5 c6da735134 docs/RUST.md: the ABI-cache half of the build-apk.sh box is done (4f6ec3a)
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-08 11:17:15 -04:00
irisandClaude Opus 5 4f6ec3a900 iris-android-app: clear Gradle's native-libs cache, so an ABI switch takes
`build-apk.sh` already removed `app/src/main/jniLibs` before each build,
with its own comment saying why. It does not reach Gradle's own copy:
`mergeReleaseNativeLibs` is up to date against its cached inputs, so a
build that switches ABI packages the previous one. An `--abi x86_64`
release APK containing `lib/arm64-v8a/libmain.so` installed fine and
aborted at startup with `Could not get adapter!: NotFound {
active_backends: VULKAN }` under libndk_translation -- which reads
exactly like the phone's own Vulkan problem and is nothing of the kind.
It cost an hour on 2026-09-07 and was written down rather than fixed.

Scoped to the three native-lib directories rather than all of
app/build, so an ABI change costs the native merge and not the whole
Gradle build. Verified on the case that produced it: this checkout held
an x86_64 libmain.so from emulator work, and `./build-apk.sh release
--abi arm64-v8a` produced an APK whose only .so is
lib/arm64-v8a/libmain.so (7,518,840 bytes) -- that APK is
ai-app-bench a012ff9.

docs/RUST.md's queue box keeps its second half open: the 648 MB debug
bench APK still will not install.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-08 11:17:07 -04:00
irisandClaude Opus 5 38bf6309cb iris: a mask is a shape, not a rectangle -- .masked_by, and touch obeys it
Iris, on the code fence: "the code block scrolling currently masks in an
inner rectangle. Ideally masks should have a shape associated with them,
rounded rectangle being one of them ... so that the mask becomes the
parent container with rounded edges. Make sure alpha works properly with
it, eg. on the corners where alpha should be decreased / multiplied."

`Mask` is now `{ primitive, parent }` -- the slot of a primitive already
written, plus the mask this one nests inside. The fragment stage
evaluates that primitive's own coverage at the masked pixel, through the
same `rounded_rect_coverage` a drawn rect goes through, and multiplies it
into the alpha along the whole `parent` chain. Nothing about the shape is
copied, so a rounded container's corner and its children's clipped corner
are one piece of arithmetic and cannot drift; two nested feathers dim a
pixel twice, which is the multiply she asked for.

`.masked()` is unchanged for callers: it writes an undrawn rect
(`Drawn::No`/`NOT_DRAWN` -- owned, moved, resized and freed like any
other primitive, simply never rasterized) and points at that, so square
clipping is the same mechanism rather than a special case. New
`.masked_by(shape)` draws `shape` behind the content in its own layer and
clips to the first primitive it drew, with no radius written twice; it
replaces `.masked().background(w)`, which drew both and clipped to the
box. `transcript-ui`'s `BlockFrame::Verbatim` is the first caller.

Hit-testing applies the shape (`SensorUi::run_sensors` ->
`UiRenderState::mask_admits`, coverage above one half, which is where the
drawn edge is), as well as the widget's own box -- the two ask different
questions and both have to hold. `primitive_corners` is a floor-for-floor
transliteration of the shader's `corners_of`, not `region.to_px()`: the
phone's 2.55 density puts nothing on a whole pixel, and skipping the
rounding disagrees with the pixels by up to one along each edge.

A mask's shape must be a rect, asserted by name in `set_mask_to`. A glyph
would need a CPU-side alpha plane before the hit test could agree with
the shader, and a standalone image a bind-group switch the fragment stage
cannot make. So no texture mask exists; the branch where one would go is
in both copies of `mask_coverage`. docs/LAYOUT.md's section end lists this
and the three other places the code is narrower than the design.

Tests. Layer 1, `layout_tests.rs`: the child's coverage swept across the
container's corner arc equals the container's own exactly; nested masks
multiply rather than intersect, asserted where both feathers are partial,
which is the only place the two differ; a press in a rounded-away corner
misses while one inside the curve and one on a straight edge hit; and
`a_plain_mask_still_clips_to_a_square_box`, the half this had no reason to
touch. The first version of the corner test swept the straight chord
between the arc's ends, which lies inside the circle everywhere -- it
proved nothing and said so, which is why it counts both sides now.

`iris/tests/mask_sdf.rs` is the only test here that needs a GPU: it lifts
`distance_from_rect` and `rounded_rect_coverage` out of
`iris_core::SHAPE_SHADER` by name -- lifted, not copied, since a copy
would be edited alongside the shader -- and runs them in a compute pass
over ~200k points at five radii against `iris_core::rounded_rect_coverage`.
Worst disagreement under 1e-5; the negative control (`+ 0.01` inside the
shader's smoothstep) fails it at 0.03.

Layer 2 for looking: `./run-headless.sh phone --phone --shot /tmp/mask.png
--seconds 6 -- -p transcript-fixture` draws the fixture's horizontally
scrolled code fence clipped on the curve at both top corners.

Two things found on the way and fixed here:

- The winit backend had the defect the Android one was fixed for in
  85869d0 -- `Backends::PRIMARY` and an `.expect` on the adapter. This
  VM's Venus device disappears when the host runs out of virgl contexts,
  which happened mid-task, and layer 2 aborted with `Could not get
  adapter!` while GL sat there working. It probes and rebuilds the
  instance on `Backends::GL` exactly as Android does now, and the request
  names the backends it tried. The rule had been written on one member of
  a set of two.
- `active_primitive_count` counted mask shapes, so `iris::frame`'s
  `primitives=` -- a number Iris reads off a phone report as "how much is
  on screen" -- would have gained one per masked widget.

`widget_trait!` now accepts a `///` doc comment on its functions, since
`masked_by` is public API and rustdoc is where a contract is read.

docs/LAYOUT.md, docs/RUST.md (both queue boxes, the commands, and where
the GPU test sits among the three layers), docs/IRIS.md and
docs/IRIS_TODO.md ("Masks defined relative to each other", now closed).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-08 02:18:57 -04:00
irisandClaude Fable 5.1 3eb0e033d5 docs: tick report hygiene and bench header (commits 7485d78, b8ea723)
Both docs/IRIS_TODO.md's night bullets and docs/RUST.md's queue items
covered by the two client-core/iris-android-app commits above.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-07 22:26:49 -04:00
irisandClaude Fable 5.1 b8ea723718 iris-android-app: Copy report always copies; restore the header's text size
Two of the phone's 2026-09-07 night reports (docs/IRIS_TODO.md):

Copy report used to silently decline ("nothing to copy -- run the
benchmark first") whenever no benchmark had run yet, which read on the
phone as the button being unhittable until Diagnostics was pressed first
-- UI_RULES's "a failure is reported where it happened" failure, since it
declined with no visible effect. It now always copies something: with no
benchmark run yet it copies the diagnostics pane's own text instead (which
needs no prior button press either), with a first line saying so, and in
every case appends the ring's tail (LogRing::tail_text,
COPY_REPORT_TAIL_LINES lines, previous commit) instead of the whole ring,
which was the other half of "causes a lot of lag" pasting it into a
message box. app_log.rs wires iris::diagnostics::trace_enabled into the
ring filter that commit added.

The header's four controls no longer fit one row at HEADER_TEXT = 18, and
a previous agent had shrunk it to 13 to make room -- exactly what
UI_RULES forbids (never shrink text to fit a layout). Restored to 18 and
split bench_controls into two rows instead (run+copy, diagnostics+trace),
doubling the header's own height rather than the outer layout's reserved
space (top_bar already sizes to its own content). Checked on this
checkout's emulator: ui-trace's --field box shows two clean, non-
overlapping rows, and a screenshot shows the restored size reading
clearly; a Copy report tap with nothing run yet now logs "copied to
clipboard" instead of declining.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-07 22:26:27 -04:00
irisandClaude Fable 5.1 7485d78d50 client-core: filter the ring's Debug/Trace lines to iris's own targets
Iris's phone report (docs/IRIS_TODO.md, 2026-09-07 night): the ring held
1339 lines and dropped 4050 more, almost all of it naga::front/wgpu_core/
jni logging at Debug unconditionally, because RingLogger accepted every
target at whatever level `log`'s own max was set to. The trace gate added
in 992c472 only covers iris's own debug! call sites, not a dependency's.

ring_accepts() is the one filter, applied in RingLogger::log rather than
per callsite: Info and above always rings, from anywhere (a dependency's
real warning is worth keeping); Debug and Trace ring only from `iris`/
`client_core` targets, and only while tracing is on. Tracing itself is
`iris::diagnostics::trace_enabled`, passed into RingLogger as a plain
`fn() -> bool` rather than called directly, since client-core sits below
iris and must not depend on it -- the same reason `inner` (the platform
logger) is already injected rather than chosen here.

Also adds LogRing::tail_text and COPY_REPORT_TAIL_LINES (150, named and
reasoned at the constant) for the next commit's Copy report trim.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-07 22:26:17 -04:00
iris b87f5a597e iris: a finger put down on a moving list catches it at that sample
Iris, from the phone (docs/IRIS_TODO.md, 2026-09-07 night): "sometimes
when I try to catch it while it's still moving (particularly if I drag)
then it fails to stop & snap to where finger is." The fling did stop on
the down -- `Selection::drag` has cancelled it since the fling landed --
but the *gesture* then went through `DRAG_SLOP` like any other press, so
for the first few frames the finger was down and the content under it
did not move. Compose does not do that: `scrollable`'s
`startDragImmediately` is `isScrollInProgress`, and the drag starts on
the DOWN with no slop.

So `DragArbiter::press_start` takes a `PressState` -- what the target
looked like when the press landed, `already_selected` and `scrolling` --
and a press on moving content enters `Panning` immediately. A catch that
is released without ever moving is `Released(None)`: not a `Tapped`,
because Compose consumes that DOWN and no click detector under it sees
the gesture, so stopping a fling must not also follow the link it landed
on; and not a velocity, because there is none to hand on. The moment it
moves anything it is an ordinary pan release again and flings normally.

`DragGesture::starts_press` is the one rule for "this frame opens a
press", read by `handle` and by `Selection::drag` -- which has to prepare
its list (cancel the fling, report whether there was one) on exactly the
frames `handle` will call `press_start` on, including the recovery frames
where no `PressStart` ever arrived.

It deliberately does **not** special-case `PressStart` to true, which is
the defect the layer-1 test found: one touch-down reaches every sensor
under the finger, and a transcript row's block and the tool row
containing it drive the same shared `DragGesture`, so `handle` sees one
`PressStart` twice. Restarting on the second delivery re-read
`PressState` after the first had already acted on it -- the fling was
cancelled by then, `scrolling` came back false, and every catch quietly
became an ordinary slop-waiting press again.

Tests. Layer 1, `transcript-fixture/tests/catch_a_fling.rs`: the
recorded 120Hz flick, 150ms of fling, then a down and three 2px moves --
the content tracks the finger sample for sample
(`a_press_on_a_flinging_list_pins_the_content_to_the_finger`, which
fails at the parent commit with "the content 0.0px"); a catch released
without moving neither taps nor flings; and the half this had no reason
to touch, `the_same_small_drag_on_a_settled_list_moves_nothing` -- 6px
total is inside `DRAG_SLOP`, so making every press pin the content would
pass the first test and take the slop away from every ordinary one.
Unit, in `sense.rs`: the catch pans from the first sample, the same
press on settled content stays undecided, a catch that drags still
flings, and the double-delivered `PressStart` stays one press.
2026-09-07 22:18:16 -04:00
irisandClaude Fable 5.1 80a75c128e docs/RUST.md: who owns the killed agents' diff, and the stale worktree note
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-07 22:08:35 -04:00
irisandClaude Fable 5.1 50e69995b6 docs/RUST.md: the emulator crash loop was the missing GLES fallback, with the panic-hook note
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-07 21:50:06 -04:00
irisandClaude Fable 5.1 85869d02f8 iris: the Android renderer falls back to GLES, and every failure reports
The bench app crash-looped on this checkout's emulator with the default
features (RUST.md's queue item). Not the surface lifecycle and not "once
backgrounded": a build without `force-gles` never got a first frame.
`AndroidRenderer::new` asked wgpu for `Backends::PRIMARY`, which does not
contain `GL`, and this emulator advertises a Vulkan ICD with no adapter
behind it -- `NotFound { active_backends: VULKAN, no_adapter_backends:
VULKAN, supported_backends: VULKAN | GL }`, `.expect`ed, so SIGABRT, so
the launcher restarts it. iris was refusing a device whose only usable
adapter is a GLES one.

It now probes for a `PRIMARY` adapter and rebuilds the instance on
`Backends::GL` when there is none. The probe runs on an instance that
never touches the window on purpose: **an Android window can be
connected to one graphics API only**, so one instance carrying both
backends fails worse -- measured here on the way to this fix, Vulkan's
`vkCreateAndroidSurfaceKHR` claims the window in `create_surface` and
the GLES surface from the same window then reports `In
Surface::configure / Invalid surface`, aborting a frame later in
`Surface::get_current_texture_view`. Vulkan still wins wherever it has
an adapter (`PowerPreference::None` does not sort, and Vulkan is
enumerated first), so nothing changes on the phone.

Second half, the same rule applied to the whole set: the surface,
adapter and device requests all report through the `Result<Self,
String>` this function already returns, where two of the three used to
panic. `surface_changed` puts that string on screen and in the log
ring, which is what the Result was added for.

Emulator evidence (API 36 x86_64, debug): after, `iris renderer: no
Backends(VULKAN | METAL | DX12 | BROWSER_WEBGPU) adapter on this
device, falling back to GLES` then `new renderer built (Gl)` and
frames. Clean on both the default and a `force-gles` build for the
cases this had no reason to touch: two background/return cycles,
rotation there and back (the `already_live=true` reuse branch), a
background/return after the rotation, and cold starts. Vulkan could not
be exercised here -- that this emulator has no Vulkan adapter is the
defect itself.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-07 21:50:02 -04:00
irisandClaude Fable 5.1 f99ae4c366 iris-android-app: a panic hook, so an abort says something Iris can read
Checked before writing anything: under `panic = "abort"` (this crate's
Cargo.toml) a panic's message reaches the tombstone's `Abort message`
and nowhere else -- not `log`, so not `client_core::log_ring`, so not
Dev Updater's Runtime tab. That tab is the only surface Iris has on a
phone with no `adb`, so every `assert!` and `expect!` in these builds
has been failing silently as far as she is concerned; the adapter crash
fixed in the next commit looked like the app simply relaunching.

`install_panic_hook` (called from `app_log::install`) writes the
message and its location at `error` level. The ring is memory only and
the process is about to die, so it also writes `last-panic.txt` in the
app's private directory; `set_crash_dir`, called from
`nativeSetFilesDir`, replays that into the ring at `error` level on the
next start and deletes it. A crash loop therefore explains itself in
the run that is still up, which is the run somebody can look at.

Verified on this checkout's emulator against the unfixed renderer:
`iris panic at .../render.rs:140:14: Could not get adapter!: NotFound
{...}` on the run that died, and `iris app log: the previous run died
-- ...` on the next one.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-07 21:49:45 -04:00
irisandClaude Fable 5.1 203f53470c iris: one primitive arena all layers share, with placement in a storage buffer
A mask is about to reference a primitive already drawn and evaluate it at
the masked pixel (docs/LAYOUT.md's "Masks with a shape"), which the data
layout could not answer: a primitive's placement lived in its layer's
*vertex* buffer, invisible to the fragment stage, and `rects`/`glyphs`
were per layer too -- so a mask whose shape is a rounded container in one
layer, clipping content a `Stack` put in another, would have read the
wrong layer's rect with nothing on screen to say so.

So the instances and the per-primitive data become one arena
(`UiRenderState::primitives`), bound once per frame; a layer keeps only
its draw *order*, which is what its vertex buffer now is -- one `u32`
slot per instance instead of eight attributes. The vertex stage reads the
placement it is drawing from `instances[slot]`; the fragment stage can
read any other primitive's from the same buffer, which is what the mask
work needs and the reason there is no second copy for masks.

Arena slots are stable (nothing is compacted), so a `Mask` can hold one
across frames. A slot freed during a redraw is therefore not reusable
until every layer's order has been compacted around it -- otherwise the
reused slot would draw twice, once through the stale order entry -- which
is what `Primitives::freed` and `UiRenderState::apply_free` are. That
compaction moved out of `UiRenderNode::update` into `UiRenderState::
update`: it is bookkeeping over `active`, not GPU work, and the harness
(which has no renderer) needs it too.

Same 164 tests, the `--phone` screenshot unchanged.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-07 21:21:55 -04:00
irisandClaude Fable 5.1 b38e797db3 docs: phone report 2026-09-07 night -- catching a fling, silent Copy report, third-party debug flooding the ring
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-07 21:18:29 -04:00
irisandClaude Fable 5.1 92985ba8e3 iris/Cargo.lock: the log entry for iris-android-app regenerated after the uploader's removal
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-07 21:09:35 -04:00
irisandClaude Fable 5.1 181ba64606 docs/REVIEW-2026-09-07.md: every finding's status after the fix pass
13 fixed, 6 moot or deferred, 2 not done on purpose. Each finding gets its
own Status line in place rather than a summary at the end, so a reader who
arrives at a finding sees what happened to it; the header carries the
counts and the six commits.

The moot ones are all in the phone-logging route 06b8a1f deleted (D2's
unbounded `POST /client-log` body, D3's silently dropped lines, R3's three
copies of one wire contract, R4's `build.rs`, and the `client_log_time`
duplication) -- the app hands its log to Dev Updater through an on-device
ContentProvider now, so there is nothing left to bound or share. Two more
are deferred to the devlog agent because `iris/android-app/**` and
`client-core/src/log_ring.rs` were open under it this pass.

The two left undone are deliberate. R2 (a mask clips drawing but not
hit-testing) waits on docs/LAYOUT.md's mask redesign, since intersecting
a chain in `resolved_region` now would be a second mechanism to unpick.
R6 is a look-at-it-on-the-phone item and no build in this VM is evidence
about her device's font set.

Full checks on the tree as pulled: `cargo fmt --check` clean in `iris/`,
`server/`, `client-core/` and `event-model/`; `cargo clippy --workspace
--all-targets` exit 0 in `iris/` and `server/` (the only line is the
`future-incompatibilities` note about naga/wgpu/winit, which predates
this pass); `cargo test --workspace` 165 in `iris/`, 160 in `server/` and
157 in `client-core/`, no failures. The one thing not run is a real
device build -- `cargo ndk -t x86_64 -P 29 check -p iris` is clean, but
`-p iris-android-app` is the devlog agent's tree.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-07 21:09:02 -04:00
irisandClaude Fable 5.1 a6a100edc6 iris: the chain bound is named for the walk, and two nits from the review
docs/REVIEW-2026-09-07.md's rule finding on `MOVE_CHAIN_LIMIT` plus both
nits.

`MOVE_CHAIN_LIMIT` bounds two different parent walks -- move offsets in
the vertex stage and `Mask::parent` in the fragment stage -- under a name
that says one, and the shader's own comment beside it already called it
"the bound on the parent walk". Renamed to `PARENT_CHAIN_LIMIT` in both
files at once (the constant has no other users), with the doc saying
which two chains it governs.

`DragGesture`'s release computed `self.velocity.velocity()` twice, once
for the outcome and once for the `iris drag release:` line -- a full Lsq2
fit each. Once now, into a local both read.

`transcript-ui`'s `selection.rs` called `ui.ui_mut().animate(id)` even
when `List::fling` had bailed (Compose's `|v| <= 1.0`, or no anchor), so
a frame was asked to advance an animation known not to exist. It is
behind `is_scrolling()` now, which is the same answer `fling` itself
reached. `phone_screen.rs`'s recorded flick still flings, which is the
half that says the guard did not turn a working release off.

Verified: `cargo test --lib -p iris` (104) and `-p transcript-fixture`
(12), fmt and clippy clean, and layer 2 (`run-headless.sh phone --phone`)
still renders with the mask chain intact -- code fences clipped to their
rows, the list clipped at the composer -- which is what the wgsl rename
needed looking at rather than compiling.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-07 21:05:55 -04:00
irisandClaude Fable 5.1 ff1d6ea932 iris: a degenerate fit has no solution, and the desktop follows a display's density
Two of docs/REVIEW-2026-09-07.md's risks.

**R7.** `poly_fit_least_squares` clamped a near-zero basis-vector norm
(`1.0 / dot(..).sqrt().max(1e-6)`) where Compose's `polyFitLeastSquares`
bails: below `0.000001f` the vectors are linearly dependent and there is
no solution. Clamping reached the solve with a `q` row of zeros and a
zero on `r`'s diagonal, produced `[NaN, NaN, NaN]`, and was rescued only
by the caller's `is_finite` check -- working, but by accident, and not
what the source it is transcribed from does. It returns `Option` now and
`velocity()` answers 0 on `None`.
`a_fit_through_linearly_dependent_points_has_no_solution` reports
`Some([NaN, NaN, NaN])` with the clamp back in place. Three samples at
one instant is exactly what the input clock produced before 2ec0fee, so
this is the second half of the same fault.

**R5.** `WindowEvent::ScaleFactorChanged` was unhandled, so dragging the
window to a display with a different scale left every `Len::dp` and every
rasterised glyph at the density the window opened on. It now re-reads
`content_scale` -- through that function rather than off the event, so
`IRIS_SCALE` still pins `--phone`'s density instead of following the
monitor -- and sets both copies. `UiRenderState::set_density` marks the
tree for a full redraw when the value actually changes, because
`Text::shape` keys its cache on `(attrs, width, density)` and nothing
else would ask for those glyphs again. Invisible on this machine (every
display here is 1.0), which is why the review asked for it in writing.

Verified: `cargo test --lib -p iris` (104), `cargo test -p
transcript-fixture` (12), `cargo ndk check -p iris`, fmt and clippy
clean, and layer 2 (`run-headless.sh phone --phone --replay
flick-120hz.touch --shot`) still draws and still clips at the composer.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-07 21:03:43 -04:00
irisandClaude Fable 5.1 2b20bb2c91 docs/RUST.md: queue -- bench header type shrunk to fit, emulator crash loop after backgrounding
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-07 21:01:27 -04:00
irisandClaude Fable 5.1 3c80d9d696 iris bench: a Trace switch for the input/frame diagnostics, and the report says when it was on
`iris::diagnostics::set_trace` landed with nothing to press it. It is the
bench header's fourth control now, reading `Trace off` or `Trace on` --
a toggle whose own appearance never changes is a button that looks like
it did nothing. Its accessibility label stays the fixed "Trace input and
frames", because that is what `run-bench.sh` and `ui-trace --do "tap
'...'"` find it by and a control that renames itself when pressed is one
no script can find twice. Pressing it rebuilds the header and shows the
diagnostics pane, so the state is on screen at the moment of the press.

Both reports carry `trace_line`, from the flag read at the *start* of
what is being reported as well as at the end: the switch is on screen
while a benchmark runs, so "somebody moved it half way through" is a
state that happens, and reported as either "on" or "off" it would be a
confident sentence about a log covering half the run.

The row's type size is one constant for all four labels and drops from
18 to 13: with a fourth control the labels overlapped each other on a
1080px screen. Shrinking one label to fit is what the UI rules forbid;
resizing the row is a layout decision and all four still match.

Checked on the emulator: the switch flips its own text and colour, the
pane reads "input/frame trace: on", and `iris::frame`/`iris::input`
lines appear in the ring only after it is pressed.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-07 20:58:58 -04:00
irisandClaude Fable 5.1 06b8a1f4b0 The app hands its log to Dev Updater on the phone, not through ai-server
Iris's call once the upload route was working: put it in Dev Updater
properly. So the app now exposes its own ring through a ContentProvider
at `<applicationId>.devlog` -- Dev Updater's contract, written down in
that project's README, not something invented here -- and Dev Updater's
phone app reads it on the same device and forwards it to its own build
machine. No tunnel, no token, no second enrolment, and any app that
server delivers can implement the same and get the same Runtime tab.

`DevLogProvider.java` plus `devlog.rs` are the platform glue only: a flat
`String[]` across JNI, a `MatrixCursor` on the Java side, and
`nativeReady` telling Rust the authority the provider actually
registered, so the Diagnostics pane can name somewhere a reader can
query rather than composing a guess. `LogRing::newest_seq()` is the one
addition in `client-core`: an in-memory ring starts again at zero, so it
is what lets a reader notice the process restarted instead of silently
skipping everything since.

Deleted with it, so there is one mechanism: `client_core::log_upload`,
`POST /client-log` on ai-server, the `AI_APP_LOG_*` baking (which left
`build.rs` with nothing to do), and the uploader on both Android
clients. Kept: the ring, `RingLogger`, `install_process_logger`, and the
Diagnostics line -- whose second half is now `devlog provider:
content://<authority>`.

Verified end to end on this checkout's emulator: iris's own
`iris::android::view` startup lines read out of the provider by the
shell, forwarded by Dev Updater's Runtime tab, and served back from
`GET /apps/android-app/components/app/logs?kind=runtime`. A component
whose package has no provider says so in as many words.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-07 20:58:48 -04:00
irisandClaude Fable 5.1 e10582a2cd iris: three layer-1 tests that could not fail in the direction the bug goes
docs/REVIEW-2026-09-07.md's T1, T2 and T3. Each was confirmed by breaking
its subject on purpose and watching the new assertion fire, and each of
those breaks is recorded beside the assertion.

**T1** (`phone_screen.rs`) bounded the fling's duration with
`FlingCalculator::new(PHONE_SCALE).duration(velocity)` -- the calculator
under test -- and only from above, so it could fail when a fling ran too
long and never when one stopped dead, which is the symptom Iris actually
reported. The companion `assert_ne!(before, after)` passes on one pixel of
travel. It now takes both bounds from `fling_spline_reference.py`, which
gains this case's own line (`density=2.55 v=15250.0: distance=11057.424px
duration=2.0716s`), and measures travel in pixels from a row's own
on-screen extent -- 10527px against the reference's 11057, the 5%
shortfall being the frames a tracked row leaves the screen on. Scaling
`tick_fling`'s elapsed by 1000 reports "stopped after 8ms"; scaling its
delta by 0.01 reports "travelled 111px".

**T2** (`top_edge.rs`) asserted the per-row box only on the return leg,
so a regression that drew rows in the wrong place while travelling
*backwards* was checked by the row count alone. The first leg still
cannot assert it (an unmeasured row has to be drawn to be measured), so
there is now a third leg -- back again, every height known. Widening
`intersects_viewport` downwards passes all 40 forward steps and fails at
"back 6", which is the leg that did not exist.

**T3** (`top_edge.rs`) asserted a mask exists and sits inside the list's
box, never that any row primitive references it, so a broken
`Mask::parent` chain -- what d507ae4 introduced -- left it green while a
code fence drew unclipped. It now walks every row primitive's chain and
requires the list's own mask slot on it (and rejects a chain that loops).
Forcing `Painter::set_mask`'s `parent` to `NONE` fails it with "clips to
[Id(1)], a chain that never reaches the list's own mask Id(0)".

Verified: `cargo test -p transcript-fixture` (12) and `cargo test --lib -p
iris` (103) pass, fmt and clippy clean.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-07 20:58:35 -04:00
irisandClaude Fable 5.1 551c01398f iris: the guards against silently wrong output survive into release
docs/REVIEW-2026-09-07.md's R1. Every invariant guard added on 2026-09-07
was a `debug_assert!`, and every build anybody runs on this project is
release -- the bench APK must be (the debug `libmain.so` is 325 MB and
will not install) and Iris's phone gets release too. So a `List` drawn
without a mask painted over its surroundings again, in exactly the build
the fault was found in, with nothing saying so.

Promoted to `assert!`, each O(1) or a handful per *draw* and each
protecting against output that is wrong on screen with no other symptom:
`List::draw`'s `painter.is_masked()`, `List`'s `extents`-are-on-screen
check, `Painter::set_mask`'s doubled-call check (the second call replaces
rather than nests, i.e. an unclipped widget), `Painter::glyphs`'s atlas
generation (glyphs sampled from coordinates now holding other letters),
and `List::fling`'s finiteness (one comparison per gesture; NaN
propagates into `deceleration_for`'s `ln()` and the fling never settles).

Left as `debug_assert!` and now saying so in a comment: `List::place`'s
slot-exists precondition (once per row placed per frame, and its release
failure is the `.expect` below rather than something wrong on screen) and
`poly_fit_least_squares`'s two preconditions (run on every velocity query,
with `MIN_SAMPLE_SIZE` and the `is_finite` check giving release a defined
outcome either way). `PointerClock::sample`'s ordering assert was already
annotated in 2ec0fee for the same reason.

Verified: `cargo test --lib -p iris` (103) and `cargo test -p
transcript-fixture` (12) pass in both debug *and* `--release`, which is
what says the promoted asserts do not fire on a real replayed flick;
fmt, clippy and `cargo ndk check -p iris` clean.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-07 20:52:45 -04:00
irisandClaude Fable 5.1 7e79ec11e0 docs: the fling's "before" velocity is what velocity_reference.py prints, 12250 and 12500
docs/REVIEW-2026-09-07.md's D5. Four places quoted 11750 px/s as the old
average estimator's answer -- for `flick-120hz.touch` *and* for the
press-plus-one-move-frame set, which are different sample sets, and one
number in both rows is the tell. `iris/benches/velocity_reference.py`,
which the same section says every number below it comes from, prints
12250 for the recording and 12500 for the two-sample set, and
`sense.rs:1406` already had the 12250.

Half of where 11750 came from is recoverable and is written down beside
the table: it is the recording's 196 px over 16.68 ms, a 60 Hz frame
rather than the 16 ms span the file itself records. That explains the
flick row; the other row was copied from it. The 1.30x ratio derived from
it becomes 1.24x.

Also settles the second disagreement about the same experiment (the
review's rule finding on the negative control): `sense.rs`'s doc comment
claimed reverting `velocity` to total-over-span fails "exactly this one,
the flick recording, and phone_screen.rs" while RUST.md said seven. Run
again today with the revert in place: seven in `-p iris` (the flick
recording, the accelerating flick, the horizon, the stopped finger, the
minimum sample count, both `drag_gesture` flick tests) plus
`phone_screen.rs`'s flick, everything else green. RUST.md was right and
the comment now says the same thing.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-07 20:43:23 -04:00
irisandClaude Fable 5.1 2ec0fee84c iris: the input clock anchors on the first event's oldest sample, not its own time
docs/REVIEW-2026-09-07.md's D4. `on_touch_event` took its one anchor as
`(Instant::now(), event.event_time_nanos())` from the first MotionEvent the
view ever sees, and dated every later sample as `anchor_at + (sample -
anchor).max(0)`. An event's historical samples are by definition *older*
than its own event_time, so if that first event is a Move -- the Down went
to another view, or the view was attached mid-gesture -- its whole batch
clamps onto one instant: three samples at the same time make the Lsq2 fit
degenerate and the flick reads 0 px/s. In a debug build the ordering
debug_assert fired first, and it was comparing against `anchor_nanos`,
a value from a different event, so it was also the wrong comparison for
the first sample of every later event.

The arithmetic moves into `sense::PointerClock`, which anchors at
`now - (event_time - oldest_sample)` and carries the last sample seen
across events, so `sample()`'s ordering assert compares against the
previous event's last sample. It lives in `sense` rather than in the
android backend because `iris::android` is cfg'd out everywhere but the
device, and this is exactly the arithmetic that wanted a test off one:
`the_first_events_batched_samples_are_dated_apart` reports [0ns, 0ns, 0ns]
against the old anchoring.

The assert stays a `debug_assert!` and now says why in a comment: it runs
once per touch sample, hundreds a second on a batching 120Hz screen, and a
mis-ordered sample degrades a velocity rather than drawing something wrong.

Also drops the stale reference to `VelocityTracker::add_sample` in the
comment above it (the review's rule finding); the method is `add_position`.

Verified: `cargo test --lib -p iris` and `cargo ndk -t x86_64 -P 29 check
-p iris` clean, fmt and clippy clean.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-07 20:36:41 -04:00
irisandClaude Fable 5.1 992c472975 iris: iris::input/iris::frame diagnostics, and gating the four debug! lines that already drowned the ring
Iris asked for a button to copy raw input events and per-frame timings
through the same report Copy report already produces. sense::log_input_event
(one line per platform pointer sample, historical samples inline on
Android) and diagnostics::log_frame (one line per frame: frame number,
frame clock, time since last input, layout/draw durations, redraw kind,
primitives on screen, animating) both land under iris::diagnostics's
trace_enabled() gate, off by default since the ring is 2000 lines/256KiB
and either target at 120Hz fills it in seconds. report_to_touch.py turns
a report's iris::input lines back into a .touch file for harness/desktop
replay, round-tripped in transcript-fixture's input_log_roundtrip test.

Folds in docs/REVIEW-2026-09-07.md's D1: four older per-frame debug!
lines (android::view's two render() lines, list.rs's fling tick,
text/mod.rs's text render) were unconditional at Debug and, with the
ring's RingLogger recording everything the app's Debug install lets
through regardless of target, filled it before Copy report ever saw
anything else. All four (and sense.rs's drag-release-samples line) are
now behind the same gate. The same test proves both directions: tracing
off leaves zero Debug lines from a replayed flick, tracing on produces
the expected iris::input/iris::frame lines with real durations.

Not wired to a Diagnostics-pane button: bench_client.rs is open under
another agent. set_trace(bool) is the whole surface a control needs.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-07 16:48:48 -04:00
irisandClaude Fable 5.1 729098756d docs: the CA travels in the enrol link, and why not the two alternatives
DECISIONS.md gets the decision with both rejected options and what the
longer link measures (89 -> 652 bytes, a 45x23 QR -> 93x47), RUST.md ticks
the enrolment queue item and marks the log-upload route superseded rather
than editing it, and IRIS.md says what changed for anyone building the
Android app.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-07 16:47:48 -04:00
irisandClaude Fable 5.1 d8562d96a3 iris android app: told which server by an enrol link, not by its build
The APK is cross-compiled here and run against the server on the host, so
everything build.rs baked in (AI_APP_TRANSCRIPT_HOST/_PORT/_TOKEN and this
machine's CA) was good for exactly the pair that built it -- and a token in
a delivered artifact besides. MainActivity registers aiapp://enroll, hands
the URI and the app's private files directory to Rust, and
client_core::config stores it 0600; transcript_client reads it afresh per
transport, so opening a new link repoints a running app.

Diagnostics says which of three things is true, because they want different
actions: 'enrolled: host:port', 'not enrolled -- open the enrol link from
Dev Updater', and 'enrolment unreadable: ...' for the case nothing could be
found out. The last is why status() has an Unknown arm at all.

ui-sandbox.sh's printed enrol command now carries the CA, which is what
makes it work for an app with no baked copy.

Verified on this checkout's emulator: fresh install reads 'not enrolled',
the intent enrols (log: 'enrolled with 10.0.2.2:8519', enrollment.json
-rw-------), Diagnostics then reads 'enrolled: 10.0.2.2:8519', and the CA
reconstructed from that link is byte-identical to the machine's ca.pem and
validates the server over curl. Android offered the chooser between this
app and the Compose one, which is the intended behaviour.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-07 16:45:51 -04:00
irisandClaude Fable 5.1 22210a42f5 docs: review of 2026-09-07's work -- 5 defects, 7 risks, 3 tests that cannot fail
Read-only review of ba2afba..origin/rustify (the fling spline and Lsq2
velocity, list culling/clamp/anchor re-homing, nested masks, the headless
harness, insets/targetSdk, platform fonts, and the client-core log ring
with POST /client-log).

The three that matter most: the app's own log ring is installed at
LevelFilter::Debug while the same day added three ungated per-frame
`log::debug!` callsites, so the 2000-line ring wraps in under ten seconds
and the route built to get Iris's logs to her carries frame spam instead;
POST /client-log inherits the router's 32 MiB body limit with no
per-message or rate cap, so an authenticated client can fill the host's
disk through ai-server's runtime log; and every invariant added today is
a `debug_assert!` while the phone and the bench APK are both release
builds, so none of the new guards can fire where the defects were found.

Also: the input clock anchors on the first MotionEvent's own event_time,
so that event's historical samples date before the anchor and are
silently clamped onto one instant; the "before" fling velocity quoted in
four docs (11750 px/s) is not what velocity_reference.py prints (12250);
masks clip drawing but not hit-testing, so a straddling row is now
invisible above the list and still tappable through the header.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-07 16:38:43 -04:00
irisandClaude Fable 5.1 ade572973a enrolment carries the CA, and one store holds it on every platform
An APK built in this VM pins this VM's CA, so it can never reach the
host's ai-server -- which is exactly the iris Android client's situation
(cross-compiled here, run against the host). So ai-server now puts the CA
in every enrollment link it mints, base64url of its DER under the 'ca'
parameter wg-app-link just learned to add, and client_core parses it back
out as PEM. Nothing has to be built on the machine it talks to.

Refused rather than ignored where 'ca' does not decode: a link that named
a certificate and then pinned nothing is the one outcome nothing
downstream could notice.

EnrollmentStore moves out of desktop-app into client_core::config, since
the Android client needs the same file for the same reason and only the
directory differs by platform (AGENTS.md's sharing rule). desktop-app's
--ca becomes the override for a link that carried none.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-07 16:34:12 -04:00
irisandClaude Fable 5.1 9b27e858b5 docs/RUST.md: APK runtime logs in Dev Updater via an on-device ContentProvider (Iris, 2026-09-07); supersedes the ai-server client-log route
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-07 16:30:08 -04:00
irisandClaude Fable 5.1 7e4e26a335 iris: resolve fontique's Android monospace generic family ourselves
fontique 0.11.1's Android backend never resolves GenericFamily::Monospace
(mono=None in the startup diagnostic, RUST.md's 2026-09-07 "Platform
fonts" gap): DEFAULT_GENERIC_FAMILIES looks up "monospace" against
name_map before fonts.xml is parsed into it, and even after parsing,
AOSP's fonts.xml names it with a <family name="monospace"> element whose
<font> children the backend's own parser never reads (a TODO left in
place) -- so the name gets a FamilyId with no font data behind it, and
family_by_name("monospace") comes back empty too. Confirmed still present
on linebender/parley's main branch, so there is no newer release to bump
to.

TextData::patch_android_monospace (Android-only, called from
TextData::default) reads fonts.xml's own "monospace" declaration for the
font filename it names, then finds which of fontique's actually-scanned
families owns a font file with that name and registers it as the
Monospace generic directly -- the same authority Compose's
Typeface.MONOSPACE resolves through, without pinning an OEM-specific
family name. Verified on this checkout's emulator:
mono=Some("Droid Sans Mono") in the startup log, and a screenshot showing
the bench-fixture's code block and tool-card values in a visibly
monospaced face beside sans body/heading text. Desktop's fontconfig
backend is unaffected.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-07 16:25:27 -04:00
irisandClaude Fable 5.1 84a13e806b iris: a fling starts at Compose's velocity, which is a curve fit and not an average
Iris, from the phone on the 4274b8b build: "flinging now actually works
but is slower than Compose's immediately after releasing the flick (the
slow down seems correct)." The spline was already AOSP's; the initial
velocity was not.

`VelocityTracker` held per-frame pan deltas and answered their sum over
the sample span -- an average, which cannot tell an accelerating flick
from a steady drag. Ported from the `-sources.jar` of
androidx.compose.ui:ui-android:1.12.0 and
androidx.compose.foundation:foundation-android:1.12.0 (the versions the
Compose app builds against) rather than from memory, and the reading
corrected the plan twice:

  * The touch path is not `Strategy.Impulse`. `scrollable`/`draggable`
    release through the 2D `VelocityTracker`, which on Android is two
    `VelocityTracker1D(strategy = Lsq2)` over absolute positions -- a
    degree-2 least-squares fit differentiated at the newest sample.
    Impulse is reached only by `DifferentialVelocityTracker`, whose one
    caller is `NonTouchScrollingLogic`: wheel and trackpad.
  * There is no minimum fling velocity. `ViewConfiguration`'s 50dp/s is
    used only by `NestedScrollInteropConnection`; `DefaultFlingBehavior`
    skips `abs(v) <= 1f`, and says in its own comment that this is to
    dodge a NaN out of the spline. So `List::fling` caps at 8000dp/s
    against its own density and floors at 1px/s, and no threshold
    Compose does not have was added.

So the tracker holds positions rather than deltas (Lsq2 refuses
differential data in Compose too), 20 of them, with Compose's 100ms
horizon and 40ms stopped-gap; `DragGesture` feeds the raw window
coordinate along the drag axis at the press and every `Pan` frame.

`iris/benches/velocity_reference.py` is the independent transcription
the checked-in numbers come from, as `fling_spline_reference.py` is for
the curve. On `flick-120hz.touch`: 11750px/s before, 15250px/s after. On
an accelerating flick -- the shape a real finger makes, which that 16ms
recording is too short to show -- 1080 before, 2445 after. An average
also flings from a standstill (2533px/s where Compose says 0) and flings
from two points that describe no curve.

Negative control: reverting `velocity` to `total / span` fails exactly
seven tests, all of them about the estimator, and leaves the steady
drag, the tap, the selection release, the sixteen arbiter tests and the
rest of phone_screen.rs passing.

`iris drag release:` keeps its info line and gains a debug
`iris drag release samples:` with every held sample as `t_ms:position`,
so a flick that felt wrong on a phone with no logcat can be replayed at
layer 1 or pasted into the reference script.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-07 16:24:29 -04:00
irisandClaude Fable 5.1 452c44249f docs/RUST.md: queue -- logging landed; iris app enrolment replaces the build-time log destination; build-apk.sh traps
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-07 16:22:29 -04:00
irisandClaude Fable 5.1 238057ad5e docs: the phone-logging decision, how to use it, and two build-apk traps
DECISIONS.md gets the route and both rejected alternatives with what each
would have cost; RUST.md gets a "Phone logging" section with the build
command, where to read it on the phone, the end-to-end verification, and
the two rig traps that cost an hour -- Gradle's merged-native-libs cache
surviving build-apk.sh's `rm -rf jniLibs` (a --abi x86_64 APK packaged
arm64 and aborted with what reads exactly like a Vulkan fault), and the
648 MB debug bench APK that cannot be installed at all. IRIS.md gets the
client-core logging API with a before/after.

Queue item ticked.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-07 16:21:23 -04:00
iris 896c93a59a iris: drop bundled Noto Sans, match Compose's platform-font fonts
Iris's call: "remove the font for now; just match what compose does."
Removes the six embedded Noto Sans/Noto Sans Mono TTFs (3.6 MB) that
TextData::default used to register ahead of the platform's own fonts;
fontique's system font discovery was already on by default and now
runs unshadowed (Roboto/Roboto Flex on Android, fontconfig on the
desktop). .so -3,748,136 bytes (11,193,608 -> 7,445,472), matching the
estimate. Verified fallback still lands on visible tofu for CJK/emoji
rather than blank, and flagged (not fixed) a fontique Android backend
gap that leaves Monospace unresolved -- see RUST.md's "Platform fonts
(2026-09-07)" and DECISIONS.md/IRIS.md's dated entries.
2026-09-07 16:14:34 -04:00
iris 690161e5e9 docs: the transcript's edges were three faults, and what the rig found
IRIS_TODO's 2026-09-07 top-edge entry closed with the root cause of
each, the six layer-1 test names, and what was suspected and turned out
not to be it -- no culling test compared a row's top against the
viewport's, and 03c6be8's header duplicate is untouched and still open.
The later report's "you shouldn't be able to scroll below the bottom (or
above top)" is ticked with why the clamp is a correction measured from
the layout walk rather than a clamp inside the scroll setter: nothing at
the moment of a scroll knows where the content ends.

RUST.md gains the same account in "Where things stand", plus the three
things this said about the new test rig -- layer 1 found all of it in
seconds and the emulator was not used; layer 2 is where the missing clip
is visible, with the command; and an assertion that reads the wrong
thing hides the bug it is for, which is how a list resting 1398px past
its own first row passed a test about stopping at that row.

Also the last of the six tests, the bottom end of the clamp
(`scrolling_past_the_last_row_settles_on_it`) -- the same rule at the
edge the top-edge work had no reason to touch.
2026-09-07 16:07:37 -04:00
iris e922b73d7a iris: a transcript row is drawn if it overlaps the viewport, and clipped to it
Iris's phone, 2026-09-07, two screenshots of the transcript at its top
edge wrong in opposite directions: rows already scrolled past still
drawn, over the header bar (`version = "0.1.0"` behind "Run benchmark"),
and a blank band where the row straddling the edge should be. Three
faults, one rule -- `List::intersects_viewport`: a row is drawn if any
part of it is inside the list's own box, and nothing outside that box
reaches the screen.

1. **The walk drew everything between the anchor and the viewport.**
   `scroll` moves the anchor's offset and nothing else, so panning leaves
   the anchor's own row further and further outside the viewport, and
   every row in between was placed *and drawn* on every frame. Measured
   on the bench fixture: 8 scrolls of 3000px left 64 rows drawn for a
   2012px viewport, ~59 of them off screen. `place` now skips a row whose
   height is already known and whose box does not overlap; `rehome_anchor`
   moves the anchor onto a visible row each frame, without moving
   anything drawn, so the walk is O(visible) again whatever distance was
   travelled. `extents` holds only what is on screen, which is what
   `key_at` already claimed of it, asserted at the end of every draw.

2. **Nothing clipped the list.** A straddling row is drawn in full --
   that is the rule -- so the part above the list was on screen. The
   transcript's list is `.masked()` now (the mechanism `examples/
   message_list.rs` and the composer already use, and one that nests as
   of the previous commit), and `List::draw` asserts it has a mask rather
   than leaving that to each caller to remember.

3. **A fling past the first row stayed past it.** `tick_fling` stops a
   fling that has reached an end, wherever the spline's last step had put
   it: `fling_toward_the_start_stops_at_the_first_row` was leaving the
   first row 1398px below a 600px viewport -- a blank screen -- and its
   assertion could not see it, since `extents` then held off-screen rows
   too and `top >= -0.5` is satisfied by +1398. `clamp_to_content` gives
   the gap back from the ends the walk already placed. Only when the
   opposite end is not also in the viewport, so a list shorter than its
   viewport stays bottom-anchored as before.

Layer 1 of the test rig throughout (`transcript-fixture/tests/
top_edge.rs`, the real screen under a bench-app-shaped header): each of
the five fails on its own subject and no other -- culling on the row's
top instead of its bottom fails only `the_row_across_the_top_edge_is_
drawn`, the pre-fix walk fails only the two about what is placed,
dropping `.masked()` fails only `the_list_is_clipped_to_its_own_box`,
dropping the clamp fails only `scrolling_past_the_first_row_settles_on_
it`. The bottom edge and a list shorter than the viewport are the ends
none of this had a reason to touch and are covered too.
2026-09-07 16:05:31 -04:00
iris d507ae4c96 iris-core: masks nest instead of aborting, and a widget can ask to be drawn again
`Painter::set_mask` refused a widget any mask of its own once an
ancestor had set one -- `assertion failed: self.mask == MaskIdx::NONE`
-- so clipping was one level deep wherever it was used at all. That is
what stopped the transcript's `List` from being clipped to its own box:
its rows already use `.masked()` themselves (a code fence, a tool card's
one-line title), and giving the list one aborted on the first fence
drawn.

A mask now carries the mask it was set inside (`Mask::parent`) and the
fragment stage walks that chain, so a pixel has to be inside every mask
on it. Chained rather than intersected on the CPU because each mask
moves with its own widget: a fence inside a transcript row carries the
row's scroll and the list's box does not, and one region resolved when
the fence was last drawn gets the second of those wrong as soon as the
row is moved rather than redrawn -- which is every scroll frame. The
child holds one ref on its parent's slot, released where the child's own
slot is, so a chain cannot outlive what it points at. The old assert
survives as the case that is still wrong: the same widget setting two
masks, which since a mask now chains would be a clip loop.

Also `Painter::draw_again`, for a layout that can only discover a
correction to itself by laying out once -- `List::clamp_to_content`, in
the commit after this -- and `Painter::is_masked`, which is how a widget
that draws outside its own box can require something to be clipping it.
2026-09-07 16:05:13 -04:00
irisandClaude Fable 5.1 9ed01e2812 docs: phone report 2026-09-07 later -- overscroll, low initial fling velocity, input/timing report; queued
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-07 16:04:34 -04:00
irisandClaude Fable 5.1 5be9f1baac iris-android-app: keep the app's own log, put it in Copy report, upload it
`app_log` is the platform half: `android_logger` as the logger the ring
forwards to, and an optional destination baked in by `build.rs` from
`AI_APP_LOG_HOST`/`_PORT`/`_TOKEN` plus the pinned CA -- the same
build-time trust boundary the transcript config and the Compose APK's CA
already use, so no token is committed and an APK is good for the server
that built it. All three or none: two of the three would be a build with
nowhere to send its log and no way to say so.

`Copy report` now appends the ring to what goes on the clipboard (not to
the pane, which is on screen and would be buried) and flushes the
uploader first, so the lines are on the server by the time the message
describing them arrives. The Diagnostics pane gains two lines: how many
lines are held and when the last arrived, and what the uploader last did
-- "not tried yet", "failing -- <why>", and "no server configured" are
each their own wording, because "nothing is arriving" has three causes
that look identical otherwise.

Also: the re-emitted lines carry the target `ai_server::client_log`, not
a bare `client_log`. `RUST_LOG=ai_server=debug` -- the filter AGENTS.md
tells people to run with -- drops a bare target, so every line a phone
sent vanished with nothing saying so. Found by running it.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-07 16:01:30 -04:00
irisandClaude Fable 5.1 977bdb9ee0 client-core: the app's own log ring, and POST /client-log to get it off a phone
Iris tests iris builds on a phone with no adb, and Android forbids one
app reading another's logcat, so a `log::info!` in the app can only reach
her if the app carries its own copy and sends it somewhere.

`client_core::log_ring` is that copy: a bounded ring (2000 lines / 256
KiB, whichever bites first) behind a `log::Log` backend that forwards to
whichever real logger the platform installed, so `logcat` and the desktop
terminal see exactly what they saw before. Reading does not consume --
the report and the uploader are two readers of one ring.

`client_core::log_upload` drains it into ai-server's new `POST
/client-log`, which re-emits each line into the server's own tracing
output. Dev Updater already shows that as ai-server's runtime log, so
nothing new is built there. A failed batch is retried from the same
cursor, and nothing in the upload path calls `log!` -- it would land in
the ring it is draining.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-07 15:56:20 -04:00
irisandClaude Fable 5.1 9cd1263080 docs/RUST.md: queue -- APK size done, the embedded-fonts question left for Iris
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-07 15:47:53 -04:00
iris 42af780639 iris android-app: strip+LTO+cgu1+opt-level=s halve libmain.so, no feature trim needed
Baseline had panic=abort only. Measured each setting in order (docs/RUST.md's
new "APK size (2026-09-07)" subsection has the full table and crate
breakdown): strip=true, lto="fat", codegen-units=1, opt-level="s" take
libmain.so from 18,546,488 to 11,193,608 bytes (-39.7%) and the release APK
from 20,678,956 to 13,326,076 bytes (-35.5%), arm64-v8a. opt-level="z" was
measured (another ~800KB) but not adopted without a frame-time check.

Investigated naga/wgpu backend features and tabs-ui/tabs-screen as trim
candidates; both are already fully eliminated by the linker on Android
(0 symbols in `llvm-nm` on the baseline .so), so no Cargo feature change
would shrink the binary -- left as documented findings rather than a diff.

Embedded Noto Sans fonts (3.6 MB) and the wgpu/naga/font-shaping stack
account for most of what remains vs. Compose, which borrows the platform's
own renderer and fonts for free; recorded honestly in the doc rather than
trimmed, since subsetting fonts or dropping a backend would change what
iris can render.
2026-09-07 15:46:40 -04:00
iris 4274b8b8d0 Merge remote-tracking branch 'origin/rustify' into worktree-agent-ace98b0bdaf33ffff
# Conflicts:
#	docs/IRIS.md
#	docs/RUST.md
2026-09-07 15:33:25 -04:00
irisandClaude Fable 5.1 73f956f8e0 iris: the fling curve was the identity function, and the keyboard was a targetSdk
Iris's 2026-09-07 phone report on ed04d4c: the resume glyph corruption is
fixed (item 4 closed with her evidence), flinging "seems to just be linear
velocity with an abrupt stop", and the keyboard still does not push
anything up. docs/RUST.md's new "The 2026-09-07 phone report" section has
the derivation and every number.

**The fling was arithmetically linear.** `android_fling_spline::
distance_fraction(t)` returned `t` for every `t`. Two halves of AOSP's
`SplineOverScroller` static initialiser had been transposed -- the
bisection solved the tension curve and the sample evaluated the P1/P2 one,
where AOSP does the opposite -- which made SPLINE_POSITION and SPLINE_TIME
identical; the lookup then bracketed `t` between SPLINE_TIME entries
instead of between even time steps, and the two cancelled to the identity.
Ported exactly now from OverScroller.java and androidx.compose.animation
1.12.0's SplineBasedDecay.kt, which agree line for line, as one table
indexed by even steps of time (AOSP's second table serves only
`adjustDuration`, which nothing here has, so it is deliberately not built
-- one array, one indexing rule). `FlingCalculator::velocity_at` is new
beside `position_at`, and `List::tick_fling` logs `iris fling tick:` with
the per-frame delta and speed.

Every existing test compared the calculator with itself -- monotonic,
signed, integrates to the closed form, deltas non-increasing -- and all of
them pass on a straight line. iris/benches/fling_spline_reference.py is an
independent hand transcription of both sources and supplies the numbers
now checked into `the_spline_matches_aosps_own_table` and
`a_flick_decelerates_the_way_aosp_says_it_does`;
`tick_fling_applies_shrinking_incremental_deltas` went from
"non-increasing" to "the last delta is under 80% of the first". Negative
control: with `sample` forced back to `t`, exactly those three fail.

Emulator (API 36, debug, force-gles): a released v=3750 decelerates
3746 -> 2624 -> 1834 -> 1144 -> 752 -> 449 -> 243 -> 83px/s over 32 frames
to t=0.664s; a flick into the end of the list stops there in one tick with
no overshoot; a tap 200ms into a fling ends it at 11 ticks.

**The keyboard: `targetSdk = 34`** in iris/android-app/app/build.gradle,
against compileSdk 37 and the Compose app's 37 -- and that app's keyboard
does push up on her phone. Below target 35 a window keeps the legacy
behaviour where adjustResize shrinks it for the IME, so
getInsets(ime()).bottom measures an already-shrunk window and is zero;
setDecorFitsSystemWindows(false) opts out of that and still takes on the
API 36 emulator here, which is why every test run passed. Now targetSdk 37.

That is a reading and not a measurement, so the other half is making the
phone able to answer it. MainActivity also registers a
WindowInsetsAnimation.Callback (onEnd re-reads getRootWindowInsets, so an
interrupted animation cannot freeze a value), which delivers the height
where only the animation path carries it and makes the push-up animate:
ime_bottom now arrives 509, 663, 833, 881, 883 instead of one jump.
`insets::Shared::updates` counts every dispatch and
`AndroidUiState::insets_report()` puts it in the Diagnostics pane --
screenshot-verified, `insets: dispatches=27 left=0 top=142 right=0
bottom=63 ime_bottom=0 ime_visible=false`. Iris has no logcat, and "the
listener never fired" and "it fired with a zero height" are otherwise the
same picture; dispatches=0 says so in words rather than showing defaults.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-07 12:44:01 -04:00
irisandClaude Fable 5.1 038f6a3832 docs: the test rig's layers 1 and 2, with their commands and their limits
RUST.md's "Three test layers" section rewritten in place with what was
built: the `cargo test -p transcript-fixture` command and the five
assertions with the mutation that fails each, the `run-headless.sh
--phone [--replay …]` commands and the 15s/18s they take, and a
paragraph on what still cannot be answered below layer 3 (anything about
pixels, any frame time, anything JNI). Also the two traps that cost time
-- `swaymsg seat - cursor` reaching nothing on a compositor with no
input devices, and a leftover window tiling beside the new one so a
screenshot looks like a duplicated-primitive bug.

IRIS.md gains the public surface: `iris::harness`, `TouchScript`,
`List::fling_velocity`, the fling's clock, and the desktop backend's
move to physical-pixel layout with `content_scale`/`IRIS_SCALE`.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-07 12:39:53 -04:00
irisandClaude Fable 5.1 1121d7cc83 docs/LAYOUT.md: masks reference a drawn primitive instead of copying a shape, and hit-testing applies the shape (Iris, 2026-09-07)
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-07 12:38:55 -04:00
irisandClaude Fable 5.1 232de0ec53 iris: a phone-shaped desktop window, driven by the same touch recordings
Layer 2 of docs/RUST.md's "Three test layers":

    ./run-headless.sh phone --phone --shot /tmp/p.png -- -p transcript-fixture

opens `transcript-fixture`'s screen -- the same fixture and the same
fold the headless tests and the Android bench use -- in a window at the
phone's own 1080x2424 and `content_scale` 2.55, and screenshots it. 15
seconds, warm. `--replay FILE` drives one of the `.touch` recordings
into it and writes `<shot>-before.png` too, so "the list moved" is two
pictures: the flick carries it back about seven turns of the fixture.

Two things this needed.

**The desktop backend now lays out in physical pixels with a density,
exactly as Android does** (`default::content_scale`, overridable with
`IRIS_SCALE`, which is how `--phone` hands it the phone's). It used to
divide winit's coordinates into a separate "logical" space, which left
`UiRenderState::resize` (physical, from `WindowEvent::Resized`) and the
window uniform (logical) disagreeing on any display whose scale factor
is not 1.0, and rasterised glyphs at one resolution to show them at
another. At 1.0 -- every display here -- the numbers are unchanged, and
the `tabs` screenshot is identical.

**`rig-input`'s `replay-touch`** puts a gesture on screen. This
machine's compositor has no pointer to move: sway runs on the headless
backend with no input devices, so `swaymsg seat - cursor press` reports
success and `swaymsg -t get_seats` shows `capabilities: 0`. wlroots 0.19
dropped `WLR_HEADLESS_INPUTS` and ydotool's uinput device would be
ignored by a compositor not reading libinput, so the virtual-pointer
protocol is what is left. It parses the *same* `TouchScript` the
harness does, so one recording drives both layers.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-07 12:38:19 -04:00
irisandClaude Fable 5.1 e430880cde docs: phone report 2026-09-07, rows at the transcript's top edge culled early or drawn through the header
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-07 12:35:20 -04:00
irisandClaude Fable 5.1 a999bd106a docs: masks with a shape (LAYOUT.md, decided 2026-09-07) and the orchestrator queue in RUST.md
Iris: masks should carry a shape, rounded rectangle first, or take a
container widget as the mask, with corner alpha multiplied rather than
cut. Design: the mask evaluates the same SDF draw_rounded_rect uses,
nested masks chain and multiply like moves, and a rounded Rect's
.masked() makes the container the mask with one radius by construction.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-07 12:34:19 -04:00
irisandClaude Fable 5.1 6840edf61e iris-android-app: the bench's fixture half comes from transcript-fixture
The fixture bytes, the backlog/tail split and the fold into a screen
were `bench_client.rs`'s alone; they are `transcript-fixture`'s now, so
the Android bench, the headless harness and the phone-shaped desktop
window open one screen from one copy (AGENTS.md: nothing UI-shaped in a
platform crate). What stays here is the JNI half -- clipboard, battery,
IME, the report and the four phases.

Built with `cargo ndk -t arm64-v8a -P 29 build --features
"transcript-screen bench"`; the two warnings it prints (bench_jni's
unused overlay methods, the unused `tabs-ui` dependency under this
feature set) predate this change.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-07 12:27:04 -04:00
irisandClaude Fable 5.1 333220196e iris: a headless in-process harness, and the bench fixture as a shared crate
Layer 1 of docs/RUST.md's "Three test layers": `iris::harness` opens a
real screen with no window, no compositor and no GPU, on an explicit
clock and a replayed touch stream -- a trivial `t_ms action x y` file,
so the batched 120Hz flick shape from Iris's phone report is
reproducible as a test. The emulator cannot produce that shape at all:
a `ui-trace` swipe is many evenly-spaced events, a finger is five
samples in 20ms.

`transcript-fixture` is the fixture-loading and fold-driving half of
`iris-android-app`'s `bench_client.rs`, moved out of the platform crate
so the harness, a desktop window and the Android bench open the same
screen from the same bytes (AGENTS.md's sharing rule).

Two supporting changes in iris itself, both about reading a clock that
was not handed in: `Fling::started_at` is now set on the first
`tick_fling` rather than at the release, so a driver running frames on
its own clock does not start every fling at the wall clock and advance
it on a different one; and `List::fling_velocity` exposes what the
release measured, which is where `Released(Some(v))` lands.

Four tests, each confirmed to fail without its subject: dropping
`animate(id)` from `Selection::drag` (the phone's own "fling does
nothing" defect) and reverting `started_at` each fail the flick test
alone; flinging on `Tapped` fails only the tap test; a 5s `LONG_PRESS`
fails only the selection test; a `set_bottom_inset` that ignores its
argument fails only the composer/IME test.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-07 12:24:54 -04:00
irisandClaude Fable 5.1 7f4ea7e8fd docs/TODO.md: Compose app crash from Iris's phone log export, reversed AnnotatedString range in ToolInput.highlighted
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-07 12:22:47 -04:00
irisandClaude Fable 5.1 591128eef1 AGENTS.md: the phone app and the planned desktop app share widgets and styling; only screen layout differs
Iris, 2026-09-07. The second central design point beside the driver
rule, so a platform crate growing a widget or a colour reads as a
defect to move. docs/RUST.md carries the detail.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-07 12:13:54 -04:00
irisandClaude Fable 5.1 ba0f2ea93f docs: the 22:16 report reconciled with what was actually run
RUST.md's "Shell lost" section and IRIS_TODO.md's matching paragraph both
said item 4's fix was written but never built or tested. It was committed
in ba2afba with its test passing, so both were stale the moment that
landed and read as if nothing had been run at all.

Replaced with one section per item, saying what was fixed, what was
measured on this checkout's emulator and what the phone still has to
settle: items 2 and 3 ticked with their numbers, item 4 ticked on the code
with phone confirmation still owed (no Vulkan adapter here), item 1 left
open with the exact logcat line for Iris to look at. The two pre-existing
faults found on the way -- the 16-deep move chain and the API-29 JNI calls
-- are recorded where the next reader will hit them.

IRIS.md gains the public-surface entry: `Widget::tick`,
`UiData::animate`/`tick_animations`, `FlingCalculator`'s density and
coefficient, and `MOVE_CHAIN_LIMIT`.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-07 12:12:03 -04:00
irisandClaude Fable 5.1 ed04d4c735 iris: the keyboard reopens, the IME's height reaches the layout, and a fling actually moves
Items 1-3 of Iris's 22:16 phone report, plus the two defects that were
hiding behind item 1 and only became visible once the first one was
fixed. Emulator evidence and the numbers are in docs/RUST.md.

**Keyboard reopen.** `attr.rs`'s already-focused branch calls
`focus_gained` on a tap that stays inside `DRAG_SLOP` -- what Android's
own `EditText` does, `showSoftInput` being idempotent. Dismissing the IME
leaves the field focused, so the only branch that requested it never ran
again. Negative control run: without this one call the second tap leaves
`mInputShown=false`. Swipes across and out of the focused field still
summon nothing.

**IME height.** `MainActivity` sends `getInsets(ime()).bottom` and
`isVisible(ime())` as two values; the height used to be sent *as* the
boolean, so nothing had a number to pad by. `Insets`/`WindowInsets` carry
both, `bench_client` reads the boolean for its state machine and the
height for `Composer::set_bottom_inset`, and the list follows because it
is `rest(1)` in the same `Span`.

**Fling.** Three defects, in the order they were found:

1. `on_touch_event` read only each `MotionEvent`'s final position, so a
   batched 120Hz flick fed the tracker one sample and `velocity()`
   answered 0.0. Historical samples are replayed through the sensor pass
   now, `CursorState::time` carries each sample's own time (so a replay
   loop's speed cannot become the measured velocity -- the winit backend
   sets it too), the press is a sample as AOSP's own tracker does, and
   `iris drag release:` logs the decision for the phone's logcat.
2. Nothing advanced a fling between input events: `tick_fling`'s only
   caller was the benchmark's own loop, so the bench flung and a finger
   never did. iris has one animation mechanism now -- `Widget::tick`,
   `UiData::animate`/`tick_animations`, called by both backends before
   the draw and re-requesting a frame while it answers true.
3. With flings finally animating, one lasted 45 seconds: `List::fling`
   hardcoded density 1.0 against physical-pixel velocities, and
   `FlingCalculator`'s coefficient used the scroll friction where AOSP
   uses its 0.84 tuning constant -- 56x, inside an exponential. Emulator:
   1.62s for v=11064, against AOSP's own 1.586s.

**Two pre-existing faults found on the way.** `MOVE_CHAIN_LIMIT` was 16
and the composer's chain is 17, so every debug build aborted on a tap of
the composer and every release build silently drew and hit-tested that
subtree short; it is 64 in both the CPU walk and shader.wgsl, and the
assert prints the chain so a cycle and a deep tree can be told apart. And
`minSdk` is 29, since `getEventTimeNanos` is API 29 and a missing JNI
method is a crash rather than a degraded fling.

Every new invariant carries its guard: sample times non-decreasing in
`on_touch_event`, and tests confirmed to fail without their fix for the
press-seeded velocity, the animation registration and the AOSP
magnitudes.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-07 12:11:55 -04:00
irisandClaude Fable 5.1 ba2afbaedb iris: a cleared glyph atlas must un-cache every RenderedText, not just empty itself
Iris's phone, 2026-09-06 22:16: after leaving the app and returning,
every glyph drawn *before* the resume came back as fragments of other
letters, while the diagnostics text drawn after it was perfect.

The renderer rebuild does force a full redraw -- `surface_changed` calls
`render.resize(...)`, which sets `UiRenderState::resized`, which makes
the next `update` take `redraw_all`. What survives that is one cache
further in: `TextView::render` returns its cached `RenderedText`
whenever the wrap width, buffer and attrs are unchanged, so
`TextData::place` is never reached, nothing is re-rasterised into the
fresh atlas, and the *previous* atlas's uv_min/uv_max/layer go straight
back to the GPU. Only text whose content changed after the resume
re-shapes -- exactly the split in the screenshot.

One mechanism rather than a per-holder invalidation path: `GlyphAtlas`
carries a `generation`, bumped by `clear`; a `RenderedText` records the
one it was placed against; and `TextView::render`'s cache key includes
it, so clearing the atlas makes every cached render un-reusable at once.
`Painter::glyphs` debug-asserts that a submitted quad's generation is
the live one, catching the fault at the submission instead of on screen.

Test `clearing_the_atlas_re_renders_cached_text_instead_of_reusing_it`
(iris/src/widget/text/mod.rs): draw, clear the atlas, resize, draw
again, and assert the atlas holds the same glyph count. Confirmed to
fail without the cache-key line -- it trips the new debug_assert with
"glyphs placed against atlas generation 0 submitted against 1".

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-06 23:22:40 -04:00
iris 10267dec27 Merge branch 'worktree-agent-a673ba12761c025d9' into rustify 2026-09-06 23:20:30 -04:00
irisandClaude Fable 5.1 7e7cbb5402 Tool-call cards and grouping, with the state a result never arrived in
P1b (docs/RUST.md). `transcript-ui/src/tool.rs` draws a card per tool
call and a group per run: collapsed, a card is its name and the one-line
summary `parse_tool_input` derives; open, it is the description, the
input (highlighted, on the verbatim surface) and the output, capped with
a "Show all N lines". A run is one surface with a heading and a chevron
bar at its foot, so it closes from either end.

Three things worth knowing.

**A collapsed card lays out its summary line and nothing else.** The
fixture's tool outputs are tens of kilobytes and a collapsed card never
builds a widget for one -- `collapsed_cards_shape_only_their_summary_
lines` opens a three-card group over 88 kB of output each and asserts the
text-shape count equals the same group's over three bytes (17 either
way; 17 against 20 when the discipline is deliberately broken, so the
test is real).

**A result arriving replaces one card.** `ToolRow::apply_calls` is the
group's half of `RowBlocks::apply_delta`'s rule, and `build_row` now
hands back one `TailRow` -- blocks for a message, cards for a run --
rather than two mechanisms chosen at each call site.

**Every tap is a tap**: `GestureOutcome::Tapped` out of the `DragArbiter`
`Selection` already owns, so a drag that started on a card scrolls the
transcript instead of opening it.

Three defects found by looking at the render, all recorded with their
repro in docs/IRIS_TODO.md: a `Span` of padded children inside another
`Span` places them a slot out of step (worked around by building the
group as one span, which costs the 4dp inset); `scrollable_on(Axis::X)`
on a non-editable text draws nothing, so a card's command is clipped
rather than pannable; and `NotoSans-Regular` has no U+25B8/25BE/25B4 at
all, so the expander mark is set in the monospace face.

Screenshots: docs/bench/p1b-2026-09-06/.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-06 22:49:51 -04:00
iris a200ddbddd docs/IRIS_TODO.md: Iris's 22:16 phone report on the 20303e0 build, four open items with the reading of each 2026-09-06 22:31:07 -04:00
iris b332873894 Merge remote-tracking branch 'origin/rustify' into worktree-agent-a673ba12761c025d9 2026-09-06 21:33:28 -04:00
irisandClaude Fable 5.1 a4809b3026 WIP: tool-call cards and grouping (P1b)
`transcript-ui::tool` draws a card per call and a group per run, with
the states, the collapsed-lays-out-nothing discipline and the
one-card-per-result update. Screenshots in docs/bench/p1b-2026-09-06/.

Includes a local fix to `List::place`'s reposition-vs-mov clash, which
is about to be dropped for rustify's own.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-06 21:33:24 -04:00
iris 1ad2f9ec6e docs/RUST.md: phone delivery is a push to ai-app-bench, not ~/host/bench 2026-09-06 20:04:26 -04:00
iris 33e8ab83a2 docs/RUST.md: the two 2026-09-06 fixes under P1a, with the emulator's first legible screenshot 2026-09-06 19:59:57 -04:00
iris f5b88932b4 iris: a widget's move slot has one owner -- move_applied + repositioned
`mov` accumulates a delta onto the slot and `reposition` overwrote it, and
both legitimately land on one widget in one frame: `List::place`'s
Bottom-known branch offers a row a same-size box that has moved (`mov`),
then corrects the placement inside it when the row's cached height no
longer matches what the row reports (`reposition`). That is what a wrapped
transcript row hit, and what the `move_applied == ZERO` debug assert was
standing in for -- an assert against a case that happens is not a
guarantee, it is a crash.

The slot means `move_applied + repositioned` now, both halves recorded on
`ActiveData`, so `reposition` adds the move rather than dropping it and
stays idempotent. The assert it replaces is a `debug_assert_eq!` that the
slot still holds that sum on entry -- i.e. that nothing but those two ever
wrote it.

Test: `a_widget_moved_by_its_parent_and_then_placed_inside_it_lands_at_the_placement`,
which draws the child at the offered position (-100px) rather than the
placement (100px) without the fix. Verified against the `.wrap(true)`
repro from docs/IRIS_TODO.md (draws correctly, no panic) and an emulator
bench run with assertions live.
2026-09-06 19:59:39 -04:00
irisandClaude Fable 5.1 9079276ec8 A tool call can say it failed, and what it is for, without a renderer
P1b's pure half (docs/RUST.md). Three pieces, all testable with no
widget in sight:

- `event_model::Event::ToolEnd` gains `is_error`, read from the CLI's own
  `tool_result` field by both the live translator and the import replay
  (`import::tool_result_is_error`, one reader so the two cannot disagree
  about the same conversation). Without it a result is all a card has,
  and a broken call draws exactly as confidently as one that worked --
  the missing state, not a wrong one. `#[serde(default)]`, so an older
  transcript reads back as "not reported to have failed".
- `client_core::transcript_fold::ToolState`: Running, Deciding,
  Succeeded, Failed, NoResult. The pair it exists for is the last two
  against Succeeded-with-empty-output -- a call that printed nothing and
  a call whose result never arrived leave the same empty string, and only
  the session's status separates "still going" from "nobody found out".
- `client_core::tool_summary::parse_tool_input` and
  `client_core::durations`: `ToolInput.kt`'s subject/description/timeout
  split and `Durations.kt`'s span formatting, ported with their tests.

The echo driver's three-call run now has a failing middle call, so the
failed appearance is reachable from `ui-sandbox.sh` at all.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-06 19:46:21 -04:00
iris 3cb18ac5c2 iris: a one-layer glyph atlas is a GL_TEXTURE_2D, so every glyph drew as a box
The emulator was blamed for two days for what is iris's own defect on any
GL adapter. `GpuTextures::new` created the atlas `texture_2d_array` with
one layer; wgpu-hal picks the GL target from the descriptor alone
(`gles::Texture::get_info_from_desc`, `(false, 1) => TEXTURE_2D`), so the
shader's `sampler2DArray` was handed a `GL_TEXTURE_2D`, the unit was
incomplete, every `textureSample` returned (0,0,0,1), and `draw_glyph`'s
`color.a *= texel.a` painted the whole glyph quad.

`MIN_ARRAY_LAYERS = 2`, with the account at `create_array_texture` and a
`debug_assert!` there. Vulkan -- the phone's build and the desktop's
default backend -- was never affected.

`force-gles` now switches the desktop backend too, so the GLES path is
reproducible on a machine with a real GPU in seconds rather than only
through an APK: that is how this was found, with two shader probes
showing the sample was exactly (0,0,0,1).
2026-09-06 19:41:40 -04:00
irisandClaude Fable 5.1 69525bd131 iris: a Rect is not size-independent, and P1a's block appearance verified
The defect P1a's screenshots found, and the one that mattered:
`Rect::is_size_independent()` answered `true`. A `Rect` fills whatever
region it is handed, so its content *is* the region -- and
`draw_inner`'s fast path, which rewrites a widget's primitives with
`r.outside(&from).within(&region)` instead of redrawing it, cannot
reproduce that once a region carries both `rel` and `abs`. What it
looked like: a fenced code block's background kept the height of the
provisional full-region draw `Span` does in its first phase, so one
fence's panel covered every block below it and every row below that,
with the text underneath laid out correctly. Likely the same cause as
RUST.md's older "the composer bar's grey background is not drawn".

Also here: a quote's bar is a `Stack` background behind padded text
rather than a two-child `Span(Dir::RIGHT)` (one widget fewer and no
provisional pass), and `transcript-ui`'s `transcript` example gains a
row holding one of every block kind -- the fixture's own heading,
paragraph, fence and table source, plus a list and a quote, which the
fixture has neither of.

docs/bench/p1a-2026-09-06/ has the pairs and docs/RUST.md's P1a box
names what still differs. The iris half is from the desktop backend
because this emulator cannot draw iris's glyphs at all (solid boxes,
reproduced on the previous commit, with Compose drawing text correctly
on the same AVD); both routes to Vulkan on this AVD were tried and both
fail. Bench stream phase, assertions live, no abort: p50 53.0ms p90
108.6ms p99 132.0ms against 52.8/108.1/137.3 before -- unchanged.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-06 19:30:39 -04:00
irisandClaude Fable 5.1 64f64b54e5 iris: per-block markdown appearance, syntax-highlighted fences, tappable links
P1a (docs/RUST.md). A transcript row's blocks are drawn the way
Markdown.kt draws them rather than as one flat span list:

- transcript-ui/src/markdown.rs is a *block* renderer now.
  `BlockFrame` is the whole widget vocabulary -- Plain, Verbatim (a
  dark rounded panel that pans sideways) and Quote (a bar and an
  indent) -- so a new markdown feature costs spans, not widgets.
  `frame_of` is the one place the BlockKind -> appearance mapping is
  written.
- Fences take `client_core::highlight`'s spans by language, in the
  same Catppuccin palette Theme.kt's `catppuccinSyntax()` uses, with
  the char->byte offset conversion the two index spaces need.
- Lists get the bullet ladder and coloured markers MarkdownPieces.kt
  draws, ordered lists count from the number they were written with,
  headings take Material's own ladder (24/22/16/14/12/11).
- Tables are padded monospace columns measured from the cells, with
  the header bold and a rule under it -- see docs/DECISIONS.md for
  what that trades against a real grid.
- Links carry their URL through to a tap. `GestureOutcome::Tapped`
  is new: a press that never committed to a pan or a selection, so a
  finger that flung the list past a link does not also open it.
  `iris::platform::OpenUrl` is the capability, implemented by each
  backend (xdg-open/open/start on the desktop, an ACTION_VIEW intent
  deferred to `after_input` on Android, the same shape
  `pending_show_keyboard` uses).
- `DragArbiter`/`DragGesture` take an axis, so a code fence pans
  across its own long lines through the same machine a list pans
  down its rows -- and a vertical drag starting on a fence still
  reaches the list.
- `TextEditCtx::byte_at` answers which byte a tap landed on without
  exposing the parley layout; `Rect::radius` takes a `Len`, so a
  corner can be written in dp.

Tests: 31 in transcript-ui (11 new, covering the frame mapping,
highlighting including a multibyte fence and an unknown language,
list markers, table padding and wrapping, link hit-testing), 85 in
iris (4 new on the tap-vs-drag rule and the two axes).
cargo fmt clean, clippy warning-free.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-06 18:55:46 -04:00
irisandClaude Fable 5.1 20303e0b4c IRIS.md: take_counters gained a fourth number, text shapes
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-06 18:40:38 -04:00
irisandClaude Fable 5.1 6973a89815 docs: the verification pass over Tasks A and B, and the composer background withdrawn
RUST.md gains the pass's findings with their commits and the numbers:
the block model held under a per-character prefix property, the
size-independent hit-box defect and its fix, the tail-rebuild selection
gap, why the three new debug_asserts are whole-set, the text-shape
counter that turns "a delta costs one block" into a measurement, and the
verification bench run.

IRIS_TODO.md's "the bar's own grey background is not drawn" is
withdrawn: decoding the screencap puts it at rgb(41,40,49), full width,
y2245..y2365 -- drawn, and dark on black, which is most likely what the
earlier reading was.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-06 18:40:25 -04:00
irisandClaude Fable 5.1 c3cfc67bb3 iris: count text layouts, so "a delta shapes one block" is measured rather than argued
take_counters gains a fourth counter, text shapes, bumped in
Painter::render_text -- which TextView::render only reaches on a cache
miss, so it counts shapes and not requests. A draw counter cannot stand
in for it in either direction: a widget can be redrawn without
re-shaping (the layout is memoized by width) and re-shaped without any
extra draw, and re-shaping is the whole thing the per-block transcript
row exists to avoid.

With it, a_delta_into_a_long_reply_redraws_the_same_widgets_as_a_short_one
asserts the number docs/DECISIONS.md's 2026-09-06 entry actually claims:
one delta into a 100-paragraph reply shapes exactly one text layout, the
same as into a one-paragraph one. Before the split that was necessarily
O(message), since the reply was one buffer.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-06 18:32:31 -04:00
irisandClaude Fable 5.1 155d899e55 transcript-ui: pin the tail rebuild's unregister with the case that broke it
e1030d6 made Selection's key (RowKey, u32) and changed apply's
ReplaceLast arm to unregister unconditionally rather than only when the
key changed -- correctly, but with nothing exercising it. The case is a
tail row rebuilt under the *same* key with fewer blocks than it had: the
blocks that no longer exist keep pointing at widgets replace_back's drop
frees, and Selection::begin resolves every registered handle on an
ordinary press, so the next tap anywhere in the transcript panics. The
old `if new_key != old_key` guard could not see it, because nothing
about the key changed.

Selection::registered_blocks (test-only) is what lets the test assert the
contract unregister states -- every block of the row, not the first --
instead of only that nothing panicked.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-06 18:32:05 -04:00
irisandClaude Fable 5.1 e63e923d44 iris: a size-independent widget's hit box lands where it is drawn
draw_inner's third fast path -- offered region changed shape, widget's
output does not depend on it -- rewrites the widget's own primitives in
place and writes no move-slot delta at all. 167862c added a
move_applied increment there, copied from mov, where region and the slot
delta really do move together. Here only region moves, so resolved_region
subtracted a distance the chain never held and every such widget's hit
box sat short of its drawing by exactly the last step it took.

Span reaches this on the first frame of any tree it is in: it measures
each child at the full region and then places it, which for a Rect (the
.background(rect(..)) idiom, list row tints) is a size change through this
branch. So the hit box was wrong from the start, with the drawing correct
-- nothing on screen to say so.

a_size_independent_widget_moved_by_its_parent_has_the_hit_box_it_is_drawn_at
is the sibling of a_panned_widgets_own_hit_box_moves_exactly_once on the
branch that fix had no reason to touch; it fails on both frames without
this.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-06 18:31:58 -04:00
irisandClaude Fable 5.1 a56a928b0c client-core: the transcript's own markdown shapes, and the streaming property as a property
split_blocks was tested on the shapes it was written against. These are
the ones a real reply contains -- a fence with blank lines in it, a `---`
inside a fence, a nested list, a fence directly under a heading, a table,
a quote -- plus the property RowBlocks::apply_delta actually depends on,
checked at every character boundary of a message that has all of them:
growing a message may rewrite its last block and never an earlier one, or
common_prefix must say so. No defect found; the split already held.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-06 18:28:57 -04:00
iris 0449a324ef docs/RUST.md: P1 started on Iris's word, sub-order P1a-P1e by what makes the bench fair 2026-09-06 18:28:48 -04:00
irisandClaude Fable 5.1 e1030d69f6 iris: a transcript row is a column of markdown blocks, so a streamed delta costs one block
A row was one TextEdit holding the whole message, so every delta
re-shaped every paragraph of a long reply through parley -- the one phase
where iris trails Compose on the phone (p50 18.2ms vs 13.4ms, bench v2).

- client-core/src/markdown_blocks.rs: split a message into its top-level
  blocks with their source, through the same pulldown-cmark the renderer
  parses with so the two cannot disagree about where a block starts, plus
  common_prefix. Appending markdown can rewrite an earlier block (a
  trailing --- turns the paragraph above into a heading), so the fast
  path compares the prefix it keeps rather than assuming it -- with the
  test that says so.
- transcript-ui: a row is a Span of one TextEdit per block;
  RowBlocks::apply_delta replaces the block a delta lands in;
  TranscriptScreen keeps the tail row's blocks, seeded in build_tree as
  well as push_row (a screen opened onto a streaming reply took the
  rebuild path for its first delta otherwise, with nothing to say so).
- A block is the selection unit: Selection is keyed by (RowKey, u32),
  which is reading order at both levels, and the pointer-captured half of
  a drag resolves the block under the finger from its drawn box
  (Selection::locate) instead of from the row's extent.

Pass condition: a_delta_into_a_long_reply_redraws_the_same_widgets_as_a_short_one
drives a real UiRenderState and asserts the draw count for a delta into a
100-paragraph (3,000+ char) reply equals the count for a one-paragraph
one. 30 either way; it read 630 against 30 twice on the way there.

Emulator stream phase, same AVD before and after: p50 61.5 -> 54.5ms,
p90 211.7 -> 113.1ms, p99 342.6 -> 137.4ms, worst 403.6 -> 143.0ms, 202
-> 293 frames in the same 21 seconds. Selection across blocks verified
with a real long-press drag.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-06 17:33:37 -04:00
irisandClaude Fable 5.1 167862ca1b iris: the composer scrolls on a finger -- a dp cap worth zero, a stale mask slot, a hit box moved twice
Wrapping the composer's field in .scrollable().masked() needed three
layout defects fixed first, each with a headless regression test that was
confirmed to fail without its fix:

- MaxSize/Sized reported a caller's declared dp length unresolved, and
  Span places a child from the abs/rel of what it reported, so dp(168)
  was worth zero: the bar got a slot of nothing the moment its content
  passed six lines and the Scroll inside measured its container at -63px
  (container=-63 content=415.8 amt=478.8 on the emulator). Len::fold_dp,
  used on the way out, plus a debug_assert in draw_inner that a reported
  Size carries no dp -- the rule is about every widget, not those two.
- Masked allocated a fresh mask slot per draw, and draw_inner's
  unchanged-region fast path does not revisit descendants, so they kept
  clipping against a box the bar had moved away from: four live mask
  entries, none of them current, and the field drew nothing.
  ActiveData::own_mask, allocated once and rewritten in place.
- mov updates active.region and accumulates the same delta on the move
  slot, and resolved_region added both, so a panned widget's own hit box
  sat at twice the pan -- the composer's field was untappable after a
  drag. ActiveData::move_applied.

Scroll itself measured the right number by a misleading route; it is
written against painter.px_size() now and still reports its content's
size, since reporting the container makes the answer a function of
itself.

Verified on this checkout's emulator: swipe 540 1200 -> 540 1460 moved
the field's Message box 31,1041..1048,1509 -> 31,1131..1048,1651 with its
height unchanged at 468px.

run-bench.sh polled logcat for a prefix copy_report also logs at startup,
so it printed a report that had never been run.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-06 17:17:42 -04:00
iris d73db97629 iris/android-app/build-apk.sh: clear jniLibs before building, so only the requested ABI is packaged 2026-09-06 16:47:54 -04:00
irisandClaude Fable 5.1 fb6b459c2c iris: Scroll pans on a finger drag; a vertical drag in a focused field scrolls rather than selects
IRIS_TODO.md's "the composer has no touch-drag scroll". `Scroll::drag`
takes its pan from the same `sense::DragGesture` `List` is driven by --
arbitration, DRAG_SLOP, velocity and pointer capture all stay in sense.rs
and only what a committed pan *means* is decided per caller -- and
`WidgetLike::scrollable()` registers it beside the wheel handler it already
registered, so every scroll area pans on a finger with nothing added at the
call site. No fling: `Scroll` has no per-frame tick to animate one and the
areas it wraps are at most a screenful. `Scroll::amt()` exposes the pan
position.

`attr.rs`'s `on_press` treated an already-focused field as the plain
click_or_drag case, so every Pressing frame extended a selection. It now
applies the same DRAG_SLOP rule its unfocused branch already did: a press
past the slop vertically abandons its pending selection for the rest of the
gesture, so the scroll area around the field wins it. That is Android
EditText's own behaviour and it is what lets a swipe up over the composer
scroll instead of dragging a highlight through what you typed.

Also fixed, found doing it: `ActiveData::mask` stored the mask a widget
*set* rather than the one it was drawn *under*, and `redraw` feeds that
field back in as the inherited mask -- so a targeted redraw of any `Masked`
handed it its own mask and aborted on `set_mask`'s nested-mask assert. A
real abort on the emulator, `assertion failed: self.mask == MaskIdx::NONE`.

And the per-frame orphan guard from 76b1f99 is now a count comparison
(O(active widgets)); the O(primitives) walk only runs to build the failure
message, because running it per frame made a debug build on the emulator too
slow to finish a bench run at all.

Tests: four in scroll.rs (pan past the slop, a tap inside it, a horizontal
drag, the end clamp), `a_finger_drag_over_a_scroll_area_pans_it` in
sense_tests.rs driving the whole registration/dispatch/capture path (fails
with "got 0" without the new registration), and
`redrawing_a_masked_widget_does_not_nest_its_own_mask` in layout_tests.rs
(aborts on the pre-fix code).

The composer itself is deliberately still not `.scrollable()`: `Scroll`
measures against the window rather than its own offered box, so inside the
`MaxSize` capping it at six lines it pans the field out of the bar --
measured, reverted and written down in RUST.md and DECISIONS.md.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-06 16:45:56 -04:00
irisandClaude Fable 5.1 76b1f99277 iris: a dirty widget redrawn by its ancestor never freed its old primitives
`draw_inner` read `needs_redraw` without consuming it, and used it to skip
the whole `if let Some(active)` block -- including the `remove(id, false)`
that frees a redrawn widget's previous primitives. So a widget that was
both already active and marked dirty, and was reached by an *ancestor's*
draw rather than by `redraw_updates` picking it first, drew a second full
set of primitives and then had `active.insert` overwrite the only handles
that could ever have freed the first set. Those primitives stay in the
layer's instance buffer for the life of the process, with a leaked move
slot and leaked mask refs, drawn every frame at whatever region they last
had -- and `List` sets no mask, so a row measured at `GENEROUS_PADDING`
leaves its ghost outside the list's own box.

That is the doubled `Compacted:` row in docs/bench/iris-phone-v2-2026-09-06.md:
overlapping copies inside the transcript and one more below the composer.

Fixed by consuming the mark (`needs_redraw.remove`) at the top of
`draw_inner` -- this call *is* the redraw it asked for -- and freeing the
old primitives on the dirty path too.

Guarded so it cannot come back silently: `UiRenderState::orphaned_primitives`
walks every layer's live instances and names any whose owner is no longer
active or no longer holds a handle to them, and `update` `debug_assert!`s it
empty every frame (debug builds only). New regression test
`an_ancestor_redrawing_a_dirty_row_leaves_no_stale_copy` in list.rs fails on
the pre-fix code with "1 primitive(s) survived their own widget's redraw".

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-06 13:59:53 -04:00
iris 3e72a4ef19 docs: the defect pass's findings -- RUST.md boxes, IRIS_TODO ticks, DECISIONS and IRIS entries 2026-09-06 13:47:28 -04:00
iris c02152a4f4 iris: a tap on an empty text field left no caret, so typing was silently dropped
TextEditCtx::select compared the tap against the laid-out text's own box
and cleared the selection for anything outside it. An empty field lays
out to a zero-width box, so tapping the composer granted focus and opened
the keyboard with no caret, and insert_str returns early without one --
every keystroke went nowhere and no glyph was ever emitted. Parley clamps
a point outside the layout by itself, and a press reaching select() has
already been hit-tested to the widget, so there was nothing for the
'outside' branch to mean.

insert_str now debug_asserts rather than dropping input silently, and
UiRenderState::draw_started -- a re-entrancy guard whose test was written
after its own remove(), so it could never fire, and which grew by one
entry per widget ever drawn -- is restored to what it was meant to be:
inserted around Widget::draw, removed when it returns, asserted empty at
the top of every update.
2026-09-06 13:43:56 -04:00
iris d9872989fa iris/android: the composer's launch position was the bench report pane, plus surface/insets lifecycle logging
The empty benchmark-report TextEdit held .height(rest(1)) beside
content.height(rest(2)), so it reserved a third of the window at every
launch and pushed the composer two thirds down -- Iris's 11:39 phone
report. It is sized to its content now, capped and scrollable, and sits
above the transcript rather than under the composer.

New log::info! lines for one insets change, one surface_changed, one
renderer build and one surface_destroyed, each with the glyph/atlas
counts, so a phone's adb logcat can answer the app-switch text loss the
emulator cannot reproduce.
2026-09-06 13:26:34 -04:00
iris 2fed8b34b3 Merge branch 'worktree-agent-a6e37a2335f436d08' into rustify 2026-09-06 13:17:22 -04:00
irisandClaude Fable 5.1 1f379e8384 docs/REVIEW-2026-09-06.md: fix all ten review findings; RUST.md/IRIS_TODO.md: DragGesture merge checks
Finding 1 (the real crash): Selection::clear() drops rows and anchor,
called from TranscriptScreen::apply's Rebuild arm right before
List::clear() -- push_row re-registers survivors as it rebuilds each row.
Fixes a WeakWidget outliving the row group_tool_runs regrouped away,
which panicked the next long-press anywhere. New apply_tests test builds
a real TranscriptScreen, forces the regroup, and confirms no panic.

Findings 2-5: debug_assert!s on List::place's slot, List::fling and
FlingCalculator's velocity finiteness, VelocityTracker::add_sample's
chronological order, and FrameReport::mark_phase's non-decreasing
start_index. Finding 7: bench_client.rs's battery_line guard restructured
so the empty check can't be separated from its unwraps by a future edit.
Findings 9/10: new List tests pinning tick_fling's per-tick deceleration
and replace_back's evicted-key cleanup with a different key than the
existing tests use. IRIS.md's replace_back/clear/apply entry gained the
side-table-clearing note the Docs finding asked for.

Also records this pass's DragGesture-merge verification in RUST.md (tap
stays vs swipe doesn't, a real fling keeps moving after release, keyboard
cycles confirmed via on_insets_changed) and annotates the two IRIS_TODO.md
phone-report items it targets.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-06 13:16:16 -04:00
iris bf3479f5c4 client-core: an unasked page is not an empty one, and two guarded invariants
Review of 73251d6's port of TranscriptSource/joinPages.

`TranscriptSource::page` answered `before == 0` with an empty `Vec`, which
is the same value it answers "this conversation has no more history" with.
That is the state the Kotlin keeps apart: `loadOlderPage` returns false at
`oldestSeq == 0` *without* touching `moreHistory`, and returns false on an
empty page *by latching it*. Collapsing the two moved AGENTS.md's paging
bug one layer down rather than fixing it. `page` returns `OlderPage` now --
`Events(vec![])` is the start of the conversation, `NothingLoaded` is not
an answer about the conversation at all.

`join_pages`' `debug_assert!` on seq ordering across the boundary is not a
true invariant: a peer note carries the seq its turn began at, which can be
older than the page it arrived in, so an ordinary transcript would have
panicked a debug build there. Replaced with the one the function exists to
enforce -- no tool id surviving in both halves.

`fetch_transcript_lines` stores `RawValue`'s exact server bytes, so the
"neither source can produce a newline" comment in `SessionCache::append`
now rests on the server's serializer staying compact rather than on a
local normalization. Checked with a `debug_assert!` in `append` and
`store_page` rather than trusted.

Tests for the failure half, which the port had none of: a 500 mid-page, a
cached line this build cannot read, and the `after` bound in the case that
actually carries one (the existing test asserted only the case with no
bound). `cargo fmt`, `cargo clippy --all-targets`, `cargo test` (112) clean
in client-core; `cargo check -p desktop-app` clean.
2026-09-06 13:00:37 -04:00
iris 312455956d Merge remote-tracking branch 'origin/rustify' into worktree-agent-a6e37a2335f436d08 2026-09-06 12:39:22 -04:00
irisandClaude Fable 5.1 73251d6b8b client-core: port TranscriptSource and joinPages page-boundary healing
Closes docs/RUST.md's "client-core prerequisites for P1" box: the
cache-vs-server stitching TranscriptSource.kt does, and the
joinPages/healSplitMessage/adoptRun page-boundary healing
TranscriptItems.kt does, both ported into client-core with no UI
framework dependency.

Neither Kotlin file had a JVM unit test of its own, so the port used the
Kotlin source and AGENTS.md's "things that have bitten" paging incidents
as the spec instead of a test-for-test transcription. Both regressions
get a dedicated test: TranscriptSource::page refuses before == 0 before
touching the cache or the network (loadOlderPage's incident), and
adopt_run now runs on every page join rather than only the one where a
split call was found (the "one run drawn as two" incident).

fetch_transcript_lines (api.rs, additive) pairs each transcript line with
the exact server bytes via serde_json::value::RawValue rather than
re-serializing a parsed Value, so a cached line and a live SSE frame for
the same event agree byte-for-byte -- the fetch_transcript_page other
callers under iris/ depend on is untouched.

client-core: 85 -> 109 tests. cargo test/clippy --all-targets/fmt clean.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-06 12:38:59 -04:00
iris 2e00e71552 docs: Iris's 11:39 phone report on the 02:07 build, four open items 2026-09-06 11:42:21 -04:00
irisandClaude Fable 5.1 f802de94b5 Merge worktree-agent-a754368325fa06839 into rustify: DragGesture, pointer capture, edge-to-edge insets
Generalizes drag arbitration into a default-input DragGesture with
pointer capture and CursorSense::Drop, and opts MainActivity into
edge-to-edge so IME insets are redelivered. See e12c708.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-06 11:38:24 -04:00
iris 9717d1c4b0 docs/RUST.md: 2026-09-06 orchestrator plan for the P0 defects and the P1 prerequisites 2026-09-06 11:37:32 -04:00
iris 9458f443ad Merge remote-tracking branch 'origin/rustify' into worktree-agent-a754368325fa06839 2026-09-06 02:10:55 -04:00
irisandClaude Fable 5.1 e12c708246 iris: generalize drag arbitration into a default-input DragGesture, with pointer capture and Drop
Iris asked (2026-09-06) that dragging be part of iris's default input
system rather than duplicated per app: "anything that provides good
performance and can be generalized well is part of iris rather than the
app." DragArbiter and VelocityTracker (both already in iris::sense) are
now bundled into a new DragGesture, which also takes exclusive pointer
capture (UiRenderState::capture_pointer/release_pointer/captured_pointer)
the moment a gesture commits to panning or selecting, and delivers a new
CursorSense::Drop -- not PressEnd -- to the captured widget when the
button lifts, wherever on screen that happens to be.

This directly targets the phone bench's "finger flings do nothing":
per-widget hit testing silently drops a gesture the instant the pointer
moves off every registered region, which a fast pan/fling does routinely
(crossing several virtualised rows, or ending off the loaded content
entirely) -- so PressEnd, and the velocity/fling-start decision hanging
off it, was frequently never delivered at all. Capture targets List's own
stable id (List::key_at resolves the row-under-pointer from its
extents), not a row's, since List retires rows mid-drag as content
scrolls.

transcript-ui::Selection::drag now only decides pan-vs-select from
DragGesture's outcome; row.rs's per-row registration is only ever a
gesture's first frame, with lib.rs registering the List-level
continuation once. New tests: sense_tests.rs's two pointer-capture
regressions, list.rs's replacing_the_last_row_many_times_does_not_leak_primitives
(a P0 stale-primitives diagnostic -- passes, pinning the widget-arena
layer as not the leak). MainActivity.java opts into edge-to-edge
(Window::setDecorFitsSystemWindows(false), API 30+, no new dependency)
so window insets are redelivered on every change including a pure IME
toggle -- the named-but-untried fix for the phone bench's "keyboard:
could not be shown" and the emulator's identical non-confirmation.

cargo fmt/clippy/test clean across the iris workspace.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-06 02:10:48 -04:00
iris 543f6d92f0 Merge worktree-agent-a9002910a315fe719 into rustify: composing text, tap-vs-swipe focus, composer rebuild, atlas reset 2026-09-06 02:08:12 -04:00
iris 27ca5b2349 Merge remote-tracking branch 'origin/rustify' into worktree-agent-a9002910a315fe719 2026-09-06 02:03:04 -04:00
irisandClaude Fable 5.1 20b12255e1 iris/android: composing text sync, tap-vs-swipe focus, composer rebuild, atlas reset on app-switch
Four fixes from Iris's phone report on the dc01f88 build, plus her same-day
follow-up on swipe-vs-tap:

- android/ime.rs: InputConnection now calls InputMethodManager.updateSelection
  after every edit (new update_ime_selection, called from after_input) -- Gboard
  was holding keystrokes back with nothing telling it the app's selection/
  composing region had moved, which read as "doesn't enter it until I hit
  space, doesn't move the caret". New unit tests in widget/text/edit.rs cover
  the buffer-level composing/commit/delete/selection operations directly.

- attr.rs: Selector/Selectable rewritten around a shared on_press dispatcher
  over PressStart/Pressing/PressEnd instead of click_or_drag(), so a field
  that isn't already focused only grants focus (and requests the IME) on a
  completed tap -- press and release with no frame past DRAG_SLOP. A drag
  is never consumed, so whatever is behind the field still sees it. New
  FocusHost::is_focused (both platform impls) and TextEdit::press_origin
  back this. Verified on the emulator: dumpsys input_method's mInputShown
  stays false after a swipe over the composer, true after a tap.

- iris_core: GlyphAtlas::clear()/Textures::reset(), called together from
  android/view.rs's surface_changed exactly when a genuinely new renderer is
  built (app-switch, not the keyboard-resize path that already reuses the
  renderer) -- both CPU-side caches otherwise kept pointing at the old,
  destroyed device's textures. Verified on the emulator: home, reopen, every
  glyph still on screen.

- transcript-ui/composer.rs: rebuilt as one widget (unchanged Stack{rect,
  span} idiom, capped at ~6 lines via MaxSize + .scrollable(), wrapped in one
  Pad whose bottom Composer::set_bottom_inset rewrites in place so the bar
  sits on the IME or nav-bar inset with no rebuild -- rebuilding would drop
  focus/selection/in-progress text). Wired from bench_client.rs's existing
  on_insets_changed.

A second, deeper bug found while verifying the composing fix is NOT fixed
this pass: composed text never becomes visible at all. A new layout_tests.rs
test proves the widget tree's own region math is correct across a keyboard
resize, ruling that out; RUST.md's P0 box has the full writeup and what to
check next (UiRenderState::redraw's single-widget path, or something
force-gles-specific -- this AVD has no Vulkan adapter to rule that out with).

cargo fmt/clippy/test --workspace and cargo ndk clippy all clean.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-06 02:03:00 -04:00
irisandClaude Fable 5.1 71a3fae655 IRIS_TODO.md: streaming re-lays out the whole message, from the phone's bench v2
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-06 01:35:26 -04:00
irisandClaude Fable 5.1 c3984da623 docs/bench: iris bench v2 report from Iris's phone, verbatim, with her observations
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-06 01:34:39 -04:00
irisandClaude Fable 5.1 2e3f4ada38 Merge iris fling/jitter fix + Benchmark v2 + header/ime follow-ups into rustify
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-06 01:23:50 -04:00
irisandClaude Fable 5.1 03c6be80a3 iris android-app: header-duplicate investigation, ime-inset fix for keyboard confirmation
Two follow-ups after the keyboard/dp/header pass, both requested against
the P0 box:

(a) The header row rendering a second time inside the transcript area
after a keyboard-triggered resize: reproduced reliably (tap the composer,
screenshot after the keyboard opens). Ruled out one concrete hypothesis --
on_insets_changed rebuilding top_bar on every ime_bottom change, unrelated
to the header's own status-bar padding -- with a guard (last_top_pad) that
reproduced the identical duplicate afterward, so repeated rebuilding is
not the cause. Kept the guard as a real (if insufficient) fix for needless
rebuilds. Not root-caused: Span's two-phase provisional/real draw and the
redraw_all-vs-redraw_updates split are the two live suspects, but pinning
which one (or something else) produces the duplicate needs instrumenting
draw_inner directly or the phone. Full writeup in RUST.md's P0 box.

(b) Why on_insets_changed's ime_bottom never confirmed the keyboard being
shown, on either the auto-diagnostics or the new bench keyboard phase:
MainActivity.java uses windowSoftInputMode="adjustResize", under which
WindowInsets.Type.ime()'s own inset amount is defined to read zero (the
window already resized to avoid the overlap that inset would describe) --
the same trap AGENTS.md already names for the Compose side. Fixed to read
insets.isVisible(ime()) instead, a boolean unaffected by resize-vs-pan.
This alone did not make the callback re-fire on this emulator, which
still shows no insets callback after the initial one at attach -- named
but unconfirmed hypothesis: a non-edge-to-edge Activity may not get insets
redelivered for a pure IME toggle handled via resize, needing an edge-to-
edge opt-in this pass did not attempt given the risk to adjustResize's
own behavior.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-06 01:23:36 -04:00
iris 4afc453faa Merge remote-tracking branch 'origin/rustify' into worktree-agent-a16b22e34539b810e
# Conflicts:
#	iris/android-app/src/bench_client.rs
#	iris/android-app/src/bench_jni.rs
2026-09-06 01:05:18 -04:00
irisandClaude Fable 5.1 1aab61bf26 iris android-app: Benchmark v2 -- fling, type and keyboard phases
Implements RUST.md's "Benchmark v2" spec in bench_client.rs: fling (8 out
+ 8 back at 12,000px/s through List::fling, waits for !is_scrolling()
capped 3s, reports travel as row index + offset via List's new
anchor_position_display), stream (unchanged), type (the 600-char P0
constant, one char per 50ms into the composer's real TextEdit via .set(),
then deleted), and keyboard (5 show/hide cycles via bench_jni.rs's new
InputMethodManager calls, confirmed from on_insets_changed's real
ime_bottom transitions rather than assumed from the JNI call returning).

FrameReport gained mark_phase/phase_stats/late_at_hz (iris/core) so the
report can show a per-phase block (frames, late%, p50/p90/p99, worst)
against the display's real refresh rate (bench_jni's new
refresh_rate_hz), matching the shape docs/bench/compose-phone-v2 uses.
RING_CAPACITY bumped 4096->16384 since a full v2 run is ~3,000+ frames.

Found and fixed a real deadlock while wiring this up: read_from_state
(a new helper that gets a value back out of a spawned task's ctx.update,
which has no return channel of its own) only worked for its first call in
a chain, because nothing called redraw.request_redraw() after enqueueing
later ones -- nothing then drains the task channel to run them. Every
call now triggers its own redraw.

Verified end to end on this checkout's x86_64 emulator (force-gles, cold
boot): fling/stream/type all report populated phase blocks; keyboard's
show never got a real on_insets_changed confirmation this run (see
follow-up work). Full report and travel numbers go in RUST.md's P0 box
next.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-06 01:02:04 -04:00
iris dc01f88d75 Merge branch 'worktree-agent-a1ff0294b6c29127e' into tmp-merge 2026-09-06 00:54:21 -04:00
iris c589a75fa0 Merge remote-tracking branch 'origin/rustify' into worktree-agent-a1ff0294b6c29127e
# Conflicts:
#	docs/RUST.md
2026-09-06 00:54:06 -04:00
irisandClaude Fable 5.1 4b62cc642e docs/RUST.md: emulator verification results for the keyboard/dp/header fixes
run-bench.sh end to end clean (24/24 swipes, 400/400 events); header
background confirmed by screenshot; the keyboard wipe fix confirmed two
ways (a forced wm size resize and an actual soft-keyboard open, both real
surface_changed triggers, text intact both times).

Also records two things found during this verification and not fixed:
the top button row appears to render a second time, out of place, after
a keyboard-triggered resize, and a tap aimed at the field below can land
on it instead -- and the keyboard diagnostics auto-capture never fired in
this session. Neither is root-caused; explicitly not attributed to this
pass's changes without more evidence, per the standing rule against
blaming ambient failures on your own code without measuring first.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-06 00:51:27 -04:00
irisandClaude Fable 5.1 80c2eadec9 docs: record the keyboard-wipe fix, the dp unit and the header fix
docs/IRIS.md's 2026-09-06 entry (public API), docs/LAYOUT.md's "Density:
Len::dp" design section, IRIS_TODO.md's density-unit item ticked, and
docs/RUST.md's P0 box gets the investigation: the keyboard-wipe
hypothesis and confirmation, the blur root cause and why the dp unit
turned out to be the same fix, the header cause, and what remains
unverified (an emulator screenshot of the keyboard fix, and Iris's real
phone).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-06 00:40:59 -04:00
irisandClaude Fable 5.1 0b587629e6 iris/android-app bench: auto-capture diagnostics when the keyboard opens
So Iris can get a report off the phone even if the keyboard wipe (or
some other keyboard-triggered regression) is still present on whatever
build she is holding, independent of whether the on-screen Diagnostics
button itself is drawing.

on_insets_changed edge-triggers on ime_bottom becoming non-zero, waits
KEYBOARD_DIAGNOSTICS_DELAY_MS (500ms, long enough for the resize and a
couple of frames to settle) via a spawned task, then
capture_keyboard_diagnostics reuses show_diagnostics's exact report text,
logs it, copies it to the clipboard unprompted, and shows it through a
new PlatformHandle::show_diagnostics_overlay call into
IrisView.showDiagnosticsOverlay -- a plain TextView + Copy/Close panel
added over the existing IrisView (not replacing it, unlike
showRendererError's one-way trip) so it draws independently of whatever
iris's own renderer is doing, and Close returns to the still-running
session underneath.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-06 00:39:14 -04:00
irisandClaude Fable 5.1 3163256d2c iris/android-app: opaque header background, header sizes onto dp
Iris's phone report (build a9232ac): "the header buttons have nothing
behind them and overlap the transcript text." Only each button's own
rect painted anything, so the gaps between and around them (and the
status-bar strip above) showed CLEAR_COLOR (black) one layer back, and
the row's reserved height was three abs (physical-pixel) button boxes --
smaller, on a dense phone, than the dp-correct size the transcript below
now uses post the previous two commits, which is what reads as overlap
once the two disagree.

Fixed with a HEADER_SURFACE rect stacked behind the whole button row
(not just behind each button), and every non-text size in the header
(button padding, row height, the report field's padding) moved from a
bare number to dp(...), so the row's reserved height in the outer
Span::DOWN matches what is actually painted. The list/report field
already sit below the header in that same Span::DOWN, not behind it --
no stacking change needed there.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-06 00:35:45 -04:00
irisandClaude Fable 5.1 6102e0d4d9 iris: a dp length unit, resolved against density; crisp glyphs at physical size
Iris asked for this 2026-09-06 (IRIS_TODO.md, "a third length kind beside
relative and pixels ... a unit resolved against the display's density at
layout time"): before this, a Len was abs (physical pixels) or rel/rest
(a fraction of the parent), and the only way to make a design size look
the same physical size on a denser display was a single global multiply
applied after layout -- which the previous commit found is also what
made text blurry.

Len gains a `dp` field, resolved against a `density: f32` (physical
pixels per dp) now carried on UiRenderState/Painter
(`UiRenderState::set_density`/`density()`, `Painter::density()`) and
threaded through every `apply_rest`/`to_uivec2` call site. `len_fns::dp`
/ `Len::dp` construct one, exactly parallel to the existing `abs`/`rel`/
`rest`. A bare number is unaffected (still `abs`, physical pixels) --
`dp` is opt-in.

Text: `TextBuffer::shape` now takes `density` and multiplies
`font_size`/`line_height` (and any span override) by it before handing
them to parley, so the size that reaches the shaper and the rasteriser
(`TextData::place`) is the display's real physical size -- the atlas
holds a bitmap at the resolution it is actually shown at, instead of a
low-resolution one stretched afterward. `GlyphKey.size` already keys on
the resolved `font_size`, so a cache entry is naturally per physical size
with no further change. `TextData` also carries its own `density` copy
for `TextEditCtx::layout` (cursor movement/hit-testing), which shapes
text from an input callback with no `Painter` to read it from.

`Span::gap` and `Padding`'s four sides move from bare `f32` to `Len`, so
`.gap(dp(4))`/`.pad(dp(10))` work the same way any other size does; a
bare number still means physical pixels, unchanged.

Migrated transcript-ui's non-text sizes (row gap/padding, composer
padding) and one example to the new unit, per IRIS_TODO.md's "done when"
list. Android's own density (`DisplayMetrics.density`) is wired to both
copies in `new_peer`; the winit backend has no per-monitor density wired
up yet and stays at the default (1.0).

docs/IRIS.md, docs/LAYOUT.md and IRIS_TODO.md updated next.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-06 00:35:38 -04:00
irisandClaude Fable 5.1 f0da383e28 iris/android: reuse the renderer across a surface resize, fix the keyboard glyph wipe
Hypothesis confirmed by reading the path end to end before changing
anything: surface_changed fires on every SurfaceView size/format change,
not only a genuinely new Surface -- showing the IME under adjustResize
resizes the same surface through this exact callback. The handler
unconditionally dropped AndroidRenderer and rebuilt it via
AndroidRenderer::new, which allocates a brand-new, empty glyph atlas and
fresh GPU buffers, while iris_core's CPU-side glyph cache kept the UV
coordinates it had already handed out against the *old* atlas -- so every
glyph drew from a rectangle pointing into a texture that had just been
recreated empty. Rects never go through the atlas, so they kept drawing:
exactly Iris's report ("rectangles stay; only text disappears").

Fixed by reusing the existing AndroidRenderer (device, atlas, buffers,
bind groups) and only reconfiguring the surface + window uniform via its
existing resize() when a renderer is already live; AndroidRenderer::new
now runs only when surface_changed finds `renderer` already None (a
genuinely new surface, e.g. after surface_destroyed/backgrounding).

While in this path, removed the global logical/physical scale stopgap
(dividing window size, touch coordinates and insets by content_scale)
that the P0 "text too small" fix had added: it is what made text blurry
next (a glyph rasterised small then stretched by the NDC mapping onto the
real physical framebuffer). Window size, touch and insets are physical
pixels throughout now, matching AndroidRenderer's own swapchain
resolution; density is resolved per-length instead (next commit).
LogicalInsets renamed to WindowInsets to match.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-06 00:35:22 -04:00
irisandClaude Fable 5.1 2d3695a1d3 Merge iris fling/jitter fix into rustify
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-06 00:20:39 -04:00
irisandClaude Fable 5.1 f06ee259b4 iris: List::fling with Android's spline physics, and fix the drag-slop scroll jitter
Adds VelocityTracker and a port of AOSP SplineOverScroller's fling curve
(FlingCalculator, cited at the definition) to iris::sense, and wires
List::fling/is_scrolling/cancel_fling/tick_fling through
Selection::drag's release path -- a pan's release now decelerates instead
of stopping dead on the finger lifting, matching IRIS_TODO.md's "swiping
has no momentum" ask. Clamped at the loaded content's start/end and
cancelled by the next touch-down.

Also fixes the scroll jitter DragArbiter's slop release caused: crossing
DRAG_SLOP applied the whole pre-threshold drag (measured from press_start)
in one step, since nothing pans while a gesture might still resolve to a
selection. Now only the excess past DRAG_SLOP is applied on that frame,
the same way Android's own touch handling consumes touch slop rather than
replaying it.

Root-caused by reading DragArbiter's state machine and covered by new
unit tests (fling distance against the closed-form spline result within
1%, cancel-on-touch, start/end clamp, the slop-crossing regression); no
emulator was used this pass, so an on-device trace/feel-check is still
open, and Benchmark v2's four-phase bench_client.rs spec was not
attempted. docs/IRIS.md, docs/IRIS_TODO.md and docs/RUST.md's P0 box
record what's done and what's left.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-06 00:20:24 -04:00
irisandClaude Fable 5.1 560a74caf8 docs: record the phone-report fixes, follow-ups and the bundled-font API
RUST.md's P0 box gets Iris's first real-phone report (no crash) and the
four defects it found (glyph-wipe-on-first-touch, missing bold glyphs,
text far too small, status-bar inset not applied), what was fixed and
how it was verified on the emulator, and what's still open (item 1's
root cause, and the top-row height anomaly noted in the last commit).

IRIS_TODO.md gets a new "From the phone, 2026-09-06" section for the two
items explicitly deferred to a follow-up agent: no scroll momentum/fling,
and occasional jitter scrolling down.

IRIS.md gets the public-API entry for TextData's bundled fonts/
font_diagnostics, UiRenderNode::new/resize's new window_size parameter,
AndroidUiState::content_scale, AndroidAppState::on_insets_changed, and
iris_core::WgpuErrorLog.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-05 23:59:40 -04:00
irisandClaude Fable 5.1 fd7e17523d iris/android: fix layout/shader unit mismatch left by the density-scale commit
surface_changed's self.render.resize(...) -- UiRenderState::output_size,
what every widget's absolute PixelRegion (a fixed .height(56), notably)
is computed against -- was still being handed raw physical width/height
after the previous commit switched AndroidRenderer's own size()/resize()/
new() to logical (physical / content_scale) for the shader's window
uniform. That split layout and the shader into two different units:
layout placed a "56"-unit row inside a ~2219-physical-unit-tall canvas
(an absolute box, still exactly 56 units), the shader then divided that
same 56 by a ~845-unit *logical* window dimension -- found on the
emulator by measuring a fresh install's top button row at ~40 physical
px against the ~147px `56 * content_scale` predicts. Proportional
(rest(n)) sizes hid the mismatch by adapting to whichever total they were
given; only fixed sizes exposed it. Now divides by content_scale here
too, matching every other call site.

Verified on this checkout's emulator (EMU_GPU default, force-gles):
run-bench.sh completes end to end (frames=691, 24/24 swipes streamed
400/400 events) and a fresh-install screenshot shows visibly larger
text than before this and the previous commit, with the top row's own
sizing still worth a closer look on a real device -- see RUST.md's P0
box for what remains unverified there.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-05 23:57:24 -04:00
irisandClaude Fable 5.1 c7682297fa docs/bench: Compose bench v2 report from Iris's phone, verbatim
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-05 23:49:02 -04:00
irisandClaude Fable 5.1 27511302f2 iris/android-app: Diagnostics control, top-bar status-bar padding, cargo fmt
Adds a third "Diagnostics" button to the bench screen's top row, filling
the existing benchmark-report TextEdit (so the existing "Copy report"
button and clipboard path work on it unchanged) with adapter identity,
font resolution, atlas view count, wgpu errors seen so far and the frame
report -- RUST.md's P0 box, "a named Diagnostics control ... copy this
and send it to Iris." Logs the same font-resolution summary once at
startup too.

Wires BenchClient::on_insets_changed (the new AndroidAppState hook) to
rebuild the top button row with Padding::top(insets.top), through a
WidgetPtr slot (top_bar) so it can be swapped once the status-bar inset
is known -- fixes RUST.md's P0 box, "the status-bar inset is not
applied," where the two top buttons sat directly under the status bar
because nothing in this file read insets().top at all.

cargo fmt --all across the touched files.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-05 23:44:43 -04:00
irisandClaude Fable 5.1 184a6c5b33 IRIS_TODO.md: a density-independent length unit, asked for by Iris
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-05 23:43:16 -04:00
irisandClaude Fable 5.1 5b2ca039f1 docs/RUST.md: bench v2 spec and the emulator smoke run
Iris's ask (2026-09-06): the fling should travel much faster for
stress-testing, plus typing and keyboard phases. Written once into the
P0 box so the iris agent implements the identical four-phase spec --
constants, ordering and report shape -- rather than a second one that
looks the same but isn't.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-05 23:41:03 -04:00
irisandClaude Fable 5.1 a8d24553d5 app: bench v2 -- a real fling, typing and keyboard phases
Iris's ask after using the Compose bench build on her phone: the old
scroll phase used animateScrollBy, which can only ever cover the fixed
distance/time it's given, so it never flings the way a real fast swipe
does. BenchRun.run now has four phases: fling (8 flings out + 8 back
through the list's own FlingBehavior at 12,000px/s), stream (unchanged),
type (600 fixed characters into the real composer TextFieldValue, then
deleted, to exercise wrapping and the transcript being pushed upward),
and keyboard (five show/hide cycles via WindowInsetsControllerCompat,
each confirmed by isImeVisible rather than assumed).

FrameStats.markPhase/phaseLines slice the same FrameMetrics recording
by phase rather than running a second recorder; debugReport gains a
phaseFrames section ahead of the existing whole-run frames/accounting/
work sections, which are otherwise unchanged.

Also fixes a pre-existing, unrelated break in MainActivity.kt's
benchSessionSummary() -- missing several SessionSummary constructor
arguments from an earlier change -- since it blocked compileBenchKotlin
outright.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-05 23:40:59 -04:00
irisandClaude Fable 5.1 3b80a88f3b iris/android: content_scale (density), per-frame diagnostics, insets hook
Threads DisplayMetrics.density (read once in new_peer, via the Context
android-view already hands the JNI entry point) through AndroidUiState
as content_scale, and divides by it everywhere a raw device-pixel number
used to reach layout unscaled: AndroidRenderer::size()/resize()/new() now
report logical (physical / density) dimensions to UiRenderNode and to
UiRenderState's own root-layout size, and on_touch_event divides the
incoming MotionEvent coordinates the same way, so touch and layout agree
on units again. This is the fix for RUST.md's P0 box, "text is far too
small" -- a font_size: 16.0 was 16 raw device pixels on a ~3x-density
phone, identical to the desktop fix in the previous commit.

Installs Device::on_uncaptured_error on the Android device (wgpu's
default handler is an unconditional panic outside UiRenderNode::new's
own error scopes) into a new iris_core::WgpuErrorLog, and adds
AndroidRenderer::diagnostics_report() combining adapter identity, font
resolution, atlas view count and the error log into one string for a
future Diagnostics screen. render() now logs a one-line diagnostic
(masks/moves resized, atlas pages grown, image bind-group creates, wgpu
error count) for the first 10 frames after each surface_changed -- the
window RUST.md's P0 box says the glyph-wipe-on-first-touch happens in.

Adds AndroidAppState::on_insets_changed(rsc, LogicalInsets), called from
render() exactly when AndroidUiState::insets() changes (once at startup
for the status bar, again on rotation/IME) -- nothing previously read
insets().top at all, which is why RUST.md's P0 box found the bench
screen's top buttons sitting under the status bar.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-05 23:40:30 -04:00
irisandClaude Fable 5.1 d8e6bc6e9b iris: bundle Noto Sans for text rendering, apply density scale on both backends
Bundles Noto Sans/Noto Sans Mono (regular/bold/italic/bold-italic, OFL
licensed) into iris-core and registers them ahead of the platform's own
fonts in the SansSerif/Monospace generic-family fallback lists, so text
no longer depends on the platform's font enumeration succeeding or
resolving weight/style correctly. Iris's phone report showed bold spans
rendering as blank gaps of the correct advance width -- the glyph simply
wasn't rasterised -- while the emulator's system fonts happened to
resolve every style; a bundled static-per-style family removes that
platform-dependent step entirely. TextData::font_diagnostics() reports
what was found/resolved, for the startup log and the Diagnostics page.

Also applies a content/device-pixel scale that neither backend had
before: UiRenderNode::new/resize now take the window size explicitly
(logical units) rather than deriving it from the surface's physical
config, so a 16.0 font size is 16 logical units rather than 16 raw
device pixels. Wired on desktop via window.scale_factor() (input events,
window_size, and the render node's own seed); the Android side (density
via DisplayMetrics, touch coordinates, layout root size) is the next
commit.

Also adds WgpuErrorLog and a per-frame atlas-grow counter
(GpuTextures::take_pages_grown), both plumbing for the Android
diagnostics page in the next commit.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-05 23:36:29 -04:00
irisandClaude Fable 5.1 b887a96765 docs/bench: iris's first phone report, before the phone fixes
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-05 23:26:09 -04:00
irisandClaude Fable 5.1 2aaa3733c3 docs/bench: the Compose P0 report from Iris's phone, verbatim
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-05 23:20:05 -04:00
irisandClaude Fable 5.1 46246ea511 iris: turn the phone bind-group-layout crash into a diagnostic, drop force-gles from phone builds
UiRenderNode::new used to let a wgpu validation error reach the default
uncaptured-error handler and panic, which is what aborted the P0 bench APK
on Iris's phone in AndroidRenderer::new with only "wgpu error: Validation
Error" surviving into the truncated crash report. It now wraps creation in
wgpu error scopes and returns Result<Self, String>; the Android backend
turns a failure into the adapter's identity, the limits/downlevel flags a
layout validates against, and wgpu's own error chain, logged as one logcat
line and shown on screen (IrisView.showRendererError) instead of crashing.

Auditing every bind-group-layout entry against wgpu-core's own validation
source names the likely cause: masks_layout's move_offsets storage buffer
is visible to the vertex stage, which Vulkan grants unconditionally but
GLES gates on the driver's own vertex-stage SSBO support -- and the
delivered APK was built with force-gles, a flag meant only to force the
*emulator* onto GLES for one frame-time measurement, that build-apk.sh's
default feature list applied to every arm64 build regardless of target.
Its default no longer includes force-gles.

Testing the diagnostic (by inducing an artificial validation error) also
found and fixed a real reentrancy bug: calling Activity.setContentView
synchronously from inside a ViewPeer callback re-enters the same peer's
RefCell borrow through onFocusChanged, aborting with "RefCell already
borrowed". Deferred through the same push_dynamic_deferred_callback
mechanism raise_if_enabled already uses.

Full audit, verification, and the named hypothesis are in RUST.md's P0
box ("iris bench crash on the phone, 2026-09-06"); the API change is in
IRIS.md. Nobody on this session has the phone, so this is unconfirmed
against real hardware -- the point of (1) is that the next run says so
either way.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-05 23:05:07 -04:00
irisandClaude Fable 5.1 a27fbdb029 docs: close I5's three blocked verifications (24/24-swipe, backend isolation, cold-boot bench)
Ran three clean iris-scroll.sh passes on a cold -gpu host boot (all
24/24 swipes confirmed scrolling via clustered render() timestamps, not
inferred from frame count) and retook the host-GPU table's iris row as
a best-of-three. EMU_GPU=software + force-gles still cannot produce a
GLES number on this hardware -- after the earlier compute-limit crash
was fixed, device creation now aborts on max_storage_buffer_binding_size
instead (SwiftShader ES 3.0 has no SSBOs, and shader.wgsl reads four
var<storage> buffers unconditionally), so the SwiftShader-Vulkan-vs-GLES
question is closed as structurally unanswerable rather than answered.
A fresh cold-boot run-bench.sh reading for P0's bench build is in line
with the earlier warm-AVD readings, closing that box's own caveat too.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-05 22:39:27 -04:00
irisandClaude Fable 5.1 c07d544aeb event-model, client-core, transcript-ui: carry main's LimitReached event
The merge that brought main into rustify added Event::LimitReached to the
server's drivers, but on this branch the enum lives in event-model, which
the merge left without it, so ai-server (and ui-sandbox.sh) did not build.
Definition copied from main's driver.rs; the fold mirrors TranscriptItems.kt's
LimitNote; the iris row shows the epoch until P1 brings a time formatter.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-05 22:18:38 -04:00
irisandClaude Fable 5.1 46d3a6fd41 docs: record the streaming-rebuild fix, its numbers, and the new scripts
RUST.md's P0 box gets the fix, the before/after streaming-phase numbers
(with their caveats), the build-apk.sh/run-bench.sh scripts, and what the
dropout-fix pass's three remaining verifications are blocked on (the
sandbox ai-server currently fails to build, unrelated to this change).
IRIS.md gets the List::replace_back/clear and TranscriptScreen::apply
API entries. AGENTS.md's rigs section gets one sentence on each script.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-05 22:14:35 -04:00
irisandClaude Fable 5.1 5655fa8093 iris-android-app: build-apk.sh and run-bench.sh
Wraps the cargo-ndk/Gradle/keystore/apksigner build and the
install/tap-by-label/read-report cycle that P0's work had been retyping
by hand, so it stops costing time and mistakes.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-05 22:14:35 -04:00
irisandClaude Fable 5.1 b3b1d47dd6 iris: streaming a transcript event no longer rebuilds the whole screen
Every client (bench_client, transcript_client, desktop-app) refolded and
rebuilt the ~3,200-row widget tree from scratch per SSE event, which is
the streaming-phase cost the P0 benchmark gate would otherwise measure
against a Compose app that updates one row. iris::widget::List gains
replace_back (swap the last row's widget in place, keeping its slot so a
pinned list stays pinned) and clear (the full-rebuild fallback);
transcript_ui::TranscriptScreen::apply diffs the folded row lists and
picks the cheapest update -- unchanged, append, replace-the-last-row, or
(rare regroup) a full rebuild, counted. TextEditCtx::set_with_spans lets a
row's text and span list land together on a streamed update.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-05 22:14:27 -04:00
iris 50fe4828a2 Merge branch 'worktree-agent-a27094a7db775552a' into tmp-merge 2026-09-05 21:37:12 -04:00
iris 800da46188 Merge remote-tracking branch 'origin/rustify' into worktree-agent-a27094a7db775552a
# Conflicts:
#	docs/IRIS.md
2026-09-05 21:37:04 -04:00
irisandClaude Fable 5.1 00767eed4d docs: P0's iris half done -- bench feature, emulator smoke run, APK
RUST.md's P0 box gets the iris-half account: the fixture, the scroll/stream
mechanism, the report fields, build commands (all clean), packaging (no
cargo xtask apk yet, so a new Gradle release build type on top of cargo
ndk), and the emulator smoke run's report next to Compose's own. Used a
second, differently-named AVD rather than contend with the session already
on this checkout's own emulator.

DECISIONS.md's P0 entry gets a matching summary bullet. IRIS.md records
AndroidAppState::platform_ready. IRIS_TODO.md notes the one gap found:
no read-only selectable text primitive, so the bench report's TextEdit
picks up a keyboard on tap it has nothing to type into.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-05 21:35:51 -04:00
irisandClaude Fable 5.1 683db4908a iris-android-app: a bench feature, P0's iris half
A third AndroidAppState (BenchClient) on top of transcript-screen: embeds
app/bench-fixture/assets/transcript.jsonl with include_str! (no server, no
enrollment), folds the first 3,200 lines through client_core's real
fold_page as the opening backlog, and holds the rest back as a streaming
tail. "Run benchmark" resets FrameReport, animates the same 24-swipe/
6-cycle scroll BenchRun.kt drives (List::scroll in ~60Hz steps, since iris
has no built-in tween), then replays the tail at 20/s through fold_event --
the same fold path a live SSE reply takes -- and shows a report in a
selectable TextEdit. "Copy report" puts it on the clipboard.

The report adds process CPU time (libc::getrusage), peak RSS (/proc/self/
status's VmHWM) and battery current (BatteryManager.getIntProperty via
direct JNI, bench_jni.rs's PlatformHandle) to FrameStats's existing
frames/janky%/percentiles/CPU-GPU-split line -- "unavailable" rather than a
fabricated number wherever the platform can't answer.

build.rs now exits early under the bench feature before requiring a live
server's host/port/token/CA: BenchClient never calls build_transport().
app/build.gradle gains a signed `release` build type (previously only
debug) so the cdylib cargo ndk builds can be packaged for a phone, the same
key app/build-apk.sh generates.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-05 21:35:44 -04:00
irisandClaude Fable 5.1 8d23a20792 iris: AndroidAppState::platform_ready, a JavaVM+View handle for later JNI calls
Default no-op lifecycle hook, called once from new_peer right after new.
P0's bench build needs to call BatteryManager/ClipboardManager through the
view's own Context from a background thread as well as the UI thread, and
neither a JavaVM nor a GlobalRef to the view was reachable from
AndroidAppState::new before this. Existing implementors (Client,
TranscriptClient) are unaffected.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-05 21:35:33 -04:00
irisandClaude Fable 5.1 d01c105037 iris: stop requesting compute-shader limits nothing uses
adapter.request_device asked for Limits::default(), which requests
desktop-tier compute-shader limits unconditionally even though nothing in
iris/iris-core creates a ComputePipeline or writes a @compute stage. That
crashed device creation outright on a downlevel GL adapter reporting
OpenGL ES 3.0 (no compute at all) -- the Android emulator's
EMU_GPU=software/force-gles path, and any real GLES-3.0-only device.

New iris_core::device_limits(), shared by both platform backends, zeros
exactly the six max_compute_* fields rather than switching to a downlevel
Limits preset -- downlevel_webgl2_defaults() also zeros
max_storage_buffers_per_shader_stage, which shader.wgsl's vertex stage
needs. rigs/gpu-probe's own mirrored limits were updated to match.

Not verified against the actual SwiftShader-ES-3.0 crash on-device this
pass: the cold boot needed would have force-restarted this checkout's
emulator while another session had its own app running on it.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-05 21:18:30 -04:00
iris 88631f5e8b Merge remote-tracking branch 'origin/rustify' into worktree-agent-a27094a7db775552a
# Conflicts:
#	AGENTS.md
#	server/src/session/driver.rs
2026-09-05 21:10:19 -04:00
irisandClaude Fable 5.1 e6924298bc iris: fix the intermittent touch-scroll dropout (missed ACTION_DOWN hit-test)
Root-caused via temporary logcat tracing (touch events, DragArbiter state,
Selection::drag dispatch), reproduced against a real sandbox session: a
gesture's ACTION_DOWN can land on a row's own padding/gap or its header,
which CursorSense has no sensor over, so the widget that ends up handling
the gesture only ever sees Pressing frames and DragArbiter never gets
press_start -- leaving it stuck in Idle (answers Undecided forever) for the
rest of that gesture. Not the previously-suspected coalesced first
ACTION_MOVE, which is now ruled out.

DragArbiter::is_idle() lets Selection::drag notice a Pressing frame with
no matching press_start and recover the press there instead. Four new unit
tests, one of which fails on the pre-fix code.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-05 21:05:03 -04:00
irisandClaude Fable 5.1 68b48cfd14 docs: record P0's Compose half (bench build, fixture, smoke run)
RUST.md's P0 box gets the emulator smoke run's report and what's done vs.
left; DECISIONS.md gets a dated summary entry; AGENTS.md's "Checking your
work" and "The rigs" get one paragraph each on the bench build type and
app/bench-fixture/.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-05 21:04:07 -04:00
irisandClaude Fable 5.1 e6c884a0cd app: fixture-mode session screen and a "Run benchmark" control
BenchFixture.kt/BenchNetwork.kt fake the backend for the bench build: a
URLStreamHandlerFactory installed only under BuildConfig.FIXTURE_MODE
answers TranscriptSource/EventStream's requests from an in-memory copy of
the bundled fixture instead of opening a socket, so the fold, the paging
and uniqueItems under test are the screen's real ones rather than a
shortcut built for this. MainActivity opens straight onto that session
when FIXTURE_MODE is set, with no enrollment and no permission prompts.

BenchRun.kt drives the same scroll loop and streaming phase
transcript-bench.sh/stream-bench.sh drive over ui-trace, but in-process
(24 swipes through the real LazyListState, then 400 fixture events
appended at 20/s through the real live-fold path), and adds process CPU
time, peak RSS and battery current to the render report -- "unavailable"
rather than a fabricated number where the device can't answer.

"Run benchmark" sits beside the existing "Copy" in session settings,
found by that exact label the way every other control here is
(SessionSettingsDialog's onRunBenchmark, null on every build but bench).
debugReport gained an optional `extra` section for this; empty and
invisible on every other build's report.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-05 21:03:57 -04:00
irisandClaude Fable 5.1 a6cb9a9082 app: a bench build type for P0's benchmark gate
Own application id (.bench suffix) and label ("AI Sessions bench" via a
build-type resValue over the new @string/app_name), release
optimisations, signed with the same key build-apk.sh already generates,
FIXTURE_MODE=true wired through BuildConfig. Its asset source set points
straight at app/bench-fixture/assets rather than a copy under androidApp,
so there is one file to keep in sync with the generator, not two.

build-apk.sh bench builds it; the CA-pinning step is untouched and still
requires a real ca.pem to exist, even though this build never connects --
simplest to let it pin whatever is there rather than special-casing it.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-05 21:03:44 -04:00
irisandClaude Fable 5.1 0be6a571c4 app/bench-fixture: the synthetic transcript P0's benchmark opens in both apps
Deterministic (seeded), in the app's own event model rather than a real
transcript: 3,601 events split into a 3,200-event opening backlog and a
400-event tail both bench harnesses replay as the streaming phase, with
headings, inline markdown, fenced code in six languages, a table, tool
calls with kilobyte-scale input/output, and two embedded PNGs.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-05 21:03:34 -04:00
irisandClaude Fable 5.1 bfe93c4188 RUST.md, DECISIONS.md: P0, the phone benchmark gate Iris asked for before P1
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-05 20:31:20 -04:00
irisandClaude Fable 5.1 5b7dc0e4e2 RUST.md, DECISIONS.md, IRIS_TODO.md: the port plan, P1-P7, after iris-over-Masonry
Adds "The port, in order (decided 2026-09-05)" to RUST.md: seven ordered
steps building the app on iris now that the framework is decided, each
naming the Kotlin files it replaces, the client-core pieces it needs
(and which are not yet covered and must be ported first), the missing
iris widgets it needs (recorded in IRIS_TODO.md's new "Build (for the
port)" section), and a pass condition a later agent can run. Ordered by
risk to the daily-use path: session screen parity, then the shell merge
and a real phone install, then root tabs, the file explorer,
settings/enrolment, desktop parity, and the cutover itself.

Crate-shape decision recorded in DECISIONS.md: one UI crate, app-ui,
grown out of transcript-ui rather than started beside it, with
desktop-app/android-app as thin entry points over it and platform-only
code staying in the E3/E5 Java shell.

Updates RUST.md's "Where things stand" and "For the next agent" to point
at P1 rather than the now-closed framework decision.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-05 20:30:02 -04:00
irisandClaude Fable 5.1 621f08d725 DECISIONS.md, RUST.md: Iris decided iris over Masonry, 2026-09-05
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-05 20:25:43 -04:00
irisandClaude Fable 5.1 e49d0e606f RUST.md, DECISIONS.md, IRIS.md: iris's host-GPU frame time, 2026-09-05
Takes the -gpu host pair the earlier software-mode comparison flagged as
missing. Under real GPU rendering (--features force-gles: the default
Vulkan backend has no adapter at all under plain host-GPU boot, confirmed
by the exact wgpu error), iris's median frame (15.0ms) is faster than
Compose's (20.0ms) on the same session content -- the opposite shape from
the software-mode table. The new redraw-to-submit/submit-to-present split
shows iris's own CPU work is a median 0.2ms per frame; almost the whole
frame is time handing off to the driver, consistent with (but not proof
of) the software-mode gap being mostly SwiftShader's CPU rasterisation
cost rather than iris-specific slowness.

A same-mode software force-gles run, meant to isolate the backend, hit a
third distinct crash instead (SwiftShader's GL path reports itself as
OpenGL ES 3.0, which has no compute shaders, and iris's device request
assumes them unconditionally) -- real scope to fix, not done here, so the
software-mode question stays open. A real intermittent touch-scroll
dropout was also reproduced (six consecutive swipes produced zero
redraws while taps kept working; an identical retry then succeeded) and
is not explained. The idle-redraw and virtualised-culling findings from
the software-mode pass were confirmed to hold under real GPU rendering
too.

DECISIONS.md's DEFERRED item carries the updated table; the iris-vs-
Masonry choice itself is still Iris's to make. IRIS.md records the
FrameReport::record_split/FrameStats::cpu_p50/gpu_wait_p50 API from the
prior commit (e2a1fad), which this pass's measurement used.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-05 20:07:20 -04:00
irisandClaude Fable 5.1 e2a1fadbec iris: FrameReport CPU/GPU split, force-gles backend switch, iris-scroll.sh rig
Splits each frame sample at queue.submit into redraw-to-submit (iris's own
CPU work) and submit-to-after-present (driver/GPU wait), so RUST.md's I5
"where does iris's frame time go" question can be answered with a number
per half instead of a single total. Adds a force-gles Cargo feature that
switches the Android wgpu::Instance from Backends::PRIMARY to Backends::GL
at compile time (no runtime env-var path exists into an already-launched
Android process on this machine), for isolating SwiftShader-Vulkan vs.
GLES/virgl as the software-mode gap's cause. app/iris-scroll.sh extracts
transcript-bench.sh's exact 24-swipe/6-cycle gesture loop for iris's own
demo app, which transcript-bench.sh cannot drive directly since it opens a
session through the Compose app's own UI.

Verification (this pass, on a disk-pressure-limited host running low on
space): cargo fmt --all clean, no diff. cargo clippy --workspace
--all-targets: no warnings from this diff (pre-existing future-incompat
notices from wgpu/winit/naga only). cargo test --workspace and cargo ndk
for iris-android-app --features transcript-screen were verified clean by
the previous pass on this identical diff (fmt/clippy/test/ndk all clean,
per that pass's own report); not re-run here because the host's disk was
93% full and a concurrent ai-server rebuild (stable toolchain moved to
1.98.1, rebuilding aws-lc-sys from scratch) had driven I/O pressure to
~60%, so a repeat cargo test --workspace sat 50+ minutes doing no useful
work and was stopped rather than left to make the disk situation worse.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-05 18:47:01 -04:00
irisandClaude Fable 5.1 0e4629361b docs: I5's clean scroll comparison between Compose and iris, one session
Same sandbox session content, same emulator, EMU_GPU=software: Compose
(debug, in-app report) 1102 frames/99.0% late/p50 33.8ms/p99 79.5ms vs
iris (release -- debug SIGSEGVs on this emulator) FrameReport 299
frames/94.65% janky/p50 79.1ms/p99 117.8ms (repeat: 233/94.42%/p50
109.3ms). Ticks I5 [x]; states plainly what's not comparable (build
profile forced asymmetric, three different jank definitions, both are
software-rasterised emulator numbers). The two "zero frames" attempts
that preceded the clean runs traced to this session's own script bug
(a cd into /tmp changed which emulator ui-trace targeted), not a
reproduction of the previously-suspected touch-delivery dropout; a
sampler ran the whole session and saw load rise during the gesture
without correlating to any failure. DECISIONS.md's DEFERRED item gets
the same table so Iris can decide iris-vs-Masonry from it -- that
choice is left to her.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-05 15:21:45 -04:00
irisandClaude Fable 5.1 1e7b1cddb7 RUST.md, IRIS.md, IRIS_TODO.md, DECISIONS.md: record I5's frame report and holddrag results
FrameReport gave a real, measured on-device number (frames=34,
janky%=61.76, p50=26.5ms p90=48.0ms p99=98.1ms worst=98.1ms) and
long-press-then-drag-to-select is now confirmed on-device (logcat plus a
screenshot of the highlighted selection). Neither closes I5's box to [x]
yet: the frame number is real but not the clean single 24-swipe loop
comparable to Compose's, because gestures against this checkout's
EMU_GPU=software emulator intermittently delivered zero touch input this
session -- a new, separately named finding (candidate cause: the
emulator's own software rasterisation measured at ~78% of a CPU core
continuously), not yet root-caused. DECISIONS.md's DEFERRED item is
updated with these numbers rather than a decision made here.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-05 15:00:29 -04:00
irisandClaude Fable 5.1 470f8e5019 transcript-ui: log selection begin/extend, for on-device verification
Selection has no accessibility label of its own yet, so a logcat line at
begin/extend is the smallest way to confirm a real long-press-then-drag
reached DragArbiter/Selection on-device. Driven with the new ui-trace
holddrag action against iris-android-app's transcript screen: produced
"iris selection: begin at row ..." then a sequence of "... extend to row
..." lines, and a screenshot right after shows the expected highlighted
selection spanning multiple rows.

New `log = "0.4.28"` dependency (matching iris-android-app's own pin) --
transcript-ui had no logging facility before this.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-05 15:00:18 -04:00
irisandClaude Fable 5.1 cf10b17c5b End a subagent on its own end_turn, not the parent's tool_result, and give the expander a touch-sized row
The Agent tool runs subagents in the background, so the parent's result
arrives at launch while the subagent works on for minutes; finishing on it
read a running agent as finished with a transcript cut off at launch. A
subagent now ends on its own message_delta end_turn, and a later line for a
finished one reopens it, since a background agent can be messaged again.

The card's expander row was only the chevron's height, so a tap for it
landed on the first subcard; it is the platform's 48dp minimum now.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-05 14:44:00 -04:00
irisandClaude Fable 5.1 7ae53ad797 iris: FrameReport, a per-frame wall-time report of iris's own render path
dumpsys gfxinfo cannot see a SurfaceView's own GPU-drawn frames at all
(RUST.md's I5 box), so iris needs its own equivalent of Compose's
render-report button before item 3 of the recommendation can be decided
by a number. FrameReport (iris/core/src/render/frame_report.rs) records
each frame's wall time -- from render()'s redraw start to after
queue.submit + present() -- into a fixed 4096-entry ring, and reports
total frames, janky % (>16.7ms, gfxinfo's own budget), P50/P90/P99 and
the worst. Wired into AndroidUiState and android/view.rs's render(), and
exposed as two named controls ("Frame report", "Reset frame report") on
iris-android-app's transcript screen, logged under the crate's fixed tag
so a script can grep "iris frame report" the way transcript-bench.sh
greps "ai-app render report".

6 new unit tests for the ring/percentile math. cargo fmt/clippy/test
--workspace clean; cargo ndk (iris, transcript-ui, and
iris-android-app --features transcript-screen) all clean.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-05 14:23:51 -04:00
irisandClaude Fable 5.1 d17040b601 RUST.md, IRIS.md, IRIS_TODO.md, DECISIONS.md: record I5's Android integration and measurements
I5's transcript screen now runs on-device against a real ai-server on
iris-android-app's new transcript-screen feature (extends I2's shell
rather than a third one), with real scrolling, real touch-drag panning
and tap-by-name accessibility all confirmed by screenshot/log evidence.
I4's own emulator-side check (tap-by-name on the tabs demo) closed the
same session, so its box ticks [x] now.

Still [~], not [x]: the render-time number RUST.md's recommendation
wants for iris couldn't be produced this pass, for a precise and
recorded reason rather than a vague one -- dumpsys gfxinfo cannot see a
SurfaceView's own GPU-drawn frames at all (0 frames reported across a
gesture loop that visibly scrolled), and a SurfaceFlinger --latency
fallback gave no per-frame history either on this Android version. The
Compose side of the same loop did produce a real number under identical
conditions (8.96% janky, 99th percentile 150ms), so this is now a
one-sided number rather than a missing one on both sides.

Also found and recorded: the AVD's saved snapshot carries a GPU config
across restarts, so switching between the documented Vulkan boot
recipes needs a cold boot (clearing snapshots/) that the emu wrapper
does not force -- cost three different-looking crashes before the
pattern was the snapshot, not the code.

DECISIONS.md's DEFERRED item is updated with the numbers Iris needs to
weigh the iris-vs-Masonry call; the call itself stays hers.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-05 14:09:04 -04:00
irisandClaude Fable 5.1 bf5087a598 iris/android: fix background-thread redraw crash, add missing INTERNET permission
Two real bugs found bringing up I5's Android transcript client, neither
specific to that screen -- any future caller of Tasks::redraw_handle()
from a background thread would hit the first one.

AndroidRedrawHandle::request_redraw called View::post_frame_callback from
a tokio worker thread; its Java side calls Choreographer.getInstance(),
which throws IllegalStateException unless the *calling* thread already
has a Looper, and a JNI-attached background thread has none. That crashed
the whole process (SIGABRT, unwrap() on a JavaException) the first time a
background fetch asked for a second frame. Fixed by routing through
View::post_delayed(0) instead, Android's own thread-safe way to queue
work onto a View's UI thread, landing on a new
IrisViewPeer::delayed_callback override that drains tasks and renders --
same body as do_frame, now running safely on the UI thread.

iris-android-app's manifest never needed INTERNET before (the tabs demo
makes no network call); its absence read as EPERM ("Operation not
permitted") from UreqTransport::new's connect, not the
ECONNREFUSED/ENETUNREACH a dead server would give.

Full account in RUST.md's I5 box and IRIS.md's Tasks::redraw_handle entry.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-05 14:08:40 -04:00
irisandClaude Fable 5.1 9fa09b0af1 Show a session's subagents as subcards, each with a read-only transcript
A subagent is a second transcript owned by a session, in the same event
model, with no process and no controls. The claude translator routes lines
carrying parent_tool_use_id to a per-subagent translator and transcript
under <session>/subagents/<tool_use_id>; three routes expose the list, a
transcript page and the SSE stream. Echo grows /subagent [n] as the rig.

On the phone a card with subagents ends in a chevron expander, collapsed by
default, opening to outlined subcards styled like dev-updater's components;
a subcard opens SessionScreen in read-only form, addressed through
TranscriptAddress so paging, cache and stream are shared.

Design in SUBAGENTS.md; choices awaiting review in DECISIONS.md.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-05 13:41:15 -04:00
irisandClaude Fable 5.1 aa3d11471f iris-android-app: transcript-screen feature -- I5's Android integration
Extends the existing tabs demo shell (I2/E5's Gradle project, JNI
registration) with a second, mutually-exclusive AndroidAppState rather than
building a third shell -- it already has the working IrisView/MainActivity
Java and the register_view_class wiring, and the only thing a transcript
screen needs on top is a different Client type (the same axis
tabs_ui::build vs. transcript_ui::build already varies along on winit).

`--no-default-features --features transcript-screen` builds
transcript_client::TranscriptClient instead of the plain tabs Client:
fetches the sandbox's session list, opens the first one, and follows it
live, reusing desktop-app's app.rs shape (fold_event/group_tool_runs/
fold_page/raw_seq, a generation counter) almost verbatim. The one real
difference is the redraw path -- android-view has no winit::EventLoopProxy,
so Tasks gained redraw_handle() (iris/src/task.rs) to let a caller request
a frame after each TaskCtx::update from inside a still-running task, not
just once when the whole future completes.

Deliberate simplification, not a template: there is no session list or
enrollment UI here. build.rs bakes the sandbox's host/port/token plus the
pinned CA in at build time from AI_APP_TRANSCRIPT_HOST/_PORT/_TOKEN and
AI_APP_CA, the same trust-boundary reasoning as the Compose app's
GeneratePinnedCert Gradle task, extended to also bake the enrollment since
building a real one (Keystore-sealed storage, a QR/link scanner) is E3's
scope, not this box's. Recorded in RUST.md's I5 box.

tabs-ui and the transcript-screen deps are now both optional, gated by
mutually exclusive tabs-screen (default) / transcript-screen features --
building one screen with the other's default deps still active tripped
Cargo's unused_dependencies lint.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-05 13:26:10 -04:00
irisandClaude Fable 5.1 78aff64844 client-core: hoist transcript_fold::{fold_page,raw_seq} out of desktop-app
Both desktop-app's app.rs and the new Android transcript client (RUST.md's
I5) need the same page-fold and live-stream resume-cursor logic; per
CODE_RULES's "write the logic once" it now lives in client-core alongside
fold_event/group_tool_runs instead of being duplicated. desktop-app calls
the shared functions; its own copies and their tests moved with them.

Also fixes a clippy::collapsible_if in config.rs's percent_decode, found
while re-running clippy after this change (let-chains are stable now).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-05 13:25:46 -04:00
irisandClaude Fable 5.1 9a33cb5384 docs/: move the design and working documents out of the repo root (CLAUDE.md and AGENTS.md stay, harnesses read them there)
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-05 13:03:42 -04:00
iris 45ced405f3 Merge branch 'worktree-agent-afe80868604fef704' into tmp-merge 2026-09-05 12:59:54 -04:00
iris eff5c8b0c0 Let the machine's own CLI refresh an expired token, and retry once
A 401 from the usage endpoint means the stored access token has expired.
Refreshing it here is not an option: Anthropic's OAuth rotates the refresh
token, so a second refresher invalidates the CLI's copy and forces a
re-login on a machine that usually has a live session on it. So run the CLI
there instead and re-read what it wrote.

`doctor` rather than `auth status`: probed against 2.1.258 with an invalid
token, `auth status` answers loggedIn:true from the file alone and never
reaches the network. The same probe showed a failed refresh blanks both
tokens, which is why this stays on the 401 path.

Also gives ProviderConfig one program() so the CLI's default path is not
written down twice.
2026-09-05 12:07:34 -04:00
iris 7b63330aaa Say when a usage 401 is an expired login, not an unreachable endpoint
A 401 is the endpoint answering and refusing the stored OAuth token, which
Claude Code refreshes as it runs -- so a machine whose CLI has been idle
hands us a stale one. Reporting it as "usage endpoint unreachable" pointed
at the network instead of at the one thing that fixes it.
2026-09-05 11:58:17 -04:00
irisandClaude Opus 5 6bdec6e785 Let a session resume itself when its usage limit lifts
Off by default and per session: it spends quota the moment quota exists,
with nobody watching, which is not a thing a default may decide. Switched
on from the session settings dialog, with the message it sends editable
("continue" unless something else is typed).

Running out of quota becomes a state rather than an error. The Claude
driver recognises its dialect's sentence -- `Claude AI usage limit
reached|1788546972` -- and reports `LimitReached` with the reset time it
gave; nothing above a driver matches on a string. The transcript draws it
as a divider, like a clear or a compaction.

The schedule is a plan to *ask*, never a plan to send. Both reset times
available are untrustworthy in the direction that matters -- the dialect's
is written when the turn fails, the endpoint's moves when the window does
-- so the wait ends in a question to the usage meter, and only `ok` with
no window at 100% sends anything. A window still spent reschedules to its
own reset time, which is what makes a limit that lifts late wait longer
and one that lifts early resume sooner. A meter that cannot be asked is a
longer wait too, never a send. A day after the limit was hit the wait
gives up and says so in the transcript, so a machine that can never be
asked is not retried for ever.

The schedule is persisted on the session: a five-hour window outlasts a
backend restart, and a wait forgotten across one never comes back.

Driven end to end with echo, never a real account: `/limit [minutes]`
reports the same event a real driver does and `/usage` sets what the meter
answers, deliberately separate so the two can disagree. The wait moved
from the dialect's two minutes to the meter's seven when the meter changed
its mind, and the message went out on the first check after the meter came
back under the limit.

Also makes the settings dialog scrollable, which these two controls made
necessary: at a 1.5x system font it clipped the last of them with nothing
on screen to say so.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-05 05:21:43 -04:00
irisandClaude Opus 5 4821a02bd3 Default thinking level for new sessions, and move the rigs out of AGENTS.md
`Config::default_effort` is what a session starts at when nothing chose one,
applied in `spawn_session` rather than filled in by the spawn screen so it
holds for an import and a bare API call too. It is set by the spawn screen's
own picker, whose label says so: one control, where new sessions are made,
rather than a settings page for a single value. Not on a provider, because
providers are discovered and the next rediscovery would erase it; not on the
phone, because a second device would then spawn at a level nobody there
chose. `GET`/`POST /defaults` carry it as a struct, so the permission mode --
still hardcoded to `auto` on the spawn screen -- can move there later without
a second route.

Only drivers that read a level are given one: an echo session was storing a
`--effort` it never passes to anything, which is a config file answering a
question about itself wrongly.

Separately, `AGENTS.md` is 35 KB sent with every request in this repo, and 12
KB of it was rigs and reference measurements that only matter once you are
running one. Those are the `ai-app-rigs` skill now -- the same text, still the
only copy, read when the work touches it. 35,198 -> 20,813 chars.

Verified on the emulator against the sandbox: the spawn screen pre-fills from
the server, picking `low` spawned a session at `low` and left `/defaults` set
to it, and an echo session spawned afterwards took no level at all.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-04 21:42:05 -04:00
irisandClaude Opus 5 1ff662c7c3 Let a session choose how hard it thinks
Output is about an eighth of what a session costs and thinking is nearly
all of it -- prose is ~1.5% of output tokens, measured over 27,015 requests
of this account's own transcripts -- so the level is the largest saving
available short of shortening the conversation itself.

Shaped like the working directory rather than like the model: the CLI's
only two setting control requests are `set_model` and `set_permission_mode`
(checked against the 2.1.258 binary), so `--effort` is read when the process
launches and cannot be asked of a running one. `set_session_effort` records
the level and stops the process; the next message or Start launches one that
has it. That is also why the picker is in the session settings dialog beside
Move, and not on the bar beside the model and the mode, which take effect
mid-turn.

`None` is a level in its own right -- the CLI's own default -- so the picker
can return to it, and a blank is normalized to it at the boundary rather
than stored as a level the CLI would reject.

Offered only where it means something: `DriverKind::takes_effort` reports
the capability and the phone leaves the row out entirely, rather than the
session-type branch this app does not have anywhere else. A llama session
would otherwise get a control whose only effect is stopping its process.

Verified on the emulator against the sandbox's fake CLI: the picker sets it,
the server reports it, and an echo session's dialog is unchanged.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-04 21:23:58 -04:00
irisandClaude Opus 5 e4f0935f98 Keep the second auth test under a subscriber, so the tripwire is not flaky
`gates_every_route_and_never_logs_the_token` failed about one full-suite
run in ten, on the assertion that a rejection *was* logged. Its sibling
ends with an unauthenticated request of its own, made with no subscriber
on that thread -- and tracing caches a callsite's interest process-wide
the first time it is reached, so whichever test got there first decided
whether the warning would ever be recorded.

That is the rule already written at the top of "Things that have bitten",
applied to one member of a set: the combined gating+logging test exists
because of it, and the enrollment test added later did not get it.
Twenty runs clean since.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-04 17:58:11 -04:00
iris 3c0214ece8 Merge branch 'main' of git.arirex.me:iris/ai-app
# Conflicts:
#	AGENTS.md
#	PLAN.md
#	app/androidApp/src/main/kotlin/com/example/aiapp/SessionUsageBar.kt
#	app/androidApp/src/main/kotlin/com/example/aiapp/SpawnScreen.kt
#	server/src/config.rs
#	server/src/main.rs
#	server/src/routes.rs
#	server/src/session/echo.rs
#	server/src/session/llama.rs
#	server/src/session/transport.rs
#	server/src/ssh.rs
#	server/src/usage.rs
2026-09-04 17:56:50 -04:00
irisandClaude Opus 5 127b25e60a Meter a session by its provider, and let llama.cpp run over ssh
The rate-limit bar answered a question about an account, and picked the
answer by machine. One machine runs echo, the Claude CLI and a local
model side by side, so every echo session on it drew the CLI's five-hour
window: a quota that session cannot spend and could never run down. A
session now names its meter (`usageProvider`, from
`DriverKind::usage_provider`, which `usage::providers_for` reads too so
the two lists cannot disagree), and the phone matches on machine *and*
provider. Nothing meters echo or llama, and nothing at all is drawn --
including while the first fetch is out, since "checking" under a session
that turns out to meter nothing is a row the screen then withdraws.

Echo gets a meter it can be *told* about instead: `/usage 42`,
`/usage 95 20`, `/usage 42 never`, `/usage notloggedin`,
`/usage unreachable`, `/usage failed`, `/usage off`. Those states cost
real quota to arrange, which is why none of them had been looked at.

And llama.cpp runs wherever a setup says, which was the last of phase 5.
`Transport::reserve_port` is the second half of what a transport is --
"run this" plus "reach this port" -- returning the port the server binds
there and the port that reaches it here, and `Launch::reaching` puts the
`-L` tunnel on the connection that already carries the command. Three
things that came out of building it:

- A forwarded launch gets a pty and every other one keeps `-T`. Killing
  the ssh client ends a CLI by closing the stdin it reads; llama-server
  never reads its stdin, so the same kill left it running on the far
  machine with the model loaded -- one orphan per stopped session.
- The model is looked for on the machine that will serve it, at that
  machine's own models directory, so `GET /setups/{id}/models` is what
  the spawn screen offers rather than the backend's own downloads.
- The readiness poll watches the process, not only the port: a model
  that will not load exits in a second and would otherwise have been
  reported as "gave up after 300s". The failure carries the log's tail.

Exercised end to end against this VM over ssh to itself: spawn, load,
answer, outlive a backend restart, be adopted, answer again, and stop --
with both the ssh client and the far llama-server gone afterwards. The
local path, the Claude bar and the spawn screen checked on the emulator.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-04 17:45:32 -04:00
222 changed files with 50777 additions and 5978 deletions

No files matched your search

+244
View File
@@ -0,0 +1,244 @@
---
name: ai-app-rigs
description: ai-app's test rigs, harness scripts and reference measurements - ui-sandbox.sh, debug-transcript.sh, transcript-bench.sh, stream-bench.sh, trace-draw.sh, the /usage fixture vocabulary, the fake CLI, the rule that no UI-driving script may tap a coordinate, how to test llama.cpp and ssh on this machine, how importing behaves, and the scroll/stream/explorer numbers not worth re-measuring. Read before running or writing a benchmark, driving the app's UI from a script, exercising the session lifecycle, testing a llama or remote session, or touching the import screen.
---
# ai-app: rigs, harnesses and measurements
Moved out of `AGENTS.md` on 2026-09-04 so it is read when it is relevant
rather than sent with every request in this repo -- it was 12 KB of the 35 KB
that file cost on every one. Unchanged in the move, and still the only copy.
## The rigs
Each exists because something was invisible without it.
- **`app/ui-sandbox.sh`** — a second `ai-server` with its own `$HOME`, config
and data directory, holding eight invented Claude Code transcripts and a
`claude` that is two lines of shell. **That isolation is the point**: the
import screen lists whatever is in `~/.claude/projects`, which in this VM is
real agent transcripts, so exercising *delete* against the ordinary server
deletes somebody's conversation and exercising *import* starts a real
`--resume` on the owner's account.
Its port and root derive from the checkout's name, so two checkouts'
sandboxes cannot reach each other, and its token is generated once into
`~/.config/ai-app/sandbox-token` and carried across restarts along with any
the enrolment flow appended — so the emulator app is enrolled **once** (the
start banner prints the command) and stays enrolled. It shares the real TLS
certificates, because the installed APK pins that CA.
Driving verbs, so none of this is re-derived per session:
`./ui-sandbox.sh spawn [title]` (an echo session, prints its id),
`./ui-sandbox.sh send SID text|@file`, and
`./ui-sandbox.sh api /path [curl args]`.
`./ui-sandbox.sh keep` restarts the server without wiping the sessions and
enrolment already there — for when the fixture under test was expensive to
build; plain `start` wipes them, which is right for the list-screen
fixtures and wrong for that.
It passes `--delay` by default, and `AI_SANDBOX_BIG_MB` puts one large
transcript among the small ones while `AI_SANDBOX_SPAWN_DELAY` makes the
fake CLI slow to start. Both exist because operations that finish in
milliseconds have states on the way that nothing can observe, and an
unobservable state is one where broken and working look identical.
It also builds a fixture tree at the sandbox home's `~/files` for the
explorer, holding the states otherwise only reachable by finding a real
machine in one: an empty directory, a name with a tab and one with an
apostrophe, a binary file, one over `FILE_LIMIT`, one `chmod 000`, a
symlink to a directory and a broken one, a source file per language, and
the three sizes the limits were measured against (`edit-32k.rs`,
`edit-128k.rs`, `big-source.rs`). Point a session at it with
`./ui-sandbox.sh api /sessions/<id>/cwd -X POST -H 'content-type: application/json' -d '{"cwd":"~/files"}'`.
The explorer's 409 is produced by editing the file on the machine
(`printf … > file`) between pressing the pencil and pressing save.
- **`app/debug-transcript.sh`** — a real conversation on the emulator. The
echo driver is the right rig for most things and the wrong one for anything
whose cost scales with what was actually written: a real reply is longer,
is real markdown, and carries tool calls whose input and output are
kilobytes. Two faults were invisible until a real transcript was loaded — a
page of history landing mid-fling threw the reader back to the newest end,
and parsing one real reply took 51ms against 4.6ms for a synthetic one.
`-b` takes the biggest conversation on the machine rather than the newest,
which is what a scrolling test wants; `--stop` takes it down.
It copies the transcript into `/tmp` and gives the server a `HOME` of its
own, so the import can only see the copy — importing spawns `claude
--resume`, and against the real file that is a second CLI writing to a
conversation somebody may still be in. **A transcript never goes in this
repository**: they hold whatever was said, read and written in that
session, and `~/repos` is shared with the host besides.
- **`/usage` in an echo session puts up an invented meter**, which is how the
rate-limit screens' states are reached without spending quota: `/usage 42`,
`/usage 95 20` (minutes left), `/usage 42 never` (the between-blocks window
with no reset time), `/usage 42 unreadable`, `/usage notloggedin`,
`/usage unreachable`, `/usage failed`, `/usage off`. The vocabulary is
`usage::Fixture`'s, since those are its states. With none set an echo
session meters nothing, which is the ordinary case and draws no bar.
- **A fake CLI exercises the process lifecycle without a token.** Point a
`claude_cli` provider's `command` at a two-line script — `#!/bin/sh` and
`cat > /dev/null` — and it behaves the way the lifecycle code cares about:
it holds the fifo open, records a real pid, writes nothing, and dies on a
signal. So adopt, stop, restart and start are all drivable without a real
`--resume` and without spending a turn on somebody's account. Reach for
this when what is under test is *whether a process is running*, and for
`debug-transcript.sh` when it is *what the transcript draws*.
- **`app/transcript-bench.sh`** is the standard scroll measurement: it opens
the first session (or `-k` keeps the current screen), scrolls a fixed
gesture loop, and prints the app's render report — the same one the in-app
copy button produces, whose `on screen:` line names what the viewport was
holding. Compare two runs with the same gestures; the emulator's absolute
frame times transfer nothing, the report's accounting does. Run it either
side of any change under `Markdown*.kt`, `Transcript*.kt` or
`SessionScreen.kt`'s list, and put the report in the commit. The numbers
that move first are the worst `record: one block`, the reparse mean while
streaming, and the draw phase's accounting line.
- **`app/stream-bench.sh [-k] FILE`** is that measurement for a reply still
arriving. It taps "Jump to latest" so the list is pinned to the newest end,
resets the report, sends FILE, waits for the transcript to stop growing,
and prints. Both of those are corrections to a first version that measured
nothing: a transcript parked further back never redraws while a reply
streams into it, and a session is idle at *both* ends of a turn, so polling
for idle answers before the turn has started.
- **`app/trace-draw.sh`** names what a scrolling frame spends inside the
framework, from `atrace` text output with no trace processor needed. It is
how the cost of a layout node per link was attributed to the framework
rather than guessed at.
### Driving the UI
**No script that drives this app's UI presses a coordinate.** Every control
is found by the name it already carries for assistive technology —
`ui-trace record --do "tap 'Session settings'"` — which resolves the label
against the screen at the moment of the gesture and fails the whole run when
it is not there. `app/bench-lib.sh` is what the bench scripts share for it. A
coordinate is a position measured once by hand, and anything that moves the
control makes the tap land on whatever now sits there — the bench then
reports a number that was never measured, which reads exactly like a result.
Both bench scripts pressed the render report at `tap 723 205` until that
button moved into the session settings dialog on 2026-09-03. The check that
none has crept back:
grep -n "tap [0-9]" app/*.sh
Swipes are still coordinates, deliberately: a gesture across a scrolling area
is a distance rather than a control.
**Two traps in the emulator bench loop**, each of which cost a run.
`adb shell pm clear` removes the enrolment and the notification permission
along with the saved anchors, so the next run measures a permission dialog —
re-enrol with the command `ui-sandbox.sh` prints, and
`pm grant … POST_NOTIFICATIONS`. And a saved scroll anchor is per session id,
so the only way two builds start a scroll from the same place is a *fresh
session for each*.
**The emulator is `~/repos/emulator-tools`' business, not this repo's.**
`emu up` creates and boots the AVD named after this checkout — whatever `emu
name` prints, never a name typed out here, since this file is the same in
every clone. `run-android.sh` is that plus a build and an install. The `adb`
on `PATH` after sourcing `android-env.sh` is that repo's wrapper, which fills
in `-s` from the same rule. Gradle does not go through it, so a Gradle init
script from `emulator-tools` runs `emu check` before `installDebug`,
`uninstallDebug` and `connectedAndroidTest` and fails rather than fanning out
to every attached device; when it refuses, say which device you mean at the
moment you use it — `ANDROID_SERIAL=$(emu serial) ./gradlew …`.
### Testing llama.cpp and ssh here
**Both are set up here as of 2026-09-04** and need nothing typed. The
prebuilt CPU llama.cpp lives outside the repo at `~/.local/opt/llama.cpp`
(the 15 MB `ubuntu-x64` release asset) and is symlinked as
`/usr/local/bin/llama-server`, which is what makes **discovery find it over
ssh**: `~/.local/bin` is not on the PATH a non-interactive ssh session gets.
It resolves its own libraries through `$ORIGIN`, so no `LD_LIBRARY_PATH` is
needed. One model is downloaded — `unsloth/Qwen3-0.6B-GGUF/Qwen3-0.6B-Q8_0.gguf`,
639 MB under `~/.local/share/ai-app/models` — and answers at usable speed on
this VM's 8 cores. **Do not test with a 2-bit quant**: the
IQ2_XXS of that model produces fluent nonsense, which reads exactly like a
broken driver — `llama-cli` produces the same from the file directly, which
is how to tell the two apart in a hurry.
There is no second machine, so **ssh this VM to itself**. That is set up
too: the key is `~/.config/ai-app/ssh-self` (its public half is in
`~/.ssh/authorized_keys`, labelled removable), and the real config carries a
setup called **"this vm over ssh"** — `bob@127.0.0.1` with that
`identityFile` plus
`options: ["StrictHostKeyChecking=no", "UserKnownHostsFile=/tmp/ai-app-known-hosts"]`
so it touches nothing real — offering `claude-cli` and `llama-cpp`. It is the
whole rig for "does a remote llama session work", since the far machine is
this one and the model file is the same file. For a throwaway setup of your
own, point a provider's `command` at something harmless like `/bin/echo`
rather than at `claude`: the transport is what is under test, the process
exiting immediately is the signal, and it costs no tokens. The remote login
shell here is **fish**; the
remote script and `ssh.rs`'s POSIX quoting happen to mean the same thing in
both, but that is luck rather than design, and a shell that is neither is the
thing to suspect first if a remote spawn ever mangles an argument.
## Importing
The import list reports each session's **size as well as its line count**,
because the two disagree in the way that matters: these transcripts embed
screenshots as base64, so one line can be a megabyte. On this machine a 69 MB
session has 3,427 lines and a 44 MB one has 6,792 — nothing about a line
count tells you what continuing a session will cost. Shown, not warned about;
importing a large session is a choice somebody is entitled to make.
**Never import a Claude Code session that is open in a terminal.** The app
refuses it — see PLAN.md for the incident that made that a refusal rather
than a warning.
**One Claude Code session id can name two files, and the listing offers it
once.** Resuming from a different working directory makes the CLI write a
second transcript with the same id under that directory's project folder — an
ordinary state of a machine, not corruption. Everything downstream addresses
a session by id, and the phone keyed its list on it, so two rows sharing one
**closed the app** on a Compose duplicate-key throw. `parse_listing` keeps
the copy with the most lines, because the other is usually a few-hundred-byte
stub and is often the *newer* of the two, so recency is the wrong key.
Deleting removes every copy rather than the first, or the row came back after
a delete that reported success. The phone's half is `uniqueItems`, which
every list keyed on a server-chosen id goes through: a repeat there must
never be able to close the app, whatever produced it.
**Deleting a session offers to take the machine's own transcript with it**
`DELETE /sessions/{id}?deleteForeign=true`, behind a switch in the
confirmation, and only where the driver keeps a record of its own
(`keepsOwnTranscript`, which today means Claude Code). Off by default,
because leaving that copy is what makes an ordinary delete recoverable — and
the dialog's paragraph is rewritten when it is on rather than appended to,
since the sentence promising the conversation "should still be there to
import again" is exactly the one the switch makes false. The server deletes
the machine's copy *first*, so a machine it cannot reach leaves the session
where it was instead of half-deleted.
## Measurements worth not re-taking
- **What the transcript screen costs to scroll.** Taken 2026-08-30 on the GPU
emulator against a real imported transcript with the server at
`--delay 120`. Settled and flinging fast, both into fresh history and back
through rows already drawn: **5.25.9% janky frames, 99th percentile
2932ms, 02 slow UI-thread frames.** The stock Settings app on the same
device is 3.3% and 38ms, so this is at the platform floor. The number that
is *not* at the floor is the first few seconds after opening a session,
where every row on the way is being composed for the first time; that is
inherent to a lazy list and it is why a measurement taken before the screen
settles reads three times worse. **Settle first, then reset `gfxinfo`.**
- **The reset path is not reachable by reopening a session.** Measured
2026-09-04 against a session streaming at 20 events a second: reopening one
with an anchor 1,800 events back connects **87119 events behind**, well
under `CATCH_UP_LIMIT`'s 200, because the restore is two requests — the
opening page, then one span covering the whole distance. To exercise the
reset at all you have to lower `CATCH_UP_LIMIT` in a throwaway build; at 5
the app takes the reset on a live connection, clears, refills and carries
on without reconnecting.
- **The session screen's stream survives backgrounding here** — 20 seconds at
the launcher while 415 events were produced brought no reconnect at all,
which is not what the comment above that loop expects, and is most likely
this emulator being headless rather than the phone's behaviour.
- **Reopening a cached session costs one request for one event** (the probe),
and scrolling the whole conversation back costs nothing more; a cold open
of the same 500-event session is two pages, 100 events. Measured
2026-09-04 on the emulator against the sandbox.
- **Reading is cheap and editing is not.** The viewer handles a 1 MiB,
28,000-line file because it draws one row per line; the editor is one
`BasicTextField`, which costs two seconds a frame at 128 kB and stops the
app at 1 MiB, so `EDIT_LIMIT` caps it at 32 kB with the reason said on
screen. If you make the editor faster, that number is what to move.
EXPLORER.md's "What the measurements said" has the rest.
+109 -14
View File
@@ -5,24 +5,42 @@ replacing the Claude app for daily use. Rust/Axum backend on the desktop,
Kotlin/Compose Android app, WireGuard + pinned self-signed TLS + bearer token
between them.
**`PLAN.md` is the design source of truth** — every decision with its date,
its rationale, and what was rejected. Read it before changing anything
**`docs/PLAN.md` is the design source of truth** — every decision with its
date, its rationale, and what was rejected. Read it before changing anything
structural, and update it in place when a decision changes rather than
letting this file and the plan become two versions of the truth. This file is
the working notes layer: layout, commands, rigs, and things that have bitten.
The design and working documents live under `docs/` — everything except this
file and `CLAUDE.md`, which stay at the root because that is where Claude
Code and other agent harnesses look for them.
The central design point, worth not undoing by accident: **a session is a
child process, translated into one common event model.** A new session type
is a new driver — never a session-type branch in shared code (routes,
transcript, app screens).
The second one, for the Rust port on the `rustify` branch: **the phone app
and a planned desktop app share almost all of their code.** Screens, widgets,
folding, paging, config and the network client live in the shared crates
(`iris`, `client-core`, `transcript-ui`, `tabs-ui`); `android-app` and
`desktop-app` are thin entry points that own only what the platform forces
(JNI and the IME on one side, winit and argv on the other). The two
*layouts* will differ, to suit a phone's screen and a finger against a
desktop's screen and a mouse -- but the widgets a layout is made of (a
button, a text field, a list, a card) and the styling (colours, spacing,
type) are one implementation with no per-platform copy. Anything that could
work on both goes in a shared crate the first time it is written, and a
platform crate growing a widget or a colour is a defect to move, not a
convenience to keep. Iris said this on 2026-09-07; docs/RUST.md carries the
details.
## Layout
Mirrors `../dev-updater` deliberately: same stack (axum 0.8 +
axum-server/rustls, tokio, clap; Kotlin 2.4.x + Compose Multiplatform, single
`:androidApp` module), same cert scheme, same registry pattern. Read
dev-updater's `README.md` and `AGENTS.md` before diverging from them.
Module-by-module intent is in PLAN.md's "Backend layout".
Module-by-module intent is in `docs/PLAN.md`'s "Backend layout".
- `server/` — the Rust backend (`ai-server`). `routes.rs`'s module doc
comment is the HTTP table and the surface's source of truth.
@@ -40,16 +58,24 @@ Module-by-module intent is in PLAN.md's "Backend layout".
projects version-locked to the commit this repo pins. What deliberately did
**not** move is the API surface and the config *schema*: routes, drivers,
sessions and setups are what makes this project itself.
- `EXPLORER.md` — the file explorer's design (`server/src/files.rs` and
`FilesScreen.kt` / `FileViewer.kt` / `FileEditor.kt`).
- `TRANSCRIPT_CACHE.md` — the phone's copy of what it has been sent. Read it
before touching `TranscriptCache.kt`, `TranscriptSource.kt`, or the opening
and stream effects in `SessionScreen.kt`.
- `TODO.md` — the working list.
- `RUST.md` — the plan for moving the app to Rust (on the `rustify`
- `docs/` — every design and working document except this file and
`CLAUDE.md`:
- `docs/EXPLORER.md` — the file explorer's design (`server/src/files.rs`
and `FilesScreen.kt` / `FileViewer.kt` / `FileEditor.kt`).
- `docs/TRANSCRIPT_CACHE.md` — the phone's copy of what it has been sent.
Read it before touching `TranscriptCache.kt`, `TranscriptSource.kt`, or
the opening and stream effects in `SessionScreen.kt`.
- `docs/TODO.md` — the working list.
- `docs/RUST.md` — the plan for moving the app to Rust (on the `rustify`
branch of the `ai-app-2` clone): what has to be reproduced, the
framework decision, and the ordered experiments with their pass
conditions. Read it before touching anything under that branch.
- `docs/IRIS.md`, `docs/IRIS_TODO.md`, `docs/DECISIONS.md`,
`docs/LAYOUT.md`, `docs/TEXTURES.md`, `docs/CLIENT_CORE.md` — iris's
own build log (**any major addition or design decision, not only
public API** -- Iris, 2026-09-08), working list, decisions log,
layout/render design, and texture-atlas design, and the client-core
crate's design, respectively.
- `.dev-updater.ron` — what Dev Updater builds here: the server (run as
`service: Managed(…)`, supervised by Dev Updater's own implementation
rather than a script kept here) and the APK, in parallel. It points at
@@ -75,6 +101,18 @@ the **Mono** face, where every glyph is one em square, which is what makes
two icon buttons the same width without either being given one — and why
`GLYPH_SIZE` is smaller than it looks like it should be.
**The Rust app does the same, from its own subset**:
`iris/core/build-icon-font.sh` -> `iris/core/assets/fonts/nerd_icons.ttf`,
with the codepoints named in `iris/core/src/icon.rs` and drawn as text
with `Family::Icons`. Same rule about the two lists agreeing (there is a
test, `every_icon_is_in_the_bundled_font`), same Mono face, same Material
Design family so an icon means the same thing in both apps. Its subset is
separate rather than shared because subsetting only what one app draws is
the point. This is the **only** font iris bundles — body and monospace
text come from the platform (docs/DECISIONS.md, 2026-09-07), and an icon
is the opposite case: a small closed set of codepoints no system font is
guaranteed to have.
## Checking your work
- **Server**: `./run-tests.sh` from the repo root (or `cargo test` from
@@ -86,7 +124,11 @@ two icon buttons the same width without either being given one — and why
:androidApp:compileDebugKotlin :androidApp:lintDebug
:androidApp:testDebugUnitTest`. The unit tests are JVM-only and cover the
syntax highlighter, the ANSI parser and the transcript cache — the app's
pure logic with no Android in it.
pure logic with no Android in it. Touching anything under `BenchFixture.kt`,
`BenchNetwork.kt`, `BenchRun.kt` or the `bench` build type also needs
`:androidApp:compileBenchKotlin :androidApp:lintBench` — a second build
type compiles separately and lint has caught real bugs debug alone never
would (see "Android Lint" below).
- **Android Lint is not optional and is not run by a build.** It found a
crash that had been shipping (`java.time` on a minSdk-24 app with
desugaring off) and later a permission check that silently dropped every
@@ -146,6 +188,27 @@ two icon buttons the same width without either being given one — and why
Each exists because something was invisible without it.
- **The `bench` build type and `app/bench-fixture/`** exist for P0 (RUST.md
and DECISIONS.md's 2026-09-05 entries), the phone benchmark gate Iris
asked for before porting continues: a deterministic, checked-in synthetic
transcript (`app/bench-fixture/generate.py`, never a real one) that both
this app and iris open with no server, so a frame-time comparison
measures the renderer rather than the data. `./build-apk.sh bench` builds
it — own application id (`com.example.aiapp.bench`) and label ("AI
Sessions bench") so it installs beside a real enrollment rather than
replacing it. Opening it goes straight to a session screen holding the
fixture (no enrollment, no permission prompts) with a "Run benchmark"
control beside "Copy" in session settings: it drives the same scroll loop
and streaming phase `transcript-bench.sh`/`stream-bench.sh` drive over
`ui-trace`, but in-process, since a real phone has no usable system
tracing and no agent can drive one (this-machine-android's skill).
`BenchFixture.kt`/`BenchNetwork.kt` fake the backend by installing a
`URLStreamHandlerFactory` that answers `TranscriptSource`/`EventStream`'s
requests from an in-memory copy of the fixture instead of opening a
socket — so the fold, the paging and `uniqueItems` under test are the
screen's real ones, never a shortcut built just for this. The report
gains a `bench:` section (process CPU time, peak RSS, battery current) on
every build, empty except when `BenchRun.kt` filled it in.
- **`app/ui-sandbox.sh`** — a second `ai-server` with its own `$HOME`, config
and data directory, holding eight invented Claude Code transcripts and a
`claude` that is two lines of shell. **That isolation is the point**: the
@@ -226,6 +289,38 @@ Each exists because something was invisible without it.
framework, from `atrace` text output with no trace processor needed. It is
how the cost of a layout node per link was attributed to the framework
rather than guessed at.
- **`iris/android-app/build-apk.sh [debug|release] [--abi ...] [--features
...]`** builds iris-android-app's cdylib (`cargo ndk`) and its APK
(Gradle) in one step and verifies the result (`aapt2`/`apksigner`), and
**`iris/android-app/run-bench.sh [--apk PATH]`** installs it on this
checkout's own emulator, taps "Run benchmark" by label, and prints the
report -- written so the P0 build/install/tap/read-report cycle stops
being retyped by hand each time (docs/RUST.md's P0 box).
- **iris's three test layers** (docs/RUST.md's "Three test layers" has
the commands and what each cannot answer): test at the cheapest one
that can answer the question. `cargo test -p transcript-fixture` runs
the real transcript screen over the bench fixture with **no window, no
compositor and no GPU** (`iris::harness`), on a clock the test owns and
a gesture replayed from a `t_ms action x y` file under
`iris/transcript-fixture/touch/` -- which is how the batched 120Hz
flick a finger actually makes is testable at all, since a `ui-trace`
swipe is many evenly-spaced events. `iris/run-headless.sh phone --phone
--shot …` opens the same screen in a window at the phone's own size and
density for looking at, and `--replay FILE` drives the same recording
into it. The emulator is for JNI, the IME, insets, the surface
lifecycle and one verification run before a build goes to the phone --
not for iterating on layout.
- **The emulator is a GLES rig, deliberately** (Iris, 2026-09-08;
docs/DECISIONS.md). Its guest has no hardware Vulkan -- only SwiftShader
in software -- while its GLES *is* the host's real GPU through virgl at
ES 3.1, so an ordinary build's runtime fallback lands there by itself
and nothing should pass `force-gles` to arrange it. The Vulkan path is
verified on the desktop build and on Iris's phone. Do not boot the
emulator with SwiftShader Vulkan to "test the Vulkan path": that
measures a software rasteriser and steers iris away from the one
hardware-accelerated backend it has there. Every run says which adapter
drew it (`iris renderer:` in logcat, printed by `run-bench.sh`); read
that line before reading a number.
### Driving the UI
@@ -312,7 +407,7 @@ means here:
## Sessions outlive the backend
Since 2026-08-29 a session's process is deliberately left running when
`ai-server` stops, and adopted again when it starts. PLAN.md has the design;
`ai-server` stops, and adopted again when it starts. docs/PLAN.md has the design;
day to day:
- **Stopping the server no longer stops the sessions.** After `pkill
@@ -345,7 +440,7 @@ count tells you what continuing a session will cost. Shown, not warned about;
importing a large session is a choice somebody is entitled to make.
**Never import a Claude Code session that is open in a terminal.** The app
refuses it — see PLAN.md for the incident that made that a refusal rather
refuses it — see docs/PLAN.md for the incident that made that a refusal rather
than a warning.
**One Claude Code session id can name two files, and the listing offers it
@@ -521,4 +616,4 @@ belongs in `~/.claude/TOOLCHAIN.md` or `~/.claude/MACHINE.md` instead.
`BasicTextField`, which costs two seconds a frame at 128 kB and stops the
app at 1 MiB, so `EDIT_LIMIT` caps it at 32 kB with the reason said on
screen. If you make the editor faster, that number is what to move.
EXPLORER.md's "What the measurements said" has the rest.
docs/EXPLORER.md's "What the measurements said" has the rest.
+41 -27
View File
@@ -1,30 +1,44 @@
# Decisions taken for Iris to review
# Decisions awaiting review
Short list of design choices made by the design agent without asking, so
they can be judged and reversed later. Detail lives in RUST.md (and IRIS.md
for iris API changes); this file is only the summary. Newest first. Items
marked **DEFERRED** are ones the agent chose not to decide alone.
Choices made while working autonomously, for Bryan to keep or change. Each
says what was picked and why; the detail is in the design doc it names.
Delete an entry once it has been looked at.
## 2026-09-05
## Subagent views (2026-09-05, `SUBAGENTS.md`)
- **Touch drag on a transcript row follows Android's own rule**: a vertical
drag pans the list immediately; a stationary press held 500 ms starts a
text selection which further dragging extends; a horizontal drag while
something is already selected extends that selection without the wait.
One `DragArbiter` per list decides it (`iris/src/sense.rs`). Chosen over a
"text layer always wins" or "list always wins" rule because either loses
one of the two gestures a reader expects.
- **E4's desktop shape is a new `iris/desktop-app` crate**: a winit window
holding `transcript-ui`'s screen beside a session list, talking to a real
`ai-server` through `client-core`. It enrols by pasting the same
`aiapp://enroll?…` link a phone scans (`client-core::config::EnrolledServer`)
and keeps it owner-only under `$XDG_CONFIG_HOME/ai-app-desktop/`. The
pinned CA is a path given on the command line, not baked in. Chosen so
the phone and desktop share one enrolment format and no second one is
invented.
- **Order of remaining work**: finish the two in-flight pieces above, then
the transcript screen's Android integration and the `transcript-bench.sh`
comparison against Compose — the numbers the recommendation still lacks.
- **DEFERRED — whether to commit to iris over Masonry for `ai-app`.** Waits
on the bench numbers above; RUST.md's recommendation says what the
measurements must show.
Made on my own judgement, limited blast radius:
1. **A subagent is a transcript, not a session.** It has no process,
controls or settings; it is addressed as `/sessions/{id}/subagents/{sub}`
and stored under the session's directory, so deleting the session takes
it. Alternative rejected: registering it as a session of its own, which
would give it a card in the main list and a driver that can do nothing.
2. **Read-only view is the session screen minus its controls**, rather than
a second, simpler transcript screen. Keeps paging, caching, selection
and rendering in one place. Cost: a `readOnly` mode threaded through
`SessionScreen`.
3. **The list only carries a count.** Each session row says how many
subagents it has; their titles and statuses are fetched when the card is
expanded. Keeps `GET /sessions` from reading every subagent transcript.
Consequence: an expanded card's statuses refresh with the list, not live.
4. **Expanded/collapsed is remembered per session on the phone**, not on
the server. Collapsed by default, per the transcript convention that new
things arrive collapsed.
5. **Subagents of imported sessions are not shown.** The import path still
skips `isSidechain` records; the CLI's own `subagents/agent-*.jsonl` files
are not read. Only subagents run while this backend was watching exist.
6. **Echo grows `/subagent [n]`** as the test rig, so nothing here needs a
paid turn to exercise.
Deferred, because they reach further than this feature:
- **Live status on the list.** Whether the session list should follow a
stream at all (it refreshes on demand today) decides whether subagent
status can ever be live there. Not changed.
- **Nested subagents.** A subagent's own Task calls are shown as tool calls
in its transcript and are not given transcripts of their own. Supporting
that is the same mechanism one level down, but the UI would need nested
expanders.
- **The subagent status row says "context unknown".** Nothing measures a
subagent's context; the row could leave it out rather than admit it.
-277
View File
@@ -1,277 +0,0 @@
# iris: notable public API changes
For Iris to read on her own time. Each entry is a change to iris's public
surface that a widget author or app author would notice: a trait method
added, removed or re-shaped; a type that callers construct differently; a
capability that moved. Small and trivial changes do not go here.
An entry gives the date, what changed, why, and a short before/after where
it helps judge the change without the session that made it. Newest first.
## 2026-09-05: `transcript_ui::build_tree` (RUST.md's E4)
`transcript_ui::build` claimed the whole window (`ui_state.set_root(tree)`)
as its last step, which is right for a window that *is* the transcript
screen (the winit example, an eventual Android cdylib) and wrong for the
desktop app, which puts a session list beside it. `build_tree` is `build`
minus that last step: it returns `(TranscriptScreen, StrongWidget)` instead
of just `TranscriptScreen`, and the caller decides where the tree goes —
into `ui_state.set_root`, or into a `WidgetPtr` alongside something else
(`iris/desktop-app`'s `rebuild_transcript`). `build` is now one line calling
`build_tree` and doing the `set_root` itself, so existing callers are
unaffected.
```rust
// before, and still available, for a caller that wants to *be* the window:
let screen = transcript_ui::build(rsc, &mut ui_state, rows);
// new, for a caller embedding the screen beside something else:
let (screen, tree) = transcript_ui::build_tree(rsc, rows);
some_widget_ptr(rsc).set(tree);
```
## 2026-09-05: `DragArbiter`, pan-vs-select for one shared touch gesture (RUST.md's I5)
New public type, `iris::sense::DragArbiter`. Why: a widget author who
registers both a list-level pan and a row-level drag-to-select on the same
touch gesture has no way to arbitrate between them — `core/src/sense.rs`'s
`run_sensors` always gives the innermost layer first refusal, so the inner
one wins every frame it is pressed, not just the frame the press started
(this is exactly what left transcript-ui's touch-drag panning unreachable
until now). `DragArbiter` is one small state machine, one instance per
gesture surface (a whole list, not per row), that a caller drives with its
own `press_start`/`update`/`release` calls and a caller-supplied `Instant`
(so it is unit-testable without a real clock or a render harness). It
decides the way Android itself does: an ordinary vertical drag pans
immediately; a stationary press held `LONG_PRESS` (500ms) starts a
selection, which any further drag then extends; a horizontal drag while
something is already selected extends it immediately, skipping the wait.
```rust
// One per list, held alongside whatever state coordinates the rows:
let mut arbiter = DragArbiter::new();
// On press-down:
arbiter.press_start(pos, Instant::now(), already_selected);
// Every frame the button/finger stays down:
match arbiter.update(pos, Instant::now()) {
DragOutcome::Pan(dy) => list.scroll(-dy),
DragOutcome::SelectStart => selection.begin(...),
DragOutcome::SelectExtend => selection.extend(...),
DragOutcome::Undecided => {}
}
// On release:
arbiter.release();
```
`transcript-ui`'s `Selection::drag` (`transcript-ui/src/selection.rs`) is
the reference caller: every row's `CursorSense::click_or_drag() |
CursorSense::unclick()` handler routes through one `Selection`-owned
arbiter instead of calling `begin`/`extend` directly, so a drag that starts
on a row's own rendered text now pans the list correctly instead of
always starting a selection. 8 new unit tests in `iris/src/sense.rs`'s
`drag_arbiter_tests` module.
## 2026-09-05: `SpanStyle`, per-range text styling (RUST.md's I5)
A `TextBuffer` used to have exactly one style (`TextAttrs`: colour, size,
family, ...) for its whole string, applied via `push_default` into parley's
ranged builder. `SpanStyle` is a second, optional layer: a byte range plus
whichever of colour/family/font size/bold/italic/underline it overrides,
pushed with parley's own `push(property, range)` instead. Why: a transcript
row's markdown (a heading, **bold**, `inline code`, a link) all inside one
wrapped paragraph needs each to carry its own look while the paragraph
still wraps and selects as a single buffer — the thing `masonry`'s
`TextArea` cannot do (`StyleSet` is one style for the whole editor,
`text_area.rs:43-44`'s `// TODO: RichTextInput`), and the reason this
existed at all.
```rust
let (text, spans) = transcript_ui::markdown::render_markdown(src, 16.0);
wtext(text)
.spans(spans) // new: TextBuilder::spans, on both Text and TextEdit
.editable(EditMode::MultiLine)
.add(rsc);
```
Two things a widget author should know before reaching for it:
- **Call `.spans()` before or after `.editable()`, both work** — the field
lives on `TextBuilder` itself, not either output type, and both
`TextOutput::run` and `TextEditOutput::run` apply it to the buffer via
`TextBuffer::set_spans`. **These two call sites are a pair**: adding a
third `TextBuilderOutput` impl without also calling `set_spans` there
reproduces the exact bug this box shipped once already (spans silently
dropped for `TextEdit`, found only by screenshotting, not by any test —
`markdown.rs`'s own unit tests check string/range logic, which is
correct in isolation and proves nothing about whether the render path
ever sees it).
- **Colour is now per-glyph, not per-buffer.** `PlacedGlyph` gained a
`color: UiColor` field (from parley's own per-run `Style::brush`), and
`Painter::glyphs` draws each glyph in its own colour instead of
`RenderedText::color` uniformly. `RenderedText::color` still exists (the
buffer's *base* colour, for a caller that wants it as a whole, e.g. to
tint a cursor) but no longer drives what a glyph actually renders as.
## 2026-09-05: accessibility names via AccessKit (RUST.md's I4)
`.label()` (already in `trait_fns.rs`, previously unused anywhere in-tree)
is now load-bearing: it's the one thing that puts a widget in the AccessKit
tree `iris_core::ui::access::AccessTree` builds and both backends push
out. A widget author who wants a control to be findable by name (and
tappable by name, through `ui-trace`/a real screen reader) calls `.label()`
on it; nothing else is required, and a widget nobody labels is invisible
to this system at zero cost, not just zero UI.
```rust
let button = rect(Color::LIME)
.on(CursorSense::click(), move |_, rsc| { ... })
.label("Add task"); // now findable by uiautomator/AccessKit as "Add task"
```
Two new things a widget author might touch directly:
- **`Widget::access_role(&self) -> accesskit::Role`**, default `Unknown`.
Override it if your widget has a real platform equivalent —
`TextEdit` now returns `TextInput`/`MultilineTextInput` by `EditMode`.
Only consulted for a widget that also has a `.label()`; an unlabelled
widget's `access_role` is never called.
- **`Widgets::named() -> impl Iterator<Item = WidgetId>`** — every widget
with an explicit label, for anything else that wants to walk the same
set `AccessTree` does.
Nothing about `Painter`, `draw`, or the layout/move machinery changed —
this sits entirely beside them, reading `resolved_region`'s output rather
than participating in producing it.
## 2026-09-05: `List`, a virtualised bottom-anchored list (RUST.md's I3)
A new widget, `iris::widget::List` (`iris/src/widget/list.rs` -- read its
module doc first), for the transcript's kind of screen: variable-height
rows, keyed by a `u64`, composed only while visible, moved rather than
re-laid-out on scroll, a scroll anchor that survives a row inserted above
it, "more" sentinels at each end, and "hold the edge nearest the tap" when
a row's height changes (`note_tap`, resolved in the layout pass).
```rust
let mut list = List::new(Axis::Y);
list.push_back(ListRow::new(key, row_widget)); // O(1)
list.push_front(ListRow::new(older_key, row)); // O(1), anchor unaffected
list.set_more_before(Some(spinner_widget)); // sentinel, drawn at the edge
list.note_tap(viewport_y); // before mutating a row's height
let (top, bottom) = list.extent(key).unwrap(); // last frame's on-screen box, if visible
```
Built entirely out of existing primitives (`Painter::widget`/`widget_within`/
`reposition`/`draw_twice`, and `draw_inner`'s own old-children diffing) --
no new mechanism was added to the render core for it. One correctness
lesson worth reading even for other widgets: a row that fills whatever
region it is offered (`Rect`, `is_size_independent`) cannot be measured at
a throwaway oversized region and then merely `reposition`ed into place --
`reposition` only ever writes an offset, never a size, so the oversized
primitive stays oversized. `List` fixes this by caching each row's real
height once measured and placing an already-known row directly at its
exact box; see `list.rs`'s `place` for the full reasoning and
`a_fill_shaped_background_is_not_left_oversized` for the regression test.
## 2026-09-05: a second backend (android-view), and what moved to make room for it
RUST.md's I2. Three changes a widget or app author would notice, all in
service of the same thing: `default` (winit) and the new `android`
(android-view) backends sharing what does not depend on windowing.
- **`Selector`/`Selectable`'s bound changed from `Rsc::State:
HasDefaultUiState` to `Rsc::State: FocusHost`** (new trait, `attr.rs`).
`HasDefaultUiState` still exists and still works — `default/attr.rs` now
implements `FocusHost` for anything that has it — so a winit app's
existing code is unaffected. An Android app implements `FocusHost` via
`HasAndroidUiState` instead. Affects only an app that referenced
`HasDefaultUiState` directly at a `Selectable`/`Selector` call site
rather than through `.attr::<Selectable>(())`, which nothing in-tree
does.
- **`Tasks::init` takes `Arc<dyn RequestRedraw>` instead of
`Arc<winit::window::Window>`.** `RequestRedraw` (`task.rs`) is one method,
`fn request_redraw(&self)`; `winit::window::Window` implements it
(`default/render.rs`), so `Tasks::init(window)` at a call site is
unchanged by inference. Only matters if something constructed a `Tasks`
directly rather than through `DefaultRsc`/`AndroidRsc`.
- **`TextEdit::apply_event`/`TextInputResult` are `#[cfg(not(target_os =
"android"))]`** — they take a `winit::event::KeyEvent`, which does not
exist on Android; `android/input.rs` drives the same primitives
(`backspace`/`delete`/`motion`/`insert`, all still unconditional) from
`ndk::event::Keycode` directly instead. New unconditional getters on the
way: `TextEdit::text()`/`selection_range()`/`caret()`, and
`TextEditCtx::delete_byte_range`/`set_cursor_byte` — the primitives
`android/ime.rs`'s `InputConnection` bridge needed and that were not
previously exposed publicly.
## 2026-09-04: `Widget::draw` reports the size it used; `desired_width`/`desired_height` are gone
A widget used to implement three methods (`draw`, `desired_width`,
`desired_height`); it now implements one, `fn draw(&mut self, painter: &mut
Painter) -> Size`, which draws into `painter.region()` and returns how much
of it was used. Why: the two extra methods routinely re-simulated what
`draw` was about to do anyway (`Span::desired_ortho` copied its own draw
loop to get cross-axis sizing right) — one visit per widget per frame
instead of up to three. A container that needs a child's size before
placing it (alignment, centering) draws the child once at a provisional
region, reads the returned `Size`, and calls the new `Painter::reposition`
to move it into its final spot — an O(1) offset write, not a second draw. A
widget whose drawn output never depends on the size it's given (a
fixed-size `Rect`, a decoded `Image`) overrides the new `fn
is_size_independent(&self) -> bool { false }` to `true`, which skips
redrawing it when only its offered region changes shape.
```rust
// before
fn draw(&mut self, painter: &mut Painter) { /* ... */ }
fn desired_width(&mut self, ctx: &mut SizeCtx) -> Len { /* ... */ }
fn desired_height(&mut self, ctx: &mut SizeCtx) -> Len { /* ... */ }
// after
fn draw(&mut self, painter: &mut Painter) -> Size { /* ... */ }
```
`SizeCtx` and `Cache` are gone with it — see `LAYOUT.md` for the full
design, the move-offset mechanism this shipped alongside, and the file
list.
## 2026-09-04: texture pipeline rebuilt off the binding array
`Textures`/`TextureHandle`, `GlyphPrimitive`, and `UiRenderNode::new` all
changed shape. Why: the old pipeline bound every texture ever drawn in one
`binding_array<texture_2d<f32>>` and asked every device, unconditionally,
for `VK_EXT_descriptor_indexing` — a real share of Android GPUs lack it,
and it failed outright on the Android emulator's software Vulkan. See
TEXTURES.md's "Recommended shape" and "Implemented, 2026-09-04".
- **`UiRenderNode::new` drops its `limits: UiLimits` parameter, and
`UiLimits` is gone.** Before: `UiRenderNode::new(&device, &queue,
&config, UiLimits::default())`. After: `UiRenderNode::new(&device,
&queue, &config)`. Nothing replaces it — there are no more
binding-array limits to size.
- **`src/default/render.rs`'s device request asks for no features and no
binding-array limits.** Before: `required_features:
Features::TEXTURE_BINDING_ARRAY | Features::PARTIALLY_BOUND_BINDING_ARRAY
| Features::SAMPLED_TEXTURE_AND_STORAGE_BUFFER_ARRAY_NON_UNIFORM_INDEXING`
plus two `max_binding_array_*` limits. After: `Features::empty()` (the
`DeviceDescriptor` default) and only `max_buffer_size` set, which was
never about the binding array.
- **`TextureHandle` has no `primitive()` method any more**; a caller
outside `iris` shouldn't have been calling it (it fed the old renderer's
internals), but if something did: use `image_index()` for a standalone
image's bind-group index. There is no equivalent for a page — a page has
no bind group of its own now, see below.
- **`GlyphPrimitive` has no public constructor from a struct literal.**
Before: `GlyphPrimitive { uv_min, uv_max, view_idx, sampler_idx, color,
flags }`. After: `GlyphPrimitive::new(uv_min, uv_max, layer, color,
flags)` — one `layer` (the shared atlas array's layer) instead of a
`view_idx`/`sampler_idx` pair, since a page is now a layer of one array
texture rather than its own bound texture.
- **A widget author drawing images is unaffected**: `Painter::texture`/
`texture_at`/`texture_within` and `Textures::add` keep their signatures.
What changed underneath is that each standalone image now gets its own
`wgpu::BindGroup` and draw call instead of a slot in the shared array —
invisible from the widget API, visible only in `UiRenderNode`'s internals
and in `iris`'s device requirements.
-254
View File
@@ -1,254 +0,0 @@
# iris: known problems and things still to build
Iris's own list for the library, recorded 2026-09-04 in her words where it
matters, so the agents working through RUST.md pick these up in a sensible
order rather than rediscovering them. Each item says where it sits in the
order and what "done" looks like. Tick and date them in place.
## Fix
- [x] **Input does not fall through by input type (2026-09-04).**
`SensorUi::run_sensors` (`src/default/sense.rs`) used to set "consumed,
stop checking lower layers" from mere hover — a widget registered for
nothing but `click()` blocked a `Scroll` meant for whatever was behind
it, since "the cursor is over this widget" and "this widget handled the
event" were the same check. Fixed by judging consumption per input
kind: with no button transition and no scroll happening this frame
("momentary" activity), the topmost hovered widget still wins, same as
before; when something momentary *is* happening, only a widget whose
registered senses actually include a matching non-hover one (checked
via a new `TypeEventManager::registered`, which lists what a widget
registered without running anything) consumes it, so a widget with only
`Hovering`/click handlers can no longer block a scroll from reaching a
list underneath. `iris/src/sense_tests.rs` builds a button-over-a-list
`Stack` with a plain `HasEvents` impl (no GPU or window) and checks both
directions: a scroll over the button reaches the list, and a real click
still reaches the button — confirmed to fail on the pre-fix code and
pass after.
- [x] **Appending one image to an already-loaded list rebuilds every other
image's bind group (2026-09-05, fixed 2026-09-05).** Found by the
benchmark below: `GpuTextures::update` (`core/src/render/texture.rs`)
triggered `rebuild_image_bind_groups` — a loop over *every live
standalone image*, rebuilding its `BindGroup` — whenever the shared
`masks` or `move_offsets` GPU buffer was resized (`masks_resized ||
moves_resized` in `UiRenderNode::update`, `core/src/render/mod.rs`), and
a widget getting its *first* move-offset slot (LAYOUT.md section 2 —
every widget gets one on first draw) could be exactly what grows that
buffer. So one new message with one new image, appended to a transcript
that already has N images loaded, did not cost O(1): it cost one
`create_image` for the new image plus one `make_image_bind_group` per
*existing* image, because the new widget's own move slot pushed the
arena past its capacity. Measured directly in
`iris/examples/bench_images.rs`: appending a 1,001st image to 1,000
already-settled ones reported **1,001** bind-group creates for that one
frame, not 1 (`./run-bench.sh images`, frame 5 in the transcript below).
**Fix**: `masks`/`move_offsets` never belonged in a standalone image's own
bind group (group 2) in the first place — the group also holds that
image's own texture view, which is the only thing that is genuinely
per-image, so a buffer shared by *everything* forced a rebuild of
*every* group the moment it moved. Gave masks/move_offsets their own
bind group (group 3 in `shader.wgsl` and `UiRenderNode`: `masks_layout`/
`masks_group`), bound once per frame in `UiRenderNode::draw` rather than
once per draw call, instead of duplicating them into every per-image
group. `GpuTextures` and its image bind groups now know nothing about
either buffer — `rebuild_image_bind_groups` is called only from
`grow_array` (the atlas array texture growing, which genuinely does
change what every image's own bind group must reference) — so a
masks/move_offsets resize now touches exactly one bind group, ever,
regardless of how many images are live. Numbers after the fix, same
benchmark and command:
./run-bench.sh images
frame=1 bind_group_creates=1000 (cold load, unchanged)
frame=2 bind_group_creates=0 (was 1000 -- see the item below)
frame=3 bind_group_creates=0
frame=4 bind_group_creates=0
(append one image here)
frame=5 bind_group_creates=1 (was 1001)
frame=6 bind_group_creates=0
`run-headless.sh tabs --shot` still 27266 bytes, byte-for-byte unchanged,
confirming the bind-group restructuring changed nothing about what is
drawn.
- [x] **Bind-group creation takes two frames to reach the steady state, not
one (2026-09-05, closed by the fix above, 2026-09-05).** Same benchmark:
loading 1,000 images cold used to report 1,000 creates on frame 1
(expected — `create_image`, one per new image) *and again* 1,000 on
frame 2, before settling to 0 from frame 3. This was `rebuild_image_bind_groups`
firing a second time for the same masks/move-offsets buffer-growth
reason as the item above, confirming the guess recorded here — the two
were exactly the same root cause measured two different ways. Frame 2
now reports 0 (see the numbers above); not a separate fix.
## Build
- [x] **Benchmarks**, not unit tests, run on demand (2026-09-05; a
`benches/` or a script under `iris/`, never in `cargo test`). The
scenario that matters most is a **message list** — chat apps and this
app's transcript alike — stressed with many messages and many images.
One case in particular: **resizing an input box** (typing enough text to
grow it) that pushes a long list of messages above it must stay very
fast and recalculate almost nothing — a move of everything above, not a
re-layout. That is exactly the O(1) move chain in LAYOUT.md; the
benchmark is what proves it. Done when the numbers are in this file with
the command, and the input-box case reports draws re-run, not just frame
time.
**Built as two rigs**, chosen per scenario by whether a real `wgpu`
device is needed (`UiRenderState`/`Widgets` touch no GPU or window, so
most of this runs as an ordinary binary — the same property
`layout_tests.rs` relies on):
- `iris/benches/message_list.rs` — a plain `Instant`-timed binary
(`[[bench]] harness = false` in `iris/Cargo.toml`), not criterion: see
the file's own header for why (short version — every scenario here
reduces to a *count* `UiRenderState::take_counters` already produces,
which criterion's statistical machinery adds nothing to and which a
new dependency is not worth pulling in for). Covers (a) first-frame
cost of a message list of N wrapped-text rows (one in 20 also carrying
a small in-memory image) for N = 100/1,000/10,000; (b) per-frame cost
of scrolling that list, 200 ticks; (c) the input-box case — a
fixed-height field at the bottom of the screen growing by a line 40
times, with the message list above it filling the rest of the screen.
Run: `cd iris && cargo bench --bench message_list` (always release —
`cargo bench` builds the `bench` profile, which is optimized).
- `iris/examples/bench_images.rs` — needs a real device, so it runs
through `iris/run-headless.sh bench_images`, printing
`UiRenderNode::take_image_bind_group_creates()` (a new counter, added
in `core/src/render/texture.rs` and `core/src/render/mod.rs`,
mirroring `UiRenderState::take_counters`) each frame. Covers (d): 1,000
image rows, checked both cold (does bind-group creation reach zero
once loaded) and after appending one more image once settled (does
*that* stay cheap) — the second question is what actually matters for
a live transcript and is what turned up the two Fix items above.
- `iris/run-bench.sh [list|images]` runs either or both and is what to
run before/after touching `Scroll`, `Span`, `Sized`, the move-offset
chain, or `GpuTextures`.
**Numbers (2026-09-05, release, `cargo bench`/`run-headless.sh`, this
VM: AMD Ryzen 7 3800X, 8 cores, rustc 1.98.0 nightly-2026-09-03):**
cd iris && cargo bench --bench message_list
(a) first frame, N=100: 30.30ms draws=227 rewrites=15 moves=0
(a) first frame, N=1000: 186.04ms draws=2252 rewrites=150 moves=0
(a) first frame, N=10000:1770.36ms draws=22502 rewrites=1500 moves=0
(b) scroll, N=100/1000/10000, 200 ticks each:
draws=200 rewrites=0 moves=200 (identical at every N)
per-tick average: 0.0002ms (identical at every N)
(c) input grows 40 lines, N=100/1000/10000 rows above it:
draws=320 rewrites=40 moves=160 (identical at every N)
per-line average: 0.0012-0.0013ms (identical at every N)
cd iris && ./run-bench.sh images (2026-09-05, before the fix)
frame=1 bind_group_creates=1000 (cold load)
frame=2 bind_group_creates=1000 (see Fix item above)
frame=3 bind_group_creates=0
frame=4 bind_group_creates=0
(append one image here)
frame=5 bind_group_creates=1001 (see Fix item above)
frame=6 bind_group_creates=0
cd iris && ./run-bench.sh images (2026-09-05, after the fix)
frame=1 bind_group_creates=1000 (cold load, unchanged -- genuine work)
frame=2 bind_group_creates=0
frame=3 bind_group_creates=0
frame=4 bind_group_creates=0
(append one image here)
frame=5 bind_group_creates=1 (one image's own create_image, O(1))
frame=6 bind_group_creates=0
**Reading it**: (a) is real, necessary work — shaping and laying out N
never-before-seen text rows — and scales with N as it must, ~10x cost
per 10x N. (b) and (c) are the pass conditions that matter: both are
**exactly flat across N = 100 to 10,000**, confirming LAYOUT.md's O(1)
move chain holds for both scrolling and for a growing input box pushing
the message list — draws/moves per tick or per line do not grow with
list size, and the per-operation cost (a fraction of a microsecond) is
nowhere near a frame budget. (d)'s cold-load and steady-state halves
behave as designed; its *append* half did not, until the fix above moved
masks/move_offsets out of the per-image bind group — now flat at O(1)
the same way (b) and (c) are.
- **I5's transcript screen (`iris/transcript-ui/`, 2026-09-05) — what it
left, each recorded at the point in the code it would go rather than
silently dropped. See RUST.md's I5 box for the full account of what
*was* built (the screen, `SpanStyle`, cross-row selection, the growing
composer).**
- [ ] **Android integration for this screen does not exist yet.** No
cdylib/Gradle shell the way `iris-android-app` wraps `tabs-ui` (I2),
so `transcript-bench.sh`'s render-number pass condition against the
Compose baseline cannot be run. Needs: real `client-core::ApiClient`/
`event_stream::follow_session_events` wiring against
`app/ui-sandbox.sh --delay` (this crate deliberately fetches nothing
itself, `transcript-ui/src/lib.rs`'s doc), a new cdylib + Gradle
module, then the bench script pointed at it.
- [x] **Touch-drag panning over a row's own rendered text — done,
2026-09-05.** `row.rs` used to register `CursorSense::click_or_drag()`
on each row's `TextEdit` for cross-row selection; `TextEdit::draw`'s
`painter.child_layer()` (`iris/src/widget/text/edit.rs:87`) meant that
registration won `core/src/sense.rs::run_sensors`'s per-layer
arbitration on every frame it was pressed, not just the frame the
press started, so a list pan gesture registered on `List` itself never
got a turn while a row was under the finger. Fixed with
`iris::sense::DragArbiter` (recorded in `IRIS.md`), one small state
machine per list deciding pan vs. select the way Android does (a
vertical drag pans immediately; a stationary press held `LONG_PRESS`
(500ms) starts a selection which further drag extends; a horizontal
drag while something is already selected extends immediately) —
`transcript-ui/src/selection.rs`'s `Selection::drag` is the one place
every row's drag now routes through. 8 new unit tests
(`iris/src/sense.rs`'s `drag_arbiter_tests`); `cargo fmt/clippy/test
--workspace` and `cargo ndk` (both `iris` and `transcript-ui`) all
clean; `run-headless.sh` screenshot byte-identical to before the
change (38578 bytes). See RUST.md's I5 box, "Gap closed, 2026-09-05".
- [ ] **Row-level accessibility names.** The composer carries
`.label("Message")`; transcript rows do not carry a `.label()` of
their own yet, so `Widgets::named()` (I4) does not include them —
`row.rs`'s `build_text_row` is where one would go, keyed to something
stable per row (its sender + a short excerpt, matching what a screen
reader announcing a chat message would say).
- [ ] **A tappable link and a background chip behind inline code.**
Both need per-range glyph geometry that `TextEditCtx` does not expose
outside `iris::widget::text` (`edit.rs`'s `layout()` helper is
private) — see `markdown.rs`'s module doc for the exact shape the fix
would take (the same primitive `TextEdit::draw`'s own selection
highlight already uses internally,
`iris/src/widget/text/edit.rs:99`).
- [ ] **`Selection`'s anchor-row shortcut.** The row a drag started in
is selected in full (`select_all`) the moment the drag leaves it,
rather than "from the click point to whichever edge points away from
the drag" — needs the same private `layout()` access as the item
above. `selection.rs`'s module doc has the exact reasoning.
- [ ] **No syntax highlighting inside a fenced code block.**
`client_core::highlight` exists (built for the file explorer) and
could feed per-token `SpanStyle`s into a code block's span; wiring it
in was not attempted this pass.
- [ ] **Masks defined relative to each other.** Wanted: mask A multiplies
by something *and also* applies mask B — a mask can reference a parent
mask, the way the move chain references a parent offset. Today masks
are independent regions. Design it beside the move chain (same shape:
a parent index and a bounded walk in the shader); do it when a real
widget needs it, not before.
- [ ] **Positions as a single float per scroll.** Iris raised, and half
rejected, letting a scroll update one float rather than positions:
input handling cares about most elements in a list, so absolute
positions must be computed on the CPU anyway. LAYOUT.md's design
already lands here (GPU walks the chain, CPU resolves on demand for
hit tests). Keep the CPU resolution lazy and per query; do not
materialise every row's absolute position per frame.
- [ ] **Animations, last.** Cosmetic, so after everything above. Must be
**modular — a piece of the library rather than a core part forced into
everything, the same way input is**. Whatever the mechanism, a widget
that does not animate must pay nothing and import nothing for it.
## Reconsider
- [ ] **`WidgetView`.** Iris is unsure of it: what she wants is an easy way
to compose a widget from others (a button is the main case). With
sizing folded into `draw`, composing may be easy enough that `View` is
redundant. Decide after the layout change lands, by writing a button
both ways and keeping the one that is shorter to explain; delete the
other rather than keeping two ways.
-2632
View File
File diff suppressed because it is too large. Load diff
+141
View File
@@ -0,0 +1,141 @@
# Subagents
A session's subagents -- the helpers a Claude Code session starts through its
Task tool -- each get a transcript of their own, listed under the session's
card and readable in the same transcript view the session has. Designed
2026-09-05; the decisions Bryan has not yet reviewed are in `DECISIONS.md`.
## What a subagent is here
**A subagent is a second transcript owned by a session, in the same event
model, with no process and no controls.** It is not a session: it cannot be
messaged, stopped or started, and it has no setup, model or usage of its
own. Everything it shares with a session -- the transcript file format, the
paging routes, the SSE stream, the phone's cache and rendering -- is reused
by addressing, not by copying.
The CLI reports a subagent's messages on the parent's own stream-json
output, each carrying `parent_tool_use_id` = the id of the Task `tool_use`
that started it. Before this the translator dropped those lines
(`subagent_events_are_not_duplicated_into_the_transcript`); now it routes
them to that subagent's own translator and transcript. The parent's
transcript still shows only the Task call itself.
## Storage
Under the session directory:
```
<session>/subagents/<tool_use_id>/meta.json {title, created}
<session>/subagents/<tool_use_id>/transcript.jsonl same SeqEvent lines as the session's
```
The id is the Task tool_use id (`toolu_…`), which is unique, stable across a
backend restart, and already the key everything on the parent side uses.
Only ids matching `[A-Za-z0-9_-]+` are ever created or looked up, since the
id becomes a path.
The transcript's sequence numbers are its own, starting at 1. `Transcript`,
`read_window`, `catch_up` and `read_after` work on it unchanged.
Its path out: deleting the session deletes its directory, subagents included.
There is no separate delete.
## Lifecycle, as events in the subagent's transcript
1. Created on the first child line for an unseen parent id (or, when the
parent Task call was seen, at that call). First lines written:
`Status Running`, then `UserMessage { text: <the Task's prompt> }` when
the prompt is known -- it genuinely is the subagent's first user turn.
2. Every child line is translated by that subagent's own `Translator`
(one per subagent: tool ids are unique but streaming deltas are by
content-block index, and parallel subagents interleave).
3. **The parent's `tool_result` never finishes a subagent.** The Task tool
runs in the background by default: the `tool_result` -- "Async agent
launched..." -- arrives the moment it *starts*, while the subagent goes
on working for however long its own turn takes, sometimes minutes. What
ends it is its own turn ending: the raw API's `message_delta` on its
stream carrying `stop_reason: "end_turn"` (a `stop_reason` of `tool_use`
is the model about to call one, not an end), or a `result` line for its
own turn if a future CLI version ever sends one. Either maps to
`Status Exited`; the subagent's vocabulary has no `Idle`, so the
equivalent event `dispatch` produces for an ordinary session is dropped
rather than written. A shipped version of this finished on the
`tool_result` instead, which read a running background agent as
"finished" with its transcript truncated at the moment it launched.
4. **A child line for a subagent that already finished reopens it**
(`Status Running`) rather than being dropped: a background Task can be
sent another message long after its first turn ended, and that is
exactly what a further line for it means. Same transcript, same child
`Translator`, just picking back up.
5. When the parent session's process exits (`Status Exited` on the
session), every subagent still `Running` gets `Status Exited` too: its
process was the parent's.
A subagent that was mid-flight when the backend restarted keeps working:
the registry reopens the existing transcript on the next child line, and
the file continues its sequence -- the same reopening #4 describes, whether
what closed it was a restart or its own `end_turn`. If its turn ended while
the backend was down nothing recorded that until the next line arrives, so
its last status stays `Running`, which the list reports as **unknown**
rather than as running (see the wire shape) until then.
Title: the Task call's `description` input, then ` (<subagent_type>)` when
one is given; falling back to the tool's name when the child arrives before
(or without) the parent call being seen.
## Server layout
- `session/subagent.rs` -- the registry: `Subagents` (per session, in
`Shared`), `Subagent` (its `Transcript` behind a mutex plus a
`broadcast::Sender<SeqEvent>`), `record(id, event)`, `start(id, title,
prompt)`, `finish(id)`, `reopen(id)`, `finish_all()`, `list()` from disk. Drivers get an
`Arc<Subagents>` beside their `EventSink`; llama ignores it.
- `session/claude/translate.rs` -- routes child lines by parent id, holds
one child `Translator` per subagent, remembers pending Task calls'
description/prompt/subagent_type.
- `session/echo.rs` -- `/subagent [n]`: the test rig. Starts *n* (default 1)
subagents at once, each named "helper k". Each writes the prompt as its
user message, streams a few words of text, runs one `Bash` tool call, then
finishes about three seconds after starting, and the parent's Task calls
end when their subagent does. Three seconds so the running state can be
seen on the phone.
- `routes.rs` -- three routes, in the doc table.
## Wire shape
```
GET /sessions/{id} SessionInfo gains `subagents: N` (count, 0 when none)
GET /sessions same field on each row
GET /sessions/{id}/subagents [{id, title, status, created, lastActivity}], oldest first
GET /sessions/{id}/subagents/{sub}/transcript exactly the session transcript's query and answer
GET /sessions/{id}/subagents/{sub}/events?after=N exactly the session events stream
```
`status` is the transcript's last `Status` event, serialised like a session's
(`running`, `exited`), except that a subagent whose session is not itself
running cannot be running: the list answers `unknown` for that one. The
phone words these as *running*, *finished* and *unknown* on the subcard.
The count on `SessionInfo` is a directory listing, so the list stays cheap.
The per-subagent status is only read when the list route is asked for.
## Phone
- `SessionSummary.subagents: Int`. A card with a non-zero count ends in an
expander row -- a full-width `Chevron(Pointing.Down)` row that flips to
`Pointing.Up` -- collapsed by default. Expanding fetches
`/sessions/{id}/subagents` and draws one `OutlinedCard` per subagent,
indented inside the session card, the way dev-updater draws a project's
components: title, then the status word and a relative time. The
expansion state is per session id and survives a refresh of the list.
- Tapping a subcard opens `Screen.Subagent`, which is `SessionScreen` in
**read-only** form: the same transcript, paging, cache, selection,
images and status row, with the composer, the process button, the model
picker, the files button, the settings cog and the usage bar left out.
The header shows the subagent's title with the session's title beneath
it. Back returns to the list.
- Addressing: `fetchTranscript`, `EventStream`, `TranscriptSource` and the
cache take a transcript address rather than a session id --
`sessions/{id}` or `sessions/{id}/subagents/{sub}` -- so the cache nests a
subagent's copy under its session's and the same code serves both.
+55 -6
View File
@@ -50,6 +50,12 @@ version = "0.23.1"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "ac07cdecf99051d9a5238b80f35af32cdeba5b336e55d957b318b50137e18da5"
[[package]]
name = "bitflags"
version = "2.13.1"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "b588b76d00fde79687d7646a9b5bdf3cc0f655e0bbd080335a95d7e96f3587da"
[[package]]
name = "bytes"
version = "1.12.1"
@@ -76,7 +82,10 @@ checksum = "9330f8b2ff13f34540b44e946ef35111825727b38d33286ef986142615121801"
name = "client-core"
version = "0.1.0"
dependencies = [
"base64",
"event-model",
"log",
"pulldown-cmark",
"serde",
"serde_json",
"ureq",
@@ -206,6 +215,15 @@ dependencies = [
"percent-encoding",
]
[[package]]
name = "getopts"
version = "0.2.24"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "cfe4fbac503b8d1f88e6676011885f34b7174f46e59956bba534ba83abded4df"
dependencies = [
"unicode-width",
]
[[package]]
name = "getrandom"
version = "0.2.17"
@@ -490,6 +508,25 @@ dependencies = [
"unicode-ident",
]
[[package]]
name = "pulldown-cmark"
version = "0.13.4"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "e9f068eba8e7071c5f9511831b44f32c740d5adf574e990f946ddb53db2f314e"
dependencies = [
"bitflags",
"getopts",
"memchr",
"pulldown-cmark-escape",
"unicase",
]
[[package]]
name = "pulldown-cmark-escape"
version = "0.11.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "007d8adb5ddab6f8e3f491ac63566a7d5002cc7ed73901f72057943fa71ae1ae"
[[package]]
name = "quote"
version = "1.0.47"
@@ -553,9 +590,9 @@ dependencies = [
[[package]]
name = "rustls"
version = "0.23.43"
version = "0.23.44"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "0283386ce02abc0151e1761d08802dfe86c173b0b494af5cbc086574e453da06"
checksum = "6725596c3f2c3a0aef021139e145d4eafe314a6623e4680ca83852b2c67ab2ba"
dependencies = [
"log",
"once_cell",
@@ -783,12 +820,24 @@ dependencies = [
"zerovec",
]
[[package]]
name = "unicase"
version = "2.9.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "dbc4bc3a9f746d862c45cb89d705aa10f187bb96c76001afab07a0d35ce60142"
[[package]]
name = "unicode-ident"
version = "1.0.24"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "e6e4313cd5fcd3dad5cafa179702e2b244f760991f45397d14d4ebf38247da75"
[[package]]
name = "unicode-width"
version = "0.2.2"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "b4ac048d71ede7ee76d585517add45da530660ef4390e49b098733c6e897f254"
[[package]]
name = "untrusted"
version = "0.9.0"
@@ -797,9 +846,9 @@ checksum = "8ecb6da28b8a351d773b68d5825ac39017e680750f980f3a1a85cd8dd28a47c1"
[[package]]
name = "ureq"
version = "3.4.0"
version = "3.4.1"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "972d7902c8735f2695410b8aed7df6ed12a47394aa1c8d7af49f0497b731a94d"
checksum = "af5546be8f5378d5414f83733f5c9a2526f4645829edbc1c41790aeef1b38e8b"
dependencies = [
"base64",
"cookie_store",
@@ -817,9 +866,9 @@ dependencies = [
[[package]]
name = "ureq-proto"
version = "0.6.1"
version = "0.6.2"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "da5f78b09e6941e1a0f2e30e695e4b120377b54d5e0aec11b594bb57b3971613"
checksum = "fabc3e92916c89c95b20eef7b06b00b066bc217ef9ea3a4ac9bf1a7e35261e10"
dependencies = [
"base64",
"http",
+39
View File
@@ -102,6 +102,16 @@ android {
targetSdk = 37
versionCode = 1
versionName = "1.0"
// Read by MainActivity to decide, at startup, whether this is the P0 benchmark build
// (docs/RUST.md's P0 box) rather than the app somebody enrolled. False everywhere except
// the `bench` build type below, which overrides it.
buildConfigField("boolean", "FIXTURE_MODE", "false")
}
buildFeatures {
// Only for FIXTURE_MODE above; nothing else here reaches for generated BuildConfig fields.
buildConfig = true
// Only for the bench build type's resValue("string", "app_name", ...) below.
resValues = true
}
packaging {
resources { excludes += "/META-INF/{AL2.0,LGPL2.1}" }
@@ -131,6 +141,35 @@ android {
isMinifyEnabled = false
if (keystore != null) signingConfig = signingConfigs.getByName("release")
}
// P0's benchmark build (docs/RUST.md, docs/DECISIONS.md's 2026-09-05 entry): release
// optimisations so a frame time measured here means what release means everywhere else in
// this project, its own application id so it installs beside a real enrollment rather than
// replacing it, and FIXTURE_MODE so MainActivity opens straight onto the fixture session
// instead of asking to be enrolled. Signed with the same key as release -- it never talks
// to a real backend, so there is no CA of its own to mismatch, and a second keystore would
// be one more secret to keep off this machine's shared mount for no benefit.
create("bench") {
initWith(getByName("release"))
// :link (wg-app-link) has no "bench" build type of its own -- it is a library shared
// with dev-updater and has no reason to know this project invented one -- so this says
// which of its build types to link against instead.
matchingFallbacks += listOf("release")
applicationIdSuffix = ".bench"
// "AI Sessions bench" everywhere the OS shows the app's name (launcher, recents,
// Settings): this resValue overrides res/values/strings.xml's app_name for this
// build type alone, and AndroidManifest.xml's android:label reads @string/app_name
// rather than a literal so a build type can override it without touching the
// manifest.
resValue("string", "app_name", "AI Sessions bench")
buildConfigField("boolean", "FIXTURE_MODE", "true")
if (keystore != null) signingConfig = signingConfigs.getByName("release")
}
}
sourceSets {
// The fixture both bench builds (this one and iris's) open with; see
// app/bench-fixture/README.md. Read directly from its own directory rather than copied
// into androidApp/src -- one file to keep in sync with the generator, not two.
getByName("bench").assets.directories.add("../bench-fixture/assets")
}
compileOptions {
sourceCompatibility = JavaVersion.VERSION_21
+1 -1
View File
@@ -25,7 +25,7 @@
the fix is a judgement about how this app should look. Drop this
suppression when a real icon lands. -->
<application
android:label="AI Sessions"
android:label="@string/app_name"
android:allowBackup="true"
android:theme="@android:style/Theme.Material.Light.NoActionBar"
tools:ignore="MissingApplicationIcon">
@@ -142,6 +142,20 @@ data class SessionSummary(
val keepsOwnTranscript: Boolean,
/** How much the session asks before acting; null when it was never set. */
val permissionMode: String?,
/**
* How hard the model thinks, or null for the CLI's own default.
*
* Null is a level somebody can choose, not only one to start in -- see [EFFORT_LEVELS]. It is
* reported rather than assumed for the same reason [permissionMode] is.
*/
val effort: String?,
/**
* Whether a thinking level does anything here -- a Claude CLI session, not a llama or echo one.
*
* Asked of the server rather than worked out from the provider's name, because this is a
* property of the driver's *kind* and the phone only has the name.
*/
val takesEffort: Boolean,
/**
* Whether this continues a session the machine already had, which changes what deleting means.
*/
@@ -153,6 +167,24 @@ data class SessionSummary(
* itself from a default is one you can turn off while believing you are reading it.
*/
val notify: Boolean,
/**
* Whether this session sends itself a message once its account's usage limit lifts, and what
* that message says.
*
* The message is what the server would actually send, with its own default already filled in,
* so the field shows the words rather than an empty box standing for them.
*/
val autoResume: Boolean,
val autoResumeMessage: String,
/**
* When the server next intends to check whether the limit has lifted, in epoch seconds, or null
* when nothing is waiting.
*
* A time to *ask*, not a time to resume: the server checks the meter at that moment and waits
* again if the limit is still on. Worded that way wherever it is shown, because a promise this
* app cannot keep is worse than no time at all.
*/
val resumeAt: Double?,
/**
* The directory the session works in, or null where it was never given one.
*
@@ -178,8 +210,27 @@ data class SessionSummary(
* server because that is where a provider's kind is known.
*/
val maxImageEdge: Int?,
/**
* Which of `GET /usage`'s snapshots is about this session, and null where nothing meters it.
*
* The rate-limit bar answers a question about an *account*, and what decides which account --
* if any -- is the provider this session runs, not the machine it runs on. Pairing by machine
* alone drew the Claude CLI's five-hour window under every echo session on a machine that also
* has the CLI: a quota that session cannot spend and could never run down. Decided by the
* server for the same reason [maxImageEdge] is -- it is a fact about the provider's kind, and
* this app has only its name.
*/
val usageProvider: String?,
val status: String,
val lastActivity: Double,
/**
* How many subagents this session has, however their own status now reads.
*
* A directory listing on the server rather than a status read per subagent, so the list stays
* cheap; the per-subagent state is only fetched when the card is expanded. Zero on a server
* that predates subagents, so this app still opens against one.
*/
val subagents: Int,
)
private fun parseSession(session: JSONObject) =
@@ -192,14 +243,24 @@ private fun parseSession(session: JSONObject) =
title = session.getString("title"),
model = session.optString("model").ifEmpty { null },
permissionMode = session.optString("permissionMode").ifEmpty { null },
effort = session.optString("effort").ifEmpty { null },
takesEffort = session.optBoolean("takesEffort", false),
imported = session.optBoolean("imported", false),
notify = session.optBoolean("notify", true),
autoResume = session.optBoolean("autoResume", false),
// The server sends its own default rather than nothing, so an empty answer means an older
// server -- and this app's word for it is the same word.
autoResumeMessage =
session.optString("autoResumeMessage").ifEmpty { DEFAULT_RESUME_MESSAGE },
resumeAt = if (session.has("resumeAt")) session.getDouble("resumeAt") else null,
cwd = session.optString("cwd").ifEmpty { null },
contextTokens =
if (session.has("contextTokens")) session.getLong("contextTokens") else null,
maxImageEdge = session.optInt("maxImageEdge", 0).takeIf { it > 0 },
usageProvider = session.optString("usageProvider").ifEmpty { null },
status = session.getString("status"),
lastActivity = session.getDouble("lastActivity"),
subagents = session.optInt("subagents", 0),
)
fun fetchSessions(settings: ServerSettings): List<SessionSummary> =
@@ -215,6 +276,35 @@ fun fetchSessions(settings: ServerSettings): List<SessionSummary> =
fun fetchSession(settings: ServerSettings, sessionId: String): SessionSummary =
requestFromServer(settings, "/sessions/$sessionId") { parseSession(it.jsonObject()) }
/**
* One row of `GET /sessions/{id}/subagents`, oldest first.
*
* A subagent is a second transcript owned by a session -- no process, no controls of its own -- so
* this carries only what a card needs to draw and to open it; see SUBAGENTS.md. [status] is
* "running", "exited" or "unknown": a subagent whose session is not itself running cannot be
* running, and the list says so rather than reporting a state that cannot hold.
*/
data class SubagentSummary(
val id: String,
val title: String,
val status: String,
val created: Double,
val lastActivity: Double,
)
fun fetchSubagents(settings: ServerSettings, sessionId: String): List<SubagentSummary> =
requestFromServer(settings, "/sessions/$sessionId/subagents") {
it.jsonObjects { row ->
SubagentSummary(
id = row.getString("id"),
title = row.getString("title"),
status = row.getString("status"),
created = row.getDouble("created"),
lastActivity = row.getDouble("lastActivity"),
)
}
}
// What the server offers, so the spawn screen has no hardcoded lists: a setup added to the server's
// config.ron appears here with no app rebuild.
//
@@ -375,6 +465,12 @@ data class SshDetails(
* Where files attached from here land on that machine; null for the session's own directory.
*/
val attachmentsDir: String? = null,
/**
* Where that machine keeps its GGUF models; null for the same place the backend keeps its own
* (`~/.local/share/ai-app/models`, read on that machine). A llama.cpp session serves the file
* from the machine it runs on, so this is where its models are looked for and listed.
*/
val modelsDir: String? = null,
)
private fun SshDetails.toJson() =
@@ -382,6 +478,7 @@ private fun SshDetails.toJson() =
if (port != null) put("port", port)
if (!identityFile.isNullOrBlank()) put("identityFile", identityFile)
if (!attachmentsDir.isNullOrBlank()) put("attachmentsDir", attachmentsDir)
if (!modelsDir.isNullOrBlank()) put("modelsDir", modelsDir)
}
/** What a machine turns out to have, without saving anything. */
@@ -450,6 +547,8 @@ fun spawnSession(
model: String? = null,
cwd: String? = null,
permissionMode: String? = null,
/** Null for whatever the server's default is; see [fetchDefaultEffort]. */
effort: String? = null,
params: Map<String, String> = emptyMap(),
/** Continue this Claude Code session instead of starting an empty one. */
import: String? = null,
@@ -467,6 +566,7 @@ fun spawnSession(
if (!model.isNullOrBlank()) put("model", model)
if (!cwd.isNullOrBlank()) put("cwd", cwd)
if (!permissionMode.isNullOrBlank()) put("permissionMode", permissionMode)
if (!effort.isNullOrBlank()) put("effort", effort)
if (!import.isNullOrBlank()) put("import", import)
if (params.isNotEmpty()) {
put("params", JSONObject(params.toMap<String, Any>()))
@@ -906,7 +1006,7 @@ fun startImport(
*/
fun fetchTranscript(
settings: ServerSettings,
sessionId: String,
address: TranscriptAddress,
before: Long? = null,
limit: Int = 80,
// Count [limit] in rows, not events, joining a reply's streamed deltas into one -- so a page of
@@ -925,7 +1025,7 @@ fun fetchTranscript(
if (coalesce) append("&coalesce=true")
if (after != null) append("&after=").append(after)
}
return requestFromServer(settings, "/sessions/$sessionId/transcript$query") { connection ->
return requestFromServer(settings, "/${address.urlPath}/transcript$query") { connection ->
val body = JSONArray(connection.inputStream.bufferedReader().readText())
// The text as well as the event: the transcript cache stores the one and the fold needs the
// other, and they have to be the same line.
@@ -973,6 +1073,56 @@ fun setSessionModel(settings: ServerSettings, sessionId: String, model: String)
*/
val PERMISSION_MODES = listOf("manual", "acceptEdits", "auto", "bypassPermissions", "plan")
/**
* What a new session's thinking level is when nothing chose one, or null for the CLI's own.
*
* Held by the server rather than by this phone, because a second device would otherwise spawn
* sessions at a level the first one's owner never picked.
*/
fun fetchDefaultEffort(settings: ServerSettings): String? =
requestFromServer(settings, "/defaults") {
it.jsonObject().optString("effort").ifEmpty { null }
}
/** Sets what new sessions start at. Nothing already running changes. */
fun setDefaultEffort(settings: ServerSettings, level: String?) {
requestFromServer(
settings,
"/defaults",
method = "POST",
jsonBody = JSONObject().put("effort", level ?: JSONObject.NULL).toString(),
) {}
}
/**
* How hard the model thinks, as `claude --effort` takes them, cheapest first.
*
* Not offered alongside the model and the permission mode on the session's own bar, because it does
* not behave like them: the CLI has a control request for those two and none for this (checked
* against 2.1.258), so a level is settled when the process is launched. Changing it therefore stops
* the process, which is what the working directory beside it in this dialog does, and why it is
* here rather than on a bar whose other controls take effect mid-turn.
*/
val EFFORT_LEVELS = listOf("low", "medium", "high", "xhigh", "max")
/** What the picker shows, and sends as null, for a session that has chosen no level. */
const val DEFAULT_EFFORT = "default"
/**
* Records how hard a session thinks and **stops its process**, since the level is read when the
* process is launched. The next message, or Start, runs one that has it.
*
* [level] is null for the CLI's own default.
*/
fun setSessionEffort(settings: ServerSettings, sessionId: String, level: String?) {
requestFromServer(
settings,
"/sessions/$sessionId/effort",
method = "POST",
jsonBody = JSONObject().put("effort", level ?: JSONObject.NULL).toString(),
) {}
}
/** Switches how much a running session asks before acting, also in place. */
fun setSessionPermissionMode(settings: ServerSettings, sessionId: String, mode: String) {
requestFromServer(
@@ -984,6 +1134,36 @@ fun setSessionPermissionMode(settings: ServerSettings, sessionId: String, mode:
}
/** Turns this session's notifications on or off. Stored on the backend -- see `SessionConfig`. */
/**
* What an auto-resume says when nothing else was typed. Mirrors the server's own default, so a
* cleared field shows the word that would actually be sent instead of going blank.
*/
const val DEFAULT_RESUME_MESSAGE = "continue"
/**
* Turns auto-resume on or off and sets what it would say, in one request because they are one
* decision -- see the server's `/sessions/{id}/auto-resume`.
*/
fun setSessionAutoResume(
settings: ServerSettings,
sessionId: String,
autoResume: Boolean,
message: String?,
) {
requestFromServer(
settings,
"/sessions/$sessionId/auto-resume",
method = "POST",
jsonBody =
JSONObject()
.put("autoResume", autoResume)
// Empty means the server's default rather than a session poked with nothing to
// read, which is the same rule the server applies to the field.
.put("message", message?.trim()?.ifEmpty { null } ?: JSONObject.NULL)
.toString(),
) {}
}
fun setSessionNotify(settings: ServerSettings, sessionId: String, notify: Boolean) {
requestFromServer(
settings,
@@ -1074,6 +1254,26 @@ private fun parseDownload(o: JSONObject) =
error = if (o.has("error")) o.getString("error") else null,
)
/**
* The models on one machine, which is the list a llama.cpp session there can choose from.
*
* Not [fetchModels], which is what the *backend* has downloaded. A session serves its model from
* the machine it runs on, so for a machine reached over ssh those are two different lists -- and
* offering the backend's would name files that are not there, turning a choice that cannot work
* into a session that fails when it tries to load one.
*/
fun fetchSetupModels(settings: ServerSettings, setupId: String): List<LocalModel> =
requestFromServer(settings, "/setups/${setupId.urlEncoded()}/models") { connection ->
JSONArray(connection.inputStream.bufferedReader().readText()).mapObjects { m ->
LocalModel(
key = m.getString("key"),
repo = m.getString("repo"),
file = m.getString("file"),
bytes = m.getLong("bytes"),
)
}
}
fun fetchModels(settings: ServerSettings): Models =
requestFromServer(settings, "/models") { connection ->
val body = JSONObject(connection.inputStream.bufferedReader().readText())
@@ -1,7 +1,9 @@
package com.example.aiapp
import androidx.activity.compose.BackHandler
import androidx.compose.foundation.background
import androidx.compose.foundation.layout.Box
import androidx.compose.foundation.layout.fillMaxSize
import androidx.compose.foundation.layout.imePadding
import androidx.compose.foundation.layout.padding
import androidx.compose.material3.AlertDialog
@@ -34,7 +36,23 @@ import kotlinx.coroutines.withContext
* session, spawning one, and settings.
*/
private sealed class Screen {
data object Main : Screen()
/**
* The session list, with a subagent's own transcript over it when [subagent] is set.
*
* A layer on this screen rather than a screen of its own, for the same reason [Session.files]
* is: [SessionListScreen] owns which cards are expanded and what each expansion fetched, kept
* in `remember`, and a subagent is opened from a card's expander. As a sibling `Screen` it was
* disposed and recreated on every return, which lost that state -- an expanded card collapsed
* itself the moment its own subagent's view was closed.
*/
data class Main(val subagent: SubagentTarget? = null) : Screen()
/**
* One subagent's own transcript, read-only. See [SessionScreen]'s `subagent` parameter and
* SUBAGENTS.md's "Phone". Closing it returns to [Main] under it, not to [Session]: a subagent
* is opened from the session list's card rather than from inside the session it belongs to.
*/
data class SubagentTarget(val summary: SessionSummary, val subagent: SubagentSummary)
/**
* One session, with the file explorer over it when [files] is set.
@@ -81,7 +99,7 @@ fun AppRoot(
val context = LocalContext.current
val scope = rememberCoroutineScope()
var settings by remember(settingsVersion) { mutableStateOf(loadServerSettings(context)) }
var screen by remember { mutableStateOf<Screen>(Screen.Main) }
var screen by remember { mutableStateOf<Screen>(Screen.Main()) }
// A notification tap this could not follow, and why. Null both before one is asked for and
// after one succeeds, since success is a screen rather than a message.
var failedOpen by remember { mutableStateOf<FailedOpen?>(null) }
@@ -96,7 +114,7 @@ fun AppRoot(
share = shareRequest
// A session already open takes it. Otherwise the list is where the choice is made,
// whatever screen was showing: Spawn and Settings have nowhere to put a file.
if (screen !is Screen.Session) screen = Screen.Main
if (screen !is Screen.Session) screen = Screen.Main()
}
}
@@ -123,7 +141,7 @@ fun AppRoot(
existing = null,
onSaved = { saved ->
settings = saved
screen = Screen.Main
screen = Screen.Main()
},
onBack = null,
)
@@ -136,7 +154,7 @@ fun AppRoot(
// shows, so it always refetches.
val goToMain = {
reloadToken++
screen = Screen.Main
screen = Screen.Main()
}
if (screen !is Screen.Main) {
BackHandler(onBack = goToMain)
@@ -185,6 +203,9 @@ fun AppRoot(
reloadToken = reloadToken,
share = share,
onOpen = { screen = Screen.Session(it) },
onOpenSubagent = { summary, subagent ->
screen = here.copy(subagent = Screen.SubagentTarget(summary, subagent))
},
onSpawn = { screen = Screen.Spawn },
onImported = { imported ->
reloadToken++
@@ -192,6 +213,27 @@ fun AppRoot(
},
onSettings = { screen = Screen.Settings },
)
// Its own back handler is registered after MainScreen's, so it is the one the
// platform asks first while a subagent is open -- the same rule the files
// explorer's handler follows over its session, below.
here.subagent?.let { target ->
BackHandler { screen = here.copy(subagent = null) }
// Its own opaque background: this screen was always the sole content under
// the theme's own Surface before, so it never had to paint one -- stacked over
// the list here, the space between its own cards let the list underneath show
// through without this. The same fix FilesScreen needed over its session.
Box(Modifier.fillMaxSize().background(MaterialTheme.colorScheme.background)) {
key(target.summary.id, target.subagent.id) {
SessionScreen(
settings = current,
summary = target.summary,
onBack = { screen = here.copy(subagent = null) },
onFiles = {},
subagent = target.subagent,
)
}
}
}
}
is Screen.Session ->
// Keyed on the id, because a different session is a different screen rather than this
@@ -0,0 +1,108 @@
package com.example.aiapp
import android.content.Context
import java.util.concurrent.CopyOnWriteArrayList
/**
* P0's benchmark gate (see docs/RUST.md and docs/DECISIONS.md's 2026-09-05 entry): an in-process
* fake of the backend, so the `bench` build type can drive a real session screen -- the real
* [TranscriptSource], the real fold, the real paging -- with no server and no network permission.
*
* Only ever installed when [BuildConfig.FIXTURE_MODE] is true (see [MainActivity]); everything else
* in this build compiles it in but never calls it, since Kotlin has no per-build-type source set
* that both [MainActivity] (which every variant compiles) and this can share without one.
*
* The design: [requestFromServer] and [Sse] talk to `https://$FIXTURE_HOST:$FIXTURE_PORT` through
* ordinary `java.net.URL`, exactly as they would talk to a real server. A
* [java.net.URLStreamHandlerFactory] registered once for the whole process intercepts every
* `https://` connection to that host and answers from this object's in-memory event log instead of
* opening a socket -- see BenchNetwork.kt. Everything above that (TranscriptSource, SessionScreen,
* the fold, uniqueItems) never learns the difference.
*/
object BenchFixture {
const val FIXTURE_HOST = "bench.fixture.invalid"
const val FIXTURE_PORT = 1
/** How many of the fixture's events are the opening backlog; see bench-fixture/README.md. */
private const val BACKLOG_COUNT = 3200
val settings = ServerSettings(FIXTURE_HOST, FIXTURE_PORT, "bench")
/** The session id every bench run opens; nothing else in this build ever mints one. */
const val SESSION_ID = "bench-fixture-session"
/**
* The whole transcript, seq order, growing as [pushLive] is called during the streaming phase.
* Read by both the REST page handler and the SSE handler, so a page requested mid- stream and a
* live frame agree on what has "already happened" -- the same thing a real server's own
* transcript file guarantees.
*/
private val log = CopyOnWriteArrayList<Pair<String, SeqEvent>>()
/** The events not yet appended to [log] -- the streaming phase's own source. */
private var streamTail: List<Pair<String, SeqEvent>> = emptyList()
private val images = mutableMapOf<String, ByteArray>()
@Volatile private var loaded = false
/**
* Parses the bundled fixture once. Safe to call more than once; only the first does anything.
*/
@Synchronized
fun ensureLoaded(context: Context) {
if (loaded) return
val lines =
context.assets.open("transcript.jsonl").bufferedReader().readLines().filter {
it.isNotBlank()
}
val parsed = lines.map { it to parseSeqEvent(it) }
log.addAll(parsed.take(BACKLOG_COUNT))
streamTail = parsed.drop(BACKLOG_COUNT)
for (name in listOf("bench1.png", "bench2.png")) {
images[name] = context.assets.open(name).readBytes()
}
loaded = true
}
/** The events the streaming phase has left to send. */
fun remainingStreamEvents(): Int = streamTail.size
/** Sends the next fixture event onto the live log, as a real SSE frame would arrive. */
fun pushNextLiveEvent(): Boolean {
val next = streamTail.firstOrNull() ?: return false
streamTail = streamTail.drop(1)
log.add(next)
return true
}
/** Undoes [pushNextLiveEvent] and reloads the opening backlog, for running the bench twice. */
@Synchronized
fun resetToBacklog(context: Context) {
loaded = false
log.clear()
ensureLoaded(context)
}
fun fileBytes(name: String): ByteArray? = images[name]
/**
* Raw JSON lines with seq > [after], in order -- what an `/events?after=` connection replays.
*/
fun linesAfter(after: Long): List<String> =
log.filter { it.second.seq > after }.map { it.first }
/**
* One REST page: [fetchTranscript]'s `before`/`limit`/`after`, against the growing log. Ignores
* `coalesce` -- the fixture's own deltas are already split the way a real reply streams, and
* what the benchmark exercises is the fold and the paging, not the server's row-joining, which
* client-core's own port tracks separately (CLIENT_CORE.md).
*/
fun page(before: Long?, limit: Int, after: Long?): List<String> {
val upper = before ?: (log.lastOrNull()?.second?.seq?.plus(1) ?: 1L)
val candidates = log.filter {
it.second.seq < upper && (after == null || it.second.seq > after)
}
return candidates.takeLast(limit).map { it.first }
}
}
@@ -0,0 +1,181 @@
package com.example.aiapp
import java.io.ByteArrayInputStream
import java.io.IOException
import java.io.InputStream
import java.io.PipedInputStream
import java.io.PipedOutputStream
import java.net.HttpURLConnection
import java.net.URL
import java.net.URLStreamHandler
import java.net.URLStreamHandlerFactory
import java.security.Principal
import java.security.cert.Certificate
import javax.net.ssl.HttpsURLConnection
import javax.net.ssl.SSLPeerUnverifiedException
import org.json.JSONArray
/**
* Installs the process-wide interception [BenchFixture] needs. Idempotent and safe to call more
* than once; the JDK only allows [URL.setURLStreamHandlerFactory] to be called successfully once
* per process, and a second real call throws -- so this guards it rather than relying on every
* caller to remember.
*
* Scoped to [BenchFixture.FIXTURE_HOST]: any other `https://` URL falls through to the platform's
* ordinary handler, so this only ever changes behaviour for the one host the bench build invents.
*/
@Synchronized
fun installFixtureNetworkOnce() {
if (installed) return
installed = true
URL.setURLStreamHandlerFactory(
URLStreamHandlerFactory { protocol ->
if (protocol != "https") null
else
object : URLStreamHandler() {
override fun openConnection(url: URL): HttpURLConnection =
if (url.host == BenchFixture.FIXTURE_HOST) FixtureConnection(url)
else
// The bench build makes no other https call -- this factory is
// installed only in FIXTURE_MODE (MainActivity) -- so there is
// deliberately no delegate to a platform handler here: once a
// URLStreamHandlerFactory is installed there is no supported way to
// ask the JDK for its own default handler back, and re-entering this
// same factory for the fallback would recurse forever rather than
// reach one.
throw java.io.IOException(
"bench build's fixture network has no route to https host " +
"${url.host} -- only ${BenchFixture.FIXTURE_HOST} is served"
)
}
}
)
}
private var installed = false
/**
* Answers one request against [BenchFixture] instead of opening a socket. Implements just enough of
* [HttpsURLConnection] for [requestFromServer] and [Sse] to work unmodified: both only call
* `connect`/`disconnect`, set a handful of request properties they never need answered, and read
* `responseCode` and `inputStream`.
*/
private class FixtureConnection(url: URL) : HttpsURLConnection(url) {
private var input: InputStream? = null
private var writer: Thread? = null
override fun connect() {
if (input != null) return
input = route(url.path, url.query)
}
override fun disconnect() {
writer?.interrupt()
try {
input?.close()
} catch (_: IOException) {}
}
override fun usingProxy() = false
override fun getResponseCode(): Int {
connect()
return 200
}
override fun getInputStream(): InputStream {
connect()
return input!!
}
override fun getErrorStream(): InputStream? = null
// Nothing here reads any of these; implemented only because HttpsURLConnection declares them
// abstract. A fixture never negotiates real TLS, so each says exactly that rather than
// fabricating a plausible-looking certificate.
override fun getCipherSuite() = "none (bench fixture, no TLS)"
override fun getLocalCertificates(): Array<Certificate>? = null
override fun getServerCertificates(): Array<Certificate> =
throw SSLPeerUnverifiedException("bench fixture connection presents no certificate")
override fun getPeerPrincipal(): Principal =
throw SSLPeerUnverifiedException("bench fixture connection presents no certificate")
override fun getLocalPrincipal(): Principal? = null
/**
* [path] is `/sessions/{id}/...`; everything else this build's fixture is asked for is a bug.
*/
private fun route(path: String, query: String?): InputStream {
val params =
(query ?: "")
.split("&")
.filter { it.contains('=') }
.associate {
val (k, v) = it.split("=", limit = 2)
k to java.net.URLDecoder.decode(v, "UTF-8")
}
return when {
path.endsWith("/transcript") -> {
val lines =
BenchFixture.page(
before = params["before"]?.toLongOrNull(),
limit = params["limit"]?.toIntOrNull() ?: 80,
after = params["after"]?.toLongOrNull(),
)
val body = JSONArray(lines.map { org.json.JSONObject(it) })
ByteArrayInputStream(body.toString().toByteArray())
}
path.endsWith("/events") -> openEventsStream(params["after"]?.toLongOrNull() ?: 0L)
path.contains("/files/") -> {
val name = path.substringAfterLast("/files/")
val bytes =
BenchFixture.fileBytes(name)
?: throw IOException("bench fixture has no file named $name")
ByteArrayInputStream(bytes)
}
else -> throw IOException("bench fixture has no route for $path")
}
}
/**
* A live SSE body: [BenchFixture.linesAfter] replayed immediately, then polled every 50ms for
* anything [BenchFixture.pushNextLiveEvent] has added since -- the same shape a real backend's
* backlog-then-follow gives [Sse], just polled instead of woken, which is a fixture's business
* rather than something worth a condition variable for.
*/
private fun openEventsStream(after: Long): InputStream {
val pipeIn = PipedInputStream(1 shl 16)
val pipeOut = PipedOutputStream(pipeIn)
var sent = after
val thread = Thread {
try {
while (!Thread.currentThread().isInterrupted) {
val fresh = BenchFixture.linesAfter(sent)
for (line in fresh) {
pipeOut.write("data: $line\n\n".toByteArray())
pipeOut.flush()
sent = org.json.JSONObject(line).getLong("seq")
}
Thread.sleep(50)
}
} catch (_: InterruptedException) {
// disconnect() -- the ordinary way this ends.
} catch (_: IOException) {
// The reader side (Sse) closed its end.
} finally {
try {
pipeOut.close()
} catch (_: IOException) {}
}
}
.also {
it.isDaemon = true
it.start()
}
writer = thread
return pipeIn
}
}
@@ -0,0 +1,325 @@
package com.example.aiapp
import android.content.Context
import android.os.BatteryManager
import android.os.Process
import android.view.View
import androidx.compose.foundation.gestures.FlingBehavior
import androidx.compose.foundation.lazy.LazyListState
import androidx.compose.ui.focus.FocusRequester
import androidx.core.view.ViewCompat
import androidx.core.view.WindowInsetsCompat
import androidx.core.view.WindowInsetsControllerCompat
import java.io.File
import kotlinx.coroutines.CoroutineScope
import kotlinx.coroutines.delay
import kotlinx.coroutines.isActive
import kotlinx.coroutines.launch
/**
* P0's scripted benchmark, run in-process instead of by a shell script: the phone has no usable
* system tracing (this-machine-android's skill) and no agent can drive it, so the same scroll loop
* and streaming phase `transcript-bench.sh`/`stream-bench.sh` drive over `ui-trace` are reproduced
* here against [LazyListState] and [BenchFixture] directly. Only reachable from the `bench` build
* (see [SessionSettingsDialog]'s `onRunBenchmark`), but compiled into every build for the reason
* [BenchFixture]'s doc comment gives.
*
* **v2 (2026-09-06)**, asked for by Iris because the v1 fling was too gentle to stress-test the
* scroll path and said nothing about typing or the keyboard. Four phases now, each a slice of the
* same [FrameStats] recording ([FrameStats.markPhase]/[FrameStats.phaseLines] -- one recorder, not
* two): **fling** (real `FlingBehavior`, not `animateScrollBy`), **stream** (unchanged from v1),
* **type** (600 fixed characters into the real composer `TextFieldValue`, then deleted), and
* **keyboard** (five show/hide cycles). The exact constants below are also written into
* `docs/RUST.md`'s P0 box, "Benchmark v2 (2026-09-06)", so the iris half implements the identical
* spec -- changing a number here without updating that box makes the two apps measure different
* things while looking like the same benchmark.
*/
object BenchRun {
/** transcript-bench.sh's default: 6 cycles of 4 swipes each, kept as the pre-v2 comparison. */
private const val CYCLES = 6
private const val SWIPE_PX = 900f
private const val SWIPE_MS = 200
private const val SWIPE_PAUSE_MS = 500L
/**
* Fling phase (v2): a real fling through the list's own [FlingBehavior], not `animateScrollBy`
* -- Iris's ask was that it "travel way faster" than the old tween-based swipe, and a tween can
* never exceed the distance it is told to cover in the time it is given, while a real fling
* decays from an initial velocity the way a finger flick does. 12,000 px/s is roughly a hard,
* fast flick on a ~420dp/in device (about 30 dp/ms-equivalent initial speed); chosen well above
* the ~4,500 px/s a moderate `animateScrollBy` swipe implies, so this phase exercises the fast
* end of what the platform's fling decay produces rather than the gentle one v1 measured.
*/
private const val FLING_VELOCITY_PX_S = 12_000f
private const val FLING_COUNT = 8
private const val FLING_SETTLE_CAP_MS = 3_000L
private const val FLING_PAUSE_MS = 300L
/** stream-bench.sh's shape: a real reply arrives as many small deltas, not one big write. */
private const val STREAM_EVENTS_PER_SEC = 20
private const val STREAM_SECONDS = 20
/**
* Type phase (v2): sentences built from long, multisyllabic words so the composer actually
* wraps across lines rather than fitting one, and long enough (600 chars) that the composer's
* own height grows over several frames, pushing the transcript above it upward the same way a
* real long message does. Exactly this string is also in `docs/RUST.md`'s P0 box so the iris
* half types the identical content.
*/
const val TYPE_TEXT =
"Benchmarking this transcript screen requires unusually long, multisyllabic words so " +
"wrapping and reflow are properly exercised: internationalization, " +
"counterproductiveness, disproportionately, incomprehensibility, " +
"deinstitutionalization, uncharacteristically, overenthusiastically, " +
"misunderstanding, straightforwardness, telecommunications, and interdisciplinary " +
"collaboration all push a narrow composer field to wrap across several lines while " +
"the transcript above is pushed upward by the growing keyboard-adjacent box, which " +
"is exactly what a real reader typing a long message sees happening now!!!"
private const val TYPE_CHAR_DELAY_MS = 50L
/**
* Keyboard phase (v2): five show/hide cycles, a second apart, is enough to see whether the
* transition is ever actually observed rather than being a one-off fluke either way.
*/
private const val KEYBOARD_CYCLES = 5
private const val KEYBOARD_SHOW_WAIT_MS = 1_000L
private const val KEYBOARD_HIDE_WAIT_MS = 1_000L
/**
* Scrolls, flings, streams, types and toggles the keyboard, then returns the extra report lines
* P0 asked for (per-phase travel/typing/keyboard counts, plus CPU time, peak RSS, battery
* current) -- [FrameStats] and [DebugStats] are reset first, exactly as `copyRenderReport`
* resets them, so the two accountings cover the same stretch of work.
*/
suspend fun run(
context: Context,
scope: CoroutineScope,
listState: LazyListState,
flingBehavior: FlingBehavior,
composerFocus: FocusRequester,
setComposerText: (String) -> Unit,
view: View,
): List<String> {
FrameStats.reset()
DebugStats.reset()
val cpuStartMs = Process.getElapsedCpuTime()
val battery = BatterySampler(context)
// Launched in the caller's scope rather than a fresh coroutineScope{} here, which would
// suspend this function until the sampler job ended -- and it only ends when told to.
val samplerJob = scope.launch {
while (isActive) {
battery.sample()
delay(1000)
}
}
val travel = runFlingPhase(listState, flingBehavior)
val sent = runStreamPhase()
runTypePhase(listState, composerFocus, setComposerText, view)
val keyboard = runKeyboardPhase(context, view)
samplerJob.cancel()
val cpuMs = Process.getElapsedCpuTime() - cpuStartMs
val rssLine = peakRssLine()
val batteryLine = battery.finish()
return listOf(
" fling: $FLING_COUNT flings out + $FLING_COUNT back at" +
" ${FLING_VELOCITY_PX_S.toInt()}px/s, travel $travel",
" scroll: $CYCLES cycles (${CYCLES * 4} swipes, legacy tween), " +
"streamed $sent/${STREAM_EVENTS_PER_SEC * STREAM_SECONDS} fixture events",
" type: ${TYPE_TEXT.length} characters inserted then deleted, one per" +
" ${TYPE_CHAR_DELAY_MS}ms",
keyboard,
" process CPU time over this run: ${cpuMs}ms",
rssLine,
batteryLine,
)
}
/**
* Phase 1: starting pinned at the newest end, [FLING_COUNT] flings away from it (toward older
* messages) through the list's real fling path, then [FLING_COUNT] back. Positive velocity here
* matches this list's existing scroll-offset convention (`TranscriptList`'s `reverseLayout`
* pins index 0 -- the newest item -- at the bottom; a positive scroll offset moves the viewport
* toward higher indices, i.e. away from the newest end and toward older content), the same sign
* the pre-v2 swipe loop below already used for its first two swipes.
*/
private suspend fun runFlingPhase(
listState: LazyListState,
flingBehavior: FlingBehavior,
): String {
FrameStats.markPhase("fling")
listState.scrollToItem(0)
val start = position(listState)
repeat(FLING_COUNT) {
listState.scroll { with(flingBehavior) { performFling(FLING_VELOCITY_PX_S) } }
waitForSettle(listState)
delay(FLING_PAUSE_MS)
}
val outward = position(listState)
repeat(FLING_COUNT) {
listState.scroll { with(flingBehavior) { performFling(-FLING_VELOCITY_PX_S) } }
waitForSettle(listState)
delay(FLING_PAUSE_MS)
}
val back = position(listState)
return "start=$start outward=$outward end=$back"
}
private fun position(listState: LazyListState) =
"idx=${listState.firstVisibleItemIndex}/off=${listState.firstVisibleItemScrollOffset}px"
/** Belt-and-suspenders on top of `performFling` already suspending until its own decay ends. */
private suspend fun waitForSettle(listState: LazyListState) {
val startedAt = System.currentTimeMillis()
while (
listState.isScrollInProgress &&
System.currentTimeMillis() - startedAt < FLING_SETTLE_CAP_MS
) {
delay(16)
}
}
/**
* Phase 2 (unchanged from v1): pinned to the newest end before streaming starts, the way
* stream-bench.sh's "Jump to latest" tap is -- a reply streamed into a list parked further back
* arrives off-screen and the report would show nothing happened.
*/
private suspend fun runStreamPhase(): Int {
FrameStats.markPhase("stream")
var sent = 0
val total = STREAM_EVENTS_PER_SEC * STREAM_SECONDS
while (sent < total && BenchFixture.remainingStreamEvents() > 0) {
BenchFixture.pushNextLiveEvent()
sent++
delay(1000L / STREAM_EVENTS_PER_SEC)
}
// Lets the last few deltas land and draw before the next phase starts.
delay(300)
return sent
}
/**
* Phase 3: focuses the real composer, shows the keyboard if the platform allows it, then types
* [TYPE_TEXT] one character at a time through the same `TextFieldValue` state a real keystroke
* updates, and deletes it the same way -- this is what exercises wrapping and the transcript
* being pushed upward, not a single big write.
*/
private suspend fun runTypePhase(
listState: LazyListState,
composerFocus: FocusRequester,
setComposerText: (String) -> Unit,
view: View,
) {
FrameStats.markPhase("type")
listState.scrollToItem(0)
composerFocus.requestFocus()
showIme(view.context, view)
// Lets focus and the keyboard's opening animation land before typing starts, so the frames
// this phase records are the wrap/reflow it is measuring, not the keyboard opening.
delay(300)
var typed = ""
for (ch in TYPE_TEXT) {
typed += ch
setComposerText(typed)
delay(TYPE_CHAR_DELAY_MS)
}
delay(200)
while (typed.isNotEmpty()) {
typed = typed.dropLast(1)
setComposerText(typed)
delay(TYPE_CHAR_DELAY_MS)
}
}
/**
* Phase 4: [KEYBOARD_CYCLES] show/hide cycles through the same [WindowInsetsControllerCompat]
* path a real IME toggle goes through, reporting how many of each were actually confirmed by
* [android.view.WindowInsets.isVisible] rather than assumed from having asked -- UI_RULES:
* never present an inferred value as a measured one. If the platform never shows it even once,
* this says so in words rather than reporting a phase with no keyboard in it.
*/
private suspend fun runKeyboardPhase(context: Context, view: View): String {
FrameStats.markPhase("keyboard")
var shown = 0
var hidden = 0
repeat(KEYBOARD_CYCLES) {
showIme(context, view)
delay(KEYBOARD_SHOW_WAIT_MS)
if (imeVisible(view)) shown++
hideIme(context, view)
delay(KEYBOARD_HIDE_WAIT_MS)
if (!imeVisible(view)) hidden++
}
return if (shown == 0) {
" keyboard: could not be shown ($KEYBOARD_CYCLES attempts, 0 confirmed visible)"
} else {
" keyboard: shown $shown/$KEYBOARD_CYCLES, hidden $hidden/$KEYBOARD_CYCLES" +
" (confirmed via isImeVisible)"
}
}
private fun controller(context: Context, view: View): WindowInsetsControllerCompat? {
val window = context.activity()?.window ?: return null
return WindowInsetsControllerCompat(window, view)
}
private fun showIme(context: Context, view: View) {
controller(context, view)?.show(WindowInsetsCompat.Type.ime())
}
private fun hideIme(context: Context, view: View) {
controller(context, view)?.hide(WindowInsetsCompat.Type.ime())
}
private fun imeVisible(view: View): Boolean =
ViewCompat.getRootWindowInsets(view)?.isVisible(WindowInsetsCompat.Type.ime()) ?: false
/** VmHWM from /proc/self/status: the process's high-water mark, in kB, since it started. */
private fun peakRssLine(): String {
val kb =
try {
File("/proc/self/status")
.readLines()
.firstOrNull { it.startsWith("VmHWM:") }
?.trim()
?.removePrefix("VmHWM:")
?.trim()
?.removeSuffix("kB")
?.trim()
?.toLongOrNull()
} catch (_: Exception) {
null
}
return " peak RSS: " +
(kb?.let { "${it}kB" } ?: "unavailable (/proc/self/status unreadable)")
}
}
/**
* Samples [BatteryManager.BATTERY_PROPERTY_CURRENT_NOW] (microamps) once a second for the length of
* a run. The property returns `Int.MIN_VALUE` on hardware that does not support it -- most
* emulators -- and that is reported as "unavailable" rather than folded into an average with the
* real samples, which would silently understate every number after it. See UI_RULES: never present
* an inferred value as a measured one.
*/
private class BatterySampler(context: Context) {
private val manager = context.getSystemService(BatteryManager::class.java)
private val samples = mutableListOf<Int>()
fun sample() {
val value = manager?.getIntProperty(BatteryManager.BATTERY_PROPERTY_CURRENT_NOW)
if (value != null && value != Int.MIN_VALUE) samples.add(value)
}
fun finish(): String {
if (samples.isEmpty()) return " battery current: unavailable on this device"
val meanUa = samples.sum() / samples.size
return " battery current: mean ${meanUa}µA over ${samples.size} samples" +
" (min ${samples.min()}, max ${samples.max()})"
}
}
@@ -127,6 +127,20 @@ fun debugReport(
frames: List<String>,
accounting: List<String>,
crash: String?,
/**
* P0's benchmark-only measurements (process CPU time, peak RSS, battery current) -- empty on
* every path but [BenchRun.runP0Benchmark], which is the only caller that has them. A section
* heading only appears when there is something to put under it, so an ordinary copy from the
* render-report button reads exactly as it did before this existed.
*/
extra: List<String> = emptyList(),
/**
* Bench v2's per-phase frame accounting ([FrameStats.phaseLines]) --
* fling/stream/type/keyboard, each a slice of the same frames the whole-run sections below
* still cover in full. Empty on every path but the scripted bench run, same reasoning as
* [extra].
*/
phaseFrames: List<String> = emptyList(),
): String = buildString {
appendLine("ai-app render report")
appendLine(device)
@@ -141,6 +155,11 @@ fun debugReport(
appendLine("transcript:")
transcript.forEach { appendLine(it) }
appendLine()
if (phaseFrames.isNotEmpty()) {
appendLine("per phase:")
phaseFrames.forEach { appendLine(it) }
appendLine()
}
appendLine("frames:")
frames.forEach { appendLine(it) }
appendLine()
@@ -152,6 +171,11 @@ fun debugReport(
appendLine("work since this was last copied:")
val work = DebugStats.lines()
if (work.isEmpty()) appendLine(" nothing recorded") else work.forEach { appendLine(it) }
if (extra.isNotEmpty()) {
appendLine()
appendLine("bench:")
extra.forEach { appendLine(it) }
}
}
/** Puts [text] on the clipboard under [label], which is what the system offers as its name. */
@@ -12,6 +12,10 @@ import androidx.compose.ui.Alignment
import androidx.compose.ui.Modifier
import androidx.compose.ui.graphics.Color
import androidx.compose.ui.unit.dp
import java.time.Instant
import java.time.ZoneId
import java.time.format.DateTimeFormatter
import java.time.format.FormatStyle
/**
* A line across the transcript saying what left the session's context.
@@ -47,3 +51,40 @@ fun TranscriptDivider(text: String, color: Color, modifier: Modifier = Modifier)
fun ClearedRow(modifier: Modifier = Modifier) {
TranscriptDivider("Context cleared", clearedColor, modifier)
}
/**
* The mark running out of quota leaves.
*
* The same red the usage bar takes when a window is spent, because it is the same fact in a second
* place: colour by consequence, so "there is nothing left to spend" is learned once.
*
* A time rather than a countdown. The row is folded once and never re-measured, so a span would go
* stale on screen the moment it was drawn; and this is when the *account* said it would reset,
* which is not a promise about when the session picks back up. A limit the session was told no
* reset time for says nothing about one -- that state has its own words rather than a plausible
* number.
*/
@Composable
fun LimitRow(item: TranscriptItem.LimitNote, modifier: Modifier = Modifier) {
TranscriptDivider(limitSummary(item.resetsAt, ZoneId.systemDefault()), overLimitColor, modifier)
}
/**
* What the row says. Split out so the wording is testable without a screen, since the two states it
* has to keep apart -- a reset time that arrived and one that never did -- are exactly the pair
* that reads the same when it goes wrong.
*
* [zone] is a parameter rather than read here so a test says the same thing wherever it runs.
*/
fun limitSummary(resetsAt: Double?, zone: ZoneId): String {
val at = resetsAt?.let {
try {
DateTimeFormatter.ofLocalizedTime(FormatStyle.SHORT)
.withZone(zone)
.format(Instant.ofEpochSecond(it.toLong()))
} catch (_: Exception) {
null
}
}
return if (at == null) "Usage limit reached" else "Usage limit reached • resets $at"
}
@@ -14,7 +14,7 @@ private const val RESET_EVENT = "reset"
* mean. [close] from any thread ends it, and the caller owns reconnecting -- with the last seq it
* saw as the new cursor.
*/
class EventStream(settings: ServerSettings, private val sessionId: String) {
class EventStream(settings: ServerSettings, private val address: TranscriptAddress) {
private val stream = Sse(settings)
fun close() = stream.close()
@@ -35,7 +35,7 @@ class EventStream(settings: ServerSettings, private val sessionId: String) {
// one and the screen folds the other, and they have to be the same line.
onEvent: (raw: String, event: SeqEvent) -> Unit,
) {
stream.run("/sessions/$sessionId/events?after=$after", onOpen) { name, data ->
stream.run("/${address.urlPath}/events?after=$after", onOpen) { name, data ->
// A named frame carries no payload and a data frame has no name.
if (name == RESET_EVENT) onReset()
else if (data.isNotEmpty()) onEvent(data, parseSeqEvent(data))
@@ -157,6 +157,18 @@ sealed class SessionEvent {
*/
data object Cleared : SessionEvent()
/**
* The session stopped because its account's usage limit was reached.
*
* Its own event rather than an [Error] carrying the CLI's sentence, because it is a state
* rather than something that went wrong -- and because the raw sentence is `Claude AI usage
* limit reached|1788546972`, which is not readable by the person it is shown to.
*
* [resetsAt] is epoch seconds and null where the session was told nothing. Only the server acts
* on it; what this draws it as is a time, not a countdown, because nothing here re-measures it.
*/
data class LimitReached(val resetsAt: Double?) : SessionEvent()
data class Error(val message: String) : SessionEvent()
/**
@@ -261,6 +273,10 @@ fun parseSeqEvent(json: String): SeqEvent {
trigger = body.optString("trigger").ifEmpty { null },
)
"cleared" -> SessionEvent.Cleared
"limitReached" ->
SessionEvent.LimitReached(
if (body.has("resetsAt")) body.getDouble("resetsAt") else null
)
"error" -> SessionEvent.Error(body.getString("message"))
else -> SessionEvent.Unknown(type)
}
@@ -42,6 +42,21 @@ object FrameStats {
private val gpu = ArrayList<Long>()
private var since = System.currentTimeMillis()
/**
* Where a named phase of a scripted run (bench v2's fling/stream/type/keyboard) started, as an
* index into [total] and a wall-clock time -- not a second recorder, just a mark on this one,
* so a phase's frames are the same [FrameMetrics] the whole-run report already has, sliced.
*/
private data class PhaseMark(val name: String, val startIndex: Int, val startMs: Long)
private val phaseMarks = ArrayList<PhaseMark>()
/** Call at the start of each named phase of a scripted run; see [BenchRun]. */
@Synchronized
fun markPhase(name: String) {
phaseMarks += PhaseMark(name, total.size, System.currentTimeMillis())
}
@Synchronized
fun add(metrics: FrameMetrics) {
// The first frame after a window opens includes inflating it and is nobody's scroll.
@@ -69,6 +84,7 @@ object FrameStats {
listOf(total, waited, input, animation, layout, draw, sync, issue, swap, gpu).forEach {
it.clear()
}
phaseMarks.clear()
since = System.currentTimeMillis()
}
@@ -95,6 +111,38 @@ object FrameStats {
) + if (gpu.isEmpty()) emptyList() else listOf(phase("gpu ", gpu))
}
/**
* One block per [markPhase] call: how many frames landed between that mark and the next (or the
* end of the run, for the last one), how many were late, the total/p50/p90/p99, the worst
* single frame, and how long the phase actually ran. Marks with no frames between them (a phase
* that finished before a frame was drawn) still get a line rather than being silently dropped
* -- UI_RULES' "say what you don't know" applies to a phase as much as to a single number.
*/
@Synchronized
fun phaseLines(refreshHz: Float): List<String> {
if (phaseMarks.isEmpty()) return emptyList()
val budget = if (refreshHz > 0) 1000.0 / refreshHz else 16.7
val lines = ArrayList<String>()
phaseMarks.forEachIndexed { i, mark ->
val endIndex = if (i + 1 < phaseMarks.size) phaseMarks[i + 1].startIndex else total.size
val endMs =
if (i + 1 < phaseMarks.size) phaseMarks[i + 1].startMs
else System.currentTimeMillis()
val samples = total.subList(mark.startIndex, endIndex)
val seconds = (endMs - mark.startMs) / 1000.0
lines += " ${mark.name}: ${samples.size} frames over ${"%.1f".format(seconds)}s"
if (samples.isEmpty()) {
lines += " no frames recorded in this phase"
} else {
val late = samples.count { it / 1_000_000.0 > budget }
lines += " late: $late (${percent(late, samples.size)})"
lines += " " + phase("total ", samples)
lines += " worst ${"%.1fms".format(samples.max() / 1_000_000.0)}"
}
}
return lines
}
/** How long the frames recorded here spent in their draw phase, and how many there were. */
@Synchronized fun drawPhase(): Pair<Long, Int> = draw.sum() to draw.size
@@ -66,6 +66,35 @@ class MainActivity : ComponentActivity() {
// Transparent status bar on every version; the Surface below paints through underneath it
// and content insets itself. Same reasoning as dev-updater's MainActivity.
enableEdgeToEdge()
// The `bench` build's entire purpose (P0, docs/RUST.md): open straight onto the session
// screen against BenchFixture's in-process fake backend, with no enrollment, no network
// permission, and no notification prompt -- none of them mean anything with no server and
// no real device to notify. See BenchFixture.kt and BenchNetwork.kt for how a screen built
// to talk to a real backend is made to talk to this instead. Still needs the same
// status/navigation-bar padding the ordinary flow below applies: edge-to-edge is the
// platform's own default from Android 15 on this app's targetSdk, with or without the call
// above, so skipping the padding here put the header's own buttons under the status bar --
// there to look at, but not there for `ui-trace`'s tap-by-label to land on.
if (BuildConfig.FIXTURE_MODE) {
installFixtureNetworkOnce()
BenchFixture.ensureLoaded(this)
setContent {
MaterialTheme(colorScheme = AiAppColors) {
Surface(modifier = Modifier.fillMaxSize()) {
Box(Modifier.fillMaxSize().statusBarsPadding().navigationBarsPadding()) {
SessionScreen(
settings = BenchFixture.settings,
summary = benchSessionSummary(),
onBack = { finish() },
onFiles = {},
)
}
}
}
}
return
}
// Dark status-bar icons only over a light background, decided from the scheme rather than
// fixed. It was hardcoded to `true`, which was right against the default light surface and
// became unreadable the moment the app wore Catppuccin Mocha.
@@ -147,6 +176,33 @@ class MainActivity : ComponentActivity() {
}
}
/** The one session the `bench` build ever shows -- BenchFixture's session id, nothing else. */
private fun benchSessionSummary() =
SessionSummary(
id = BenchFixture.SESSION_ID,
setup = "bench",
setupName = "bench",
provider = "bench",
title = "P0 benchmark",
model = null,
keepsOwnTranscript = false,
permissionMode = null,
effort = null,
takesEffort = false,
imported = false,
notify = false,
autoResume = false,
autoResumeMessage = "",
resumeAt = null,
cwd = null,
contextTokens = null,
maxImageEdge = null,
usageProvider = null,
status = "idle",
lastActivity = 0.0,
subagents = 0,
)
// launchMode="singleTop": an enrollment scan, or a notification tapped while the app is open,
// lands here rather than in a second activity instance.
override fun onNewIntent(intent: Intent) {
@@ -48,6 +48,8 @@ fun MainScreen(
/** What another app shared in and no session has taken yet; see [ShareRequest]. */
share: ShareRequest? = null,
onOpen: (SessionSummary) -> Unit,
/** Opens one session's subagent, from the expander under its card. */
onOpenSubagent: (SessionSummary, SubagentSummary) -> Unit,
onSpawn: () -> Unit,
onImported: (SessionSummary) -> Unit,
onSettings: () -> Unit,
@@ -139,6 +141,7 @@ fun MainScreen(
settings = settings,
reloadToken = token,
onOpen = onOpen,
onOpenSubagent = onOpenSubagent,
onSpawn = onSpawn,
)
MainTab.Import ->
@@ -1,7 +1,9 @@
package com.example.aiapp
import androidx.compose.foundation.ExperimentalFoundationApi
import androidx.compose.foundation.clickable
import androidx.compose.foundation.combinedClickable
import androidx.compose.foundation.layout.Arrangement
import androidx.compose.foundation.layout.Box
import androidx.compose.foundation.layout.Column
import androidx.compose.foundation.layout.Row
@@ -9,6 +11,7 @@ import androidx.compose.foundation.layout.Spacer
import androidx.compose.foundation.layout.fillMaxSize
import androidx.compose.foundation.layout.fillMaxWidth
import androidx.compose.foundation.layout.height
import androidx.compose.foundation.layout.heightIn
import androidx.compose.foundation.layout.padding
import androidx.compose.foundation.layout.width
import androidx.compose.foundation.lazy.LazyColumn
@@ -17,6 +20,7 @@ import androidx.compose.material3.Card
import androidx.compose.material3.CircularProgressIndicator
import androidx.compose.material3.FloatingActionButton
import androidx.compose.material3.MaterialTheme
import androidx.compose.material3.OutlinedCard
import androidx.compose.material3.Switch
import androidx.compose.material3.Text
import androidx.compose.material3.TextButton
@@ -30,6 +34,8 @@ import androidx.compose.runtime.setValue
import androidx.compose.ui.Alignment
import androidx.compose.ui.Modifier
import androidx.compose.ui.platform.LocalContext
import androidx.compose.ui.semantics.contentDescription
import androidx.compose.ui.semantics.semantics
import androidx.compose.ui.unit.dp
import kotlinx.coroutines.Dispatchers
import kotlinx.coroutines.launch
@@ -47,12 +53,40 @@ fun SessionListScreen(
settings: ServerSettings,
reloadToken: Int,
onOpen: (SessionSummary) -> Unit,
/** Opens one session's subagent, from the expander under its card. */
onOpenSubagent: (SessionSummary, SubagentSummary) -> Unit,
onSpawn: () -> Unit,
) {
val scope = rememberCoroutineScope()
var listState by remember { mutableStateOf<LoadState<List<SessionSummary>>>(LoadState.Loading) }
var confirmingDelete by remember { mutableStateOf<SessionSummary?>(null) }
// Which session cards are expanded to show their subagents, and what each expansion fetched.
// Ids rather than a flag on the row for the same reason `deleting` is: the rows are rebuilt
// from
// whatever the server last said, and this belongs to the reader's own choice, which survives a
// refresh.
var expandedSessions by remember { mutableStateOf(setOf<String>()) }
var subagentLoads by remember {
mutableStateOf(mapOf<String, LoadState<List<SubagentSummary>>>())
}
fun loadSubagents(sessionId: String) {
subagentLoads = subagentLoads + (sessionId to LoadState.Loading)
scope.launch {
subagentLoads =
subagentLoads +
(sessionId to
try {
LoadState.Loaded(
withContext(Dispatchers.IO) { fetchSubagents(settings, sessionId) }
)
} catch (e: ApiException) {
LoadState.failed(e)
})
}
}
// Failures that belong to one session rather than to the list, keyed by its id and shown on its
// own card. The two scopes are decided by whether the server answered: it answered and refused,
// so this says nothing about the other rows.
@@ -84,6 +118,14 @@ fun SessionListScreen(
withContext(Dispatchers.IO) {
transcriptCache.retainOnly(loaded.value.map { it.id }.toSet())
}
// A session gone from this answer cannot still be expanded, and an expanded one
// that is still here asks again -- its subagents may have changed since the
// last
// fetch.
val ids = loaded.value.map { it.id }.toSet()
expandedSessions = expandedSessions intersect ids
subagentLoads = subagentLoads.filterKeys { it in ids }
expandedSessions.forEach(::loadSubagents)
loaded
} catch (e: ApiException) {
LoadState.failed(e)
@@ -127,6 +169,17 @@ fun SessionListScreen(
deleting = session.id in deleting,
onOpen = { onOpen(session) },
onLongPress = { confirmingDelete = session },
expanded = session.id in expandedSessions,
subagents = subagentLoads[session.id],
onToggleSubagents = {
if (session.id in expandedSessions) {
expandedSessions = expandedSessions - session.id
} else {
expandedSessions = expandedSessions + session.id
loadSubagents(session.id)
}
},
onOpenSubagent = { subagent -> onOpenSubagent(session, subagent) },
)
Spacer(Modifier.height(12.dp))
}
@@ -225,7 +278,7 @@ fun SessionListScreen(
deleteSession(settings, session.id, alsoDeleteForeign)
// After it succeeded, not before: a refused delete leaves the
// session exactly as it was, and its transcript with it.
transcriptCache.session(session.id).purge()
transcriptCache.session(TranscriptAddress(session.id)).purge()
}
// Only this row, and only what changed. Refetching the list instead
// put every other session back through loading and handed the
@@ -276,6 +329,12 @@ private fun SessionCard(
deleting: Boolean,
onOpen: () -> Unit,
onLongPress: () -> Unit,
/** Whether the expander below is open. Collapsed by default; see [SessionListScreen]. */
expanded: Boolean,
/** What the expander's own fetch answered, or null before it has been asked. */
subagents: LoadState<List<SubagentSummary>>?,
onToggleSubagents: () -> Unit,
onOpenSubagent: (SubagentSummary) -> Unit,
) {
BusyItem(label = if (deleting) "deleting" else null) {
Card(
@@ -332,10 +391,104 @@ private fun SessionCard(
color = MaterialTheme.colorScheme.error,
)
}
// Nothing at all for a card with no subagents: a disabled expander here would be
// noise on every ordinary session's card. Its own row at the bottom rather than
// beside the title or the machine line, so opening it never displaces text that was
// already on screen -- see UI_RULES on a control not displacing the text beside it.
if (session.subagents > 0) {
Spacer(Modifier.height(8.dp))
// The platform's minimum touch height, not the chevron's own ten or so dp:
// at the chevron's height a tap meant for it landed on the first subcard
// beneath and opened a subagent instead.
Row(
horizontalArrangement = Arrangement.Center,
verticalAlignment = Alignment.CenterVertically,
modifier =
Modifier.fillMaxWidth()
.heightIn(min = 48.dp)
.clickable(enabled = !deleting, onClick = onToggleSubagents)
.semantics {
contentDescription =
if (expanded) "Collapse subagents" else "Expand subagents"
},
) {
Chevron(if (expanded) Pointing.Up else Pointing.Down)
}
if (expanded) {
Spacer(Modifier.height(4.dp))
Column(verticalArrangement = Arrangement.spacedBy(8.dp)) {
when (subagents) {
null,
is LoadState.Loading ->
CircularProgressIndicator(
modifier = Modifier.width(20.dp).height(20.dp),
strokeWidth = 2.dp,
)
is LoadState.Error ->
// Said here rather than left silent: a fetch that failed and an
// expander that simply found nothing must not look the same --
// see UI_RULES on designing the unknown state first.
Text(
subagents.message,
style = MaterialTheme.typography.bodySmall,
color = MaterialTheme.colorScheme.error,
)
is LoadState.Loaded ->
subagents.value.forEach { subagent ->
SubagentCard(
subagent,
onClick = { onOpenSubagent(subagent) },
)
}
}
}
}
}
}
}
}
}
/**
* One subagent, indented inside its session's card -- the way dev-updater draws a project's
* components (`ComponentCard`, `UpdaterScreen.kt`): an outlined card, not the session card's own
* filled one, so the nesting reads as one step rather than as another session.
*/
@Composable
private fun SubagentCard(subagent: SubagentSummary, onClick: () -> Unit) {
OutlinedCard(Modifier.fillMaxWidth().clickable(onClick = onClick)) {
Column(Modifier.padding(horizontal = 12.dp, vertical = 8.dp)) {
Text(subagent.title, style = MaterialTheme.typography.titleSmall)
Spacer(Modifier.height(2.dp))
Row(modifier = Modifier.fillMaxWidth()) {
Text(
subagentStatusLabel(subagent.status),
style = MaterialTheme.typography.bodySmall,
color = MaterialTheme.colorScheme.onSurfaceVariant,
modifier = Modifier.weight(1f),
)
Text(
relativeTime(subagent.lastActivity),
style = MaterialTheme.typography.bodySmall,
color = MaterialTheme.colorScheme.onSurfaceVariant,
)
}
}
}
}
/**
* The subcard's word for a subagent's status -- see SUBAGENTS.md's "Wire shape". Its own function
* rather than a branch inside [StatusText], because a subagent's three states are not that
* composable's five: "exited" reads as "finished" here, since its process was always its parent's
* and never something of its own to have merely stopped.
*/
private fun subagentStatusLabel(status: String) =
when (status) {
"running" -> "running"
"exited" -> "finished"
else -> "unknown"
}
@Composable
fun StatusText(status: String) {
@@ -12,6 +12,7 @@ import androidx.activity.result.PickVisualMediaRequest
import androidx.activity.result.contract.ActivityResultContracts
import androidx.compose.foundation.background
import androidx.compose.foundation.clickable
import androidx.compose.foundation.gestures.ScrollableDefaults
import androidx.compose.foundation.gestures.awaitEachGesture
import androidx.compose.foundation.gestures.awaitFirstDown
import androidx.compose.foundation.layout.Box
@@ -64,12 +65,15 @@ import androidx.compose.runtime.snapshots.Snapshot
import androidx.compose.ui.Alignment
import androidx.compose.ui.Modifier
import androidx.compose.ui.draw.drawWithContent
import androidx.compose.ui.focus.FocusRequester
import androidx.compose.ui.focus.focusRequester
import androidx.compose.ui.graphics.graphicsLayer
import androidx.compose.ui.input.pointer.PointerEventPass
import androidx.compose.ui.input.pointer.pointerInput
import androidx.compose.ui.layout.onSizeChanged
import androidx.compose.ui.platform.LocalContext
import androidx.compose.ui.platform.LocalDensity
import androidx.compose.ui.platform.LocalView
import androidx.compose.ui.semantics.contentDescription
import androidx.compose.ui.semantics.semantics
import androidx.compose.ui.text.TextRange
@@ -216,16 +220,33 @@ fun SessionScreen(
share: ShareRequest? = null,
/** Said once [share] has been attached here, so it is not attached again. */
onShareTaken: () -> Unit = {},
/**
* Draws this screen read-only, on a subagent's own transcript instead of the session's.
*
* A subagent has no process and no controls of its own -- see SUBAGENTS.md's "Phone" -- so
* every gate below keyed on this switches off the composer, the files button, the settings cog,
* the usage bar and notifications, while everything that draws a transcript (paging, cache,
* selection, images, the status row, stream reconnects) is reused unchanged, pointed at
* [address] instead of the session's own.
*/
subagent: SubagentSummary? = null,
) {
DebugStats.count("session screen recomposed")
val isSubagent = subagent != null
val address = TranscriptAddress(summary.id, subagent?.id)
val scope = rememberCoroutineScope()
val topEdgeHeld = remember { TopEdgeHold() }
var items by remember { mutableStateOf(listOf<TranscriptItem>()) }
var status by remember { mutableStateOf(summary.status) }
var status by remember { mutableStateOf(subagent?.status ?: summary.status) }
// Seeded from the row this screen was opened from, so a conversation already under way says how
// much it is holding before any turn happens here. Null is "nobody has measured it", which is a
// different answer from an empty context and is drawn differently.
var contextTokens by remember(summary.id) { mutableStateOf(summary.contextTokens) }
//
// A subagent has no context measurement of its own, so it always starts unmeasured rather than
// borrowing the parent session's figure -- see UI_RULES on not showing an inferred value as one
// that was measured.
var contextTokens by
remember(address) { mutableStateOf(if (isSubagent) null else summary.contextTokens) }
// When the current compaction started. The moment comes off the `compacting` status event
// itself -- the server timestamps every transcript line -- rather than off this device noticing
// one, which is what makes it survive leaving the session and reopening it.
@@ -241,7 +262,13 @@ fun SessionScreen(
val context = LocalContext.current
// Seeded from what was left in the box last time and written back on every keystroke, so
// leaving the screen does not throw away a half-typed message. See `Drafts.kt`.
var input by remember(summary.id) { mutableStateOf(atEnd(loadDraft(context, summary.id))) }
//
// A subagent has no box to type into, so it never touches a draft at all -- not this session's,
// which is what reading one keyed only by `summary.id` would do here.
var input by
remember(summary.id) {
mutableStateOf(if (isSubagent) atEnd("") else atEnd(loadDraft(context, summary.id)))
}
// A model the reader has chosen and not yet confirmed. See [ModelSwitchWarning]: switching
// makes the session re-read the whole conversation.
var pendingModel by remember { mutableStateOf<String?>(null) }
@@ -294,27 +321,26 @@ fun SessionScreen(
// Reload throws away what it was reading from.
val cache = remember(settings) { TranscriptCache(cacheRoot(context, settings)) }
val source =
remember(summary.id, epoch) {
TranscriptSource(settings, summary.id, cache.session(summary.id))
}
remember(address, epoch) { TranscriptSource(settings, address, cache.session(address)) }
// Whether the cached tail has been shown to still be the server's own line. Nothing is resumed
// from a cached cursor until it has, and a probe that could not be made leaves this false for
// the stream loop to try again.
var probePassed by remember(summary.id, epoch) { mutableStateOf(false) }
var probePassed by remember(address, epoch) { mutableStateOf(false) }
// Whether the opening effect is still settling that question. It draws the cached rows and
// lifts [ready] before the answer arrives, which is the point of the cache -- so the stream
// below waits for this rather than for `ready`, or it asks the same question twice.
var probing by remember(summary.id, epoch) { mutableStateOf(true) }
var probing by remember(address, epoch) { mutableStateOf(true) }
// The oldest sequence number loaded, and whether there is more behind it. Paging backwards is
// what keeps opening a long session cheap.
var oldestSeq by remember { mutableLongStateOf(0L) }
// Where this session was last being read, from this device's own store. Read once, because the
// answer stops being interesting the moment the list is on screen.
val savedAnchor = remember(summary.id, epoch) { loadScrollAnchor(context, summary.id) }
// Where this transcript was last being read, from this device's own store, keyed by the address
// rather than the session id so a subagent's saved position cannot collide with its session's.
// Read once, because the answer stops being interesting the moment the list is on screen.
val savedAnchor = remember(address, epoch) { loadScrollAnchor(context, address.cachePath) }
// Whether the saved position is still being put back. Nothing is drawn while it is: opening at
// the newest end and then travelling to the anchor is exactly the journey a reader must never
// see.
var restoring by remember(summary.id, epoch) { mutableStateOf(savedAnchor != null) }
var restoring by remember(address, epoch) { mutableStateOf(savedAnchor != null) }
// Messages the server has taken and the session has not read yet, by the id that will resolve
// them. From the event stream rather than from what this screen sent, so they survive leaving
// the session -- and a message sent from another device is drawn waiting on this one too.
@@ -327,11 +353,22 @@ fun SessionScreen(
var loadingHistory by remember { mutableStateOf(false) }
var ready by remember { mutableStateOf(false) }
// Replies parsed ahead of the rows that draw them; see [ParsedReplies].
val replies = remember(summary.id) { ParsedReplies() }
// Keyed like everything else describing one session's transcript. `rememberLazyListState` saves
// through `rememberSaveable`, and this screen restores by its own anchor instead -- two
// restores would fight over the first frame.
val listState = remember(summary.id) { LazyListState() }
val replies = remember(address) { ParsedReplies() }
// Keyed like everything else describing one transcript. `rememberLazyListState` saves through
// `rememberSaveable`, and this screen restores by its own anchor instead -- two restores would
// fight over the first frame.
val listState = remember(address) { LazyListState() }
// The list's own fling path -- what a real flick decays through -- captured here so BenchRun's
// fling phase can drive `LazyListState.scroll` through exactly the `FlingBehavior` this
// screen's
// `TranscriptList` already uses by not overriding it (its `LazyColumn` takes no `flingBehavior`
// argument, so this is the same default it gets).
val flingBehavior = ScrollableDefaults.flingBehavior()
// Where BenchRun's type phase focuses before it types, and the view it toggles the keyboard on
// -- both bench-only, but cheap enough (a remembered object, a CompositionLocal read) to hold
// unconditionally rather than behind a second code path only the bench build compiles.
val composerFocus = remember { FocusRequester() }
val view = LocalView.current
// Whether the newest message is on screen right now. The list is reversed, so the newest end is
// the scrolling start: nothing behind you is exactly being at the bottom. Asked of the scroll
// state rather than of item indices, because a zero-height first item makes an index ambiguous.
@@ -637,7 +674,7 @@ fun SessionScreen(
// ended and carries live events only. The window comes from this phone's own copy when there is
// one, and then costs a single request to check that the server's transcript is still the one
// it came from. See TRANSCRIPT_CACHE.md.
LaunchedEffect(summary.id, epoch) {
LaunchedEffect(address, epoch) {
/**
* One opening window onto the screen, whichever side it came from.
*
@@ -667,11 +704,16 @@ fun SessionScreen(
// A replay is as old as the last visit; the row this screen was opened from was
// fetched moments ago. So the transcript comes from the cache and everything that
// is not the transcript comes from the summary -- otherwise a session that finished
// an hour ago opens saying "working" until the stream connects.
status = summary.status
// an hour ago opens saying "working" until the stream connects. A subagent's status
// comes from its own summary, never the parent session's: they are two different
// things running or not, and the parent's model and permission mode do not apply to
// it at all.
status = subagent?.status ?: summary.status
if (!isSubagent) {
model = summary.model
permissionMode = summary.permissionMode ?: "auto"
if (summary.status != "compacting") compactingSince = null
}
if (status != "compacting") compactingSince = null
// Nothing to put back, so these rows are the screen and the probe can return under
// them. A restore still has history to fetch and is gated below.
if (savedAnchor == null) ready = true
@@ -798,7 +840,7 @@ fun SessionScreen(
// at the top on their return. Switching apps is a choice somebody made, not a fault to report.
// Stopping the stream deliberately makes the drop a close rather than an error, and resuming
// reconnects from the same cursor.
LaunchedEffect(summary.id, ready, epoch, lifecycleOwner) {
LaunchedEffect(address, ready, epoch, lifecycleOwner) {
if (!ready) return@LaunchedEffect
// The opening effect draws cached rows and lifts `ready` *before* it has checked that the
// cursor under them is still the server's, so `ready` is no longer the whole gate. Without
@@ -868,10 +910,14 @@ fun SessionScreen(
// The screen going away entirely, which the lifecycle scope above does not cover: a composable
// can leave the composition while the activity stays started. Keyed on the epoch as well, so
// Reload's replacement source is the one a later disposal closes.
DisposableEffect(summary.id, epoch) { onDispose { source.close() } }
DisposableEffect(address, epoch) { onDispose { source.close() } }
// Nothing gets announced about the session somebody is reading; see NotificationService.
// RESUMED rather than STARTED because "looking at it" means the foreground.
//
// Not for a subagent: it has no notifications of its own, and it is not the session this would
// otherwise mark as being read.
if (!isSubagent) {
LaunchedEffect(summary.id, lifecycleOwner) {
lifecycleOwner.repeatOnLifecycle(Lifecycle.State.RESUMED) {
NotificationService.showing(context, summary.id)
@@ -882,6 +928,7 @@ fun SessionScreen(
}
}
}
}
// Back at the newest end, so the backlog [apply] held can land. Everything at once rather than
// paced out: they are at the bottom, which is the one place the list is allowed to follow new
@@ -924,7 +971,7 @@ fun SessionScreen(
val (index, offset, awayFromNewest) = settled
saveScrollAnchor(
context,
summary.id,
address.cachePath,
// Nothing to restore at the newest end, which is where a session with no anchor
// opens anyway. One *before* the index, because item zero is the "below" slot.
if (!awayFromNewest) null
@@ -947,7 +994,7 @@ fun SessionScreen(
//
// There is no correction beside this one. Following the newest message is not an effect: the
// list is reversed, so an arriving message extends the end the viewport is pinned to.
val unitSizes = remember(summary.id) { HashMap<Any, Int>() }
val unitSizes = remember(address) { HashMap<Any, Int>() }
LaunchedEffect(listState, moreHistory) {
snapshotFlow { listState.layoutInfo }
.collect { info ->
@@ -983,6 +1030,8 @@ fun SessionScreen(
}
}
// Only for the model picker, which a subagent does not have.
if (!isSubagent) {
LaunchedEffect(summary.setupName, summary.provider) {
offeredModels =
try {
@@ -995,10 +1044,12 @@ fun SessionScreen(
.orEmpty()
}
} catch (_: Exception) {
// Not worth reporting: the picker simply has nothing to offer, which is visible.
// Not worth reporting: the picker simply has nothing to offer, which is
// visible.
emptyList()
}
}
}
/**
* Asks the server to take back a message the session has not read yet.
@@ -1148,8 +1199,9 @@ fun SessionScreen(
}
// One poll for the machines' limits, read by everything on this screen that reports them.
val usageFeed = rememberUsageFeed(settings)
val usage = usageFeed.forSetup(summary.setup)
// Nothing meters a subagent -- it has no account of its own -- so it never starts this poll.
val usageFeed = if (isSubagent) null else rememberUsageFeed(settings)
val usage = usageFeed?.forSession(summary) ?: SessionUsage.NotMetered
RecordFrames()
var usageOpen by remember { mutableStateOf(false) }
var settingsOpen by remember { mutableStateOf(false) }
@@ -1191,7 +1243,12 @@ fun SessionScreen(
// the bench scripts keep working when this moves again. They pressed it at a hand-measured
// coordinate until 2026-09-03, and anything that moved the header made that tap land on
// whatever now sat there -- reporting a number that was never measured.
val copyRenderReport = {
// Shared by the ordinary "Copy" button and (bench build only) "Run benchmark": what differs
// between them is only whether there is a [extra] section, built by BenchRun.run beforehand --
// everything about assembling, copying and logging the report is exactly the same act either
// way, and a second copy of it beside `onRunBenchmark` below would be the two silently
// disagreeing about what "the report" contains the first time either one changed.
fun buildAndCopyReport(extra: List<String> = emptyList()) {
val report =
debugReport(
device =
@@ -1213,6 +1270,11 @@ fun SessionScreen(
accounting =
FrameStats.drawPhase().let { (nanos, count) -> drawAccounting(nanos, count) },
crash = lastCrash(context),
extra = extra,
// Empty outside a BenchRun.run pass -- copyRenderReport's own reset below clears
// the
// marks along with everything else, so an ordinary copy never has any to show.
phaseFrames = FrameStats.phaseLines(context.refreshHz()),
)
context.copyToClipboard("ai-app render report", report)
// Also to the log, so a session driving the app over adb can read the same report the
@@ -1226,6 +1288,29 @@ fun SessionScreen(
DebugStats.reset()
Toast.makeText(context, "Copied render report", Toast.LENGTH_SHORT).show()
}
val copyRenderReport = { buildAndCopyReport() }
// Bench build only: P0's scripted fling/stream/type/keyboard benchmark (BenchRun.kt), against
// the fixture session opened below instead of a real server. Null everywhere else -- see
// [SessionSettingsDialog]'s onRunBenchmark.
val runBenchmark: (() -> Unit)? =
if (BuildConfig.FIXTURE_MODE) {
{
settingsOpen = false
scope.launch {
val extra =
BenchRun.run(
context = context,
scope = scope,
listState = listState,
flingBehavior = flingBehavior,
composerFocus = composerFocus,
setComposerText = { text -> input = atEnd(text) },
view = view,
)
buildAndCopyReport(extra)
}
}
} else null
Box(Modifier.fillMaxSize()) {
Column(Modifier.fillMaxSize()) {
Row(
@@ -1236,20 +1321,35 @@ fun SessionScreen(
// A ring's worth, which is what the arrow already keeps on its other three sides.
Spacer(Modifier.width(GLYPH_BUTTON_MARGIN))
Column(Modifier.weight(1f)) {
// A subagent's own title, with the session's beneath it in a smaller style --
// the header says whose conversation this is as well as what it is. Otherwise
// just the session's title, as before.
if (subagent != null) {
Text(subagent.title, style = MaterialTheme.typography.titleMedium)
Text(
title,
style = MaterialTheme.typography.bodySmall,
color = MaterialTheme.colorScheme.onSurfaceVariant,
)
} else {
Text(title, style = MaterialTheme.typography.titleMedium)
// Machine first, then what runs on it -- the same order and the same wording
// everywhere this pair appears, so it reads as one fact rather than two
// sentences with different grammar.
// Machine first, then what runs on it -- the same order and the same
// wording everywhere this pair appears, so it reads as one fact rather than
// two sentences with different grammar.
//
// No model. The picker in the footer already shows what this session is set to,
// and showing it twice means two things to keep in step -- they disagreed for a
// moment on every model change.
// No model. The picker in the footer already shows what this session is set
// to, and showing it twice means two things to keep in step -- they
// disagreed for a moment on every model change.
Text(
"${summary.setupName} · ${summary.provider}",
style = MaterialTheme.typography.bodySmall,
color = MaterialTheme.colorScheme.onSurfaceVariant,
)
}
}
// None of this is a subagent's: it has no files of its own to browse, no settings,
// and nothing meters it -- see SUBAGENTS.md's "Phone".
//
// Beside the provider it reports on, which is the line directly to its left. Its
// real home is this provider's settings, which do not exist yet. A session on a
// provider with no such service gets an honest "unavailable" rather than a hidden
@@ -1264,6 +1364,7 @@ fun SessionScreen(
// Usage, files, settings -- widest scope first, narrowing to the right, so the cog
// stays at the end where every other screen keeps it. Asked for in this order by
// Iris on 2026-09-03.
if (!isSubagent) {
Row {
GlyphButton(
USAGE_GLYPH,
@@ -1281,25 +1382,29 @@ fun SessionScreen(
FilesTarget(
setup = summary.setup,
setupName = summary.setupName,
// Where this session works, and the machine's own home when it
// was never given a directory -- resolved there rather than
// guessed at here, since this app does not know that home.
// Where this session works, and the machine's own home when
// it was never given a directory -- resolved there rather
// than guessed at here, since this app does not know that
// home.
start = summary.cwd?.takeIf { it.isNotBlank() } ?: "~",
)
)
},
)
// What it opens is about this session, so it sits at the end of the session's
// own row. A cog and not a word because there will be more, and a bar of words
// has nowhere to put it.
// What it opens is about this session, so it sits at the end of the
// session's own row. A cog and not a word because there will be more, and a
// bar of words has nowhere to put it.
GlyphButton(SETTINGS_GLYPH, "Session settings", { settingsOpen = true })
}
}
}
// Under the header, above everything the session itself says: it is a fact about the
// machine rather than a turn in the conversation, and it is the number that decides
// whether to keep going.
// whether to keep going. Nothing meters a subagent.
if (!isSubagent) {
SessionUsageBar(usage)
}
(streamError ?: actionError)?.let { message ->
Text(
@@ -1545,6 +1650,7 @@ fun SessionScreen(
is TranscriptItem.ClearedNote -> ClearedRow()
is TranscriptItem.CompactedNote ->
CompactedRow(item)
is TranscriptItem.LimitNote -> LimitRow(item)
// Never reached: a peer message is flattened into
// its own units. Here because a `when` over the
// item kinds has to stay exhaustive.
@@ -1643,24 +1749,34 @@ fun SessionScreen(
)
}
// Kept for a subagent -- see SUBAGENTS.md's "Phone" -- with the wording that turns
// "exited" into "finished" for one, since it has no process to leave running or stop.
SessionStatusRow(
status = status,
compactingFor = compactingFor,
contextTokens = contextTokens,
subagent = isSubagent,
)
// Between the transcript and the box: above what is being typed, so the list does not
// cover the thing the command is about, and below everything that explains it.
// Everything from here down is the composer: a subagent cannot be messaged, so none of
// it applies -- see SUBAGENTS.md's "Phone".
if (!isSubagent) {
// Between the transcript and the box: above what is being typed, so the list does
// not cover the thing the command is about, and below everything that explains it.
CommandSuggestions(
// Nothing to suggest about a suggestion that was just taken. `/compact` is a whole
// command *and* a prefix of itself, so picking it left the list standing there with
// the one row already chosen. Held by what was picked rather than by a flag, so
// typing anything else brings the list back without a second thing to reset.
commands = if (input.text == picked) emptyList() else suggestedCommands(input.text),
// Nothing to suggest about a suggestion that was just taken. `/compact` is a
// whole command *and* a prefix of itself, so picking it left the list standing
// there with the one row already chosen. Held by what was picked rather than by
// a flag, so typing anything else brings the list back without a second thing
// to
// reset.
commands =
if (input.text == picked) emptyList() else suggestedCommands(input.text),
onPick = { command ->
// At the end of what was inserted, which is where the reader carries on typing:
// a command with an argument is put in the box half-written, and a cursor left
// at the front makes the next keystroke the first character of "/rename".
// At the end of what was inserted, which is where the reader carries on
// typing: a command with an argument is put in the box half-written, and a
// cursor left at the front makes the next keystroke the first character of
// "/rename".
input = atEnd(command.typed())
picked = command.typed()
},
@@ -1669,11 +1785,14 @@ fun SessionScreen(
// Always enabled -- a send while the session is running becomes a steering message
// injected at the next tool boundary, which is the point of the whole app.
//
// The field gets a row of its own, above the buttons: sharing one put the full width
// behind three controls, so the thing being typed into was the narrowest on the row.
// The field gets a row of its own, above the buttons: sharing one put the full
// width
// behind three controls, so the thing being typed into was the narrowest on the
// row.
Column(Modifier.fillMaxWidth().padding(8.dp)) {
// Directly above the box they will be sent from, so what is attached is visible
// rather than counted: the "+2" on the button below said how many and never which.
// rather than counted: the "+2" on the button below said how many and never
// which.
PendingAttachments(
settings = settings,
sessionId = summary.id,
@@ -1686,9 +1805,12 @@ fun SessionScreen(
input = it
saveDraft(context, summary.id, it.text)
},
modifier = Modifier.fillMaxWidth(),
// No longer "(+image)": the images are on screen above this, and a placeholder
// saying so said it in words beside the thing itself.
// BenchRun's type phase requests focus on this exact field
// (`composerFocus`)
// so it types through the real composer rather than a stand-in.
modifier = Modifier.fillMaxWidth().focusRequester(composerFocus),
// No longer "(+image)": the images are on screen above this, and a
// placeholder saying so said it in words beside the thing itself.
placeholder = { Text("Message") },
maxLines = 4,
)
@@ -1696,17 +1818,19 @@ fun SessionScreen(
verticalAlignment = Alignment.CenterVertically,
modifier = Modifier.fillMaxWidth(),
) {
// Photo or file, asked here rather than by two buttons: the row is full, and
// Photo or file, asked here rather than by two buttons: the row is full,
// and
// attaching is one action whichever picker answers it.
var attaching by remember { mutableStateOf(false) }
Box {
// Just "+". The count it used to carry was standing in for showing them.
// Just "+". The count it used to carry was standing in for showing
// them.
BubbleButton(onClick = { attaching = true }) { Text("+") }
DropdownMenu(
expanded = attaching,
onDismissRequest = { attaching = false },
// See PickerButton: without this the menu opens a status bar's height
// away from the button in an edge-to-edge activity.
// See PickerButton: without this the menu opens a status bar's
// height away from the button in an edge-to-edge activity.
properties = PopupProperties(clippingEnabled = false),
shape = BubbleMenuShape,
) {
@@ -1730,11 +1854,11 @@ fun SessionScreen(
)
}
}
// The settings share what is left after the actions have taken what they need.
// A Row hands out intrinsic widths in order and clips whatever runs past the
// edge, so with these laid out first the arrival of Stop pushed Send off the
// screen entirely -- the app's central control, gone at the moment it is most
// in use.
// The settings share what is left after the actions have taken what they
// need. A Row hands out intrinsic widths in order and clips whatever runs
// past the edge, so with these laid out first the arrival of Stop pushed
// Send off the screen entirely -- the app's central control, gone at the
// moment it is most in use.
Row(
verticalAlignment = Alignment.CenterVertically,
modifier = Modifier.weight(1f),
@@ -1742,16 +1866,18 @@ fun SessionScreen(
if (offeredModels.isNotEmpty()) {
PickerButton(
current = modelLabel(model),
// What the machine offers, plus the state a session is in when it
// has chosen none of them. The button has always been able to say
// "default"; until this the list could not, so leaving it was a
// one-way trip.
// What the machine offers, plus the state a session is in when
// it has chosen none of them. The button has always been able
// to
// say "default"; until this the list could not, so leaving it
// was a one-way trip.
options = listOf(DEFAULT_MODEL) + offeredModels,
// Not set here. The button follows what the session reports it is
// set to, which arrives a moment later and is sometimes a different
// answer -- a name the CLI resolved, or no change at all on a
// provider whose model is fixed. Asked about first, unless there is
// nothing to lose by it -- see [ModelSwitchWarning].
// Not set here. The button follows what the session reports it
// is set to, which arrives a moment later and is sometimes a
// different answer -- a name the CLI resolved, or no change at
// all on a provider whose model is fixed. Asked about first,
// unless there is nothing to lose by it -- see
// [ModelSwitchWarning].
onPick = { chosen ->
if (
modelLabel(chosen) == modelLabel(model) ||
@@ -1772,14 +1898,16 @@ fun SessionScreen(
},
)
}
// The same filled shape as the button beside it, not an outlined one: these are
// two things you can do about the session, and weighting one as secondary said
// they were a primary action and its qualifier. What separates them is the
// colour and the mark, which is what they mean.
// The same filled shape as the button beside it, not an outlined one: these
// are two things you can do about the session, and weighting one as
// secondary said they were a primary action and its qualifier. What
// separates them is the colour and the mark, which is what they mean.
//
// Always here, rather than arriving with the turn as it used to. A control that
// comes and goes makes its own presence the signal, and a button always in the
// same place also cannot push Send off the end of the row by turning up.
// Always here, rather than arriving with the turn as it used to. A control
// that comes and goes makes its own presence the signal, and a button
// always
// in the same place also cannot push Send off the end of the row by turning
// up.
val process =
when {
running -> ProcessAction.Pause
@@ -1799,18 +1927,21 @@ fun SessionScreen(
Glyph(
process.glyph,
colour = LocalContentColor.current,
modifier = Modifier.semantics { contentDescription = process.label },
modifier =
Modifier.semantics { contentDescription = process.label },
)
}
Spacer(Modifier.width(8.dp))
// The paper plane, with a clock on it while a turn is in flight: sending then
// queues the message for the next tool boundary rather than starting a turn of
// its own, and the two have to be told apart at a glance. The label says the
// same thing to a screen reader.
// The paper plane, with a clock on it while a turn is in flight: sending
// then queues the message for the next tool boundary rather than starting a
// turn of its own, and the two have to be told apart at a glance. The label
// says the same thing to a screen reader.
//
// Disabled while there is nothing to send, rather than pressable and silent:
// `send` has always returned early on an empty composer, so the button promised
// something it would not do. Disabled and not hidden, for the reason above.
// Disabled while there is nothing to send, rather than pressable and
// silent:
// `send` has always returned early on an empty composer, so the button
// promised something it would not do. Disabled and not hidden, for the
// reason above.
Button(
onClick = { send() },
enabled = input.text.isNotBlank() || pendingAttachments.isNotEmpty(),
@@ -1827,12 +1958,13 @@ fun SessionScreen(
}
}
}
}
// Beside the other two dialogs, and outside the list for the same reason as them: what is open
// is the screen's business rather than any row's. See [SessionImageViewer].
fullImage?.let { ref -> SessionImageViewer(settings, summary.id, ref) { fullImage = null } }
if (usageOpen) {
UsageDialog(feed = usageFeed, onDismiss = { usageOpen = false })
usageFeed?.let { UsageDialog(feed = it, onDismiss = { usageOpen = false }) }
}
if (settingsOpen) {
// Measured when the dialog opens rather than kept up to date: what the reader is being told
@@ -1846,6 +1978,8 @@ fun SessionScreen(
settings = settings,
sessionId = summary.id,
title = title,
effort = summary.effort.takeIf { summary.takesEffort },
takesEffort = summary.takesEffort,
cachedBytes = cachedBytes,
// The purge finishes before the epoch moves, because the relaunched opening effect
// reads the same directory and would otherwise draw what is about to be deleted. The
@@ -1869,6 +2003,7 @@ fun SessionScreen(
},
onDismiss = { settingsOpen = false },
onCopyRenderReport = copyRenderReport,
onRunBenchmark = runBenchmark,
)
}
}
@@ -2122,6 +2257,13 @@ private fun SessionStatusRow(
/** Context the session is holding, or null where nothing has measured it. */
contextTokens: Long?,
modifier: Modifier = Modifier,
/**
* Whether this row is for a subagent rather than a session, which changes only one word:
* "exited" reads as "finished" there too, the same as the subagent list's own card -- a
* subagent's process was always its parent's, so "exited" would read as a fault rather than the
* ordinary way one of these ends.
*/
subagent: Boolean = false,
) {
DebugStats.count("status row recomposed")
Row(
@@ -2178,7 +2320,7 @@ private fun SessionStatusRow(
Text(
when (status) {
"idle" -> "idle"
"exited" -> "exited"
"exited" -> if (subagent) "finished" else "exited"
"awaitingInput" -> "your turn"
"unknown" -> "can't tell"
else -> status
@@ -2249,7 +2391,7 @@ private const val ONE_TAP_MS = 250L
* session is set to without spending a second line on saying it.
*/
@Composable
private fun PickerButton(current: String, options: List<String>, onPick: (String) -> Unit) {
fun PickerButton(current: String, options: List<String>, onPick: (String) -> Unit) {
var open by remember { mutableStateOf(false) }
// When an outside touch last closed the menu.
//
@@ -6,8 +6,10 @@ import androidx.compose.foundation.layout.Spacer
import androidx.compose.foundation.layout.fillMaxWidth
import androidx.compose.foundation.layout.height
import androidx.compose.foundation.layout.width
import androidx.compose.foundation.rememberScrollState
import androidx.compose.foundation.text.KeyboardActions
import androidx.compose.foundation.text.KeyboardOptions
import androidx.compose.foundation.verticalScroll
import androidx.compose.material3.AlertDialog
import androidx.compose.material3.CircularProgressIndicator
import androidx.compose.material3.MaterialTheme
@@ -26,6 +28,10 @@ import androidx.compose.ui.Alignment
import androidx.compose.ui.Modifier
import androidx.compose.ui.text.input.ImeAction
import androidx.compose.ui.unit.dp
import java.time.Instant
import java.time.ZoneId
import java.time.format.DateTimeFormatter
import java.time.format.FormatStyle
import kotlinx.coroutines.Dispatchers
import kotlinx.coroutines.launch
import kotlinx.coroutines.withContext
@@ -54,6 +60,16 @@ fun SessionSettingsDialog(
*/
title: String,
onRenamed: (String) -> Unit,
/**
* How hard the model thinks, as the session reports it, or null for the CLI's own default.
*
* Taken from the row this dialog was opened over rather than fetched, because unlike the
* notification switch there is nothing else that changes it: the level is this app's to set and
* the server does not resolve it into something else.
*/
effort: String?,
/** Whether a level does anything here; the row is left out entirely where it does not. */
takesEffort: Boolean,
/**
* What this phone is holding of the conversation, or null while that is being measured -- see
* the Reload row below, which is what would discard it.
@@ -66,9 +82,18 @@ fun SessionSettingsDialog(
* measures is that screen's own state.
*/
onCopyRenderReport: () -> Unit,
/**
* Runs P0's scripted scroll-and-stream benchmark and copies the extended report, or null on
* every build but `bench` -- see [BuildConfig.FIXTURE_MODE] and BenchRun.kt. Null rather than
* always-present-but-disabled: this has no meaning at all outside the bench build, and a
* control with nothing behind it on every other build is not a state worth drawing.
*/
onRunBenchmark: (() -> Unit)? = null,
) {
val scope = rememberCoroutineScope()
var name by remember(sessionId) { mutableStateOf(title) }
var level by remember(sessionId) { mutableStateOf(effort) }
var effortError by remember { mutableStateOf<String?>(null) }
var saving by remember { mutableStateOf(false) }
var error by remember { mutableStateOf<String?>(null) }
// Null until the server has been asked. The row this dialog was opened over is a snapshot of
@@ -77,6 +102,16 @@ fun SessionSettingsDialog(
// and a spinner sits beside it, which is what not knowing looks like.
var notify by remember(sessionId) { mutableStateOf<Boolean?>(null) }
var notifyError by remember { mutableStateOf<String?>(null) }
// The same three-state shape the notification switch has, for the same reason: until the
// server has answered, the switch is disabled rather than showing a position nothing confirmed.
var autoResume by remember(sessionId) { mutableStateOf<Boolean?>(null) }
var resumeMessage by remember(sessionId) { mutableStateOf(DEFAULT_RESUME_MESSAGE) }
// When the server next intends to ask whether the limit has lifted, or null when nothing is
// waiting. Read once with everything else: it moves on the server's schedule, not this
// screen's, and a figure that redrew itself here would be this app re-measuring what it was
// told.
var resumeAt by remember(sessionId) { mutableStateOf<Double?>(null) }
var resumeError by remember { mutableStateOf<String?>(null) }
// Where the session works. Null until the server has been asked, for the same reason the switch
// above is. An empty answer is a session that was never given a directory, which is not the
// same as one whose directory is unknown -- the field is only enabled once one of those is
@@ -90,6 +125,9 @@ fun SessionSettingsDialog(
try {
val fresh = withContext(Dispatchers.IO) { fetchSession(settings, sessionId) }
notify = fresh.notify
autoResume = fresh.autoResume
resumeMessage = fresh.autoResumeMessage
resumeAt = fresh.resumeAt
cwd = fresh.cwd.orEmpty()
typedCwd = fresh.cwd.orEmpty()
} catch (e: ApiException) {
@@ -97,6 +135,8 @@ fun SessionSettingsDialog(
// instead of offering a position nothing confirmed.
notifyError = e.message
notify = null
resumeError = e.message
autoResume = null
}
}
@@ -125,6 +165,26 @@ fun SessionSettingsDialog(
}
}
/**
* Chooses a thinking level, which ends the process the old level was launched with.
*
* Put back if the request is refused, for the reason the notification switch below gives: a
* control that stays where it was put after a refusal is stating something untrue.
*/
fun setEffort(chosen: String?) {
val was = level
level = chosen
effortError = null
scope.launch {
try {
withContext(Dispatchers.IO) { setSessionEffort(settings, sessionId, chosen) }
} catch (e: ApiException) {
level = was
effortError = e.message
}
}
}
// Moved optimistically so the switch answers the finger that moved it, and put back if the
// request is refused -- a switch that waits for a round trip reads as broken on a slow tunnel,
// and one that stays moved after a refusal lies.
@@ -142,6 +202,39 @@ fun SessionSettingsDialog(
}
}
/**
* Turns auto-resume on or off, or changes what it would say.
*
* One request for both, because the server takes one: switching it on and typing the message
* are two halves of the same decision, and sending them separately would leave a moment where
* the session is armed with the old words.
*
* Put back if refused, like the notification switch. Turning it off also clears what was
* scheduled -- said here rather than only on the server, or the row would go on naming a time
* that no longer exists.
*/
fun setAutoResume(on: Boolean, message: String) {
val wasOn = autoResume
val wasMessage = resumeMessage
val wasAt = resumeAt
autoResume = on
resumeMessage = message
if (!on) resumeAt = null
resumeError = null
scope.launch {
try {
withContext(Dispatchers.IO) {
setSessionAutoResume(settings, sessionId, on, message)
}
} catch (e: ApiException) {
autoResume = wasOn
resumeMessage = wasMessage
resumeAt = wasAt
resumeError = e.message
}
}
}
// Nothing to do when the name has not changed, so the button says so rather than sending a
// request whose success would look exactly like the failure of having typed nothing.
val changed = name.trim().isNotEmpty() && name.trim() != title
@@ -168,7 +261,10 @@ fun SessionSettingsDialog(
onDismissRequest = onDismiss,
title = { Text("Session settings") },
text = {
Column {
// Scrollable, because this dialog grew past a screenful: a Material dialog constrains
// its own height and clips what does not fit, so the last control on the list is one
// large system font away from being unreachable with nothing on screen to say so.
Column(Modifier.verticalScroll(rememberScrollState())) {
OutlinedTextField(
value = name,
onValueChange = { name = it },
@@ -212,6 +308,70 @@ fun SessionSettingsDialog(
)
}
Spacer(Modifier.height(8.dp))
Row(
verticalAlignment = Alignment.CenterVertically,
modifier = Modifier.fillMaxWidth(),
) {
Text("Resume after a usage limit", modifier = Modifier.weight(1f))
if (autoResume == null && resumeError == null) {
CircularProgressIndicator(
modifier = Modifier.width(16.dp).height(16.dp),
strokeWidth = 2.dp,
)
Spacer(Modifier.width(8.dp))
}
Switch(
checked = autoResume == true,
onCheckedChange = { setAutoResume(it, resumeMessage) },
enabled = autoResume != null,
)
}
// Disabled rather than hidden while the switch is off: a field that comes and goes
// makes its own presence the signal, and a visible one teaches what the switch will
// do. Committed on the keyboard's Done rather than on every keystroke, so typing a
// sentence is one request instead of one per letter.
OutlinedTextField(
value = resumeMessage,
onValueChange = { resumeMessage = it },
label = { Text("Message to send") },
// What an empty field means, in the field: the server's own word rather than a
// session poked with nothing to read.
placeholder = { Text(DEFAULT_RESUME_MESSAGE) },
singleLine = true,
enabled = autoResume == true,
modifier = Modifier.fillMaxWidth(),
keyboardOptions = KeyboardOptions(imeAction = ImeAction.Done),
keyboardActions =
KeyboardActions(onDone = { setAutoResume(true, resumeMessage) }),
)
// What it does and what it costs, in the order it happens. The last sentence is the
// one that matters: the time below is when the server will *ask*, not a promise
// about when the session speaks.
Text(
"When this session stops because the account is out of quota, the server " +
"checks the limit and sends this message once it has lifted. It checks " +
"again if the limit is still on.",
style = MaterialTheme.typography.bodySmall,
color = MaterialTheme.colorScheme.onSurfaceVariant,
)
// Only where something is actually waiting. Absent is not a state worth a row: a
// session that has not hit a limit has nothing scheduled, which the reader can see
// from the switch.
resumeAt?.let { at ->
Text(
"Waiting now -- next check ${formatCheckTime(at)}.",
style = MaterialTheme.typography.bodySmall,
color = MaterialTheme.colorScheme.onSurfaceVariant,
)
}
resumeError?.let {
Text(
it,
color = MaterialTheme.colorScheme.error,
style = MaterialTheme.typography.bodySmall,
)
}
Spacer(Modifier.height(8.dp))
Row(
verticalAlignment = Alignment.CenterVertically,
modifier = Modifier.fillMaxWidth(),
@@ -257,6 +417,44 @@ fun SessionSettingsDialog(
style = MaterialTheme.typography.bodySmall,
)
}
// Left out rather than disabled, the one place this dialog does that: a disabled
// control teaches what the thing can do, and a llama session cannot do this at all
// -- the row would be teaching something false about it.
if (takesEffort) {
Spacer(Modifier.height(8.dp))
Row(
verticalAlignment = Alignment.CenterVertically,
modifier = Modifier.fillMaxWidth(),
) {
Text("Thinking", modifier = Modifier.weight(1f))
PickerButton(
current = level ?: DEFAULT_EFFORT,
// The level the CLI picks for itself is in the list as well as in the
// button, so leaving a level is not a one-way trip -- the same
// correction the model picker carries.
options = listOf(DEFAULT_EFFORT) + EFFORT_LEVELS,
onPick = { chosen ->
setEffort(chosen.takeIf { it != DEFAULT_EFFORT })
},
)
}
// What it costs, said where it is about to be pressed, like Move above: the
// CLI reads the level when it launches and has no control request for
// changing one.
Text(
"Changing this stops the session's process. It starts again with the " +
"next message, or with Start.",
style = MaterialTheme.typography.bodySmall,
color = MaterialTheme.colorScheme.onSurfaceVariant,
)
effortError?.let {
Text(
it,
color = MaterialTheme.colorScheme.error,
style = MaterialTheme.typography.bodySmall,
)
}
}
Spacer(Modifier.height(8.dp))
Row(
verticalAlignment = Alignment.CenterVertically,
@@ -317,6 +515,21 @@ fun SessionSettingsDialog(
Text("Render timings", modifier = Modifier.weight(1f))
TextButton(onClick = onCopyRenderReport) { Text("Copy") }
}
// Bench-build only: see [onRunBenchmark]. Named exactly "Run benchmark" because
// ui-trace and the emulator smoke run find it by that label, the same way every
// other control here is found -- see AGENTS.md's "Driving the UI".
onRunBenchmark?.let { run ->
Spacer(Modifier.height(8.dp))
Row(
verticalAlignment = Alignment.CenterVertically,
modifier = Modifier.fillMaxWidth(),
) {
Glyph(SPEED_GLYPH, colour = MaterialTheme.colorScheme.onSurface)
Spacer(Modifier.width(8.dp))
Text("P0 benchmark", modifier = Modifier.weight(1f))
TextButton(onClick = run) { Text("Run benchmark") }
}
}
}
},
// Disabled rather than absent while there is nothing to save: a button that comes and goes
@@ -329,3 +542,21 @@ fun SessionSettingsDialog(
dismissButton = { TextButton(onClick = onDismiss) { Text("Close") } },
)
}
/**
* When the server will next look, as a local time.
*
* A time rather than a countdown, for the reason the transcript's own limit row gives: this screen
* reads the figure once, and a span drawn from a value nothing refreshes goes stale while somebody
* is looking at it.
*/
private fun formatCheckTime(epochSeconds: Double): String =
try {
DateTimeFormatter.ofLocalizedTime(FormatStyle.SHORT)
.withZone(ZoneId.systemDefault())
.format(Instant.ofEpochSecond(epochSeconds.toLong()))
} catch (_: Exception) {
// A time that cannot be read is not a time to show: the sentence above still says a check
// is coming, which is the part the reader can act on.
"soon"
}
@@ -70,12 +70,21 @@ class UsageFeed(
/** Ask the backend again now. The dialog's refresh button; the poll does it on its own. */
val refresh: () -> Unit,
) {
/** What [setup]'s own limits came back as. See [usageFor] for why the states are these. */
fun forSetup(setup: String): SessionUsage =
when (val state = snapshots) {
/**
* What meters [session], and what that meter came back as. See [usageFor] for the states.
*
* A session rather than a machine, because a machine is not what is metered: one machine runs
* the Claude CLI and an echo session side by side, and only the first of them spends anything.
*/
fun forSession(session: SessionSummary): SessionUsage {
// Settled without asking anybody: a session nothing meters has nothing to check, and
// "checking" is what the fetch's own states would say about it for as long as one is out.
val provider = session.usageProvider ?: return SessionUsage.NotMetered
return when (val state = snapshots) {
is LoadState.Loading -> SessionUsage.Waiting
is LoadState.Error -> SessionUsage.Unavailable(state.message)
is LoadState.Loaded -> usageFor(state.value, setup)
is LoadState.Loaded -> usageFor(state.value, session.setup, provider)
}
}
}
@@ -158,9 +167,15 @@ fun SessionUsageBar(usage: SessionUsage, modifier: Modifier = Modifier) {
}
}
// Nothing at all for a machine that meters nothing: a row saying "unknown" there would report a
// problem about a setup somebody chose, on every screen, forever.
if (usage is SessionUsage.NotMetered) {
// Nothing at all for a session that meters nothing: a row saying "unknown" there would report
// a problem about a setup somebody chose, on every screen, forever.
//
// And nothing while the first fetch is out, which is a different silence. A request in flight
// is not a state to report -- and the session that meters nothing is exactly the one this
// cannot yet tell apart, so "5-hour usage: checking" appeared under an echo session for half a
// second and was then taken away. A row that has to be withdrawn is worse than one that
// arrives late.
if (usage is SessionUsage.NotMetered || usage is SessionUsage.Waiting) {
return
}
@@ -171,9 +186,10 @@ fun SessionUsageBar(usage: SessionUsage, modifier: Modifier = Modifier) {
// Words, not a colour and not an empty bar: every one of these is a different kind of
// answer from "this much is used", and only words carry a difference in kind.
when (val state = usage) {
SessionUsage.NotMetered -> Unit
// Both handled above, before the row exists at all.
SessionUsage.NotMetered,
SessionUsage.Waiting -> Unit
is SessionUsage.Unavailable -> UsageNote("5-hour usage unknown -- ${state.why}")
SessionUsage.Waiting -> UsageNote("5-hour usage: checking")
is SessionUsage.Known -> {
val window = state.windows.firstOrNull { it.kind == "session" }
if (window == null) {
@@ -234,16 +250,22 @@ private fun fiveHourLabel(window: UsageWindow, now: OffsetDateTime): String {
}
/**
* One machine's snapshot, out of every machine's.
* One meter's snapshot, out of every machine's: [setup]'s row for [provider].
*
* Both halves are needed to pick it. A machine can hold more than one meter -- the Claude CLI's
* account and, while a test has one set, an echo session's invented one -- and a snapshot is one
* service on one machine.
*
* Every way of having *failed* to get numbers is [SessionUsage.Unavailable] with the reason in it.
* None of them may look like zero, and none may look like [SessionUsage.NotMetered], which is the
* machine having no quota rather than the question going unanswered.
*/
fun usageFor(snapshots: List<UsageSnapshot>, setup: String): SessionUsage {
// No snapshot at all means the backend never asked, which it only does for a machine with
// nothing metered on it. That is a different answer from having asked and failed.
val mine = snapshots.firstOrNull { it.setup == setup } ?: return SessionUsage.NotMetered
fun usageFor(snapshots: List<UsageSnapshot>, setup: String, provider: String): SessionUsage {
// No snapshot at all means the backend never asked, which it only does where there is nothing
// to ask about. That is a different answer from having asked and failed.
val mine =
snapshots.firstOrNull { it.setup == setup && it.provider == provider }
?: return SessionUsage.NotMetered
if (mine.state != "ok") {
return SessionUsage.Unavailable(mine.detail ?: mine.state)
}
@@ -236,6 +236,7 @@ private fun AddSetupDialog(
var address by remember { mutableStateOf("") }
var identity by remember { mutableStateOf("") }
var attachmentsDir by remember { mutableStateOf("") }
var modelsDir by remember { mutableStateOf("") }
var tested by remember { mutableStateOf<String?>(null) }
var testing by remember { mutableStateOf(false) }
@@ -250,6 +251,7 @@ private fun AddSetupDialog(
port = typedPort,
identityFile = identity.trim().ifEmpty { null },
attachmentsDir = attachmentsDir.trim().ifEmpty { null },
modelsDir = modelsDir.trim().ifEmpty { null },
)
}
@@ -293,6 +295,14 @@ private fun AddSetupDialog(
label = { Text("Folder for attached files (optional)") },
singleLine = true,
)
// Where that machine's GGUFs are, for a llama.cpp session on it. Blank means
// the same place this backend keeps its own downloads, read on that machine.
OutlinedTextField(
value = modelsDir,
onValueChange = { modelsDir = it },
label = { Text("Folder for models (optional)") },
singleLine = true,
)
tested?.let {
Spacer(Modifier.height(8.dp))
Text(it, style = MaterialTheme.typography.bodySmall)
@@ -61,18 +61,31 @@ fun SpawnScreen(
// "auto" rather than "manual": on a phone every ask is a round trip to a question card, and
// answering "allow Bash?" dozens of times per task is what this app exists to avoid.
var permissionMode by remember { mutableStateOf("auto") }
// Null until the server has been asked, and null again if it answers "no level chosen" -- the
// two are told apart by [defaultsAsked], because a picker that shows a level before the answer
// arrives is one you can spawn at without having chosen it.
var effort by remember { mutableStateOf<String?>(null) }
var defaultsAsked by remember { mutableStateOf(false) }
var busy by remember { mutableStateOf(false) }
// Only the spawn's own failure. The fetch's lives in `options`: this one leaves a filled-in
// form worth keeping, and that one leaves nothing to fill in.
var spawnError by remember { mutableStateOf<String?>(null) }
// Downloaded models, for a llama provider to choose between. Kept separate from the setups: a
// Claude session needs none, so failing to list them must not stop the screen rendering.
// The models on the *chosen machine*, for a llama provider to choose between. Kept separate
// from the setups: a Claude session needs none, so failing to list them must not stop the
// screen rendering. Refetched when the machine changes, because a model is a file on one
// machine -- see [fetchSetupModels].
var models by remember { mutableStateOf<List<LocalModel>>(emptyList()) }
var modelKey by remember { mutableStateOf<String?>(null) }
var contextSize by remember { mutableStateOf("") }
var temperature by remember { mutableStateOf("") }
LaunchedEffect(Unit) {
// Separate from the setups fetch below and deliberately not fatal: failing to learn the
// default must leave a screen you can still spawn from, so the picker stays on "default"
// and says so rather than the whole form refusing to draw.
runCatching { withContext(Dispatchers.IO) { fetchDefaultEffort(settings) } }
.onSuccess { effort = it }
defaultsAsked = true
options =
try {
val fetched = withContext(Dispatchers.IO) { fetchSetups(settings) }
@@ -83,9 +96,6 @@ fun SpawnScreen(
} catch (e: ApiException) {
LoadState.failed(e)
}
models =
runCatching { withContext(Dispatchers.IO) { fetchModels(settings).local } }
.getOrDefault(emptyList())
}
Column(Modifier.fillMaxSize().verticalScroll(rememberScrollState()).padding(16.dp)) {
@@ -115,6 +125,17 @@ fun SpawnScreen(
is LoadState.Loaded -> state.value
}
val setup = setups.firstOrNull { it.name == setupName }
// Whichever machine is chosen now, asked again when that changes. The old machine's list
// is dropped first rather than left on screen: a file name from another machine looks
// exactly like one from this one.
LaunchedEffect(setup?.id) {
models = emptyList()
modelKey = null
val id = setup?.id ?: return@LaunchedEffect
models =
runCatching { withContext(Dispatchers.IO) { fetchSetupModels(settings, id) } }
.getOrDefault(emptyList())
}
val current = setup?.providers?.firstOrNull { it.name == providerName }
// Only the Claude CLI has models, a working directory and permission modes; keying the
// extra fields on the kind rather than the provider name keeps a second Claude provider
@@ -175,12 +196,13 @@ fun SpawnScreen(
)
if (isLlama) {
// A llama session names one of the models this backend has downloaded, so the choice is
// that list rather than free text -- a name that is not on disk is a session that
// cannot start.
// A llama session names one of the models on the machine it will run on, so the
// choice is that list rather than free text -- a name that is not on that machine's
// disk is a session that cannot start.
if (models.isEmpty()) {
Text(
"No models downloaded yet. Get one from the Models screen first.",
"No models on ${setup?.name ?: "this machine"}. The Models screen downloads " +
"to the backend; another machine needs the file put there itself.",
style = MaterialTheme.typography.bodyMedium,
color = MaterialTheme.colorScheme.onSurfaceVariant,
)
@@ -251,6 +273,19 @@ fun SpawnScreen(
selected = permissionMode,
onSelect = { permissionMode = it },
)
Spacer(Modifier.height(16.dp))
// Says what it does to *later* spawns as well, because it does: the level chosen here
// is stored as the default, which is the whole way that default is set. A picker that
// quietly changed a global would be the same control with the fact left out.
ChipGroup(
label = "Thinking (kept as the default for new sessions)",
options = listOf(DEFAULT_EFFORT) + EFFORT_LEVELS,
// The CLI's own default is a level in the list, so this cannot be a one-way trip.
// Disabled-looking until the server has answered, for the reason above.
selected = if (defaultsAsked) effort ?: DEFAULT_EFFORT else null,
onSelect = { chosen -> effort = chosen.takeIf { it != DEFAULT_EFFORT } },
)
}
Spacer(Modifier.height(24.dp))
@@ -268,6 +303,13 @@ fun SpawnScreen(
try {
val spawned =
withContext(Dispatchers.IO) {
// Stored before the spawn and not after it: choosing a level is
// an intent about new sessions in general, so a spawn that then
// fails must not also lose the choice. Non-fatal for the same
// reason the fetch above is -- the session is what was asked for.
if (isClaude) {
runCatching { setDefaultEffort(settings, effort) }
}
spawnSession(
settings,
// The id, not the label: labels are editable and the server
@@ -280,6 +322,7 @@ fun SpawnScreen(
if (isLlama) modelKey else model.trim().takeIf { isClaude },
cwd = cwd.trim().takeIf { isClaude },
permissionMode = permissionMode.takeIf { isClaude },
effort = effort.takeIf { isClaude },
// Sent only when set, so blank means "whatever llama.cpp does
// by default" rather than a zero.
params =
@@ -0,0 +1,27 @@
package com.example.aiapp
/**
* Where one transcript lives: a session's own, or one of its subagents'.
*
* The single mechanism [fetchTranscript], [EventStream], [TranscriptSource] and
* [TranscriptCache.session] all take, rather than each growing its own branch between a session and
* a subagent -- see SUBAGENTS.md's "Phone" and "Wire shape". A caller that has only a session id
* builds one with the one-argument constructor; a subagent's screen supplies both ids.
*/
data class TranscriptAddress(val sessionId: String, val subagentId: String? = null) {
/** The URL segment naming this transcript, before `/transcript` or `/events`. */
val urlPath: String
get() =
if (subagentId == null) "sessions/$sessionId"
else "sessions/$sessionId/subagents/$subagentId"
/**
* Where this transcript's cache lives on the phone, relative to the cache root.
*
* A subagent's nests under its session's directory rather than sitting beside it, so deleting a
* session's cache directory takes its subagents' with it -- the same one-way door the server's
* own storage describes.
*/
val cachePath: String
get() = if (subagentId == null) sessionId else "$sessionId/subagents/$subagentId"
}
@@ -35,8 +35,15 @@ class TranscriptCache(
private val root: File,
private val warn: (String) -> Unit = { Log.w("ai-app", it) },
) {
/** The cache for one session, whether or not anything has been stored for it yet. */
fun session(id: String): SessionCache = SessionCache(File(root, id), warn)
/**
* The cache for one transcript, whether or not anything has been stored for it yet.
*
* A subagent's [TranscriptAddress.cachePath] nests it under its session's directory, so
* deleting the session (below) takes its subagents' caches with it -- there is no separate
* purge for one.
*/
fun session(address: TranscriptAddress): SessionCache =
SessionCache(File(root, address.cachePath), warn)
/**
* Deletes every session directory not in [ids], called after a successful list fetch. The path
@@ -175,6 +175,18 @@ sealed class TranscriptItem {
val preTokens: Long?,
val postTokens: Long?,
) : TranscriptItem()
/**
* The account ran out of quota, so the turn stopped here.
*
* A divider rather than an error: nothing failed, and what a reader scrolling back needs from
* it is the same thing a clear or a compaction gives them -- why the conversation stops at this
* line.
*
* [resetsAt] is epoch seconds and null where the session was told nothing, which is a state the
* row has words for rather than a time it invents.
*/
data class LimitNote(override val seq: Long, val resetsAt: Double?) : TranscriptItem()
}
/**
@@ -458,6 +470,7 @@ fun foldEvent(items: List<TranscriptItem>, entry: SeqEvent): List<TranscriptItem
} else {
items + TranscriptItem.ImageItem(entry.seq, event.ref)
}
is SessionEvent.LimitReached -> items + TranscriptItem.LimitNote(entry.seq, event.resetsAt)
is SessionEvent.Cleared -> items + TranscriptItem.ClearedNote(entry.seq)
is SessionEvent.Compacted ->
items + TranscriptItem.CompactedNote(entry.seq, event.preTokens, event.postTokens)
@@ -18,7 +18,7 @@ import java.util.concurrent.atomic.AtomicReference
*/
class TranscriptSource(
private val settings: ServerSettings,
private val sessionId: String,
private val address: TranscriptAddress,
val cache: SessionCache,
) {
private val stream = AtomicReference<EventStream?>(null)
@@ -65,7 +65,7 @@ class TranscriptSource(
val tail = cache.tail() ?: return false
// `before = seq + 1` is the newest event with seq <= the cursor, which is the event *at*
// the cursor when the server still has one there.
val answer = fetchTranscript(settings, sessionId, before = tail.seq + 1, limit = 1)
val answer = fetchTranscript(settings, address, before = tail.seq + 1, limit = 1)
val matches =
answer.size == 1 &&
try {
@@ -83,7 +83,7 @@ class TranscriptSource(
*/
suspend fun fetchOpening(): List<SeqEvent> {
DebugStats.count("transcript page from server")
val page = fetchTranscript(settings, sessionId, limit = OPENING_WINDOW)
val page = fetchTranscript(settings, address, limit = OPENING_WINDOW)
page.forEach { (line, entry) -> cache.append(line, entry.seq) }
cache.flush()
return page.map { it.second }
@@ -108,7 +108,7 @@ class TranscriptSource(
val page =
fetchTranscript(
settings,
sessionId,
address,
before = before,
limit = limit,
coalesce = coalesce,
@@ -131,7 +131,7 @@ class TranscriptSource(
* well lose.
*/
fun follow(after: Long, onOpen: () -> Unit, onReset: () -> Unit, onEvent: (SeqEvent) -> Unit) {
val opened = EventStream(settings, sessionId)
val opened = EventStream(settings, address)
stream.getAndSet(opened)?.close()
try {
opened.run(after, onOpen, onReset) { raw, entry ->
@@ -0,0 +1,6 @@
<?xml version="1.0" encoding="utf-8"?>
<resources>
<!-- Overridden by the `bench` build type's resValue (build.gradle.kts) to "AI Sessions bench",
so the two are never mistaken for each other in the launcher or in Settings. -->
<string name="app_name">AI Sessions</string>
</resources>
@@ -0,0 +1,32 @@
package com.example.aiapp
import java.time.ZoneId
import kotlin.test.Test
import kotlin.test.assertEquals
import kotlin.test.assertTrue
/**
* What the transcript says where a session ran out of quota.
*
* The pair worth a test is the one that reads the same when it goes wrong: a reset time that
* arrived and one that never did. The second must not turn into a plausible-looking time, because a
* reader has no way of telling an invented one from a reported one.
*/
class LimitRowTest {
private val utc = ZoneId.of("UTC")
@Test
fun `a reported reset time is shown as a time`() {
// 2026-09-05T12:00:00Z. Asserted as a prefix and the clock reading rather than as the
// whole string: the platform's own short-time format is what this asks for, and it
// differs by JDK and locale down to which space character separates the meridiem.
val summary = limitSummary(1_788_609_600.0, utc)
assertTrue(summary.startsWith("Usage limit reached • resets "), summary)
assertTrue(summary.contains("12:00"), summary)
}
@Test
fun `a limit with no reset time says only what is known`() {
assertEquals("Usage limit reached", limitSummary(null, utc))
}
}
@@ -23,7 +23,7 @@ class TranscriptCacheTest {
private fun cache() = TranscriptCache(File(temp, "v1/host_8443")) { said += it }
private fun session(id: String = "s") = cache().session(id)
private fun session(id: String = "s") = cache().session(TranscriptAddress(id))
private fun line(seq: Long, type: String = "toolStart") =
"""{"seq":$seq,"ts":1.5,"type":"$type","id":"x"}"""
+28
View File
@@ -0,0 +1,28 @@
# The P0 benchmark fixture
`transcript.jsonl` is a synthetic transcript in the app's own event model (the JSON lines
`GET /sessions/{id}/transcript` returns; see `Events.kt`'s `parseSeqEvent` and
`server/src/session/driver.rs`) -- never a real one. It is what both the Compose `bench` build
and iris's bench build open with no server, so the two apps draw exactly the same content and a
frame-time comparison is measuring the renderer rather than the data.
Generated by `./generate.py` (Python stdlib only, seeded -- `SEED = 20260905` -- so re-running it
reproduces the same file byte for byte). It writes into `assets/` -- a separate directory from this
script and README, because the Compose `bench` build type points its own asset source set straight
at `assets/` (`app/androidApp/build.gradle.kts`'s `sourceSets { getByName("bench") }`), and a Python
script and a markdown file have no business inside an APK:
- `transcript.jsonl` -- 3,601 events. The first 3,200 (`BACKLOG_COUNT`) are the scrolled-back
history the benchmark opens with: user turns, tool calls with kilobyte-scale input/output,
assistant replies built from headings, bold/italic/inline code, a link, fenced code blocks that
rotate through rust/kotlin/python/sh/json/toml, a markdown table, two embedded images, and
periodic `usageDelta`/`compacted` events. The remaining 400 (`STREAM_COUNT`) are not part of the
opening window -- both bench harnesses replay them at a fixed rate (20/s) through the same live
fold path a real SSE reply arrives on, which is P0's "streaming phase."
- `bench1.png`, `bench2.png` -- tiny (8x8) flat-colour PNGs, base64-free on disk but served the
same way a real attachment is (`GET /sessions/{id}/files/{name}`), referenced by the two
`"type":"image"` events in the transcript.
Regenerate after changing the shape (a new event type, a different backlog/stream split) with
`./generate.py`, and commit the result -- it is checked in rather than generated at build time so
both apps' bench builds embed the identical bytes without needing this script at build time.
Binary file not shown.

After

Width:  |  Height:  |  Size: 74 B

Binary file not shown.

After

Width:  |  Height:  |  Size: 74 B

File diff suppressed because it is too large. Load diff
+186
View File
@@ -0,0 +1,186 @@
#!/usr/bin/env python3
"""Generates transcript.jsonl -- the synthetic fixture P0's benchmark opens in both apps.
Deterministic (fixed seed), so a Compose bench APK and an iris bench APK draw byte-identical
content: the point of the fixture is a like-for-like comparison, not a realistic one.
Never a real transcript -- see AGENTS.md's ui-sandbox.sh, which this borrows its vocabulary
style from (headings, code fences, a table, a link) rather than reusing its Claude-Code JSONL
shape. This file's shape is the *app's own event model* instead: one JSON object per line,
matching what GET /sessions/{id}/transcript returns and what Events.kt's parseSeqEvent reads
(server/src/session/driver.rs is the source of truth for the field names).
./generate.py writes transcript.jsonl and bench1.png/bench2.png here
BACKLOG_COUNT events (seq 1..BACKLOG_COUNT) are the scrolled-back history the benchmark opens
with. A further STREAM_COUNT events (seq BACKLOG_COUNT+1..) are not part of the opening window;
both bench harnesses replay them at a fixed rate as the "streaming reply" phase, appended through
the same live path a real SSE reply arrives on. Keeping both halves in one file means one
generator and one seed to keep in sync, rather than two fixtures that can drift apart.
"""
import base64
import json
import random
import struct
import zlib
from pathlib import Path
SEED = 20260905
BACKLOG_COUNT = 3200
STREAM_COUNT = 400
HERE = Path(__file__).resolve().parent / "assets"
random.seed(SEED)
LANGUAGES = ["rust", "kotlin", "python", "sh", "json", "toml"]
CODE_SNIPPETS = {
"rust": '''fn fold_event(items: Vec<Item>, seq: u64) -> Vec<Item> {
// a comment worth keeping: this is the fold the app's own screen runs
let mut out = items;
out.push(Item::new(seq));
out
}''',
"kotlin": '''fun foldEvent(items: List<TranscriptItem>, entry: SeqEvent): List<TranscriptItem> {
// mirrors the server's own event model, one item per line
return items + TranscriptItem.from(entry)
}''',
"python": '''def render_report(frames, cpu_ms, rss_kb):
# printed for a human to paste back, so every number carries its unit
return f"{frames} frames, {cpu_ms}ms cpu, {rss_kb}kb peak rss"''',
"sh": '''#!/bin/sh
# scripted scroll loop, the shape transcript-bench.sh drives on a phone
for i in $(seq 1 24); do
ui-trace record --do "swipe 540 700 540 1600 200"
done''',
"json": '{"seq": 1, "type": "status", "state": "running"}',
"toml": '''[package]
name = "bench-fixture"
version = "0.1.0"''',
}
HEADINGS = [
"## Plan",
"## What changed",
"## Why this approach",
"### Open questions",
"## Results",
]
WORDS = (
"session render report frame budget scroll transcript fold event cache "
"cursor probe stream backlog swipe fixture bench compose iris widget layout "
"measure place draw tool call token context window anchor"
).split()
def paragraph(n=24):
words = [random.choice(WORDS) for _ in range(n)]
words[0] = words[0].capitalize()
text = " ".join(words) + "."
# Sprinkle markdown inline spans so the syntax highlighter/markdown parser sees a real mix.
text = text.replace(" fold ", " **fold** ", 1)
text = text.replace(" cursor ", " *cursor* ", 1)
text = text.replace(" cache ", " `cache` ", 1)
if "bench" in text:
text = text.replace(
" bench ", " [bench](https://example.com/bench) ", 1
)
return text
def make_png(rgb, size=8):
"""A tiny, valid PNG -- flat colour, no external dependency."""
def chunk(tag, data):
c = tag + data
return struct.pack(">I", len(data)) + c + struct.pack(">I", zlib.crc32(c))
sig = b"\x89PNG\r\n\x1a\n"
ihdr = struct.pack(">IIBBBBB", size, size, 8, 2, 0, 0, 0)
raw = b""
for _ in range(size):
raw += b"\x00" + bytes(rgb) * size
idat = zlib.compress(raw)
return sig + chunk(b"IHDR", ihdr) + chunk(b"IDAT", idat) + chunk(b"IEND", b"")
def main():
HERE.mkdir(exist_ok=True)
lines = []
seq = 1
ts = 1_788_000_000.0
def emit(type_, **fields):
nonlocal seq, ts
obj = {"seq": seq, "ts": round(ts, 3), "type": type_}
obj.update(fields)
lines.append(json.dumps(obj, separators=(",", ":")))
seq += 1
ts += random.uniform(0.05, 2.0)
emit("status", state="running")
emit("settings", model="bench-model", permissionMode="auto")
image_refs = []
turn = 0
while seq <= BACKLOG_COUNT:
turn += 1
emit("userMessage", text=f"Turn {turn}: {paragraph(12)}", id=None, attachments=[])
# A tool call with kilobyte-scale input/output every few turns.
if turn % 3 == 0:
tool_id = f"tool-{turn}"
big_input = json.dumps({"path": f"/repo/file_{turn}.rs", "content": paragraph(400)})
emit("toolStart", id=tool_id, tool="Edit", input=big_input)
big_output = "\n".join(paragraph(60) for _ in range(20))
emit("toolUpdate", id=tool_id, output=big_output[: len(big_output) // 2])
emit("toolEnd", id=tool_id, output=big_output)
# A reply: a heading, prose, a fenced block in a rotating language, a table, then deltas.
emit("assistantText", delta=f"{random.choice(HEADINGS)}\n\n")
emit("assistantText", delta=paragraph(30) + "\n\n")
lang = LANGUAGES[turn % len(LANGUAGES)]
emit("assistantText", delta=f"```{lang}\n{CODE_SNIPPETS[lang]}\n```\n\n")
if turn % 5 == 0:
emit(
"assistantText",
delta="| column | value |\n|---|---|\n| a | " + paragraph(3) + " |\n\n",
)
# A run of small deltas -- the shape a live reply actually streams in.
for _ in range(random.randint(3, 8)):
emit("assistantText", delta=paragraph(6) + " ")
# A couple of images, base64 PNGs, the way a real transcript embeds a screenshot.
if turn in (10, 40):
ref = f"bench{len(image_refs) + 1}.png"
image_refs.append(ref)
emit("image", ref=ref, about=None)
emit("usageDelta", tokens=random.randint(200, 4000), context=random.randint(2000, 180000))
if turn % 15 == 0:
emit(
"compacted",
preTokens=180000,
postTokens=20000,
trigger="auto",
)
# The streaming-phase tail: one long reply, built entirely from text deltas, the shape a
# bench harness replays at a fixed events/sec through the live fold path.
emit("userMessage", text="One more, streamed live for the benchmark's timing phase.", id=None, attachments=[])
while seq <= BACKLOG_COUNT + STREAM_COUNT:
emit("assistantText", delta=paragraph(5) + " ")
emit("status", state="idle")
(HERE / "transcript.jsonl").write_text("\n".join(lines) + "\n")
(HERE / "bench1.png").write_bytes(make_png((220, 90, 90)))
(HERE / "bench2.png").write_bytes(make_png((90, 150, 220)))
print(f"wrote {len(lines)} events ({BACKLOG_COUNT} backlog + {STREAM_COUNT} stream) to transcript.jsonl")
if __name__ == "__main__":
main()
+9 -3
View File
@@ -4,6 +4,11 @@
# ./build-apk.sh the release build, signed (what the phone runs)
# ./build-apk.sh debug the debug build, for reproducing something the
# emulator scripts would build anyway
# ./build-apk.sh bench P0's benchmark build (own app id, "AI Sessions
# bench" label, opens straight onto the fixture
# session -- see docs/RUST.md's P0 box and
# app/bench-fixture/README.md). Signed the same
# as release; never touches the CA it pins.
#
# Dev Updater's `.dev-updater.ron` at the checkout root spells these out as
# build modes, one command line each; it passes nothing else, so the word
@@ -25,8 +30,9 @@ VARIANT=${1:-release}
case "$VARIANT" in
release) TASK=assembleRelease ;;
debug) TASK=assembleDebug ;;
bench) TASK=assembleBench ;;
*)
echo "build-apk.sh: unknown variant '$VARIANT' (release, debug)" >&2
echo "build-apk.sh: unknown variant '$VARIANT' (release, debug, bench)" >&2
exit 2
;;
esac
@@ -81,7 +87,7 @@ fi
# uninstalling it first: the signatures differ, and Android refuses to
# update across them.
KEYSTORE="${AI_APP_KEYSTORE:-${XDG_CONFIG_HOME:-$HOME/.config}/ai-app/release.jks}"
if [ "$VARIANT" = release ] && [ ! -f "$KEYSTORE" ]; then
if { [ "$VARIANT" = release ] || [ "$VARIANT" = bench ]; } && [ ! -f "$KEYSTORE" ]; then
KEYTOOL="${JAVA_HOME:+$JAVA_HOME/bin/keytool}"
KEYTOOL="${KEYTOOL:-keytool}"
if ! command -v "$KEYTOOL" >/dev/null 2>&1; then
@@ -97,7 +103,7 @@ if [ "$VARIANT" = release ] && [ ! -f "$KEYSTORE" ]; then
-keyalg RSA -keysize 2048 -validity 10000 \
-storepass "$PASSWORD" -keypass "$PASSWORD" -dname "CN=ai-app" >/dev/null 2>&1)
fi
if [ "$VARIANT" = release ]; then
if [ "$VARIANT" = release ] || [ "$VARIANT" = bench ]; then
AI_APP_KEYSTORE="$KEYSTORE"
AI_APP_KEYSTORE_PASSWORD=$(cat "$KEYSTORE.password")
export AI_APP_KEYSTORE AI_APP_KEYSTORE_PASSWORD
+34
View File
@@ -0,0 +1,34 @@
#!/bin/sh
# RUST.md's I5 "Where iris's frame time goes" pass (2026-09-05). The same
# 24-swipe/6-cycle loop as transcript-bench.sh's, extracted for iris's own
# demo app -- transcript-bench.sh itself is Compose-specific (opens by
# session title through the Compose app's own UI) and cannot be called
# directly against dev.iris.android.demo.
#
# MUST be run from inside this checkout (not /tmp): ui-trace/adb pick which
# emulator to target from the current directory's basename (the
# per-checkout-AVD rule), and a previous pass lost two attempts to a `cd`
# into /tmp that made this resolve to a nonexistent "tmp" checkout.
set -eu
cd "$(dirname "$0")"
. ./android-env.sh >/dev/null 2>&1
cycles=${1:-6}
ui-trace record -d 3000 --do "tap 'Reset frame report'" -o /tmp/iris-bench-reset.txt >/dev/null
adb logcat -c
DO=""
i=0
while [ "$i" -lt "$cycles" ]; do
DO="$DO --do 'swipe 540 700 540 1600 200' --do 'wait 500'"
DO="$DO --do 'swipe 540 700 540 1600 200' --do 'wait 500'"
DO="$DO --do 'swipe 540 1600 540 700 200' --do 'wait 500'"
DO="$DO --do 'swipe 540 1600 540 700 200' --do 'wait 500'"
i=$((i + 1))
done
eval ui-trace record -d $((cycles * 16000 + 20000)) $DO -o /tmp/iris-bench-scroll.txt >/dev/null
ui-trace record -d 3000 --do "tap 'Frame report'" -o /tmp/iris-bench-report.txt >/dev/null
sleep 1
adb logcat -d -s iris-android-app:I | grep "iris frame report:"
+16 -1
View File
@@ -404,12 +404,27 @@ done
# Percent-encoded because the app URL-decodes the deep link's query: a
# token with '+' in it enrols as one with a space, and nothing reports it.
enc=$(python3 -c 'import sys, urllib.parse; print(urllib.parse.quote(sys.argv[1], safe=""))' "$TOKEN")
# The CA rides in the link (`wg_app_link::enroll::ca_param`: base64url of
# the DER, which needs no percent-encoding). The Compose app ignores it and
# pins the copy its APK was built with; the iris app has no baked copy at
# all -- it is cross-compiled and could be pointed at any machine -- so
# without this it enrols and then trusts nothing. Minted here rather than by
# `--enroll-link` because this token is the sandbox's own, carried across
# restarts so the emulator stays enrolled (see the top of this file).
ca=$(python3 - "$CERTS/ca.pem" <<'CA'
import base64, sys
pem = open(sys.argv[1]).read()
body = pem.split("-----BEGIN CERTIFICATE-----")[1].split("-----END CERTIFICATE-----")[0]
der = base64.b64decode("".join(body.split()))
print(base64.urlsafe_b64encode(der).decode().rstrip("="))
CA
)
cat <<INFO
sandbox: server $pid on 127.0.0.1:$PORT, log $LOG
sandbox: 9 invented Claude Code sessions under $PROJECTS (one of them ${BIG_MB}MB)
enrol the emulator (once; it survives sandbox restarts):
adb shell "am start -a android.intent.action.VIEW -d 'aiapp://enroll?host=10.0.2.2&port=$PORT&token=$enc'"
adb shell "am start -a android.intent.action.VIEW -d 'aiapp://enroll?host=10.0.2.2&port=$PORT&token=$enc&ca=$ca'"
drive it:
./ui-sandbox.sh spawn [title] an echo session; prints its id
+49 -6
View File
@@ -46,7 +46,10 @@ checksum = "9330f8b2ff13f34540b44e946ef35111825727b38d33286ef986142615121801"
name = "client-core"
version = "0.1.0"
dependencies = [
"base64",
"event-model",
"log",
"pulldown-cmark",
"serde",
"serde_json",
"tempfile",
@@ -173,6 +176,15 @@ dependencies = [
"percent-encoding",
]
[[package]]
name = "getopts"
version = "0.2.24"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "cfe4fbac503b8d1f88e6676011885f34b7174f46e59956bba534ba83abded4df"
dependencies = [
"unicode-width",
]
[[package]]
name = "getrandom"
version = "0.2.17"
@@ -425,6 +437,25 @@ dependencies = [
"unicode-ident",
]
[[package]]
name = "pulldown-cmark"
version = "0.13.4"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "e9f068eba8e7071c5f9511831b44f32c740d5adf574e990f946ddb53db2f314e"
dependencies = [
"bitflags",
"getopts",
"memchr",
"pulldown-cmark-escape",
"unicase",
]
[[package]]
name = "pulldown-cmark-escape"
version = "0.11.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "007d8adb5ddab6f8e3f491ac63566a7d5002cc7ed73901f72057943fa71ae1ae"
[[package]]
name = "quote"
version = "1.0.47"
@@ -469,9 +500,9 @@ dependencies = [
[[package]]
name = "rustls"
version = "0.23.43"
version = "0.23.44"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "0283386ce02abc0151e1761d08802dfe86c173b0b494af5cbc086574e453da06"
checksum = "6725596c3f2c3a0aef021139e145d4eafe314a6623e4680ca83852b2c67ab2ba"
dependencies = [
"log",
"once_cell",
@@ -661,12 +692,24 @@ dependencies = [
"zerovec",
]
[[package]]
name = "unicase"
version = "2.9.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "dbc4bc3a9f746d862c45cb89d705aa10f187bb96c76001afab07a0d35ce60142"
[[package]]
name = "unicode-ident"
version = "1.0.24"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "e6e4313cd5fcd3dad5cafa179702e2b244f760991f45397d14d4ebf38247da75"
[[package]]
name = "unicode-width"
version = "0.2.2"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "b4ac048d71ede7ee76d585517add45da530660ef4390e49b098733c6e897f254"
[[package]]
name = "untrusted"
version = "0.9.0"
@@ -675,9 +718,9 @@ checksum = "8ecb6da28b8a351d773b68d5825ac39017e680750f980f3a1a85cd8dd28a47c1"
[[package]]
name = "ureq"
version = "3.4.0"
version = "3.4.1"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "972d7902c8735f2695410b8aed7df6ed12a47394aa1c8d7af49f0497b731a94d"
checksum = "af5546be8f5378d5414f83733f5c9a2526f4645829edbc1c41790aeef1b38e8b"
dependencies = [
"base64",
"cookie_store",
@@ -695,9 +738,9 @@ dependencies = [
[[package]]
name = "ureq-proto"
version = "0.6.1"
version = "0.6.2"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "da5f78b09e6941e1a0f2e30e695e4b120377b54d5e0aec11b594bb57b3971613"
checksum = "fabc3e92916c89c95b20eef7b06b00b066bc217ef9ea3a4ac9bf1a7e35261e10"
dependencies = [
"base64",
"http",
+22 -2
View File
@@ -7,7 +7,7 @@ edition = "2024"
# with `server/` via `event-model`), the REST + SSE clients for its HTTP
# surface (see `server/src/routes.rs`'s module doc for the table), the
# transcript fold and cache, the markdown block model, the syntax
# highlighter and the ANSI parser. See CLIENT_CORE.md at the repo root for
# highlighter and the ANSI parser. See `docs/CLIENT_CORE.md` for
# what this holds today, what it does not yet, and how it corresponds to
# the Kotlin it replaces.
#
@@ -17,7 +17,12 @@ edition = "2024"
[dependencies]
event-model = { path = "../event-model" }
serde = { version = "1", features = ["derive"] }
serde_json = { version = "1", features = ["float_roundtrip"] }
# "raw_value" is `fetch_transcript_lines`'s reason -- it needs the exact
# bytes the server sent, not this crate's own re-serialization of a parsed
# `Value`, so a cached line and a live SSE frame for the same event agree
# byte-for-byte (see that method's doc). "float_roundtrip" is why they
# agree on a `ts` at all -- see server/Cargo.toml's identical comment.
serde_json = { version = "1", features = ["float_roundtrip", "raw_value"] }
# The blocking HTTP client for the REST calls and the long-lived SSE GETs.
# `server/` already depends on ureq for its own outbound HTTPS (the usage
# poll in usage.rs) and it is rustls-backed like the rest of this project's
@@ -27,6 +32,21 @@ serde_json = { version = "1", features = ["float_roundtrip"] }
# no need of an async runtime, and RUST.md's brief for this port is
# "lightweight" throughout.
ureq = { version = "3", features = ["json"] }
# The markdown block split (`markdown_blocks`), which has to agree with the
# renderer in `iris/transcript-ui` about where a block begins -- so it is
# the same parser at the same version, rather than a hand-written splitter
# that would drift from it.
pulldown-cmark = "0.13.4"
# The enrollment link's `ca` parameter is base64url of the CA's DER
# (`config::parse_link`). Same version `wg-app-link` already pins for the
# minting half, so a workspace that has both resolves one copy.
base64 = "0.23"
# The logging facade only -- `log_ring` implements a `log::Log` backend and
# wraps whichever real one the platform installed (`android_logger` on the
# phone, `env_logger` on the desktop), which is why neither of those is a
# dependency here. See `log_ring`'s module doc.
log = { version = "0.4.34", features = ["std"] }
[dev-dependencies]
tempfile = "3"
+67 -1
View File
@@ -10,6 +10,7 @@
use std::io::Read;
use event_model::SeqEvent;
use serde::Deserialize;
use serde_json::Value;
@@ -116,6 +117,14 @@ impl<T: Transport> ApiClient<T> {
Self { transport }
}
/// The transport underneath, for a caller that needs the raw SSE
/// stream (`event_stream::follow_session_events`) rather than one of
/// this client's typed REST calls -- `transcript_source::TranscriptSource`
/// is the one that does.
pub fn transport(&self) -> &T {
&self.transport
}
fn json_request<R: for<'de> Deserialize<'de>>(
&self,
method: &str,
@@ -266,6 +275,61 @@ impl<T: Transport> ApiClient<T> {
limit: u32,
coalesce: bool,
) -> Result<Vec<Value>, ApiError> {
self.json_request(
"GET",
&transcript_path(session_id, before, limit, coalesce, None),
None,
)
}
/// A page of transcript history, each line handed back paired with the
/// exact text it came from, and bounded below by `after` -- the shape
/// `crate::transcript_source::TranscriptSource` needs to store what it
/// fetched in the transcript cache without a second round trip to fetch
/// the raw text separately. Ported from `Api.kt`'s `fetchTranscript`.
///
/// Uses [`serde_json::value::RawValue`] rather than re-serializing a
/// parsed [`Value`], so the stored line is the exact bytes the server
/// sent (key order and float literal included) rather than this
/// crate's own idea of how to write them back out -- the cache and a
/// live SSE frame must agree byte-for-byte on the same event, which is
/// exactly what caught the `serde_json` float-rounding bug this
/// project's `AGENTS.md` records.
pub fn fetch_transcript_lines(
&self,
session_id: &str,
before: Option<u64>,
limit: u32,
coalesce: bool,
after: Option<u64>,
) -> Result<Vec<(String, SeqEvent)>, ApiError> {
let path = transcript_path(session_id, before, limit, coalesce, after);
let raw: Vec<Box<serde_json::value::RawValue>> = self.json_request("GET", &path, None)?;
raw.into_iter()
.map(|value| {
let line = value.get().to_string();
let event: SeqEvent = serde_json::from_str(&line).map_err(|e| ApiError {
message: format!(
"the server sent a transcript line this build couldn't parse: {e}"
),
status: None,
})?;
Ok((line, event))
})
.collect()
}
}
/// The query string shared by [`ApiClient::fetch_transcript_page`] and
/// [`ApiClient::fetch_transcript_lines`], so the two agree on how each
/// parameter is written rather than keeping two copies to drift.
fn transcript_path(
session_id: &str,
before: Option<u64>,
limit: u32,
coalesce: bool,
after: Option<u64>,
) -> String {
let mut path = format!("/sessions/{session_id}/transcript?limit={limit}");
if let Some(before) = before {
path.push_str(&format!("&before={before}"));
@@ -273,8 +337,10 @@ impl<T: Transport> ApiClient<T> {
if coalesce {
path.push_str("&coalesce=true");
}
self.json_request("GET", &path, None)
if let Some(after) = after {
path.push_str(&format!("&after={after}"));
}
path
}
/// The blocking [`Transport`] backed by `ureq`, the same crate `server/`
+235 -19
View File
@@ -6,33 +6,57 @@
//! the same text a phone would scan as a QR, with no second format
//! invented for it (RUST.md's E4).
//!
//! What this type deliberately does not decide: where it is persisted, and
//! under what file permissions. A phone seals its token in the Android
//! Keystore; a desktop client has its own `$XDG_CONFIG_HOME/<app>/`
//! directory and its own file-mode conventions (MACHINE.md: owner-only,
//! never in the repo). Both are caller-specific, so they stay out of this
//! crate per the code rules' "ask for the least you need" -- see
//! `iris/desktop-app/src/config.rs` for the desktop instance.
//! [`EnrollmentStore`] persists one of these as JSON, owner-only, in a
//! directory the caller names -- `$XDG_CONFIG_HOME/ai-app-desktop` for the
//! desktop app, the app-private files directory on Android. **Which**
//! directory is the only part left to the platform: the format, the file
//! mode and the "nothing saved yet is not an error" answer are the same on
//! both, and were written twice before this.
//!
//! JSON rather than the project's usual RON: `wg-app-link`'s RON house
//! rules (`format`) are for configs a person hand-edits, and this file
//! never is one -- only the app itself writes or reads it.
use base64::Engine;
use serde::{Deserialize, Serialize};
use std::io;
use std::path::{Path, PathBuf};
/// One enrolled server: reachable at `https://{host}:{port}`, authenticated
/// with `token` as a bearer header. Does not carry the pinned CA -- that is
/// a public certificate rather than a secret, and where to find it differs
/// by caller (a phone pins the one its APK was built against; a desktop
/// client is told a path).
/// with `token` as a bearer header.
///
/// `ca_pem` is the trust anchor to pin, when the link carried one (the
/// `ca` parameter, `wg_app_link::enroll::ca_param`). It is optional
/// because an app built on the machine its server runs on pins the CA at
/// build time and needs nothing from the link; one built elsewhere -- the
/// iris Android client is cross-compiled in a VM and run against the
/// host's server -- has no other way to get it. A public certificate
/// rather than a secret, so it costs the link nothing but length.
#[derive(Debug, Clone, PartialEq, Serialize, Deserialize)]
pub struct EnrolledServer {
pub host: String,
pub port: u16,
pub token: String,
/// `#[serde(default)]` so an enrollment saved before this field
/// existed still loads, as the enrolled server it always was.
#[serde(default)]
pub ca_pem: Option<String>,
}
impl EnrolledServer {
/// Parses `aiapp://enroll?host=H&port=P&token=T` (query order does not
/// matter; unrecognised keys are ignored). `token` is percent-decoded,
/// since `ui-sandbox.sh` encodes it precisely because a raw token can
/// contain `+`, which turns into a space if left to a naive splitter.
/// Parses `aiapp://enroll?host=H&port=P&token=T[&ca=B]` (query order
/// does not matter; unrecognised keys are ignored). `token` is
/// percent-decoded, since `ui-sandbox.sh` encodes it precisely because
/// a raw token can contain `+`, which turns into a space if left to a
/// naive splitter.
///
/// `ca` is base64url of the certificate's DER and is rebuilt into PEM
/// here, because that is what every consumer of it wants
/// (`UreqTransport::new`, and the file a person points `curl --cacert`
/// at). A `ca` that does not decode fails the whole link rather than
/// enrolling a server with no trust anchor: the link said which
/// certificate to pin, and quietly not pinning it is the one outcome
/// nothing downstream could notice.
pub fn parse_link(link: &str) -> Result<Self, String> {
let query = link.split_once('?').map(|(_, q)| q).ok_or_else(|| {
format!(
@@ -44,6 +68,7 @@ impl EnrolledServer {
let mut host = None;
let mut port = None;
let mut token = None;
let mut ca = None;
for pair in query.split('&') {
let Some((key, value)) = pair.split_once('=') else {
continue;
@@ -53,6 +78,7 @@ impl EnrolledServer {
"host" => host = Some(value),
"port" => port = Some(value),
"token" => token = Some(value),
"ca" => ca = Some(value),
_ => {}
}
}
@@ -63,8 +89,14 @@ impl EnrolledServer {
.parse()
.map_err(|e| format!("'{link}''s port ('{port_str}') is not a number: {e}"))?;
let token = token.ok_or_else(|| format!("'{link}' is missing 'token'"))?;
let ca_pem = ca.map(|ca| pem_from_link_param(&ca)).transpose()?;
Ok(Self { host, port, token })
Ok(Self {
host,
port,
token,
ca_pem,
})
}
/// Where a `client_core::api::UreqTransport` reaches this server.
@@ -73,20 +105,94 @@ impl EnrolledServer {
}
}
/// The `ca` parameter (base64url of DER, unpadded) as a PEM certificate.
fn pem_from_link_param(ca: &str) -> Result<String, String> {
let der = base64::engine::general_purpose::URL_SAFE_NO_PAD
.decode(ca.as_bytes())
.map_err(|e| format!("the link's 'ca' is not base64url ({e})"))?;
let body = base64::engine::general_purpose::STANDARD.encode(&der);
let mut pem = String::from("-----BEGIN CERTIFICATE-----\n");
for line in body.as_bytes().chunks(64) {
pem.push_str(std::str::from_utf8(line).expect("base64 is ASCII"));
pem.push('\n');
}
pem.push_str("-----END CERTIFICATE-----\n");
Ok(pem)
}
/// Where one client keeps the enrollment it should not have to be told
/// about a second time. `dir` is the caller's, because that is the only
/// part that differs by platform -- see this module's doc.
pub struct EnrollmentStore {
dir: PathBuf,
}
impl EnrollmentStore {
pub fn new(dir: impl Into<PathBuf>) -> Self {
Self { dir: dir.into() }
}
pub fn dir(&self) -> &Path {
&self.dir
}
fn file(&self) -> PathBuf {
self.dir.join("enrollment.json")
}
/// Writes `server` under `dir`, creating it if needed, and sets the
/// file owner-only -- it carries a bearer token, the same reason
/// `server/`'s own token store is 0600.
pub fn save(&self, server: &EnrolledServer) -> io::Result<()> {
std::fs::create_dir_all(&self.dir)?;
let path = self.file();
let json = serde_json::to_vec_pretty(server)
.expect("EnrolledServer holds nothing that fails to serialise");
std::fs::write(&path, json)?;
#[cfg(unix)]
{
use std::os::unix::fs::PermissionsExt;
std::fs::set_permissions(&path, std::fs::Permissions::from_mode(0o600))?;
}
Ok(())
}
/// `Ok(None)` when nothing has been enrolled yet, rather than an error
/// -- "not enrolled" is an ordinary first-run state, not a failure
/// (UI_RULES' "a deliberate choice is not a problem to report" applies
/// just as well to a file that simply hasn't been written yet).
pub fn load(&self) -> io::Result<Option<EnrolledServer>> {
let path = self.file();
match std::fs::read(&path) {
Ok(bytes) => {
let server = serde_json::from_slice(&bytes).map_err(|e| {
io::Error::new(
io::ErrorKind::InvalidData,
format!("{} is not a valid enrollment ({e})", path.display()),
)
})?;
Ok(Some(server))
}
Err(e) if e.kind() == io::ErrorKind::NotFound => Ok(None),
Err(e) => Err(e),
}
}
}
fn percent_decode(s: &str) -> String {
let bytes = s.as_bytes();
let mut out = Vec::with_capacity(bytes.len());
let mut i = 0;
while i < bytes.len() {
if bytes[i] == b'%' && i + 2 < bytes.len() {
if let Ok(byte) =
if bytes[i] == b'%'
&& i + 2 < bytes.len()
&& let Ok(byte) =
u8::from_str_radix(std::str::from_utf8(&bytes[i + 1..i + 3]).unwrap_or(""), 16)
{
out.push(byte);
i += 3;
continue;
}
}
out.push(bytes[i]);
i += 1;
}
@@ -108,6 +214,7 @@ mod tests {
host: "127.0.0.1".to_string(),
port: 8547,
token: "abcDEF123".to_string(),
ca_pem: None,
}
);
assert_eq!(server.base_url(), "https://127.0.0.1:8547");
@@ -141,6 +248,115 @@ mod tests {
);
}
/// The CA travels as base64url of the DER and comes back out as the
/// PEM every consumer of it wants -- the same round trip
/// `wg_app_link::enroll::ca_param` mints.
#[test]
fn a_ca_in_the_link_comes_back_as_pem() {
let der = [0x30u8, 0x82, 0x01, 0xfb, 0x3e, 0x7f];
let param = base64::engine::general_purpose::URL_SAFE_NO_PAD.encode(der);
let server =
EnrolledServer::parse_link(&format!("aiapp://enroll?host=h&port=1&token=t&ca={param}"))
.unwrap();
let pem = server.ca_pem.expect("the link carried a CA");
assert!(pem.starts_with("-----BEGIN CERTIFICATE-----\n"), "{pem}");
assert!(
pem.trim_end().ends_with("-----END CERTIFICATE-----"),
"{pem}"
);
assert_eq!(
base64::engine::general_purpose::STANDARD
.decode(
pem.lines()
.filter(|l| !l.starts_with("-----"))
.collect::<String>()
)
.unwrap(),
der
);
}
/// A link with no `ca` is an ordinary link, not a broken one: an app
/// that pins at build time mints and reads exactly these.
#[test]
fn no_ca_parameter_is_none_not_an_error() {
let server = EnrolledServer::parse_link("aiapp://enroll?host=h&port=1&token=t").unwrap();
assert_eq!(server.ca_pem, None);
}
/// The half that cannot be noticed later: a `ca` that does not decode
/// must fail the link rather than enrolling with nothing pinned.
#[test]
fn a_ca_that_does_not_decode_fails_the_link() {
let err =
EnrolledServer::parse_link("aiapp://enroll?host=h&port=1&token=t&ca=not!base64url")
.unwrap_err();
assert!(err.contains("ca"), "{err}");
}
#[test]
fn a_saved_enrollment_reads_back_the_same() {
let dir = tempfile::tempdir().unwrap();
let store = EnrollmentStore::new(dir.path());
let server = EnrolledServer {
host: "127.0.0.1".to_string(),
port: 8547,
token: "tok".to_string(),
ca_pem: Some("-----BEGIN CERTIFICATE-----\nQUJD\n-----END CERTIFICATE-----\n".into()),
};
store.save(&server).unwrap();
assert_eq!(store.load().unwrap(), Some(server));
}
#[test]
fn nothing_saved_yet_is_none_not_an_error() {
let dir = tempfile::tempdir().unwrap();
assert_eq!(EnrollmentStore::new(dir.path()).load().unwrap(), None);
}
/// An enrollment written before `ca_pem` existed still loads.
#[test]
fn an_enrollment_without_a_ca_still_loads() {
let dir = tempfile::tempdir().unwrap();
let store = EnrollmentStore::new(dir.path());
std::fs::create_dir_all(dir.path()).unwrap();
std::fs::write(
dir.path().join("enrollment.json"),
br#"{"host":"h","port":1,"token":"t"}"#,
)
.unwrap();
assert_eq!(store.load().unwrap().unwrap().ca_pem, None);
}
#[test]
#[cfg(unix)]
fn the_saved_file_is_owner_only() {
use std::os::unix::fs::PermissionsExt;
let dir = tempfile::tempdir().unwrap();
let store = EnrollmentStore::new(dir.path());
store
.save(&EnrolledServer {
host: "h".to_string(),
port: 1,
token: "t".to_string(),
ca_pem: None,
})
.unwrap();
let mode = std::fs::metadata(dir.path().join("enrollment.json"))
.unwrap()
.permissions()
.mode();
assert_eq!(mode & 0o777, 0o600);
}
#[test]
fn a_corrupt_file_is_named_in_the_error() {
let dir = tempfile::tempdir().unwrap();
std::fs::write(dir.path().join("enrollment.json"), b"not json").unwrap();
let err = EnrollmentStore::new(dir.path()).load().unwrap_err();
assert!(err.to_string().contains("enrollment.json"));
}
#[test]
fn a_non_numeric_port_is_named_in_the_error() {
let err = EnrolledServer::parse_link("aiapp://enroll?host=h&port=x&token=t").unwrap_err();
+100
View File
@@ -0,0 +1,100 @@
//! A span of milliseconds, written the way somebody reads it -- the port
//! of `Durations.kt`'s `formatMillis`/`formatMillisText`, with its tests.
//!
//! Only the tool-timeout half is here. `formatSpan` (the usage
//! countdown's rounding-up rule) belongs with whatever draws the usage
//! bar, and nothing in this crate needs it yet.
/// A span of milliseconds, written the way somebody reads it.
///
/// A tool's timeout arrives as `480000`, which nobody reads as eight
/// minutes. The rule has two halves, because a short span and a long one
/// are read for different things. Under a minute the question is "roughly
/// how long", so only the largest unit is shown and a fraction carries the
/// rest -- `2.5s`. At a minute or more the question is "how long exactly",
/// so every unit with something in it is written out -- `5d 12h 4m`. Empty
/// units are left out rather than written as zero.
///
/// Sub-second precision is dropped past a minute: nothing that takes days
/// is measured in milliseconds.
pub fn format_millis(ms: i64) -> String {
if ms < 0 {
return format!("-{}", format_millis(-ms));
}
if ms < 1000 {
return format!("{ms}ms");
}
if ms < 60_000 {
let tenths = (ms + 50) / 100;
let (whole, rest) = (tenths / 10, tenths % 10);
return if rest == 0 {
format!("{whole}s")
} else {
format!("{whole}.{rest}s")
};
}
let seconds = ms / 1000;
[
("d", seconds / 86_400),
("h", seconds / 3600 % 24),
("m", seconds / 60 % 60),
("s", seconds % 60),
]
.iter()
.filter(|(_, n)| *n > 0)
.map(|(unit, n)| format!("{n}{unit}"))
.collect::<Vec<_>>()
.join(" ")
}
/// `text` as a span when it is a whole number of milliseconds, and
/// unchanged when it is not.
pub fn format_millis_text(text: &str) -> String {
match text.trim().parse::<i64>() {
Ok(ms) => format_millis(ms),
Err(_) => text.to_string(),
}
}
#[cfg(test)]
mod tests {
use super::*;
/// The two ways a span of time is written here, and the rule each of
/// them follows -- ported from `DurationsTest.kt`, whose doc says why:
/// both are read off a screen to make a decision, so what matters is
/// that the shortest form that answers the question is what appears.
#[test]
fn under_a_minute_is_the_largest_unit_alone() {
assert_eq!(format_millis(30), "30ms");
assert_eq!(format_millis(999), "999ms");
assert_eq!(format_millis(1000), "1s");
assert_eq!(format_millis(2500), "2.5s");
// One decimal, rounded rather than cut: 2.46s is nearer two and a
// half than two and four.
assert_eq!(format_millis(2460), "2.5s");
assert_eq!(format_millis(59_900), "59.9s");
}
#[test]
fn a_minute_or_more_is_every_unit_that_has_something_in_it() {
// The figure this rule was written for: a tool timeout, which
// arrives as milliseconds and is unreadable as 480000.
assert_eq!(format_millis(480_000), "8m");
assert_eq!(format_millis(60_000), "1m");
assert_eq!(format_millis(90_000), "1m 30s");
assert_eq!(format_millis(475_440_000), "5d 12h 4m");
// Empty units are left out rather than written as zero: the labels
// say which is which, and "5d 0h 4m" is only longer.
assert_eq!(format_millis(432_240_000), "5d 4m");
}
#[test]
fn only_a_whole_number_of_milliseconds_is_rewritten() {
assert_eq!(format_millis_text(" 480000 "), "8m");
// A timeout a tool expressed some other way is its own words,
// passed through rather than guessed at.
assert_eq!(format_millis_text("2 minutes"), "2 minutes");
assert_eq!(format_millis_text(""), "");
}
}
+6 -1
View File
@@ -1,15 +1,20 @@
//! The app's pure logic, shared between the server and any Rust client --
//! see `CLIENT_CORE.md` at the repo root for what lives here and what does
//! see `docs/CLIENT_CORE.md` for what lives here and what does
//! not yet.
pub mod ansi;
pub mod api;
pub mod config;
pub mod durations;
pub mod event_stream;
pub mod highlight;
pub mod log_ring;
pub mod markdown_blocks;
pub mod notifications;
pub mod sse;
pub mod tool_summary;
pub mod transcript_cache;
pub mod transcript_fold;
pub mod transcript_source;
pub use event_model::*;
+812
View File
@@ -0,0 +1,812 @@
//! The app's own recent log, held in memory so it can be read back
//! without `logcat`.
//!
//! **Why this exists**: Iris tests iris builds on a GrapheneOS phone with
//! no `adb`, and Android forbids one app reading another's logcat, so
//! nothing outside the process can recover what it wrote. The only way a
//! line reaches her is for the app to carry its own copy. This is that
//! copy: a bounded ring every `log::info!` in the process lands in, on top
//! of whichever platform logger was already installed (`android_logger`,
//! `env_logger`) rather than instead of it -- see [`RingLogger`].
//!
//! Two consumers, both reading the same ring rather than each keeping
//! their own: the bench app's `Copy report`/`Diagnostics` (which reads
//! [`LogRing::tail_text`] and [`LogRing::summary`]) and whatever hands the
//! log out of the process -- on Android, the `DevLogProvider` Dev Updater
//! queries, which reads [`LogRing::since`] and [`LogRing::newest_seq`].
//! That is why reading does not consume: a line already handed over must
//! still be in the report, and a report taken twice must say the same
//! thing.
use std::collections::VecDeque;
use std::sync::{Arc, Mutex, OnceLock};
use std::time::{SystemTime, UNIX_EPOCH};
/// How many lines a default ring holds, and how many bytes of message.
///
/// Both bounds apply -- whichever bites first -- because the two failure
/// modes are different: a flood of short lines exhausts the count, and one
/// pathological line (a stack trace, a pretty-printed JSON body) exhausts
/// the bytes. A ring bounded only by lines can hold megabytes; one bounded
/// only by bytes can be emptied by a single line.
pub const DEFAULT_MAX_LINES: usize = 2000;
pub const DEFAULT_MAX_BYTES: usize = 256 * 1024;
/// How many of the ring's newest lines [`LogRing::tail_text`] includes.
/// Sized for a phone's share sheet rather than for the ring itself: 150
/// lines of `HH:MM:SS.mmm LEVEL target: message` is a few KiB, comfortably
/// short of whatever made pasting the full (up to 2000-line) ring into a
/// chat's message box laggy on Iris's phone. The full ring is still
/// reachable through `devlog`'s provider, so this only bounds what a
/// report inlines.
pub const COPY_REPORT_TAIL_LINES: usize = 150;
/// One recorded line. `seq` is assigned by the ring and only ever
/// increases, so a reader that remembers where it got to can ask for what
/// came after -- and a gap in the sequence is exactly the lines the bound
/// dropped.
#[derive(Debug, Clone, PartialEq, Eq)]
pub struct LogLine {
pub seq: u64,
/// Milliseconds since the unix epoch, from the app's own clock. The
/// app's rather than the receiver's: a line is timestamped when it
/// happened, and an upload can be minutes later or never.
pub at_ms: u64,
pub level: log::Level,
pub target: String,
pub message: String,
}
impl LogLine {
/// Roughly what the line costs the ring. The two `String`s dominate;
/// the fixed fields are counted as a flat overhead so a ring of empty
/// messages still has a bound.
fn weight(&self) -> usize {
self.target.len() + self.message.len() + 32
}
/// `12:34:56.789 INFO iris::android: the message`, the shape a
/// person skims. Time of day only -- the date is in the report's own
/// header, and a ring never spans one.
pub fn format(&self) -> String {
format!(
"{} {:<5} {}: {}",
clock_time(self.at_ms),
self.level,
self.target,
self.message
)
}
}
/// `HH:MM:SS.mmm` in UTC from a unix millisecond count, without a date
/// library: the only field this needs is the time of day, and dividing out
/// the day is the whole calculation. Deliberately not local time -- the
/// phone's offset is not knowable here, and a report that says UTC is
/// comparable with the server's log, which is what it gets read against.
fn clock_time(at_ms: u64) -> String {
let ms = at_ms % 1000;
let secs_of_day = (at_ms / 1000) % 86_400;
format!(
"{:02}:{:02}:{:02}.{:03}",
secs_of_day / 3600,
(secs_of_day % 3600) / 60,
secs_of_day % 60,
ms
)
}
/// Now, in unix milliseconds. Saturating rather than panicking on a clock
/// before the epoch: a wrong timestamp in a diagnostic is not worth taking
/// the app down for.
pub fn now_ms() -> u64 {
SystemTime::now()
.duration_since(UNIX_EPOCH)
.map(|d| d.as_millis() as u64)
.unwrap_or(0)
}
#[derive(Debug)]
struct Inner {
lines: VecDeque<LogLine>,
bytes: usize,
max_lines: usize,
max_bytes: usize,
next_seq: u64,
/// How many lines the bounds have discarded since the ring was made.
/// Reported rather than inferred, so "the log starts here" and "the
/// log was cut off here" are distinguishable -- the unknown state the
/// UI rules ask for.
dropped: u64,
}
/// A bounded, shareable ring of recent log lines. Cloning shares the ring;
/// there is one per process and every holder sees the same lines.
#[derive(Debug, Clone)]
pub struct LogRing(Arc<Mutex<Inner>>);
impl LogRing {
pub fn new(max_lines: usize, max_bytes: usize) -> Self {
assert!(
max_lines > 0 && max_bytes > 0,
"a ring with no room holds nothing"
);
Self(Arc::new(Mutex::new(Inner {
lines: VecDeque::new(),
bytes: 0,
max_lines,
max_bytes,
next_seq: 0,
dropped: 0,
})))
}
/// The bounds this project ships with: [`DEFAULT_MAX_LINES`] and
/// [`DEFAULT_MAX_BYTES`].
pub fn with_defaults() -> Self {
Self::new(DEFAULT_MAX_LINES, DEFAULT_MAX_BYTES)
}
/// A poisoned lock is a bug in a panicking logger, not a reason to
/// take the app down a second time -- the ring is a diagnostic, and
/// losing it must not be worse than the fault it was recording.
fn with<R>(&self, f: impl FnOnce(&mut Inner) -> R) -> R {
let mut guard = match self.0.lock() {
Ok(guard) => guard,
Err(poisoned) => poisoned.into_inner(),
};
f(&mut guard)
}
/// Records a line, evicting the oldest until both bounds hold again.
pub fn push(&self, level: log::Level, target: &str, message: String) {
self.with(|inner| {
let line = LogLine {
seq: inner.next_seq,
at_ms: now_ms(),
level,
target: target.to_string(),
message,
};
inner.next_seq += 1;
inner.bytes += line.weight();
inner.lines.push_back(line);
// `!is_empty()` rather than `len() > 1`: one line larger than
// the whole byte bound is kept, because dropping it would
// leave the ring silently empty while lines were arriving.
while inner.lines.len() > inner.max_lines
|| (inner.bytes > inner.max_bytes && inner.lines.len() > 1)
{
if let Some(evicted) = inner.lines.pop_front() {
inner.bytes -= evicted.weight();
inner.dropped += 1;
}
}
})
}
/// Every line held, oldest first.
pub fn snapshot(&self) -> Vec<LogLine> {
self.with(|inner| inner.lines.iter().cloned().collect())
}
/// The lines with a sequence number at or after `seq`, oldest first,
/// and the sequence to ask from next time. Does not consume: see this
/// module's doc for why.
pub fn since(&self, seq: u64) -> (Vec<LogLine>, u64) {
self.with(|inner| {
let lines: Vec<LogLine> = inner
.lines
.iter()
.filter(|line| line.seq >= seq)
.cloned()
.collect();
let next = lines.last().map(|line| line.seq + 1).unwrap_or(seq);
(lines, next)
})
}
pub fn len(&self) -> usize {
self.with(|inner| inner.lines.len())
}
pub fn is_empty(&self) -> bool {
self.len() == 0
}
pub fn dropped(&self) -> u64 {
self.with(|inner| inner.dropped)
}
/// The sequence number of the newest line held, or `None` for a ring
/// nothing has been written to.
///
/// What a reader needs to notice that this process **restarted**: the
/// ring is in memory, so a new process starts again at zero, and a
/// reader holding a cursor from the previous one would otherwise ask
/// for lines after a number nothing will reach for hours and see
/// nothing at all -- silently, which is worse than seeing the log
/// begin again. Answering `None` rather than 0 for an empty ring is
/// the same distinction [`Self::summary`] draws: "nothing has been
/// logged" is not a sequence number.
pub fn newest_seq(&self) -> Option<u64> {
self.with(|inner| inner.lines.back().map(|line| line.seq))
}
/// When the newest line was written, in unix milliseconds, or `None`
/// for a ring nothing has been written to.
pub fn last_at_ms(&self) -> Option<u64> {
self.with(|inner| inner.lines.back().map(|line| line.at_ms))
}
/// Every line held, formatted one per line -- what `Copy report`
/// appends.
pub fn to_text(&self) -> String {
self.snapshot()
.iter()
.map(LogLine::format)
.collect::<Vec<_>>()
.join("\n")
}
/// The newest `max_lines` lines, formatted, with a first line naming
/// how many older ones were left out of *this* text when the ring held
/// more than that -- what `Copy report` appends instead of
/// [`Self::to_text`].
///
/// Iris's own report: pasting the full ring (over a thousand lines on
/// a session that ran with tracing on) into a phone's message box was
/// what "causes a lot of lag" meant (docs/IRIS_TODO.md, 2026-09-07
/// night) -- nothing is actually lost, since `devlog`'s provider still
/// hands Dev Updater's Runtime tab the whole ring; this only caps what
/// gets inlined into a share.
pub fn tail_text(&self, max_lines: usize) -> String {
let lines = self.snapshot();
if lines.len() <= max_lines {
return lines
.iter()
.map(LogLine::format)
.collect::<Vec<_>>()
.join("\n");
}
let omitted = lines.len() - max_lines;
let tail = lines[lines.len() - max_lines..]
.iter()
.map(LogLine::format)
.collect::<Vec<_>>()
.join("\n");
format!("{omitted} earlier lines omitted; full log in Dev Updater's Runtime tab\n{tail}")
}
/// The newest `max_lines` lines, formatted, or `None` if the ring is
/// locked at this instant.
///
/// For the one caller that must not block: **the panic hook**. A panic
/// raised while this ring's own lock was held -- an allocation failing
/// inside [`Self::push`], an assertion in a `log::Log` on the way here
/// -- would deadlock the hook against the thread that is panicking,
/// and the process would hang instead of aborting, with nothing
/// written anywhere. Losing the context lines is the right trade
/// against that, and `None` says which happened rather than looking
/// like an empty log.
pub fn try_tail_text(&self, max_lines: usize) -> Option<String> {
let guard = match self.0.try_lock() {
Ok(guard) => guard,
// A poisoned lock is uncontended, so its contents are still
// readable -- the same judgement as `with`.
Err(std::sync::TryLockError::Poisoned(poisoned)) => poisoned.into_inner(),
Err(std::sync::TryLockError::WouldBlock) => return None,
};
let lines = &guard.lines;
let from = lines.len().saturating_sub(max_lines);
Some(
lines
.iter()
.skip(from)
.map(LogLine::format)
.collect::<Vec<_>>()
.join("\n"),
)
}
/// One line for a diagnostics pane: how much is held, how much was
/// dropped, and when the last line arrived. "no lines yet" is its own
/// wording rather than a count of zero with a made-up time, because
/// "nothing has been logged" and "logging is not running" would
/// otherwise look the same.
pub fn summary(&self) -> String {
let (len, dropped, last) = self.with(|inner| {
(
inner.lines.len(),
inner.dropped,
inner.lines.back().map(|line| line.at_ms),
)
});
match last {
None => "app log: no lines yet".to_string(),
Some(at) => {
let dropped = if dropped > 0 {
format!(", {dropped} dropped")
} else {
String::new()
};
format!(
"app log: {len} lines held{dropped}, last {}",
clock_time(at)
)
}
}
}
}
/// Whether a target belongs to this app's own crates (`iris` or
/// `client_core`) rather than a dependency's -- `starts_with` guarded by an
/// exact match or a `::` so an unrelated crate that merely begins with the
/// same letters (there is no such crate today, but the check should not
/// rely on that) is never mistaken for one of ours.
fn is_own_target(target: &str) -> bool {
target == "iris"
|| target.starts_with("iris::")
|| target == "client_core"
|| target.starts_with("client_core::")
}
/// Whether a line at `level` from `target` belongs in the ring, given
/// whether tracing is on right now.
///
/// This is the one filter docs/IRIS_TODO.md's "logs way too big" entry
/// asked for, applied once here rather than at each `debug!` call site:
/// Info and above always ring, from anything, because a real warning or
/// error from a dependency is worth keeping. Debug and Trace ring only
/// from this app's own targets, and only while tracing is switched on --
/// otherwise `naga::front`/`wgpu_core`/`jni` log at Debug unconditionally
/// (the process logger's own level, set once at install and unrelated to
/// tracing), which is what filled the ring with 1339 lines of it and
/// dropped 4050 more before this existed. `iris`'s own Debug lines already
/// self-gate on `iris::diagnostics::trace_enabled` at their call sites
/// (commit 992c472); this is the backstop for lines this crate does not
/// control.
fn ring_accepts(level: log::Level, target: &str, trace_enabled: bool) -> bool {
level <= log::Level::Info || (trace_enabled && is_own_target(target))
}
/// A `log` backend that records into a [`LogRing`] **and** forwards to the
/// logger the platform already installs, so nothing that reads the
/// platform's log (`logcat`, a terminal) changes.
///
/// The inner logger is passed in rather than chosen here: `client-core`
/// has no business depending on `android_logger` or `env_logger`, and
/// which one is right is exactly what differs between the two platforms
/// (the sharing rule in AGENTS.md).
pub struct RingLogger {
ring: LogRing,
inner: Box<dyn log::Log>,
/// Whether `iris::input`/`iris::frame`-style tracing is switched on
/// right now, consulted by [`ring_accepts`]. A plain fn pointer rather
/// than a dependency on `iris::diagnostics::trace_enabled` directly:
/// `client-core` sits below `iris` (AGENTS.md's "dependencies flow one
/// direction"), so the platform crate that depends on both is the one
/// that wires this closure through, the same way it already supplies
/// `inner`.
trace_enabled: fn() -> bool,
}
impl RingLogger {
pub fn new(ring: LogRing, inner: Box<dyn log::Log>, trace_enabled: fn() -> bool) -> Self {
Self {
ring,
inner,
trace_enabled,
}
}
}
impl log::Log for RingLogger {
/// True for anything `log`'s own max level lets through: the ring
/// wants everything the *inner* logger might also want, even where the
/// platform logger would filter it out. Which lines the ring itself
/// keeps is decided in [`Self::log`] by [`ring_accepts`].
fn enabled(&self, _metadata: &log::Metadata) -> bool {
true
}
fn log(&self, record: &log::Record) {
if ring_accepts(record.level(), record.target(), (self.trace_enabled)()) {
self.ring
.push(record.level(), record.target(), record.args().to_string());
}
if self.inner.enabled(record.metadata()) {
self.inner.log(record);
}
}
fn flush(&self) {
self.inner.flush();
}
}
/// Installs a [`RingLogger`] as the process logger and answers the ring it
/// records into.
///
/// Fails only if a logger is already installed, which is a programmer
/// error (two initialisation paths) rather than a recoverable condition --
/// the caller is named in the error so it is findable.
pub fn install(
ring: LogRing,
inner: Box<dyn log::Log>,
max_level: log::LevelFilter,
trace_enabled: fn() -> bool,
) -> Result<(), log::SetLoggerError> {
log::set_boxed_logger(Box::new(RingLogger::new(ring, inner, trace_enabled)))?;
log::set_max_level(max_level);
Ok(())
}
/// The one ring this process records into.
///
/// **A deliberate process-global, where this project's rules otherwise say
/// pass context explicitly.** What is being modelled is already one: `log`
/// has exactly one backend per process, set once, and every `log::info!`
/// anywhere in the binary goes to it. A ring handed around as a parameter
/// would be a *second* answer to "which lines exist" -- the report would
/// show one ring while the logger filled another, and which one a caller
/// got would depend on how far down the call tree it was. The tests above
/// all use their own [`LogRing`], so nothing here needs this to be
/// testable.
static PROCESS_RING: OnceLock<LogRing> = OnceLock::new();
/// The process's ring, created on first use with the default bounds.
/// Safe to call before [`install_process_logger`] -- it will simply be
/// empty.
pub fn process_ring() -> &'static LogRing {
PROCESS_RING.get_or_init(LogRing::with_defaults)
}
/// Installs [`process_ring`] as the recording half of the process logger,
/// forwarding to `inner` (the platform's own logger, already configured).
/// The platform half of AGENTS.md's sharing rule is `inner`; everything
/// else is shared. `trace_enabled` is the platform's own trace toggle
/// (`iris::diagnostics::trace_enabled` on Android) -- see
/// [`ring_accepts`] and the field doc on `RingLogger` for why it is
/// passed in rather than called directly.
pub fn install_process_logger(
inner: Box<dyn log::Log>,
max_level: log::LevelFilter,
trace_enabled: fn() -> bool,
) -> Result<(), log::SetLoggerError> {
install(process_ring().clone(), inner, max_level, trace_enabled)
}
#[cfg(test)]
mod tests {
use super::*;
use log::Level;
fn fill(ring: &LogRing, count: usize) {
for n in 0..count {
ring.push(Level::Info, "test", format!("line {n}"));
}
}
#[test]
fn lines_come_back_oldest_first() {
let ring = LogRing::new(10, 1 << 20);
fill(&ring, 3);
let text: Vec<String> = ring.snapshot().into_iter().map(|l| l.message).collect();
assert_eq!(text, ["line 0", "line 1", "line 2"]);
}
#[test]
fn the_line_bound_drops_the_oldest_and_says_how_many() {
let ring = LogRing::new(3, 1 << 20);
fill(&ring, 5);
let text: Vec<String> = ring.snapshot().into_iter().map(|l| l.message).collect();
assert_eq!(text, ["line 2", "line 3", "line 4"], "the newest survive");
assert_eq!(ring.len(), 3);
assert_eq!(ring.dropped(), 2, "and the loss is reported, not silent");
}
#[test]
fn the_byte_bound_bites_before_the_line_bound_when_lines_are_large() {
// Room for 1000 lines but only a few hundred bytes.
let ring = LogRing::new(1000, 300);
for n in 0..10 {
ring.push(Level::Info, "t", format!("{n}{}", "x".repeat(100)));
}
assert!(
ring.len() < 10,
"the byte bound evicted: {} held",
ring.len()
);
assert!(ring.dropped() > 0);
assert!(
ring.snapshot().last().unwrap().message.starts_with('9'),
"and it evicted from the old end"
);
}
/// The case the `len() > 1` guard exists for: one line larger than the
/// whole bound must still be readable, or a ring that is over budget
/// reads as a ring nothing was written to.
#[test]
fn one_oversized_line_is_kept_rather_than_leaving_the_ring_empty() {
let ring = LogRing::new(100, 64);
ring.push(Level::Error, "t", "y".repeat(5000));
assert_eq!(ring.len(), 1);
assert_eq!(ring.dropped(), 0);
}
#[test]
fn sequence_numbers_only_increase_and_survive_eviction() {
let ring = LogRing::new(2, 1 << 20);
fill(&ring, 5);
let seqs: Vec<u64> = ring.snapshot().into_iter().map(|l| l.seq).collect();
assert_eq!(seqs, [3, 4], "a gap is exactly what was dropped");
}
#[test]
fn since_returns_only_what_is_new_and_the_next_cursor() {
let ring = LogRing::new(100, 1 << 20);
fill(&ring, 3);
let (first, cursor) = ring.since(0);
assert_eq!(first.len(), 3);
assert_eq!(cursor, 3);
let (none, cursor) = ring.since(cursor);
assert!(none.is_empty(), "nothing new yet");
assert_eq!(cursor, 3, "and the cursor does not move");
ring.push(Level::Warn, "test", "later".into());
let (more, cursor) = ring.since(cursor);
assert_eq!(more.len(), 1);
assert_eq!(more[0].message, "later");
assert_eq!(cursor, 4);
}
/// The restart signal: a reader that saw sequence 4 and is now told
/// the newest is 0 knows the process is not the one it was reading.
#[test]
fn the_newest_sequence_says_where_the_ring_is_and_nothing_for_an_empty_one() {
let ring = LogRing::new(100, 1 << 20);
assert_eq!(ring.newest_seq(), None, "an empty ring has no newest line");
fill(&ring, 5);
assert_eq!(ring.newest_seq(), Some(4));
let restarted = LogRing::new(100, 1 << 20);
fill(&restarted, 1);
assert_eq!(
restarted.newest_seq(),
Some(0),
"a fresh ring starts again, which is exactly what a reader has to notice"
);
}
#[test]
fn reading_does_not_consume() {
let ring = LogRing::new(100, 1 << 20);
fill(&ring, 2);
let (sent, _) = ring.since(0);
assert_eq!(sent.len(), 2);
assert_eq!(ring.len(), 2, "the report still has them after an upload");
assert_eq!(ring.to_text().lines().count(), 2);
}
#[test]
fn tail_text_is_the_whole_ring_untouched_when_under_the_cap() {
let ring = LogRing::new(100, 1 << 20);
fill(&ring, 5);
assert_eq!(ring.tail_text(150), ring.to_text());
}
#[test]
fn tail_text_trims_to_the_newest_lines_and_says_how_many_were_left_out() {
let ring = LogRing::new(1000, 1 << 20);
fill(&ring, 200);
let tail = ring.tail_text(150);
let mut lines = tail.lines();
assert_eq!(
lines.next().unwrap(),
"50 earlier lines omitted; full log in Dev Updater's Runtime tab"
);
let rest: Vec<&str> = lines.collect();
assert_eq!(rest.len(), 150, "exactly the cap, after the header line");
assert!(
rest[0].ends_with("line 50"),
"the oldest line kept is the 50th, not line 0: {}",
rest[0]
);
assert!(rest.last().unwrap().ends_with("line 199"));
}
#[test]
fn try_tail_text_gives_the_newest_lines_with_no_header() {
let ring = LogRing::new(1000, 1 << 20);
fill(&ring, 200);
let tail = ring.try_tail_text(80).expect("nothing holds the lock");
let lines: Vec<&str> = tail.lines().collect();
assert_eq!(lines.len(), 80, "the cap, and no header: this is a file");
assert!(lines[0].ends_with("line 120"), "{}", lines[0]);
assert!(lines[79].ends_with("line 199"), "{}", lines[79]);
}
/// The whole point of the `try_`: the panic hook calls this from a
/// thread that may already hold the ring's lock, and a blocking read
/// there would hang the process instead of aborting it.
#[test]
fn try_tail_text_answers_none_rather_than_blocking_on_a_held_lock() {
let ring = LogRing::new(10, 1 << 20);
fill(&ring, 3);
let held = ring.0.lock().expect("fresh ring");
assert_eq!(ring.try_tail_text(80), None);
drop(held);
assert!(ring.try_tail_text(80).is_some());
}
#[test]
fn an_empty_ring_says_so_rather_than_reporting_a_time() {
let ring = LogRing::with_defaults();
assert_eq!(ring.summary(), "app log: no lines yet");
assert_eq!(ring.last_at_ms(), None);
assert!(ring.is_empty());
}
#[test]
fn the_summary_names_dropped_lines_only_when_there_are_some() {
let ring = LogRing::new(2, 1 << 20);
fill(&ring, 2);
assert!(!ring.summary().contains("dropped"), "{}", ring.summary());
fill(&ring, 2);
assert!(ring.summary().contains("2 dropped"), "{}", ring.summary());
}
#[test]
fn a_line_formats_as_time_level_target_message() {
let line = LogLine {
seq: 0,
// 1970-01-01T12:34:56.789Z, so the arithmetic is checkable by
// hand rather than against another clock.
at_ms: (12 * 3600 + 34 * 60 + 56) * 1000 + 789,
level: Level::Info,
target: "iris::android".into(),
message: "surface created".into(),
}
.format();
assert_eq!(line, "12:34:56.789 INFO iris::android: surface created");
}
/// The forwarding half: a line reaches the ring *and* the logger the
/// platform already had, and one the inner logger filters out is still
/// in the ring.
#[test]
fn the_ring_logger_forwards_to_the_inner_logger() {
use log::Log;
struct Collect(Arc<Mutex<Vec<String>>>, log::Level);
impl Log for Collect {
fn enabled(&self, metadata: &log::Metadata) -> bool {
metadata.level() <= self.1
}
fn log(&self, record: &log::Record) {
self.0.lock().unwrap().push(record.args().to_string());
}
fn flush(&self) {}
}
let seen = Arc::new(Mutex::new(Vec::new()));
let ring = LogRing::with_defaults();
// Own target, tracing on: this is the case where the ring and the
// inner logger disagree, which is the thing under test -- a
// foreign target is covered separately below.
let logger = RingLogger::new(
ring.clone(),
Box::new(Collect(seen.clone(), Level::Info)),
|| true,
);
logger.log(
&log::Record::builder()
.args(format_args!("kept"))
.level(Level::Info)
.target("iris::test")
.build(),
);
logger.log(
&log::Record::builder()
.args(format_args!("filtered"))
.level(Level::Debug)
.target("iris::test")
.build(),
);
assert_eq!(
*seen.lock().unwrap(),
["kept"],
"the inner logger's own filter still applies"
);
let held: Vec<String> = ring.snapshot().into_iter().map(|l| l.message).collect();
assert_eq!(
held,
["kept", "filtered"],
"own-target debug still rings while tracing is on"
);
}
/// The bug this filter fixes: `naga`/`wgpu_core`/`jni` log at Debug
/// unconditionally, and used to flood the ring even though nothing in
/// this app asked for their Debug output. A foreign target's Debug
/// line must not ring even while tracing is on -- tracing controls
/// this app's own diagnostics, not a dependency's chatter.
#[test]
fn a_foreign_targets_debug_line_never_rings_even_while_tracing_is_on() {
use log::Log;
struct Discard;
impl Log for Discard {
fn enabled(&self, _: &log::Metadata) -> bool {
true
}
fn log(&self, _: &log::Record) {}
fn flush(&self) {}
}
let ring = LogRing::with_defaults();
let logger = RingLogger::new(ring.clone(), Box::new(Discard), || true);
logger.log(
&log::Record::builder()
.args(format_args!("naga debug spam"))
.level(Level::Debug)
.target("naga::front")
.build(),
);
logger.log(
&log::Record::builder()
.args(format_args!("naga warning"))
.level(Level::Warn)
.target("wgpu_core::device")
.build(),
);
let held: Vec<String> = ring.snapshot().into_iter().map(|l| l.message).collect();
assert_eq!(
held,
["naga warning"],
"Info-and-above always rings; foreign Debug never does"
);
}
#[test]
fn ring_accepts_is_own_target_debug_only_while_tracing() {
assert!(
ring_accepts(Level::Info, "wgpu_core::device", false),
"Info+ from anything, tracing off"
);
assert!(
ring_accepts(Level::Warn, "jni", true),
"Info+ from anything, tracing on"
);
assert!(
!ring_accepts(Level::Debug, "jni", true),
"foreign Debug, tracing on: still excluded"
);
assert!(
!ring_accepts(Level::Debug, "iris::sense", false),
"own Debug, tracing off: excluded"
);
assert!(
ring_accepts(Level::Debug, "iris::sense", true),
"own Debug, tracing on: included"
);
assert!(
ring_accepts(Level::Trace, "client_core::api", true),
"own Trace, tracing on: included"
);
}
#[test]
fn is_own_target_matches_the_crate_or_its_modules_only() {
assert!(is_own_target("iris"));
assert!(is_own_target("iris::sense"));
assert!(is_own_target("client_core"));
assert!(is_own_target("client_core::log_ring"));
assert!(!is_own_target("iris_something_else"));
assert!(!is_own_target("naga::front"));
assert!(!is_own_target("jni"));
}
}
+325
View File
@@ -0,0 +1,325 @@
//! Split a markdown message into its top-level **blocks** -- one
//! paragraph, heading, fenced code block, list, table or quote each, as a
//! byte slice of the original source.
//!
//! This exists for streaming. A transcript row used to be one text widget
//! holding the whole message, so a single streamed delta re-shaped every
//! paragraph of it through the text engine again; the phone's bench v2 put
//! the stream phase at p50 18.2ms against Compose's 13.4ms for exactly
//! that reason (docs/IRIS_TODO.md). A row is a column of one widget per
//! block now, and a delta that lands in the last block leaves every
//! earlier block's layout alone. `docs/DECISIONS.md`'s 2026-09-06 entry has
//! what that rejected and why the split lives here rather than in the UI
//! crate: `docs/CLIENT_CORE.md` already wanted a block model for P1, and
//! keeping it here means iris stays a text renderer that knows nothing
//! about markdown.
//!
//! **Blocks only.** Inline styling (bold, links, inline code) is still the
//! renderer's own job, per block -- this deliberately does not build a
//! full AST, because nothing needs one yet.
//!
//! ## Appending is not guaranteed to leave earlier blocks alone
//!
//! It nearly always does, which is what makes the fast path worth having,
//! but markdown has no such rule: appending a "```" line can turn text
//! that was three paragraphs into one fenced block, and appending "---"
//! under a paragraph turns that paragraph into a heading. So a caller
//! taking the O(last block) path **must compare the prefix it is about to
//! keep** rather than assume it. [`common_prefix`] is that comparison, and
//! it is cheap next to laying the text out again.
use pulldown_cmark::{Event, Options, Parser, Tag};
/// What a block is, for a renderer that wants to style or space blocks
/// differently. `Other` is deliberately present rather than a panic or a
/// silent fallback to `Paragraph`: markdown has more block kinds than this
/// list and more get added, and a renderer treating an unknown one as
/// prose is right, but it should be able to *tell* that is what it is
/// doing.
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
pub enum BlockKind {
Paragraph,
Heading,
/// A fenced or indented code block.
Code,
List,
Table,
Quote,
/// A thematic break, raw HTML, a footnote -- anything with no
/// distinguished treatment here.
Other,
}
/// One top-level block: its kind and the exact source that produced it.
/// `source` is a slice of the input with trailing whitespace removed, so
/// two splits of the same prefix compare equal even when one of them had a
/// delta arriving after it.
#[derive(Debug, Clone, PartialEq, Eq)]
pub struct Block {
pub kind: BlockKind,
pub source: String,
}
fn kind_of(tag: &Tag) -> BlockKind {
match tag {
Tag::Paragraph => BlockKind::Paragraph,
Tag::Heading { .. } => BlockKind::Heading,
Tag::CodeBlock(_) => BlockKind::Code,
Tag::List(_) => BlockKind::List,
Tag::Table(_) => BlockKind::Table,
Tag::BlockQuote(_) => BlockKind::Quote,
_ => BlockKind::Other,
}
}
fn options() -> Options {
// The same set `transcript-ui`'s renderer parses with, so a block
// boundary here and the styling there cannot disagree about what the
// source means.
Options::ENABLE_STRIKETHROUGH | Options::ENABLE_TABLES | Options::ENABLE_TASKLISTS
}
/// Split `src` into its top-level blocks, in source order. An empty or
/// whitespace-only input gives no blocks; text the parser does not put
/// inside any block (a stray fence marker mid-stream) still comes back,
/// as `Other`, rather than being dropped.
pub fn split_blocks(src: &str) -> Vec<Block> {
let mut out: Vec<Block> = Vec::new();
let mut depth = 0usize;
let mut kind = BlockKind::Other;
for (event, range) in Parser::new_ext(src, options()).into_offset_iter() {
match event {
Event::Start(tag) => {
if depth == 0 {
kind = kind_of(&tag);
}
depth += 1;
}
Event::End(_) => {
depth -= 1;
if depth == 0 {
push(&mut out, kind, &src[range]);
}
}
// A top-level event that is not part of any block -- a
// thematic break, a block of raw HTML. Inside one, it is the
// enclosing block's business and this does nothing.
_ => {
if depth == 0 {
push(&mut out, BlockKind::Other, &src[range]);
}
}
}
}
out
}
fn push(out: &mut Vec<Block>, kind: BlockKind, source: &str) {
let source = source.trim_end();
if source.is_empty() {
return;
}
out.push(Block {
kind,
source: source.to_string(),
});
}
/// How many leading blocks of `old` and `new` are identical -- what a
/// caller may keep the laid-out widgets for. See the module doc for why
/// this is a comparison rather than an assumption.
pub fn common_prefix(old: &[Block], new: &[Block]) -> usize {
old.iter().zip(new).take_while(|(a, b)| a == b).count()
}
#[cfg(test)]
mod tests {
use super::*;
fn kinds(src: &str) -> Vec<BlockKind> {
split_blocks(src).into_iter().map(|b| b.kind).collect()
}
#[test]
fn a_message_splits_into_its_top_level_blocks() {
let src = "# Title\n\nFirst para.\n\n```rust\nfn main() {}\n```\n\n- a\n- b\n";
assert_eq!(
kinds(src),
vec![
BlockKind::Heading,
BlockKind::Paragraph,
BlockKind::Code,
BlockKind::List
]
);
let blocks = split_blocks(src);
assert_eq!(blocks[1].source, "First para.");
assert_eq!(blocks[2].source, "```rust\nfn main() {}\n```");
}
#[test]
fn blank_input_has_no_blocks() {
assert!(split_blocks("").is_empty());
assert!(split_blocks(" \n\n ").is_empty());
}
/// The property the streaming fast path rests on, in its ordinary
/// shape: a delta landing in the last paragraph must leave every
/// earlier block byte-identical.
#[test]
fn a_delta_into_the_last_paragraph_leaves_earlier_blocks_untouched() {
let before = split_blocks("# Title\n\nFirst para.\n\nSecond par");
let after = split_blocks("# Title\n\nFirst para.\n\nSecond paragraph now.");
assert_eq!(common_prefix(&before, &after), 2);
assert_eq!(before.len(), 3);
assert_eq!(after.len(), 3);
assert_ne!(before[2], after[2]);
}
/// A delta that starts a *new* block keeps every old block, including
/// the one that was last -- so the fast path appends rather than
/// replacing.
#[test]
fn a_delta_that_starts_a_new_block_keeps_every_old_one() {
let before = split_blocks("First para.\n\nSecond para.");
let after = split_blocks("First para.\n\nSecond para.\n\nThird");
assert_eq!(common_prefix(&before, &after), 2);
assert_eq!(after.len(), 3);
}
/// A code fence arrives one delta at a time and is unterminated for
/// most of its life. It must still be *one* block the whole way, or
/// every delta would re-split the message into a different number of
/// pieces.
#[test]
fn an_unterminated_fence_is_one_block_while_it_streams() {
for src in [
"Here:\n\n```rust\n",
"Here:\n\n```rust\nfn main() {\n",
"Here:\n\n```rust\nfn main() {\n println!(\"hi\");\n",
] {
assert_eq!(
kinds(src),
vec![BlockKind::Paragraph, BlockKind::Code],
"{src:?}"
);
}
}
/// The half the fast path had no reason to touch, and the reason
/// `common_prefix` is a comparison rather than an assumption:
/// appending can rewrite what came before. `---` under a paragraph
/// turns that paragraph into a setext heading, so the block that was
/// already laid out is not the block it is now.
#[test]
fn appending_can_rewrite_an_earlier_block_and_the_prefix_says_so() {
let before = split_blocks("Not a heading\n\nsecond");
let after = split_blocks("Not a heading\n\nsecond\n---");
assert_eq!(before[1].kind, BlockKind::Paragraph);
assert_eq!(after[1].kind, BlockKind::Heading);
assert_eq!(
common_prefix(&before, &after),
1,
"the rewritten block must not be reported as keepable"
);
}
#[test]
fn a_thematic_break_is_its_own_block() {
assert_eq!(
kinds("one\n\n---\n\ntwo"),
vec![BlockKind::Paragraph, BlockKind::Other, BlockKind::Paragraph]
);
}
/// The shapes a real transcript actually contains, each checked for
/// the one property the streaming fast path needs: the *number* of
/// blocks and every earlier block's source stay put while the message
/// grows. A fence's own blank lines, a `---` inside one, a nested
/// list and a table are all places where a naive line-based split
/// would break the message into more pieces than there are blocks.
#[test]
fn the_transcripts_own_block_shapes_survive_a_split() {
let fence_with_blanks = "Intro.\n\n```rust\nfn a() {}\n\nfn b() {}\n```\n\nAfter.";
assert_eq!(
kinds(fence_with_blanks),
vec![BlockKind::Paragraph, BlockKind::Code, BlockKind::Paragraph],
"a blank line inside a fence is not a block boundary"
);
assert_eq!(
kinds("```\n---\n```"),
vec![BlockKind::Code],
"a thematic break inside a fence is code, not a break"
);
assert_eq!(
kinds("- a\n - a1\n - a2\n- b"),
vec![BlockKind::List],
"a nested list is one top-level block"
);
assert_eq!(
kinds("## Heading\n```sh\nls\n```"),
vec![BlockKind::Heading, BlockKind::Code],
"a fence directly under a heading, with no blank line"
);
assert_eq!(
kinds("| a | b |\n|---|---|\n| 1 | 2 |"),
vec![BlockKind::Table]
);
assert_eq!(
kinds("> quoted\n> more\n\nplain"),
vec![BlockKind::Quote, BlockKind::Paragraph]
);
}
/// `apply_delta`'s precondition, stated as the property rather than
/// the arithmetic: for every prefix of a realistic streamed message,
/// the blocks before the last one must be exactly the blocks the
/// previous prefix had. Where markdown breaks that (the `---` case
/// above), `common_prefix` has to *say* so -- which is what the
/// `>= len - 1` assertion below checks: the split may rewrite the
/// last block, never an earlier one, or `RowBlocks::apply_delta`
/// would keep a widget whose text is no longer what it holds.
#[test]
fn every_prefix_of_a_streamed_message_keeps_all_but_its_last_block() {
let full = "# Report\n\nFirst finding, at some length.\n\n```rust\nfn main() {\n\n println!(\"hi\");\n}\n```\n\n- one\n - nested\n- two\n\n| a | b |\n |---|---|\n| 1 | 2 |\n\n> and a closing quote.";
// Every character boundary, so a delta landing mid-word and one
// landing exactly on a fence's closing backtick are both covered.
let mut prev = Vec::new();
for end in full.char_indices().map(|(i, _)| i).chain([full.len()]) {
let now = split_blocks(&full[..end]);
let common = common_prefix(&prev, &now);
assert!(
prev.is_empty() || common + 1 >= prev.len(),
"at {end} bytes the split rewrote block {common} of {}, not just the last one:\n before={prev:#?}\nafter={now:#?}",
prev.len()
);
prev = now;
}
}
/// The half a growing message cannot show: a fence that never closes.
/// The stream ends there and the block must still be the code block
/// it has been all along, not re-split into paragraphs.
#[test]
fn a_stream_that_ends_inside_a_fence_still_ends_with_one_code_block() {
let src = "Here is the patch:\n\n```diff\n- old line\n+ new line";
let blocks = split_blocks(src);
assert_eq!(
blocks.iter().map(|b| b.kind).collect::<Vec<_>>(),
vec![BlockKind::Paragraph, BlockKind::Code]
);
assert_eq!(blocks[1].source, "```diff\n- old line\n+ new line");
}
/// A delta that closes a fence changes the *last* block only, so the
/// fast path takes it -- the case the module doc says is the reason
/// `common_prefix` is a comparison.
#[test]
fn the_delta_that_closes_a_fence_changes_only_the_last_block() {
let before = split_blocks("Text.\n\n```\ncode\n");
let after = split_blocks("Text.\n\n```\ncode\n```");
assert_eq!(before.len(), after.len());
assert_eq!(common_prefix(&before, &after), 1);
assert_ne!(before[1], after[1]);
}
}
+244
View File
@@ -0,0 +1,244 @@
//! A tool call's input, read rather than dumped -- the port of
//! `ToolInput.kt`'s `parseToolInput`, which is what both the collapsed
//! card's one-line summary and the expanded card's key/value list are
//! derived from.
//!
//! Every tool's input arrives as JSON, and showing it raw makes the reader
//! parse `{"command":"…","timeout":120000}` themselves to find the one
//! line they care about. So the fields that carry the meaning are pulled
//! out, and anything left over is still shown, because dropping a field
//! would be claiming the tool has no other input when it might.
//!
//! Pure, and here rather than in the widget crate, for the reason the rest
//! of this crate exists: the derivation is the same on a phone and on a
//! desktop, and it is testable without a renderer.
use crate::durations::format_millis_text;
use crate::highlight::Language;
use serde_json::{Map, Value};
/// A tool call's input, split into the parts a card draws separately.
#[derive(Debug, Clone, PartialEq, Eq, Default)]
pub struct ToolInput {
/// The thing that will actually be run or read, if this tool has one.
pub subject: Option<String>,
/// The language [`ToolInput::subject`] is written in, for
/// highlighting.
pub language: Option<Language>,
/// The tool's own one-line summary, when it wrote one.
pub description: Option<String>,
/// How long the call may take, in the largest units it fits. Shown
/// apart because it is a limit on the call rather than part of what
/// the call does.
pub timeout: Option<String>,
/// Everything else, as `name: value` lines. Never dropped.
pub rest: Vec<String>,
}
impl ToolInput {
/// The one line to show when there is only room for one: what this
/// call is for.
pub fn title(&self) -> Option<&str> {
self.description
.as_deref()
.or(self.subject.as_deref())
// A subject that is only whitespace would draw as an empty
// summary line, which reads as a tool with nothing to say
// rather than as one whose subject was blank.
.filter(|t| !t.trim().is_empty())
}
}
/// Which field of which tool is the subject.
///
/// A table rather than a chain of `if`s: adding a tool is a row, and the
/// shape stops any of them from being the special case that gets its own
/// code path. Unknown tools fall through to "no subject, everything is
/// rest".
const SUBJECTS: &[(&str, &str, Option<Language>)] = &[
("Bash", "command", Some(Language::Shell)),
("Read", "file_path", None),
("Write", "file_path", None),
("Edit", "file_path", None),
("Glob", "pattern", None),
("Grep", "pattern", None),
("WebFetch", "url", None),
];
/// Fields that are the tool's own prose about itself rather than input to
/// it.
const DESCRIPTIONS: &[&str] = &["description", "prompt"];
/// One JSON value as the Kotlin's `JSONObject.optString`/`get` wrote it: a
/// string is its own characters, anything else is its JSON form.
///
/// One function rather than two, because the same coercion decides both
/// what a subject reads as and what a leftover field's value reads as, and
/// two copies would eventually disagree about a number.
fn as_text(value: &Value) -> String {
match value {
Value::String(s) => s.clone(),
other => other.to_string(),
}
}
fn non_blank(value: Option<&Value>) -> Option<String> {
let text = as_text(value?);
(!text.trim().is_empty()).then_some(text)
}
/// Split `input` (a tool call's JSON) into the parts a card draws.
///
/// Input that is not a JSON object -- older transcripts and some tools
/// send a bare string -- is still the input, so it is still shown, as the
/// whole of `rest`.
pub fn parse_tool_input(tool: &str, input: &str) -> ToolInput {
let Ok(Value::Object(json)) = serde_json::from_str::<Value>(input) else {
return ToolInput {
rest: match input.trim().is_empty() {
true => Vec::new(),
false => vec![input.to_string()],
},
..ToolInput::default()
};
};
parse_object(tool, &json)
}
fn parse_object(tool: &str, json: &Map<String, Value>) -> ToolInput {
let (subject_key, language) = SUBJECTS
.iter()
.find(|(name, ..)| *name == tool)
.map(|(_, key, language)| (Some(*key), *language))
.unwrap_or((None, None));
let subject = subject_key.and_then(|key| non_blank(json.get(key)));
let description = DESCRIPTIONS
.iter()
.find_map(|key| non_blank(json.get(*key)));
let timeout = non_blank(json.get("timeout")).map(|t| format_millis_text(&t));
// Sorted, so the leftovers are in the same order every time this call
// is drawn rather than in whatever order the JSON happened to arrive
// in. A field is left out only when it is already drawn somewhere
// else on the card.
let mut keys: Vec<&String> = json
.keys()
.filter(|k| Some(k.as_str()) != subject_key || subject.is_none())
.filter(|k| !DESCRIPTIONS.contains(&k.as_str()) || description.is_none())
.filter(|k| k.as_str() != "timeout" || timeout.is_none())
.collect();
keys.sort();
let rest = keys
.into_iter()
.map(|key| format!("{key}: {}", as_text(&json[key])))
.collect();
ToolInput {
subject,
language,
description,
timeout,
rest,
}
}
#[cfg(test)]
mod tests {
use super::*;
#[test]
fn each_tool_in_the_table_has_its_own_subject() {
// One assertion per row of `SUBJECTS`, because the table is the
// whole of the rule and a row lost in an edit would otherwise
// only show up as a card with no summary line.
let cases = [
("Bash", r#"{"command":"ls -la"}"#, "ls -la"),
("Read", r#"{"file_path":"/tmp/x.rs"}"#, "/tmp/x.rs"),
("Write", r#"{"file_path":"/tmp/y.rs"}"#, "/tmp/y.rs"),
("Edit", r#"{"file_path":"/tmp/z.rs"}"#, "/tmp/z.rs"),
("Glob", r#"{"pattern":"**/*.rs"}"#, "**/*.rs"),
("Grep", r#"{"pattern":"fn main"}"#, "fn main"),
("WebFetch", r#"{"url":"https://x/y"}"#, "https://x/y"),
];
for (tool, input, expected) in cases {
let parsed = parse_tool_input(tool, input);
assert_eq!(parsed.subject.as_deref(), Some(expected), "{tool}");
assert_eq!(parsed.title(), Some(expected), "{tool}");
assert!(parsed.rest.is_empty(), "{tool}: {:?}", parsed.rest);
}
assert_eq!(
parse_tool_input("Bash", r#"{"command":"ls"}"#).language,
Some(Language::Shell),
"a Bash command is shell, and is the one row that names a language"
);
}
#[test]
fn a_tools_own_description_is_what_the_one_line_says() {
// The description wins over the subject: it is the tool's own
// prose about what this call is for, which is what a reader
// scanning a collapsed run is looking for.
let parsed = parse_tool_input(
"Bash",
r#"{"command":"cargo test -p iris","description":"Run the iris tests"}"#,
);
assert_eq!(parsed.title(), Some("Run the iris tests"));
assert_eq!(parsed.subject.as_deref(), Some("cargo test -p iris"));
assert!(parsed.rest.is_empty(), "{:?}", parsed.rest);
}
#[test]
fn a_timeout_is_read_as_a_span_and_kept_apart_from_the_rest() {
let parsed = parse_tool_input("Bash", r#"{"command":"sleep 500","timeout":480000}"#);
assert_eq!(parsed.timeout.as_deref(), Some("8m"));
assert!(parsed.rest.is_empty(), "{:?}", parsed.rest);
}
#[test]
fn every_field_not_drawn_elsewhere_is_still_shown() {
// The half the "never dropped" promise is about: a tool this
// build has never heard of has no subject, so *everything* is
// rest -- and a known tool's extra fields are too.
let parsed = parse_tool_input(
"Edit",
r#"{"file_path":"/a.rs","old_string":"x","new_string":"y","replace_all":true}"#,
);
assert_eq!(
parsed.rest,
vec![
"new_string: y".to_string(),
"old_string: x".to_string(),
"replace_all: true".to_string(),
],
"sorted, and a non-string value written as JSON"
);
let unknown = parse_tool_input("SomeNewTool", r#"{"b":2,"a":"one"}"#);
assert_eq!(unknown.subject, None);
assert_eq!(unknown.rest, vec!["a: one".to_string(), "b: 2".to_string()]);
}
#[test]
fn input_that_is_not_an_object_is_still_the_input() {
// Older transcripts and some tools send a bare string; a card
// that dropped it would claim the call had no input at all.
assert_eq!(
parse_tool_input("Bash", "just a string").rest,
vec!["just a string".to_string()]
);
assert_eq!(parse_tool_input("Bash", " ").rest, Vec::<String>::new());
assert_eq!(parse_tool_input("Bash", "").title(), None);
}
#[test]
fn a_blank_subject_is_no_subject_rather_than_an_empty_summary_line() {
let parsed = parse_tool_input("Bash", r#"{"command":" ","other":1}"#);
assert_eq!(parsed.subject, None);
assert_eq!(parsed.title(), None);
// Not dropped just because it was blank -- it is still a field
// the call carried.
assert_eq!(
parsed.rest,
vec!["command: ".to_string(), "other: 1".to_string()]
);
}
}
+16 -4
View File
@@ -1,7 +1,7 @@
//! This phone's copy of the transcripts it has already been sent, so
//! reopening a session does not download it again. Ported from
//! `app/.../TranscriptCache.kt`; see `TRANSCRIPT_CACHE.md` at the repo root
//! for the design and `CLIENT_CORE.md` for how this file corresponds to it.
//! `app/.../TranscriptCache.kt`; see `docs/TRANSCRIPT_CACHE.md`
//! for the design and `docs/CLIENT_CORE.md` for how this file corresponds to it.
//!
//! What is stored is the server's own JSON for one event per line, in
//! transcript order. Reading the cache means running the same [`seq_of`]
@@ -361,6 +361,10 @@ impl SessionCache {
{
return Ok(false);
}
debug_assert!(
lines.iter().all(|l| !l.contains('\n')),
"a stored page's lines must each be one line"
);
fs::create_dir_all(&this.dir)?;
let kind = if rows { "rows" } else { "raw" };
let mut content = lines.join("\n");
@@ -389,8 +393,16 @@ impl SessionCache {
return Ok(());
};
// Written as it arrived. A newline inside it would split one
// event into two unreadable halves, but neither source can
// produce one.
// event into two unreadable halves. No source here can produce
// one -- an SSE `data:` field cannot hold a raw newline, and a
// fetched line is one element of a compact JSON array -- but
// that is a fact about the *server's* serializer rather than
// anything this file controls, so it is checked rather than
// trusted.
debug_assert!(
!line.contains('\n'),
"a cached transcript line must be one line: {line}"
);
use std::io::Write;
writer.write_all(line.as_bytes())?;
writer.write_all(b"\n")?;
+695 -2
View File
@@ -78,6 +78,11 @@ pub enum TranscriptItem {
input: String,
output: String,
done: bool,
/// Whether the result that arrived said the call failed
/// ([`Event::ToolEnd`]'s `is_error`). Meaningless while `done` is
/// false, and [`ToolState::of`] is the only thing that reads the
/// pair, so the two cannot be combined wrongly at a call site.
failed: bool,
asks: Vec<QuestionCard>,
images: Vec<String>,
},
@@ -112,6 +117,13 @@ pub enum TranscriptItem {
ClearedNote {
seq: u64,
},
/// The account's usage limit stopped the turn; `resets_at` is epoch
/// seconds when the dialect said when it lifts (`LimitNote` in
/// `TranscriptItems.kt`).
LimitNote {
seq: u64,
resets_at: Option<f64>,
},
CompactedNote {
seq: u64,
pre_tokens: Option<u64>,
@@ -131,6 +143,7 @@ impl TranscriptItem {
| Self::CommandRow { seq, .. }
| Self::Note { seq, .. }
| Self::ClearedNote { seq }
| Self::LimitNote { seq, .. }
| Self::CompactedNote { seq, .. } => *seq,
Self::QuestionCard(card) => card.seq,
}
@@ -286,6 +299,201 @@ fn split_run(tail: &[TranscriptItem], behind: Option<&str>) -> Vec<TranscriptIte
out
}
/// Puts a page of older items in front of the ones already loaded, healing
/// whatever the page boundary cut in two. Ported from `TranscriptItems.kt`'s
/// `joinPages`.
///
/// Two things straddle a boundary: a tool call separated from its result,
/// and a message separated from the rest of itself. Both were one thing
/// before the transcript was cut into pages.
///
/// A boundary lands wherever it lands, and roughly half the time that is
/// between a call and its result. The newer page then holds a `ToolEnd`
/// whose start it never saw, which `fold_event` draws as a row of its own
/// -- correctly, because a call that renders as nothing is indistinguishable
/// from one that never happened. When the older page arrives it brings the
/// real `ToolStart`, and concatenating the two lists left *both*: the same
/// call twice.
///
/// Merged by the call's own id rather than by position, because position is
/// exactly what a page boundary destroys. The older row wins on what a
/// start knows and the newer on what an end knows, which is the only way
/// round that loses nothing.
///
/// The third thing is the *run*, and it is the one the Kotlin original used
/// to miss (AGENTS.md's "things that have bitten"): every page ends up
/// here, but `adopt_run` must run on *every* join, not only the one where a
/// split call was found -- a boundary landing cleanly between two finished
/// calls, which is most of them, would otherwise leave the older page's
/// calls under the run name they were folded with. On screen: one run of
/// tool calls drawn as two groups, with the seam wherever the reader
/// happened to have paged.
pub fn join_pages(earlier: &[TranscriptItem], later: &[TranscriptItem]) -> Vec<TranscriptItem> {
let (older, newer) = heal_split_message(earlier, later);
let started_earlier: std::collections::HashSet<&str> = older
.iter()
.filter_map(TranscriptItem::as_tool_run)
.collect();
// Owned rather than borrowed from `newer`: `kept` below needs to consume `newer` by
// value, and a map borrowing it would keep that alive.
let ended_later: std::collections::HashMap<String, TranscriptItem> = newer
.iter()
.filter_map(|item| item.as_tool_run().map(|id| (id.to_string(), item.clone())))
.filter(|(id, _)| started_earlier.contains(id.as_str()))
.collect();
let healed: Vec<TranscriptItem> = older
.into_iter()
.map(|row| match row {
TranscriptItem::ToolRun {
seq,
id,
run_id,
tool,
input,
asks: row_asks,
images: row_images,
..
} if ended_later.contains_key(id.as_str()) => {
let &TranscriptItem::ToolRun {
ref output,
done,
failed,
asks: ref half_asks,
images: ref half_images,
..
} = &ended_later[id.as_str()]
else {
unreachable!("filtered to ToolRun above");
};
TranscriptItem::ToolRun {
seq,
id,
run_id,
tool,
input,
output: output.clone(),
done,
failed,
// Kept from both halves: a question or an image can be
// attached to either, depending on which side of the
// boundary its event fell.
asks: row_asks.into_iter().chain(half_asks.clone()).collect(),
images: row_images.into_iter().chain(half_images.clone()).collect(),
}
}
other => other,
})
.collect();
let kept: Vec<TranscriptItem> = newer
.into_iter()
.filter(|item| match item.as_tool_run() {
Some(id) => !ended_later.contains_key(id),
None => true,
})
.collect();
let mut out = adopt_run(&healed, &kept);
out.extend(kept);
// What this function exists to prevent, checked rather than assumed: the same
// call drawn twice, once from the page that saw its start and once from the page
// that saw its end. Not a seq-ordering check -- a peer note is stamped with the
// seq its turn began at, which can be older than the page it arrived in, so the
// two pages' seqs legitimately interleave at the boundary.
debug_assert!(
{
let mut ids: Vec<&str> = out.iter().filter_map(TranscriptItem::as_tool_run).collect();
let before = ids.len();
ids.sort_unstable();
ids.dedup();
ids.len() == before
},
"join_pages left the same tool call in both halves"
);
out
}
/// Rejoins a message the page boundary cut, and hands back the two pages to
/// concatenate. Ported from `TranscriptItems.kt`'s `healSplitMessage`.
///
/// `fold_event` never leaves two assistant messages next to each other
/// inside one page, so two meeting at a join are always the two halves of
/// one reply, and leaving them apart drew a single answer as two with a
/// paragraph break through the middle of a sentence.
///
/// The newer half keeps its identity, for the reason `adopt_run`'s doc
/// gives. It grows by what the older half brings, which is safe here and
/// nowhere else -- the join is at the oldest end of what is loaded, so the
/// growth extends off the top of the screen.
fn heal_split_message(
earlier: &[TranscriptItem],
later: &[TranscriptItem],
) -> (Vec<TranscriptItem>, Vec<TranscriptItem>) {
let (
Some(TranscriptItem::AssistantMsg {
text: head_text, ..
}),
Some(TranscriptItem::AssistantMsg {
seq: tail_seq,
text: tail_text,
settled: tail_settled,
}),
) = (earlier.last(), later.first())
else {
return (earlier.to_vec(), later.to_vec());
};
let merged = TranscriptItem::AssistantMsg {
seq: *tail_seq,
text: format!("{head_text}{tail_text}"),
settled: *tail_settled,
};
let mut newer = vec![merged];
newer.extend(later[1..].iter().cloned());
(earlier[..earlier.len() - 1].to_vec(), newer)
}
/// Hands the older calls at the join the name of the run they are joining.
/// Ported from `TranscriptItems.kt`'s `adoptRun`.
///
/// The two pages were folded separately, so a run split by the boundary
/// came back as two runs with two names. Naming the joined run after the
/// *older* half would be the obvious way round and is wrong: the newer half
/// is the part already on screen, and renaming it is renaming the row the
/// reader is looking at, which is how a list loses its anchor.
fn adopt_run(earlier: &[TranscriptItem], later: &[TranscriptItem]) -> Vec<TranscriptItem> {
let Some(TranscriptItem::ToolRun { run_id, tool, .. }) = later.first() else {
return earlier.to_vec();
};
// A question is in a run of its own on both sides of the join, the same as it would be
// had the two pages been folded as one. Without this the heal would merge a group
// straight through the row the reader was asked something on.
if tool == ASK_USER_QUESTION {
return earlier.to_vec();
}
let joining = run_id.clone();
let tail_len = earlier
.iter()
.rev()
.take_while(|item| matches!(item, TranscriptItem::ToolRun { tool, .. } if tool != ASK_USER_QUESTION))
.count();
if tail_len == 0 {
return earlier.to_vec();
}
let split = earlier.len() - tail_len;
let mut out = earlier[..split].to_vec();
out.extend(earlier[split..].iter().cloned().map(|mut item| {
// `take_while` above already restricted this slice to non-question tool calls;
// this just guards the invariant rather than trusting it silently.
debug_assert!(
matches!(&item, TranscriptItem::ToolRun { tool, .. } if tool != ASK_USER_QUESTION),
"adopt_run must never rename a question's own run"
);
if let TranscriptItem::ToolRun { run_id, .. } = &mut item {
*run_id = joining.clone();
}
item
}));
out
}
/// Folds one transcript event onto `items`, the way `foldEvent` does in
/// `TranscriptItems.kt`. Every wire event has a case; see the module doc
/// for the one difference from the Kotlin original (no `Unknown` fallback
@@ -361,6 +569,7 @@ pub fn fold_event(items: &[TranscriptItem], entry: &SeqEvent) -> Vec<TranscriptI
input: input.to_string(),
output: String::new(),
done: false,
failed: false,
asks: Vec::new(),
images: Vec::new(),
});
@@ -371,15 +580,23 @@ pub fn fold_event(items: &[TranscriptItem], entry: &SeqEvent) -> Vec<TranscriptI
*out = output.clone();
}
}),
Event::ToolEnd { id, output } => {
Event::ToolEnd {
id,
output,
is_error,
} => {
if items.iter().any(|i| i.as_tool_run() == Some(id.as_str())) {
update_tool(items, id, |item| {
if let TranscriptItem::ToolRun {
output: out, done, ..
output: out,
done,
failed,
..
} = item
{
*out = output.clone();
*done = true;
*failed = *is_error;
}
})
} else {
@@ -393,6 +610,7 @@ pub fn fold_event(items: &[TranscriptItem], entry: &SeqEvent) -> Vec<TranscriptI
input: String::new(),
output: output.clone(),
done: true,
failed: *is_error,
asks: Vec::new(),
images: Vec::new(),
});
@@ -449,6 +667,7 @@ pub fn fold_event(items: &[TranscriptItem], entry: &SeqEvent) -> Vec<TranscriptI
input,
output,
done,
failed,
images,
} if asks.iter().any(|a| &a.id == id) => {
for ask in asks.iter_mut() {
@@ -464,6 +683,7 @@ pub fn fold_event(items: &[TranscriptItem], entry: &SeqEvent) -> Vec<TranscriptI
input,
output,
done,
failed,
asks,
images,
}
@@ -525,6 +745,14 @@ pub fn fold_event(items: &[TranscriptItem], entry: &SeqEvent) -> Vec<TranscriptI
items.push(TranscriptItem::ClearedNote { seq });
items
}
Event::LimitReached { resets_at } => {
let mut items = items.to_vec();
items.push(TranscriptItem::LimitNote {
seq,
resets_at: *resets_at,
});
items
}
Event::Compacted {
pre_tokens,
post_tokens,
@@ -541,6 +769,74 @@ pub fn fold_event(items: &[TranscriptItem], entry: &SeqEvent) -> Vec<TranscriptI
}
}
/// What became of one tool call -- every state a card has to be able to
/// draw, including the two that are not answers.
///
/// The pair this enum exists for is [`ToolState::Succeeded`] against
/// [`ToolState::NoResult`]. A call that finished having printed nothing
/// and a call whose result never arrived both leave an empty `output`,
/// and drawing them the same way states a verdict nobody reached: "it
/// worked and said nothing" reads as a fact, where the truth is that the
/// turn ended before anything came back.
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
pub enum ToolState {
/// Started, no result yet, and the session is still working -- the
/// ordinary state of a call in flight.
Running,
/// Stopped on the reader: a permission or question this call carries
/// has not been answered, so nothing is happening until somebody
/// answers it. Distinct from [`Self::Running`] because whose move it
/// is differs, which is the Compose card's "your turn".
Deciding,
/// A result arrived and the tool did not report a failure.
Succeeded,
/// A result arrived and the tool reported that the call failed
/// (`is_error`).
Failed,
/// No result ever arrived and the session is not working any more --
/// the turn was interrupted, or the process went away. Not a verdict
/// on the call: it says only that nobody found out.
NoResult,
}
impl ToolState {
/// The state of one call. `session_working` is
/// [`session_working`]'s answer for the session this call is in --
/// the only thing here that is not a property of the call itself, and
/// what separates "still running" from "never came back".
///
/// Written once, over the fields rather than per call site, because
/// the five states are decided by four conditions and every place
/// that re-derived a subset of them got a different subset.
pub fn of(item: &TranscriptItem, session_working: bool) -> Option<Self> {
let TranscriptItem::ToolRun {
done, failed, asks, ..
} = item
else {
return None;
};
debug_assert!(
!failed || *done,
"a call cannot have failed before its result arrived"
);
Some(if asks.iter().any(|ask| ask.answers.is_empty()) {
// Ahead of `done`: a call waiting on permission has not
// finished either, and which of the two the reader is being
// told about is the one they can act on.
Self::Deciding
} else if !*done {
match session_working {
true => Self::Running,
false => Self::NoResult,
}
} else if *failed {
Self::Failed
} else {
Self::Succeeded
})
}
}
/// One row as the transcript draws it: a run of consecutive tool calls, or
/// anything else. Ported from `ToolRows.kt`'s `TranscriptRow` and
/// `groupToolRuns` -- the Compose card rendering in that file is not part
@@ -606,6 +902,37 @@ pub fn group_tool_runs(items: &[TranscriptItem]) -> Vec<TranscriptRow> {
rows
}
/// Folds a page of raw transcript lines (`ApiClient::fetch_transcript_page`'s
/// `Vec<Value>`) into the flat item list this module works over. A line
/// this build can't parse fails the whole page rather than being skipped --
/// CODE_RULES's "an enumeration must be able to say 'it broke'" -- since
/// silently dropping one event could hide, say, a user message that then
/// looks like it was never sent. Moved here from `desktop-app`'s `app.rs`
/// (RUST.md's E4) when the Android transcript client (I5) needed the same
/// fold: "write the logic once" applies to any caller embedding
/// `transcript-ui` against a live server, not just the first one.
pub fn fold_page(values: &[serde_json::Value]) -> Result<Vec<TranscriptItem>, String> {
let mut items = Vec::new();
for value in values {
let event: SeqEvent = serde_json::from_value(value.clone()).map_err(|e| {
format!("the server sent a transcript line this build couldn't parse: {e}")
})?;
items = fold_event(&items, &event);
}
Ok(items)
}
/// The wire `seq` a raw transcript line carries -- the live-stream resume
/// cursor after loading a page must be this, not a folded item's `seq()`.
/// A folded `AssistantMsg` keeps the seq of the *first* delta it
/// accumulated (`fold_event`'s own doc), so resuming from that seq would
/// re-deliver every delta already folded into it, duplicating the tail of
/// a reply that was mid-stream when the page was fetched -- found via a
/// real screenshot in E4 (RUST.md), where the assistant's line doubled.
pub fn raw_seq(value: &serde_json::Value) -> Option<u64> {
value.get("seq")?.as_u64()
}
#[cfg(test)]
mod tests {
use super::*;
@@ -749,6 +1076,7 @@ mod tests {
Event::ToolEnd {
id: "x".to_string(),
output: "done".to_string(),
is_error: false,
},
)]);
assert_eq!(
@@ -761,6 +1089,7 @@ mod tests {
input: String::new(),
output: "done".to_string(),
done: true,
failed: false,
asks: Vec::new(),
images: Vec::new(),
}]
@@ -824,4 +1153,368 @@ mod tests {
other => panic!("expected a QuestionCard, got {other:?}"),
}
}
fn line(seq: u64, json: serde_json::Value) -> serde_json::Value {
let mut obj = json;
obj["seq"] = serde_json::json!(seq);
obj["ts"] = serde_json::json!(1.0);
obj
}
/// The regression for a bug a real `run-headless.sh` screenshot found
/// in `desktop-app` (E4, RUST.md): resuming the live stream from the
/// last *item's* seq re-delivers the deltas already folded into a
/// still-open assistant message, doubling its tail. `raw_seq` of the
/// last wire line must be the true high-water mark instead, which for a
/// run of deltas is higher than every item's own `seq()`.
#[test]
fn the_resume_cursor_is_the_last_wire_seq_not_the_last_items_seq() {
let values = vec![
line(1, serde_json::json!({"type": "userMessage", "text": "hi"})),
line(
2,
serde_json::json!({"type": "assistantText", "delta": "a"}),
),
line(
3,
serde_json::json!({"type": "assistantText", "delta": "b"}),
),
line(
4,
serde_json::json!({"type": "assistantText", "delta": "c"}),
),
];
let after = raw_seq(values.last().unwrap()).unwrap();
assert_eq!(after, 4);
let items = fold_page(&values).unwrap();
let assistant_seq = items
.iter()
.find(|i| matches!(i, TranscriptItem::AssistantMsg { .. }))
.unwrap()
.seq();
assert_eq!(assistant_seq, 2);
assert_ne!(
after, assistant_seq,
"the fixed bug: these must differ here"
);
}
#[test]
fn a_page_folds_into_one_settled_assistant_message() {
let values = vec![
line(1, serde_json::json!({"type": "userMessage", "text": "hi"})),
line(
2,
serde_json::json!({"type": "assistantText", "delta": "hel"}),
),
line(
3,
serde_json::json!({"type": "assistantText", "delta": "lo"}),
),
];
let items = fold_page(&values).unwrap();
assert_eq!(
items,
vec![
TranscriptItem::UserMsg {
seq: 1,
text: "hi".to_string(),
attachments: Vec::new(),
},
TranscriptItem::AssistantMsg {
seq: 2,
text: "hello".to_string(),
settled: false,
},
]
);
}
#[test]
fn an_unparseable_line_fails_the_whole_page() {
let values = vec![serde_json::json!({"seq": 1, "ts": 1.0, "type": "not-a-real-type"})];
let err = fold_page(&values).unwrap_err();
assert!(err.contains("couldn't parse"));
}
fn tool_start(seq: u64, id: &str, tool: &str) -> SeqEvent {
event(
seq,
Event::ToolStart {
id: id.to_string(),
tool: tool.to_string(),
input: serde_json::json!({}),
},
)
}
fn tool_end(seq: u64, id: &str, output: &str) -> SeqEvent {
event(
seq,
Event::ToolEnd {
id: id.to_string(),
output: output.to_string(),
is_error: false,
},
)
}
/// AGENTS.md's "things that have bitten": `joinPages` used to run
/// `adoptRun` only on the path where a *split* call was found, so a
/// boundary landing cleanly between two already-finished calls -- most
/// of them -- left the older page's calls under the run name they were
/// folded with, drawing one run of tool calls as two groups. Two
/// finished, unrelated calls (no id in common) must still end up under
/// one run name after the join.
#[test]
fn a_clean_boundary_between_two_finished_runs_is_still_healed_into_one_run() {
let older = fold_all(&[tool_start(1, "a", "Bash"), tool_end(2, "a", "old output")]);
let newer = fold_all(&[tool_start(3, "b", "Bash"), tool_end(4, "b", "new output")]);
let joined = join_pages(&older, &newer);
let run_ids: Vec<_> = joined
.iter()
.map(|item| match item {
TranscriptItem::ToolRun { run_id, .. } => run_id.as_str(),
other => panic!("expected only ToolRun items, got {other:?}"),
})
.collect();
assert_eq!(
run_ids,
vec!["b", "b"],
"the older call must adopt the newer, already-on-screen run's name"
);
}
#[test]
fn a_call_split_across_the_boundary_merges_into_one_row() {
let older = fold_all(&[tool_start(1, "x", "Bash")]);
let newer = fold_all(&[tool_end(2, "x", "the result")]);
let joined = join_pages(&older, &newer);
assert_eq!(
joined,
vec![TranscriptItem::ToolRun {
seq: 1,
id: "x".to_string(),
run_id: "x".to_string(),
tool: "Bash".to_string(),
input: "{}".to_string(),
output: "the result".to_string(),
done: true,
failed: false,
asks: Vec::new(),
images: Vec::new(),
}],
"the older half's tool/input and the newer half's output/done must both survive"
);
}
#[test]
fn a_message_split_across_the_boundary_is_rejoined_with_the_newer_halfs_identity() {
let older = vec![TranscriptItem::AssistantMsg {
seq: 1,
text: "Hel".to_string(),
settled: false,
}];
let newer = vec![
TranscriptItem::AssistantMsg {
seq: 2,
text: "lo".to_string(),
settled: true,
},
TranscriptItem::UserMsg {
seq: 3,
text: "next".to_string(),
attachments: Vec::new(),
},
];
let joined = join_pages(&older, &newer);
assert_eq!(
joined,
vec![
TranscriptItem::AssistantMsg {
seq: 2,
text: "Hello".to_string(),
settled: true,
},
TranscriptItem::UserMsg {
seq: 3,
text: "next".to_string(),
attachments: Vec::new(),
},
]
);
}
/// A question is in a run of its own on both sides of a join -- healing
/// must never rename the run of calls the reader was asked something
/// on, the same rule `splitRun` enforces for a live turn boundary.
#[test]
fn adopt_run_never_renames_into_a_question_row() {
let older = fold_all(&[tool_start(1, "a", "Bash"), tool_end(2, "a", "done")]);
let newer = vec![TranscriptItem::ToolRun {
seq: 3,
id: "q".to_string(),
run_id: "q".to_string(),
tool: ASK_USER_QUESTION.to_string(),
input: "{}".to_string(),
output: String::new(),
done: false,
failed: false,
asks: Vec::new(),
images: Vec::new(),
}];
let joined = join_pages(&older, &newer);
match &joined[0] {
TranscriptItem::ToolRun { run_id, .. } => assert_eq!(run_id, "a"),
other => panic!("expected a ToolRun, got {other:?}"),
}
}
}
/// [`ToolState`] is what a card colours itself by, so each of its five
/// states is asserted from the events that actually produce it rather than
/// from a hand-built item -- a mapping that agreed with a fixture and
/// disagreed with the fold would be invisible until it was on screen.
#[cfg(test)]
mod tool_state_tests {
use super::*;
fn event(seq: u64, e: Event) -> SeqEvent {
SeqEvent {
seq,
ts: 0.0,
event: e,
}
}
fn fold_all(events: &[SeqEvent]) -> Vec<TranscriptItem> {
events
.iter()
.fold(Vec::new(), |items, e| fold_event(&items, e))
}
fn start(id: &str) -> SeqEvent {
event(
1,
Event::ToolStart {
id: id.to_string(),
tool: "Bash".to_string(),
input: serde_json::json!({"command": "ls"}),
},
)
}
fn end(id: &str, output: &str, is_error: bool) -> SeqEvent {
event(
2,
Event::ToolEnd {
id: id.to_string(),
output: output.to_string(),
is_error,
},
)
}
fn state_of(events: &[SeqEvent], session_working: bool) -> ToolState {
let items = fold_all(events);
ToolState::of(&items[0], session_working).expect("the fixture's first item is a tool call")
}
#[test]
fn a_result_that_arrived_is_read_from_is_error() {
assert_eq!(
state_of(&[start("a"), end("a", "ok", false)], false),
ToolState::Succeeded
);
assert_eq!(
state_of(&[start("a"), end("a", "No such file", true)], false),
ToolState::Failed
);
}
/// The pair this enum exists for. Both calls have an empty `output`
/// and nothing else distinguishes them, so a card that only looked at
/// the text would draw the interrupted one as a call that ran fine and
/// printed nothing.
#[test]
fn a_call_that_printed_nothing_is_not_a_call_that_never_answered() {
assert_eq!(
state_of(&[start("a"), end("a", "", false)], false),
ToolState::Succeeded,
"a result arrived; it was empty"
);
assert_eq!(
state_of(&[start("a")], false),
ToolState::NoResult,
"no result, and the session is not working any more"
);
}
/// The same call, mid-turn: still running rather than abandoned. The
/// only thing separating the two is the session's own status, which is
/// why `of` takes it.
#[test]
fn no_result_while_the_session_works_is_still_running() {
assert_eq!(state_of(&[start("a")], true), ToolState::Running);
}
#[test]
fn an_unanswered_ask_is_the_readers_move_whatever_else_is_true() {
let asking = event(
3,
Event::Question {
id: "q1".to_string(),
prompt: "Allow?".to_string(),
header: None,
options: vec![QuestionOption {
label: "Allow".to_string(),
description: None,
preview: None,
}],
multi_select: false,
about: Some("a".to_string()),
},
);
let answered = event(
4,
Event::Answered {
id: "q1".to_string(),
answers: vec!["Allow".to_string()],
},
);
// Ahead of both "still running" and "no result": the reader can
// act on this one, and cannot act on either of those.
assert_eq!(
state_of(&[start("a"), asking.clone()], true),
ToolState::Deciding
);
assert_eq!(
state_of(&[start("a"), asking.clone()], false),
ToolState::Deciding
);
assert_eq!(
state_of(
&[start("a"), asking, answered, end("a", "ok", false)],
false
),
ToolState::Succeeded,
"once it is answered the call is an ordinary one again"
);
}
#[test]
fn nothing_but_a_tool_call_has_a_tool_state() {
assert_eq!(
ToolState::of(
&TranscriptItem::UserMsg {
seq: 1,
text: "hi".to_string(),
attachments: Vec::new(),
},
true
),
None
);
}
}
+588
View File
@@ -0,0 +1,588 @@
//! Where a session screen gets a transcript from: this phone's copy first,
//! the server for the rest. Ported from `app/.../TranscriptSource.kt`; see
//! `docs/TRANSCRIPT_CACHE.md` for the design this implements and
//! `docs/CLIENT_CORE.md` for how this file corresponds to the Kotlin.
//!
//! One seam rather than a cache the screen has to remember to consult.
//! Everything fetched before is asked of this, and everything the server
//! sends is written into the cache on the way past, so a caller never
//! learns which side answered. The one rule worth keeping in mind: the
//! cache is never load-bearing. Every read here has a network path beside
//! it producing the same result.
//!
//! **Not ported**: `EventStream.kt`'s reconnect-with-backoff loop and the
//! ability to close a live stream from another thread. Both are wall-clock
//! and thread-lifetime concerns that belong to whatever runtime the caller
//! embeds this crate in (a Tokio task, an iris timer, a Kotlin coroutine
//! scope) rather than to this pure logic -- `follow` below is the same
//! decorator shape `iris/desktop-app/src/app.rs` and
//! `iris/android-app/src/transcript_client.rs` already hand-wrote around
//! `event_stream::follow_session_events`, just with the cache write built
//! in so a future caller does not have to repeat it a third time.
use event_model::SeqEvent;
use crate::api::{ApiClient, ApiError, Transport};
use crate::event_stream::{self, StreamItem};
use crate::transcript_cache::SessionCache;
/// How many events a session screen opens with, cached or fetched.
///
/// The server's own default page size, named here because the cached
/// opening has to be the same size as the fetched one -- a reader must not
/// get a shorter first screen for having been here before (`OPENING_WINDOW`
/// in the Kotlin original).
pub const OPENING_WINDOW: u32 = 80;
/// A transcript-line parse failure, told apart from [`ApiError`] so a
/// caller can tell "the server is unreachable" from "the server (or this
/// phone's own disk) sent something this build cannot read" -- the two
/// mean different things to a reader (retry, versus a build that is
/// behind).
#[derive(Debug, Clone)]
pub struct ParseError(pub String);
impl std::fmt::Display for ParseError {
fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result {
f.write_str(&self.0)
}
}
impl std::error::Error for ParseError {}
/// Either half of what can go wrong asking for a page: the network, or a
/// line neither the cache's nor the server's copy of `parseSeqEvent` could
/// read.
#[derive(Debug, Clone)]
pub enum PageError {
Api(ApiError),
Parse(ParseError),
}
impl From<ApiError> for PageError {
fn from(e: ApiError) -> Self {
Self::Api(e)
}
}
impl From<ParseError> for PageError {
fn from(e: ParseError) -> Self {
Self::Parse(e)
}
}
/// What [`TranscriptSource::page`] found, kept as two states rather than
/// one possibly-empty list.
///
/// The difference is the whole of AGENTS.md's `loadOlderPage` incident: an
/// empty [`Self::Events`] means "this conversation has no more history",
/// which a caller is meant to latch, and [`Self::NothingLoaded`] means the
/// question could not be asked yet, which it must not. Collapsing the two
/// into an empty `Vec` puts the bug back, because the caller cannot tell
/// them apart -- and `unwrap_or_default()` on an `Option` would do the
/// same silently.
#[derive(Debug, Clone, PartialEq)]
pub enum OlderPage {
/// The events before the cursor, oldest first. Empty means the start
/// of the conversation has been reached.
Events(Vec<SeqEvent>),
/// Nothing is loaded, so there was no cursor to page back from
/// (`before == 0`). Not an answer about the conversation at all.
NothingLoaded,
}
fn parse_line(line: &str) -> Result<SeqEvent, ParseError> {
serde_json::from_str(line).map_err(|e| ParseError(format!("{e}")))
}
/// This phone's copy of one session's transcript, plus the server it
/// falls back to. Ported from the Kotlin `TranscriptSource` class.
pub struct TranscriptSource<T: Transport> {
api: ApiClient<T>,
session_id: String,
pub cache: SessionCache,
}
impl<T: Transport> TranscriptSource<T> {
pub fn new(api: ApiClient<T>, session_id: impl Into<String>, cache: SessionCache) -> Self {
Self {
api,
session_id: session_id.into(),
cache,
}
}
/// The cached opening window, or `None` when there is nothing usable
/// to draw.
///
/// Meant to be drawn *before* [`Self::probe`] returns, which is the
/// whole point of the feature: the rows are on screen while the check
/// that they are still the server's rows is in flight, and a failed
/// check replaces them exactly as a reset does.
pub fn cached_opening(&self, limit: usize) -> Option<Vec<SeqEvent>> {
self.cache.tail()?;
let lines = self.cache.newest(limit);
if lines.is_empty() {
return None;
}
match lines.iter().map(|l| parse_line(l)).collect() {
Ok(events) => Some(events),
// A line this build cannot read at all, which the cache's own checks cannot
// see: it reads a seq off a line, not an event. Nothing to serve, so a cold
// open.
Err(ParseError(_)) => {
self.cache.purge();
None
}
}
}
/// Whether the server's event at the cached cursor is still the cached
/// one.
///
/// A caller must not resume a live stream from a cached seq unless it
/// is the same conversation: a transcript is append-only in ordinary
/// use, but the file backing it can be replaced or truncated (a
/// sandbox re-seeded with the same ids, a backup restored, a session
/// re-imported), and the server's catch-up on such a file would hand
/// this phone a continuation of a *different* conversation, spliced
/// onto the cached one with no seam. Caught with one request of a few
/// hundred bytes.
///
/// `Ok(false)` purges the cache and means "open cold". `Err` is the
/// server not being askable, which is neither: the cached rows stay
/// on screen and the caller tries again on its own reconnect schedule.
///
/// What this cannot see is a line changed in the middle of the file
/// with the tail intact -- that is what a full reload is for.
pub fn probe(&self) -> Result<bool, ApiError> {
let Some(tail) = self.cache.tail() else {
return Ok(false);
};
// `before = seq + 1` is the newest event with seq <= the cursor, which is the
// event *at* the cursor when the server still has one there.
let page = self.api.fetch_transcript_lines(
&self.session_id,
Some(tail.seq + 1),
1,
false,
None,
)?;
let matches = page.len() == 1
&& parse_line(&tail.line)
.map(|cached| cached == page[0].1)
.unwrap_or(false);
if !matches {
self.cache.purge();
}
Ok(matches)
}
/// Today's opening fetch, kept as the start of the live run. Only
/// called when the cache has nothing to open with, or when
/// [`Self::probe`] said what it had was not the server's.
pub fn fetch_opening(&self) -> Result<Vec<SeqEvent>, ApiError> {
let page =
self.api
.fetch_transcript_lines(&self.session_id, None, OPENING_WINDOW, false, None)?;
for (line, event) in &page {
self.cache.append(line, event.seq);
}
self.cache.flush();
Ok(page.into_iter().map(|(_, event)| event).collect())
}
/// The page before `before`: from the cache when it holds it,
/// otherwise from the server bounded by what the cache already has.
///
/// The server bound (`after`) is what keeps the cache worth having. A
/// coalesced page reaches back as far as its row count takes it -- a
/// single reply is hundreds of lines -- so a page fetched after the
/// reader has been away could run straight past the cached run and
/// overlap it, and an overlapping page cannot be stored. Told where
/// this phone's copy starts, the server stops there instead.
///
/// `before == 0` answers [`OlderPage::NothingLoaded`] without asking
/// the cache or the server anything -- see AGENTS.md's "things that
/// have bitten": there is no event before the first one, so the
/// request is not a harmless no-op, and its empty answer is
/// indistinguishable from having reached the start of history.
/// Guarded here rather than left to every caller, because it is a fact
/// about the question, not about who is asking it.
pub fn page(&self, before: u64, limit: u32, coalesce: bool) -> Result<OlderPage, PageError> {
if before == 0 {
return Ok(OlderPage::NothingLoaded);
}
if let Some(lines) = self.cache.page(before, limit as usize, coalesce) {
let events: Vec<SeqEvent> = lines
.iter()
.map(|l| parse_line(l).map_err(PageError::from))
.collect::<Result<_, _>>()?;
return Ok(OlderPage::Events(events));
}
let after = self.cache.covered_up_to(before).map(|v| v - 1);
let page = self.api.fetch_transcript_lines(
&self.session_id,
Some(before),
limit,
coalesce,
after,
)?;
if let Some((_, first_event)) = page.first() {
// `before` rather than the newest line's seq: a coalesced page covers
// everything up to the cursor it was asked with, and nothing in its lines
// says so.
let lines: Vec<String> = page.iter().map(|(line, _)| line.clone()).collect();
self.cache
.store_page(&lines, first_event.seq, before, coalesce);
}
Ok(OlderPage::Events(
page.into_iter().map(|(_, event)| event).collect(),
))
}
/// [`event_stream::follow_session_events`], with every frame written to
/// the cache before `on_item` sees it.
///
/// Before, so that an event held back for a reader who is scrolled
/// away is already on disk -- what the cache holds is what the server
/// sent, not what a screen has got round to drawing. Flushed on each
/// status change, which is a turn's boundary and the granularity a
/// crash may as well lose, and once more when the stream ends.
pub fn follow(
&self,
after: u64,
mut on_item: impl FnMut(StreamItem) -> bool,
) -> Result<(), ApiError> {
let cache = &self.cache;
let result = event_stream::follow_session_events(
self.api.transport(),
&self.session_id,
after,
|item| {
if let StreamItem::Event { raw, event } = &item {
cache.append(raw, event.seq);
if matches!(event.event, event_model::Event::Status { .. }) {
cache.flush();
}
}
on_item(item)
},
);
cache.flush();
result
}
/// Leaves the cache with everything it was given -- called once a
/// caller is done with this source, mirroring the Kotlin `close`'s
/// final flush (that method's stream cancellation itself is the
/// runtime concern the module doc says is not ported here).
pub fn close(&self) {
self.cache.flush();
}
}
#[cfg(test)]
mod tests {
use super::*;
use crate::api::{Body, RawResponse};
use std::collections::VecDeque;
use std::io::Read;
use std::sync::Mutex;
/// A transport that answers fixed bodies in call order, and records
/// every path it was asked for -- so a test can assert *how many*
/// requests a method made, which is the point for the `before == 0`
/// guard (AGENTS.md's regression: the guard must stop the request
/// before it happens, not merely tolerate the empty answer).
#[derive(Default)]
struct ScriptedTransport {
responses: Mutex<VecDeque<(u16, String)>>,
calls: Mutex<Vec<String>>,
}
impl ScriptedTransport {
fn respond(&self, status: u16, body: impl Into<String>) {
self.responses
.lock()
.unwrap()
.push_back((status, body.into()));
}
fn call_count(&self) -> usize {
self.calls.lock().unwrap().len()
}
}
impl Transport for ScriptedTransport {
fn request(
&self,
_method: &str,
path: &str,
_body: Option<Body>,
) -> Result<RawResponse, ApiError> {
self.calls.lock().unwrap().push(path.to_string());
let (status, body) = self
.responses
.lock()
.unwrap()
.pop_front()
.unwrap_or_else(|| panic!("ScriptedTransport got an unscripted request: {path}"));
Ok(RawResponse {
status,
body: body.into_bytes(),
})
}
fn stream(&self, path: &str) -> Result<Box<dyn Read + Send>, ApiError> {
self.calls.lock().unwrap().push(path.to_string());
let (_, body) = self
.responses
.lock()
.unwrap()
.pop_front()
.unwrap_or_else(|| {
panic!("ScriptedTransport got an unscripted stream request: {path}")
});
Ok(Box::new(std::io::Cursor::new(body.into_bytes())))
}
}
fn source(
transport: ScriptedTransport,
cache_root: &std::path::Path,
) -> TranscriptSource<ScriptedTransport> {
let api = ApiClient::new(transport);
let cache = crate::transcript_cache::TranscriptCache::new(cache_root).session("s1");
TranscriptSource::new(api, "s1", cache)
}
fn status_line(seq: u64) -> String {
format!(r#"{{"seq":{seq},"ts":1.0,"type":"status","state":"idle"}}"#)
}
#[test]
fn a_cold_cache_has_no_opening_and_fetches_from_the_server() {
let dir = tempfile::tempdir().unwrap();
let transport = ScriptedTransport::default();
transport.respond(200, format!("[{}]", status_line(1)));
let source = source(transport, dir.path());
assert_eq!(source.cached_opening(80), None);
let opening = source.fetch_opening().unwrap();
assert_eq!(opening.len(), 1);
assert_eq!(opening[0].seq, 1);
// The fetch wrote through: reopening the same cache now has something to show.
assert!(source.cache.tail().is_some());
}
#[test]
fn probe_matching_the_cached_tail_leaves_the_cache_alone() {
let dir = tempfile::tempdir().unwrap();
let transport = ScriptedTransport::default();
transport.respond(200, format!("[{}]", status_line(1)));
let source = source(transport, dir.path());
source.fetch_opening().unwrap();
let transport2 = ScriptedTransport::default();
transport2.respond(200, format!("[{}]", status_line(1)));
let cache = crate::transcript_cache::TranscriptCache::new(dir.path()).session("s1");
let source2 = TranscriptSource::new(ApiClient::new(transport2), "s1", cache);
assert!(source2.probe().unwrap());
assert!(source2.cache.tail().is_some());
}
#[test]
fn probe_mismatching_the_cached_tail_purges_the_cache() {
let dir = tempfile::tempdir().unwrap();
let transport = ScriptedTransport::default();
transport.respond(200, format!("[{}]", status_line(1)));
let source = source(transport, dir.path());
source.fetch_opening().unwrap();
// The server now answers with a different event at the same seq -- the file
// behind this session was replaced.
let transport2 = ScriptedTransport::default();
let different = r#"{"seq":1,"ts":1.0,"type":"status","state":"running"}"#.to_string();
transport2.respond(200, format!("[{different}]"));
let cache = crate::transcript_cache::TranscriptCache::new(dir.path()).session("s1");
let source2 = TranscriptSource::new(ApiClient::new(transport2), "s1", cache);
assert!(!source2.probe().unwrap());
assert!(source2.cache.tail().is_none());
}
#[test]
fn probe_finding_no_server_leaves_the_cache_untouched() {
let dir = tempfile::tempdir().unwrap();
let transport = ScriptedTransport::default();
transport.respond(200, format!("[{}]", status_line(1)));
let source = source(transport, dir.path());
source.fetch_opening().unwrap();
let transport2 = ScriptedTransport::default();
transport2.respond(500, "server on fire");
let cache = crate::transcript_cache::TranscriptCache::new(dir.path()).session("s1");
let source2 = TranscriptSource::new(ApiClient::new(transport2), "s1", cache);
assert!(source2.probe().is_err());
assert!(
source2.cache.tail().is_some(),
"an unreachable server must not be treated as a mismatch"
);
}
/// The regression this module exists to close: `before == 0` must
/// never reach the network or the cache, because an empty answer there
/// is indistinguishable from "there is genuinely no more history" --
/// AGENTS.md's `loadOlderPage` incident.
#[test]
fn paging_before_the_first_event_makes_no_request_at_all() {
let dir = tempfile::tempdir().unwrap();
let transport = ScriptedTransport::default();
let source = source(transport, dir.path());
assert_eq!(source.page(0, 80, true).unwrap(), OlderPage::NothingLoaded);
assert_eq!(source.api.transport().call_count(), 0);
}
#[test]
fn a_page_already_covered_by_the_cache_never_reaches_the_server() {
let dir = tempfile::tempdir().unwrap();
let transport = ScriptedTransport::default();
transport.respond(200, format!("[{},{}]", status_line(1), status_line(2)));
let source = source(transport, dir.path());
source.fetch_opening().unwrap();
let calls_before = source.api.transport().call_count();
let OlderPage::Events(page) = source.page(2, 10, true).unwrap() else {
panic!("a cursor of 2 is a real question about the conversation");
};
assert_eq!(page.len(), 1);
assert_eq!(page[0].seq, 1);
assert_eq!(
source.api.transport().call_count(),
calls_before,
"a cache hit must not touch the network"
);
}
/// With nothing older cached there is no floor to give the server, so
/// the request carries no `after` at all.
#[test]
fn a_server_page_with_nothing_older_cached_carries_no_bound() {
let dir = tempfile::tempdir().unwrap();
let transport = ScriptedTransport::default();
transport.respond(200, format!("[{}]", status_line(5)));
let source = source(transport, dir.path());
source.fetch_opening().unwrap();
let transport2 = ScriptedTransport::default();
transport2.respond(200, format!("[{}]", status_line(3)));
let cache = crate::transcript_cache::TranscriptCache::new(dir.path()).session("s1");
let source2 = TranscriptSource::new(ApiClient::new(transport2), "s1", cache);
source2.page(5, 10, true).unwrap();
assert_eq!(
source2.api.transport().calls.lock().unwrap()[0],
"/sessions/s1/transcript?limit=10&before=5&coalesce=true"
);
}
/// The half the test above cannot show: when the cache *does* hold an
/// older run, the fetch is floored at its end, or the page would run
/// straight past it and overlap -- which `store_page` then refuses,
/// silently costing the phone the page it just paid for.
#[test]
fn a_server_page_is_floored_at_the_end_of_the_cached_run() {
let dir = tempfile::tempdir().unwrap();
let cache = crate::transcript_cache::TranscriptCache::new(dir.path()).session("s1");
// A stored page covering [3, 6) and two live events above it, so the run this
// phone holds is [3, 8) -- the newest chunk has to be an appended one, or the
// cache reads the directory as damaged and discards it.
let lines: Vec<String> = (3..6).map(status_line).collect();
assert!(cache.store_page(&lines, 3, 6, true));
cache.append(&status_line(6), 6);
cache.append(&status_line(7), 7);
cache.flush();
let transport = ScriptedTransport::default();
transport.respond(200, format!("[{}]", status_line(9)));
let source = TranscriptSource::new(ApiClient::new(transport), "s1", cache);
source.page(10, 10, true).unwrap();
assert_eq!(
source.api.transport().calls.lock().unwrap()[0],
"/sessions/s1/transcript?limit=10&before=10&coalesce=true&after=7",
"the fetch must stop one seq below where this phone's copy ends"
);
}
/// A page the server could not answer is an error, never an empty
/// page: the caller would read the second as "this conversation has no
/// more history" and stop paging for good.
#[test]
fn a_failing_server_page_is_an_error_rather_than_an_empty_one() {
let dir = tempfile::tempdir().unwrap();
let transport = ScriptedTransport::default();
transport.respond(500, "server on fire");
let source = source(transport, dir.path());
assert!(matches!(source.page(9, 10, true), Err(PageError::Api(_)),));
}
/// A cached line this build cannot read is told apart from the network
/// failing, for the same reason: neither is "no more history".
#[test]
fn an_unreadable_cached_page_is_a_parse_error_rather_than_an_empty_one() {
let dir = tempfile::tempdir().unwrap();
let cache = crate::transcript_cache::TranscriptCache::new(dir.path()).session("s1");
cache.store_page(
&[r#"{"seq":3,"but":"not an event"}"#.to_string()],
3,
4,
true,
);
cache.append(&status_line(4), 4);
cache.flush();
let transport = ScriptedTransport::default();
let source = TranscriptSource::new(ApiClient::new(transport), "s1", cache);
assert!(matches!(source.page(4, 10, true), Err(PageError::Parse(_)),));
assert_eq!(
source.api.transport().call_count(),
0,
"a cache hit that cannot be read must not fall through to the server unnoticed"
);
}
#[test]
fn a_bad_cached_opening_line_purges_rather_than_panicking() {
let dir = tempfile::tempdir().unwrap();
let cache = crate::transcript_cache::TranscriptCache::new(dir.path()).session("s1");
cache.append("not json at all", 1);
cache.flush();
let transport = ScriptedTransport::default();
let source = TranscriptSource::new(ApiClient::new(transport), "s1", cache);
assert_eq!(source.cached_opening(80), None);
assert!(
source.cache.tail().is_none(),
"a damaged line purges the cache"
);
}
#[test]
fn follow_writes_events_to_the_cache_before_the_caller_sees_them() {
let dir = tempfile::tempdir().unwrap();
let transport = ScriptedTransport::default();
transport.respond(200, format!("{}\n\n", sse_frame(&status_line(1))));
let source = source(transport, dir.path());
let mut seen = Vec::new();
source
.follow(0, |item| {
if let StreamItem::Event { event, .. } = item {
seen.push(event.seq);
}
true
})
.unwrap();
assert_eq!(seen, vec![1]);
assert_eq!(source.cache.tail().unwrap().seq, 1);
}
fn sse_frame(data: &str) -> String {
format!("data:{data}")
}
}
+95 -25
View File
@@ -25,16 +25,19 @@ next (a Masonry or iris transcript screen, most likely).
| `sse.rs` | `Sse.kt` (the framing half) | Done, new tests (Kotlin had none of its own beyond integration) |
| `api.rs` | `Api.kt` | Partial -- see below |
| `event_stream.rs` | `EventStream.kt` | Done |
| `transcript_fold.rs` | `TranscriptItems.kt`, `ToolRows.kt` | Partial -- see below |
| `transcript_fold.rs` | `TranscriptItems.kt`, `ToolRows.kt` | Done -- see below |
| `config.rs` | `ServerConfig.kt`'s `handleEnrollment` | New, desktop-only so far -- see below |
| *(not started)* | `TranscriptSource.kt` | Not started |
| `transcript_source.rs` | `TranscriptSource.kt` | Done -- see below |
| *(not ported, and may never be)* | `TranscriptUnits.kt` | Out of scope -- see below |
Every file above whose Kotlin counterpart had a JVM unit test (`AnsiTest`,
`HighlighterTest`, `TranscriptCacheTest`) has had every one of those test
cases ported alongside it, plus new tests for the pieces that had none
(`sse.rs`, `api.rs`, `event_stream.rs`, `transcript_fold.rs`). Test count by
crate as of this writing: **85 in `client-core`**, 0 in `event-model` (its
(`sse.rs`, `api.rs`, `event_stream.rs`, `transcript_fold.rs`,
`transcript_source.rs` -- the Kotlin `TranscriptSource.kt`/`TranscriptItems.kt`
had no JVM unit tests of their own, so these were written fresh against the
Kotlin source and AGENTS.md's paging incidents as the spec). Test count by
crate as of this writing: **109 in `client-core`**, 0 in `event-model` (its
types carry no logic of their own to test -- `server/`'s own tests exercise
them via `session::transcript`'s round-trip coverage).
@@ -88,13 +91,27 @@ the full table to work from when one of these is next.
including tool-call/question/image attachment and peer-message placement.
`group_tool_runs` groups adjacent calls into `TranscriptRow::Tools`.
**Not ported:** `TranscriptItems.kt`'s `joinPages` (and its
`healSplitMessage`/`adoptRun` helpers) -- the page-boundary healing that
merges a tool call split across two fetched pages and re-merges a run a
boundary cut through. This matters the moment paging backward through
history is exercised; it is deliberately left rather than rushed, since
it is exactly the kind of boundary logic this project's own "things that
have bitten" section warns reads fine and is wrong at the edges.
`join_pages` (with `heal_split_message` and `adopt_run`, both private) is
now ported too, 2026-09-06 -- the page-boundary healing that merges a tool
call split across two fetched pages, rejoins a message a boundary cut
through, and renames a run of tool calls onto whichever name is already on
screen. Ported with AGENTS.md's "things that have bitten" incidents as the
spec rather than a JVM test file (`TranscriptItems.kt` had none of its
own): `a_clean_boundary_between_two_finished_runs_is_still_healed_into_one_run`
is the regression test for the bug that shipped -- `adopt_run` must run on
*every* join, not only the one where a split call was found, or a boundary
landing cleanly between two already-finished calls (most of them) leaves
one run drawn as two. `a_call_split_across_the_boundary_merges_into_one_row`,
`a_message_split_across_the_boundary_is_rejoined_with_the_newer_halfs_identity`,
and `adopt_run_never_renames_into_a_question_row` cover the other three
edges the Kotlin doc calls out. `join_pages` ends in a `debug_assert!`
that no tool id survives in both halves -- the duplicate row it exists to
prevent, checked rather than assumed. What it deliberately does *not*
assert is seq ordering across the boundary: a peer note carries the seq
its turn began at (`place_peer_note`), which can be older than the page
it arrived in, so the two pages' seqs legitimately interleave there. An
earlier draft asserted it and would have panicked in debug builds on an
ordinary transcript.
**Known gap, and a decision for whoever closes it:** `event_model::Event`
has no `Unknown`/catch-all variant, unlike `Events.kt`'s hand-kept mirror.
@@ -119,20 +136,72 @@ caller-specific (the code rules' "ask for the least you need"). Its only
caller today is `desktop-app`; a future Android build of this crate would
be a second one, not a reason to move the type.
## What `transcript_source.rs` covers, and what it does not
`TranscriptSource<T: Transport>` is the seam a session screen asks for a
page, ported test-for-test against the Kotlin doc rather than a JVM test
file (there wasn't one): `cached_opening`, `probe`, `fetch_opening`,
`page` and `follow`, each matching its Kotlin namesake's contract --
including `probe`'s three-way outcome (matches / cache purged /
unreachable, told apart so a caller never treats "couldn't ask" as "was
wrong") and `page`'s cache-vs-server split bounded by `covered_up_to`.
Two additions beyond a literal port, both load-bearing:
- **`page(before, ..)` refuses `before == 0` before touching the cache or
the network**, answering `OlderPage::NothingLoaded`. This is AGENTS.md's
`loadOlderPage` incident (`before = 0` is "no event before the first
one," indistinguishable from "reached the start of history" if a caller
ever asks it) moved out of the Kotlin screen and into this layer, so
every future caller gets the guard rather than having to remember it.
**The return type is `OlderPage`, not a `Vec`, and that is the guard.**
The Kotlin's two falses are different answers -- `oldestSeq == 0`
returns without touching `moreHistory`, an empty page latches it false
-- so a port that answered both with an empty list would have moved the
bug rather than fixed it, one layer down and out of sight of the screen
that used to hold the check. `OlderPage::Events(vec![])` means the start
of the conversation; `OlderPage::NothingLoaded` is not an answer about
the conversation at all. Reviewed 2026-09-06.
`paging_before_the_first_event_makes_no_request_at_all` asserts zero
transport calls, not just the variant, since a request that happens to
answer empty is exactly what caused the original bug, and
`a_failing_server_page_is_an_error_rather_than_an_empty_one` plus
`an_unreadable_cached_page_is_a_parse_error_rather_than_an_empty_one`
are the same rule for the two ways a page can fail.
- **`fetch_transcript_lines`** (new in `api.rs`) hands back each line
paired with the exact server bytes it came from, via
`serde_json::value::RawValue` rather than re-serializing a parsed
`Value` -- the cache and a live SSE frame for the same event have to
agree byte-for-byte, which is exactly what the `serde_json`
float-rounding bug (AGENTS.md) was about. The existing
`fetch_transcript_page` is untouched (other callers under `iris/`
depend on its signature); the two share a `transcript_path` helper so
the query string is written in one place.
**Not ported:** `EventStream.kt`'s reconnect-with-backoff loop, and
`TranscriptSource.close`'s ability to cancel a live stream from another
thread. Both are wall-clock/thread-lifetime policy that belongs to
whichever runtime embeds this crate (iris's own timers, a Tokio task, a
Kotlin coroutine scope), not to this pure logic -- `follow` is the same
"write to the cache, then hand the frame to the caller" decorator
`iris/desktop-app/src/app.rs` and `iris/android-app/src/transcript_client.rs`
already hand-wrote around `event_stream::follow_session_events` before this
existed; the cache write moved into one shared place so a third caller
does not repeat it again by hand.
## What is not started at all
- **`TranscriptSource.kt`** -- the layer that decides whether a page comes
from the transcript cache or the server, and stitches the two. Needs
`transcript_cache.rs` and `api.rs`'s transcript-page method, both of
which exist now, so this is unblocked whenever picked up.
- **The markdown *block* model beyond syntax spans** -- `highlight/markdown.rs`
colours a `.md` file or fence for the highlighter, but does not build the
block tree (headings, lists, tables, fences as distinct nodes) that a
renderer walks to lay out prose versus code versus a table.
`CodeFence.kt`'s use of `org.intellij.markdown` for that full CommonMark
AST is Compose rendering plumbing, not something to port as-is; a Rust
UI layer will want its own block parser or a crate for it, decided
alongside the framework choice in RUST.md.
- **A full markdown AST.** `markdown_blocks` (2026-09-06) splits a message
into its *top-level* blocks -- heading, paragraph, fence, list, table,
quote -- with each block's own source, which is what a renderer needs to
lay out prose versus code and what lets a streamed delta re-lay out one
block instead of the message (docs/RUST.md's Task B). What it
deliberately does **not** build is the tree below that: nested list
items, table cells, inline spans. Inline styling is still the renderer's
own job per block (`iris/transcript-ui/src/markdown.rs`), and nothing
has needed the rest yet. `CodeFence.kt`'s use of `org.intellij.markdown`
for a full CommonMark AST is Compose rendering plumbing, not something
to port as-is.
- **`TranscriptUnits.kt`** (see above) -- deliberately out of scope, since
it flattens a row into bounded units for a *specific* lazy-list
framework's composition cost, which is a fact about that framework
@@ -142,5 +211,6 @@ be a second one, not a reason to move the type.
`./run-tests.sh` from the repo root now runs `event-model`, `client-core`
and `server` in that order (each `cargo test`, forwarding arguments the
same way it always has). From `client-core/` directly: `cargo test`,
`cargo clippy --all-targets`, `cargo fmt` -- all clean as of this writing.
same way it always has). From `client-core/` directly: `cargo test`
(119 tests), `cargo clippy --all-targets`, `cargo fmt` -- all clean as of
this writing (2026-09-06).
+790
View File
@@ -0,0 +1,790 @@
# Decisions taken for Iris to review
Short list of design choices made by the design agent without asking, so
they can be judged and reversed later. Detail lives in RUST.md (and IRIS.md
for iris API changes); this file is only the summary. Newest first. Items
marked **DEFERRED** are ones the agent chose not to decide alone.
## 2026-09-08 (later still: the list's overscroll clamp, in frame)
Finishes the item the previous entry deferred. IRIS.md has the account.
- **`List` lays out a second time within the frame** when its walk lands
off the end of the content, instead of writing the correction to the
anchor and asking for another frame. The extra walk is paid only on an
overscrolled frame, and it is mostly O(1) moves.
- **`Painter::draw_again` is removed**, `List` having been its only
caller -- so the framework no longer offers a way to ask for a
corrective frame at all.
- **`List::place`'s top-known and bottom-known cases are one path**
(`Placement::edges`), which is the "write the logic once" rule applied
to two symmetric directions rather than a behaviour change.
## 2026-09-08 (later: a scroll area measures and places in one frame)
From Iris's phone report about the composer's padding while typing
newlines, and the rule she stated when she read the first fix: layout is
a pure function of the state, nothing self-heals, and two draws to place
something happen in the same frame. IRIS.md's entry has the account.
- **`Scroll::draw` draws its child twice** -- once at last frame's length
to measure it, once at the measured length to place it -- instead of
placing against the stale length and leaving a wrong frame on screen.
The second draw is free unless the content's length changed.
- **An end-anchored `Scroll` is at its end on its first drawn frame**, a
consequence of the above. Two layout tests now build their area with
`at_end: false`, which is what they meant: they scroll down from the
top.
- **`List::clamp_to_content`'s next-frame correction is left in place**
and written down in docs/IRIS_TODO.md instead of fixed here, because
`List::place` is a larger piece of machinery and deserves its own
before/after on the phone.
## 2026-09-08 (every crate to its latest version, wgpu 28 -> 30)
At Iris's request. RUST.md's "Every crate to its latest version" box has
the full list and the migration.
- **wgpu 30 taken now rather than pinned at 28.** Two majors of API
change, all mechanical (instance descriptor, optional bind-group and
vertex-buffer slots, `Queue::present`, a `CurrentSurfaceTexture` enum),
and one that would have been a startup abort on a device rather than a
compile error: naga now demands `@interpolate(flat)` on the shader's
integer varyings. Verified on both backends before this was called
done, since a renderer that compiles proves nothing.
- **The desktop instance now carries winit's display handle.** wgpu 30
asks for it when a GLES surface will be presented on Wayland, which is
what this machine's Vulkan-to-GLES fallback produces. Android passes
none: its surface comes from a `NativeWindow`.
- **`syn` 2 -> 3, `pollster` 0.4 -> 1.0** with no source change in
`iris/macro` or anywhere else.
## 2026-09-08 evening (the fling is shared; a cancel is not a release)
From Iris's four-item phone report; RUST.md's "2026-09-08 (evening)" box
has the reasoning and the tests, IRIS.md the summary.
- **A `Flinger` that does not know which way the content moves.** Every
scroll area flings now, on either axis, as Iris asked -- and the
physics is one type shared by `List` and `Scroll` rather than a copy
each. The choice worth reviewing is the seam: `Flinger` owns the curve
and the clock, and the *caller* owns the sign convention and where the
content ends. Rejected: teaching `Flinger` a direction, which would
have to be told to it -- and being told is the same thing as not
knowing, with an extra field to get wrong.
- **A cancel is a first-class end to a gesture, not an early release.**
`CursorState::cancelled` is new state on the pointer sample, set by
Android's `ACTION_CANCEL` and the harness's `TouchAction::Cancel`.
Rejected: mapping a cancel to `PressEnd` and having each widget decide
what to suppress, which is what shipped and is why leaving the app
flung the transcript.
- **A `DragGesture` ignores a `Cancel` it caused.** One gesture is
driven by several widgets, so the widget that was pressed can be a
"loser" on the frame its own gesture won. The test is whether the
gesture's own capture id is the holder. This is what makes it safe for
every widget driving a gesture to register the whole `drag_senses()`
set, which is now the rule without exception.
- **`List::place` draws a resized row twice in one frame.** The old
comment accepted a one-frame lag by analogy with `Scroll`'s content
length. That analogy was wrong: a stale *length* only misplaces the
next thing, while a stale *box* is drawn, because a background fills
whatever box it is handed. The extra draw is bounded to frames where a
row's height actually changed.
## 2026-09-08 (iris ships an icon font, and the drawn mark is deleted)
- **Directed by Iris.** Her question on seeing `widget::mark`: "why does
mark exist? The font should be working if it's working for compose and
nerd fonts are bundled." It was not: the Compose app draws its icons
from **its own committed Nerd Fonts subset**, while iris was setting
the disclosure mark with bare Unicode geometric codepoints
(U+25B8/25BE/25B4) out of whatever face the platform resolved -- an
empty box on her phone, a dot on this VM. The 2026-09-07 entry below,
which said "iris had no equivalent icon font to keep", is what left
that gap: iris had no icon font because it had never had one, not
because it needed none.
- **So iris now bundles the same kind of subset**:
`iris/core/build-icon-font.sh` writes
`iris/core/assets/fonts/nerd_icons.ttf` (992 bytes, three Material
Design glyphs today), `iris::icon` names the codepoints, and
`Family::Icons` draws them. This does **not** reopen the platform-fonts
decision: body and monospace text still come from the platform, and an
icon is the opposite case -- a small, closed, known set of codepoints,
which is exactly the division the Compose app already makes.
- **`iris::widget::mark` is deleted** (added earlier the same day). It
drew a correct triangle, but only a triangle, and every further icon
would have been another bespoke rasteriser. An icon as text also takes
the size, colour and baseline of the line it sits in for free.
## 2026-09-08 (the emulator is a GLES machine, and Vulkan is verified elsewhere)
- **Directed by Iris, carried out here**: "make sure the setup uses GL for
the android emulator and remove any vulkan requirements. That'll be
tested through both the desktop version as well as my phone." So the
emulator is settled as a GLES rig and nothing chases hardware Vulkan in
it any more; the Vulkan path is covered by the desktop build and by her
phone.
- **Nothing had to be forced to make that true.** Measured in the guest
the same day: the emulator has no hardware Vulkan at all (its only
Vulkan is SwiftShader, in software) and its GLES is the host's real RX
7900 XT through virgl at ES 3.1. iris's existing runtime fallback --
`Backends::PRIMARY`, no adapter, rebuild on `Backends::GL` -- already
lands there, verified end to end with an ordinary (no `force-gles`)
debug APK.
- **The emulator and the phone therefore run the same binary**, differing
only in what that binary finds. That is deliberate and worth not
undoing: a build flag that changed the backend would mean the thing
measured on the emulator is not the thing shipped. `force-gles` stays,
but only for pinning the backend on a machine that *does* have Vulkan
(the desktop), and never for a phone build.
- **Every run now says which adapter drew it.** The Android renderer logs
the full adapter line at startup the way the desktop already did -- only
the backend enum was logged before, which cannot separate `Gl` on the
host's GPU from `Gl` on SwiftShader, or a phone's real Vulkan from a
software one. `run-bench.sh` prints that line before any number.
- **No Vulkan requirement was found in iris to remove.** `device_limits()`
asks for nothing beyond wgpu's defaults (and zeroes the compute fields),
neither backend requires a feature, and both probe rather than
`.expect()` an adapter. What was removed was the *documentation* telling
people to boot the emulator with SwiftShader Vulkan.
## 2026-09-07 (platform fonts, not bundled ones)
- **Iris's own decision, carried out as directed**: removed the 3.6 MB of
bundled Noto Sans/Noto Sans Mono TTFs from `iris-core` and load text
from the platform's own font collection instead (`fontique`'s system
discovery, already on by default). Matches what the Compose app does --
it takes body text from `FontFamily.Default` and code text from
`FontFamily.Monospace`, both platform-resolved, and ships no text font
of its own. Rejected alternative (the one this pass had left open
2026-09-06): subsetting the bundled Noto Sans to Latin/common
punctuation instead of removing it outright, which would have kept
identical rendering across devices for a smaller (not zero) size cost;
Iris chose to match Compose instead.
- `.so` **-3,748,136 bytes** (11,193,608 -> 7,445,472), matching the
original 3.6 MB estimate. Fallback still lands on the platform's own
tofu for a codepoint no resolved face has (checked with CJK + emoji on
desktop) rather than blank space, so the UI_RULES unknown-glyph rule
still holds.
- **Gap found, then closed same day**: this fontique version's Android
backend never resolved the `Monospace` generic family at all (confirmed
on this checkout's emulator, `mono=None` in the startup diagnostic) --
two pre-existing bugs in fontique's own `fonts.xml` parsing stacked (an
ordering bug, and a `<family name="monospace">` declaration whose
`<font>` children the backend's parser never reads), not something this
change introduced, but this change is what stopped masking it (the
bundled mono font used to be registered ahead of the broken platform
lookup, so it always won). Checked `linebender/parley`'s `main` branch
on GitHub: neither bug is fixed there, so there was no newer release to
bump to. Fixed instead in `iris-core` itself
(`TextData::patch_android_monospace`, Android-only): reads
`/system/etc/fonts.xml`'s own `"monospace"` declaration for the font
filename it names, then registers whichever of fontique's actually-
scanned families owns that file as the `Monospace` generic -- the same
authority Compose's `Typeface.MONOSPACE` resolves through, without
pinning an OEM-specific family name. Verified on this checkout's
emulator: `mono=Some("Droid Sans Mono")`, and a screenshot showing the
bench-fixture's code block and tool-card values in a visibly monospaced
face beside sans body text; the desktop `fontconfig` backend is
unaffected (still resolves monospace correctly, confirmed unchanged).
docs/RUST.md's "Platform fonts (2026-09-07)" has the full account.
## 2026-09-07 (a phone log reaches Iris through Dev Updater's own tab)
**Supersedes the "how a phone log reaches Iris" entry below, same day.**
Iris's call once the route was working: put it in Dev Updater properly
rather than smuggling the lines through `ai-server`'s log.
- **The app exposes its own log on the device, and Dev Updater reads it
there.** A `ContentProvider` at `<applicationId>.devlog`, one table of
lines queried with `?since=<seq>` so a poll is incremental, plus a
`status` row (`held`, `dropped`, `newest_seq`). Dev Updater's phone app
polls it while the component's **Runtime** tab is open and forwards what
is new to its own build machine, into that APK component's runtime log
-- so the same tab renders both kinds and the history outlives the
phone. No tunnel, no token, no second enrolment: the two apps are on the
same phone.
**It is a contract, not a feature for iris.** Written down in
dev-updater's `README.md` ("An app's own log"), so any app that server
delivers gets the tab by implementing it; the Compose app in `app/` can
do the same later. That is the reason it beat the route below on its
second look -- the earlier one only ever worked for the one project that
had a server, and put a phone's lines under a *different component* than
the one they came from.
- **Read access is `protectionLevel="normal"`, and that is a real trade.**
`signature` is what this wants and is not available: Dev Updater and the
apps it delivers are built on one machine but signed with different
locally generated keys, so a signature permission would be held by
nothing at all. What `normal` costs is that any app on that phone which
requests `dev.updater.permission.READ_DEVLOG` by name can read another
app's dev log. Accepted because these are development builds on a
development phone and the alternative was no log; stated in the manifest
beside the declaration and in dev-updater's README so it is not
rediscovered as a surprise.
- **The provider polls rather than notifying.** `notifyChange` was not
implemented: the ring is filled by a `log::Log` backend on whatever
thread logged, and giving that a route to a `ContentProvider` means
plumbing a callback through `client-core` for every platform. Dev
Updater's contract therefore says it polls (about a second, only while
the tab is open), which is what keeps implementing the contract cheap --
a provider that does notify loses nothing.
- **What was deleted, so there is one mechanism**: `client-core`'s
`log_upload` module, `POST /client-log` on `ai-server`, the
`AI_APP_LOG_HOST`/`_PORT`/`_TOKEN` baking in `iris/android-app/build.rs`
(which left that file with nothing to do, so it is gone too), and the
uploader fields on both Android clients. Kept: the ring, `RingLogger`,
`install_process_logger`, and the Diagnostics line counting what is
held. The upload-status line there is now **"devlog provider:
content://<authority>"** -- named from what the provider registered
rather than composed from the package here, so a screenshot of that pane
is evidence the contract is live and says which package's log it is.
## 2026-09-07 (how a phone log reaches Iris) -- superseded, see above
- **The app sends its own log to `ai-server`, and Dev Updater shows it as
`ai-server`'s runtime log.** Iris has no `adb`/`logcat` on her phone, and
Android forbids one app reading another's logcat, so the app has to carry
its own copy and post it somewhere. `POST /client-log` on `ai-server`
re-emits each line into that server's own `tracing` output; Dev Updater
already runs `ai-server` as a `Managed` component, whose stdout its own
service script redirects to a file and reports through
`GET /apps/{key}/components/{name}/logs?kind=runtime`, which the phone
app's log dialog already offers as a **Runtime** tab for a `server`
component. So **no change to Dev Updater at all** -- one route on
`ai-server`, and the client in `client-core`.
**Rejected: posting to Dev Updater's own server** (the first candidate,
and what the entry above went on to build -- the estimate below was
right about the work and wrong about it being too much).
It would need a new authenticated *write* route on a TLS surface whose
module doc says every route on it "is, or decides, the bytes that get
handed to `REQUEST_INSTALL_PACKAGES` next"; a per-app device-log store;
a change to `component_logs` so an APK component can have a runtime log;
a change to the phone app's `hasBothKinds = component.kind == "server"`
gate and to what `hasRuntimeLogs` means on the wire; and -- the real
cost -- a **second** enrollment for the iris app, since it has no CA or
token for Dev Updater and Dev Updater mints tokens per device by QR.
Five changes across two repos against one route, for the same line
landing in the same viewer.
**Rejected: a share intent from a debug button** (a log file in the app's
external files dir, shared by hand). It works today and needs no server,
but every line costs Iris a manual export and a message, which is the
round trip through a person this was meant to remove. It is still the
fallback when the tunnel is down, and GrapheneOS's own per-app log export
already covers the crash case (that is how the `ToolInput.highlighted`
crash was reported).
- **The ring is in `client-core`, not in the Android crate.** A bounded
in-memory ring (2000 lines or 256 KiB, whichever bites first) behind a
`log::Log` backend that *forwards* to whichever logger the platform
already installed, so `logcat` and a desktop terminal see exactly what
they saw before. The platform supplies only its own logger and its
destination. `Copy report` appends the ring to what goes on the
clipboard, and flushes the uploader first.
- **The destination is baked in at build time, from the build machine's
own files** (`AI_APP_LOG_HOST`/`_PORT`/`_TOKEN` plus the pinned CA) --
*gone; the provider above replaced it.* What is worth keeping from it is
the reason it went: an APK good only for the server that built it cannot
be built in this VM for Iris's phone, which is the case that mattered.
all three or none, never two. The same trust boundary the transcript
config and the Compose APK's CA already use: nothing secret is
committed, and an APK is good for the server that built it. A build told
nothing still keeps its ring and still copies it; the diagnostics pane
says which of "not tried yet", "failing -- <why>" and "no server
configured" it is, because otherwise all three look like silence.
## 2026-09-06 (how a tool call looks, P1b)
- **A card that never got a result says "no result", in yellow, and it is
a state Compose cannot say.** A call that finished having printed
nothing and a call whose turn was interrupted before anything came back
both leave an empty output. Compose draws both as an ordinary finished
call, which reads as a fact somebody established. There are five states
now, each with a word and a colour: nothing at all for a call that
worked, "running" (grey), "your turn" (peach, Compose's own wording and
colour), "failed" (red), "no result" (yellow).
- **A failed call is drawn as failed, which needed a field on the wire.**
`is_error` is on the CLI's `tool_result` and was being dropped; the
server now carries it to the phone. Reversible, but the alternative is a
card that says a call succeeded because it cannot tell.
- **A group's cards do not each carry their own surface.** Compose gives
each card a fill and squares the corners where it faces a neighbour, so
a run reads as one object broken into parts. iris has no per-corner
radius, and -- more to the point -- a group built the way Compose builds
it hit a framework layout defect that drew every card's text a card
below its own box. So a group is one surface with its cards on it,
separated by a small gap, and the 4dp inset Compose holds them off the
edge by is gone. Worth revisiting once the layout defect is fixed
(docs/IRIS_TODO.md).
- **A long tool output is capped at 80 lines or 4 kB with a "Show all N
lines".** Compose draws the whole thing, and gets away with it because
its `Text` inside a `LazyColumn` lays out lazily; here the output is one
text widget and shaping a hundred kilobytes of it costs what the file
editor's 32 kB limit was measured against. If iris's text gets cheaper,
this is the number to move.
- **A card's command is clipped, not pannable, and its summary line is
clipped rather than ellipsised.** Both are framework gaps rather than
choices (`scrollable_on` on a non-editable text draws nothing; there is
no overflow ellipsis), and both are worse than Compose today. Named here
because they are visible.
## 2026-09-06 (how a markdown block looks, P1a)
- **A table is drawn as padded monospace columns, not as a grid.** Your
call to reverse. Compose draws a real grid: cells on a tint, each
column with a 136dp floor, scrolling sideways when there are too many.
iris has no grid widget, and building one would be a widget per
markdown feature -- which is the thing the block model exists to avoid.
In a monospace face a character count *is* a pixel width, so padding
each cell to its column's width is alignment, the widths are still
measured from the cells, and a table that is too wide pans sideways
through the same mechanism a code fence already uses. The header is
bold with a rule under it, and a long cell wraps inside its column
(capped at 28 characters, which is what fits three columns across a
phone). **What it trades:** no cell borders, and a table looks like
code rather than like a table. If you want the grid, it is a new widget
and it is a day's work.
- **Three block frames, and only three.** A heading, paragraph and list
are plain text with spans; a fence and a table are a rounded panel that
does not wrap; a quote is a bar with the text padded past it.
Everything else markdown says is expressed in span styles, which cost
no widgets and no layout nodes. So a new markdown feature is a span,
not a widget.
- **A list's marker is part of the text, so a wrapped item's second line
returns to the left margin.** Compose keeps it indented by giving the
marker its own column. Doing the same here needs per-line indent in
iris's text attributes; it is written down rather than done, because
the list items in a real reply are usually one line.
- **A link opens on a tap and not on the end of a drag.** A press that
panned the transcript past a link, or that held long enough to start a
selection, does not follow it -- decided by the same gesture machine
that decides pan-versus-select, so there is one rule rather than two
that can disagree.
## 2026-09-06 (composer scroll and the streaming block model)
- **A streamed message becomes a column of per-block widgets.** Decided by
the design agent; recorded here because it is the shape of every message
on screen. A transcript row is one `TextEdit` today, so a streamed delta
re-shapes the entire message through parley on every event -- the stream
phase is the one place iris is behind Compose on your phone (p50 18.2ms
vs 13.4ms). A row becomes a column of one widget per markdown block
(paragraph, heading, fence, list, table) and a delta replaces only the
last block, keeping every earlier block's layout. **Rejected:** splitting
parley's layout at block boundaries inside one text widget (couples
iris's text widget to markdown structure, and parley has no incremental
API), and caching shaped runs per paragraph inside `TextEdit` (a second
cache with its own invalidation beside the glyph cache). Chosen because
P1's markdown block model is needed anyway, so the split happens once, in
`client-core`, and iris stays a text renderer. **Status: designed, not
built** -- this pass spent its budget on the composer's three layout
defects; docs/RUST.md has the design and the pass conditions.
- **The composer's overflowing text now scrolls on a finger**, capped at
six lines and clipped to the bar. Reverses the "still does not scroll"
item below.
- **A widget may not report a `dp` length** (see IRIS.md). A rule for
widget authors, enforced by a `debug_assert!`; nothing changes for app
code.
## 2026-09-06 (stale-primitives and touch-scroll pass)
- **A vertical drag inside a focused composer now scrolls rather than
selects.** Android's own `EditText` does this -- a vertical drag scrolls
the field, and only a long press starts a selection -- so the platform
decided it. What it costs: you can no longer drag straight down inside
the composer to select several lines of what you typed; use a long press
and then drag, or drag sideways. Say if that trade is wrong for you.
- **`Scroll` gets a finger pan but no fling.** `List` flings; a scroll area
does not, because it has no per-frame tick to animate one and the areas
it wraps are at most a screenful (Android does not fling a six-line text
box either). Easy to add later if a scroll area ever wraps something long.
- **The composer still does not scroll its overflowed text**, though the
mechanism it needs is now in place. Wrapping the field in `.scrollable()`
was tried and reverted the same day: `Scroll` measures its content and
container against the *window*, so inside the `MaxSize` that caps the
composer at six lines the two are in different spaces and the field pans
itself entirely out of the bar (measured on the emulator with 474
characters in it -- the bar collapsed to its padding). Fixing that means
`Scroll` measuring against its own offered box, which is a change to a
widget the transcript and the bench shell both use, so it is its own
piece of work rather than a rider on this one.
## 2026-09-06 (defect pass)
- **The keyboard-open diagnostics overlay is gone; the capture only
logs now.** It was added when `on_insets_changed` was not firing at all
and there was no way to get a report off the phone. It fires reliably
since the activity went edge-to-edge -- and what that looks like in
use is a full-screen report covering the app **every time the keyboard
opens**, with its own Copy/Close buttons sitting underneath the
keyboard, so it cannot be dismissed (reproduced on the emulator this
pass: two `tap 'CLOSE'` runs left it up). An interruption for something
nobody asked for, over the app you are trying to type into. The named
`Diagnostics` button still shows the same text on demand, and the new
`iris surface:`/`iris insets:` log lines carry the lifecycle a `logcat`
pull needs. Reversible: `capture_keyboard_diagnostics` is still the one
place this is decided, and `PlatformHandle::show_diagnostics_overlay`
is still there.
- **The bench shell's report pane is sized to its report, not to a share
of the window.** It held `.height(rest(1))` beside the transcript's
`rest(2)`, so an *empty* `TextEdit` reserved a third of every screen --
which is what Iris's "the app does not start with keyboard spacing
correct" screenshot was showing, with the composer two thirds down and
black below it. It is `.max_height(dp(260))` now and sits above the
transcript rather than under the composer, where it was eating the
navigation-bar clearance. Cost: a filled report is clipped at 260dp
rather than scrolling (a `Scroll` there drew itself off the top of the
screen, since `Scroll` pins to the end of its content and reports its
content's full length to the parent -- worth fixing in `Scroll`, not
worked around here). "Copy report" and `logcat` still have the whole
thing.
## 2026-09-05
- **iris no longer asks every device for compute-shader limits it never
uses.** `adapter.request_device` (both `iris/src/android/render.rs` and
`iris/src/default/render.rs`) used `Limits::default()` plus an override
for `max_buffer_size`, and `Limits::default()` unconditionally requests
desktop-tier compute limits (`max_compute_workgroups_per_dimension:
65535`, per `wgpu_types`) even though nothing in `iris`/`iris-core`
creates a `ComputePipeline` or writes a `@compute` shader stage —
confirmed by grepping the whole tree, not assumed. That crashed
`request_device` outright on the Android emulator's software GL path
(`EMU_GPU=software`, `--features force-gles`): SwiftShader's GL reports
itself as OpenGL ES 3.0, which has no compute shaders at all, so the
adapter's real limit is 0 against the unconditional request for 65535 —
`RUST.md`'s "Software mode ... crashes for a third, different reason,"
2026-09-05, earlier today. The same would happen on any real
GLES-3.0-only Android device, not just the emulator. Fixed by a new
`iris_core::device_limits()` (`iris/core/src/render/mod.rs`), shared by
both platform backends so the two requests cannot drift, that zeros the
six `max_compute_*` fields explicitly rather than switching to a
downlevel `Limits` preset — `Limits::downlevel_webgl2_defaults()` was
considered and rejected: it also zeros
`max_storage_buffers_per_shader_stage`, and `shader.wgsl`'s vertex stage
reads four `var<storage>` buffers (rects, glyphs, masks, move_offsets),
so that preset would trade the compute crash for a bind-group-layout
one on the same downlevel hardware this is meant to support. No
capability check or fallback path was needed since nothing is being
disabled — the request is simply narrowed to what the pipeline actually
uses. `rigs/gpu-probe`'s own mirrored limits (it is deliberately its own
crate, not a workspace member, so it cannot call `device_limits()`
directly) were updated to match, and confirm `IRIS DEVICE: ok` against
this VM's own Vulkan and GL adapters. **Not verified this pass**: the
specific SwiftShader-ES-3.0 crash this fixes, on-device — the
`EMU_GPU=software` cold boot this needs would have force-restarted this
checkout's emulator while another session was actively running its own
app on it (`com.example.aiapp` had window focus at the time), so it was
left for a pass when the emulator is free rather than disrupting that
session. Everything reachable without the emulator is clean: `cargo
fmt`/`clippy --workspace --all-targets`/`test --workspace`, `cargo ndk
build`/`clippy` for `iris-android-app` with `force-gles`, and
`gpu-probe` against this VM's own Vulkan and GL(ES 3.2, which still has
compute and so would not have reproduced the crash even before this
fix — not a substitute for the real ES-3.0 test).
- **P0's Compose half is built and smoke-tested on the emulator** — the
`bench` build type, the shared `app/bench-fixture/` transcript, and an
in-process fake backend (`BenchFixture.kt`/`BenchNetwork.kt`) that
answers `TranscriptSource`/`EventStream` from an in-memory event log
instead of a real server, so the fold and paging under test are the real
ones. Full account, the smoke run's report, and what is deliberately
left (the iris half, the real on-phone runs) are in RUST.md's P0 box.
Not a decision to review so much as the gate itself now being runnable —
flagged here because it is the first half of something Iris explicitly
asked to see before P1.
- **P0's iris half is also built and smoke-tested on the emulator,
2026-09-05.** A new `bench` Cargo feature on `iris-android-app`, on top
of `transcript-screen`: the same checked-in fixture (`include_str!`, no
asset pipeline needed), the same 24-swipe scroll loop animated through
`List::scroll` and the same 400-event/20s streaming phase through
`fold_event`, "Run benchmark"/"Copy report" as named accessible
controls, and the same three added report fields (process CPU time,
peak RSS, battery current) via direct JNI calls
(`bench_jni.rs::PlatformHandle`) since `android_view` has no
`BatteryManager`/`ClipboardManager` wrapper of its own. One small public
API addition to get there: `AndroidAppState::platform_ready` (`IRIS.md`),
a default-no-op lifecycle hook handing an implementor a `JavaVM` +
`GlobalRef` it can call Java through from any thread. Packaged with a
new `release` build type on `iris-android-app`'s own Gradle project
(there was previously only `debug`), signed with the same key
`app/build-apk.sh` generates. Smoke run and the full report are in
RUST.md's P0 box; not attempted this pass: the real on-phone runs and
Iris's pass/fail call, which is the actual gate.
- **The intermittent touch-scroll dropout is root-caused and fixed: a
missed `ACTION_DOWN` hit-test, not the previously-suspected coalesced
first `ACTION_MOVE`.** Diagnosed by temporary logcat tracing of every
touch event, `DragArbiter` state transition and `Selection::drag`
dispatch (removed once confirmed), reproduced on this checkout's own
emulator against a real sandbox session. The trace showed the actual
mechanism: a gesture's `ACTION_DOWN` lands wherever the finger actually
is, which is not guaranteed to fall inside the same row-local sensor
region a later `ACTION_MOVE` in the same gesture lands in (a row's own
padding/gap, or its non-selectable sender-name header, is
pointer-transparent to `iris::sense::CursorSense`). When that happens,
the widget that ends up handling the gesture never saw `PressStart`, so
`DragArbiter` sits in `Idle` — which answers every subsequent frame with
`Undecided` and has no way to tell "no press is happening" from "a press
is happening but I missed its start," so it never recovers on its own
for the rest of that gesture. One real trace showed exactly this: touch
`Down`/`Move`/`Up` all delivered correctly, but zero `PressStart`
reaching the arbiter, `state=Idle` unchanged from first frame to last.
Fixed at the call site that has the context to recover
(`iris::transcript_ui::selection::Selection::drag`,
`iris/transcript-ui/src/selection.rs`): a new `DragArbiter::is_idle()`
(`iris/src/sense.rs`) lets it notice a `Pressing` frame arriving with the
arbiter still `Idle` — which can only mean a missed `PressStart`, since a
`Pressing` sense requires the button to genuinely be down — and start the
press there instead of where it was missed. Three new unit tests in
`sense.rs`'s `drag_arbiter_tests` and one in `transcript-ui`'s
`selection::tests` (the latter fails on the code before this fix).
Commit follows. Not the same failure the earlier pass's `DECISIONS.md`
DEFERRED item speculated about (a coalesced first `ACTION_MOVE` skipping
slop detection) — that hypothesis is now ruled out; the arbiter's own
slop/long-press logic was never wrong. RUST.md's I5 box,
"Touch-scroll dropout root-caused, 2026-09-05" has the full trace.
- **P0, a phone benchmark gate before any porting, asked for by Iris
2026-09-05**: "before P1 I'd like to see benchmarks & also maybe stress
test on my own phone ... If it doesn't match compose reasonably well then
I don't think I'd wanna continue." Design (RUST.md's P0 box has the
detail): the same embedded synthetic fixture in both apps with no server
needed; the same scripted scroll loop then a streaming phase, run
programmatically since the phone has no usable system tracing and no
agent can drive it; the same report from both (frames, janky %, p50/p90/
p99, process CPU time, peak RSS, battery current where readable) with a
copy button; the iris app under its own id and the Compose one as a new
`bench` build type with an id suffix, so neither replaces her production
install; two arm64 APKs plus instructions delivered under `~/host/bench/`.
The gate is hers: iris within a reasonable margin of Compose release on
p50, p99 and CPU time, no crashes, no visible stutter. If it fails, the
port stops.
- **The rest of the port is one UI crate, `iris/app-ui`, grown out of
`iris/transcript-ui` rather than started beside it.** It holds a
`Screen` enum plus a back stack — the Rust equivalent of `AppRoot.kt`'s
`when` — and `iris/desktop-app`/`iris/android-app` become thin entry
points over it. Chosen over a fresh crate because `transcript-ui`
already has the right generic shape (`Rsc: HasEvents` +
`Rsc::State: FocusHost`) and the `client-core`/`event-model` path
dependencies every later screen needs, so growing it in place is the
smaller diff. Platform-only code (notification service, share target,
QR scanner, Keystore token, deep-link enrolment) stays in the E3/E5
Java shell (`android-shell/` + `app/shellApp`) rather than moving into
this crate, since none of it is a screen. The Android APK is built by
`cargo xtask apk` (E5), merging the app-ui cdylib into the E3 shell so
there is one app rather than a demo shell plus a service shell.
`app/androidApp` (the Compose app) stays untouched and is the baseline
every step is measured against, until parity is reached (P7 decides
the switch, and is itself a load-bearing decision left to Iris). Order
is by risk to the daily-use path: session screen first (P1, where
every hard behaviour already lives), then the shell merge and a real
phone install (P2), then root tabs (P3), the explorer (P4),
settings/enrolment (P5), desktop parity (P6), and the cutover itself
(P7). Full plan: RUST.md's "The port, in order (decided 2026-09-05)".
- **iris gets its own measured frame report, rather than waiting on a
`dumpsys`/`gfxinfo` answer that cannot see a `SurfaceView`'s GPU-drawn
frames.** `iris_core::FrameReport` (`iris/core/src/render/frame_report.rs`)
times each frame's wall clock from the same point `render()`'s redraw
starts to just after `queue.submit` + `present()` — the span Compose's
own render report and `gfxinfo` both count — into a fixed 4096-entry
ring (no allocation per frame; `report()` is the only place that
allocates, and only on a button tap). The report gives total frames,
janky % over the same 16.7ms budget `gfxinfo` uses, P50/P90/P99 and the
worst, plus a reset. Exposed the way the Compose app's copy-button
report already is: two named controls ("Frame report", "Reset frame
report") on the transcript screen, tappable by accessibility name via
`ui-trace`, logging under this crate's fixed `android_logger` tag
(`iris-android-app`) so a script can grep `"iris frame report"` the way
`transcript-bench.sh` greps `"ai-app render report"`. The report's own
`Display` line says plainly that it measures up to the `present()` call
returning, not GPU/compositor completion — wgpu's `present()` is not
fenced against either, so presenting that span as "time to reach the
screen" would be a measured-looking number that is actually inferred,
which the standing UI rule forbids.
- **`ui-trace` gains a hold-then-drag gesture, additive, in
`emulator-tools`.** Neither of its two existing actions can produce
"hold stationary for `LONG_PRESS`, then move without lifting" — `tap`
has no hold and `swipe X1 Y1 X2 Y2 MS` interpolates motion across its
whole duration from t=0. A new action presses, waits, then moves to a
second point and releases as one continuous touch (raw
`sendevent`/`MotionEvent` injection, extending whatever mechanism the
existing `swipe` already uses), so `DragArbiter`'s pan-vs-select rule
(`iris/src/sense.rs`, already covered by 8 unit tests against a
synthetic clock) can finally be driven on a real device instead of only
in a test harness.
- **Touch drag on a transcript row follows Android's own rule**: a vertical
drag pans the list immediately; a stationary press held 500 ms starts a
text selection which further dragging extends; a horizontal drag while
something is already selected extends that selection without the wait.
One `DragArbiter` per list decides it (`iris/src/sense.rs`). Chosen over a
"text layer always wins" or "list always wins" rule because either loses
one of the two gestures a reader expects.
- **E4's desktop shape is a new `iris/desktop-app` crate**: a winit window
holding `transcript-ui`'s screen beside a session list, talking to a real
`ai-server` through `client-core`. It enrols by pasting the same
`aiapp://enroll?…` link a phone scans (`client-core::config::EnrolledServer`)
and keeps it owner-only under `$XDG_CONFIG_HOME/ai-app-desktop/`. The
pinned CA is a path given on the command line, not baked in. Chosen so
the phone and desktop share one enrolment format and no second one is
invented.
- **I5's Android integration extends `iris-android-app` (I2's shell)
behind a Cargo feature (`transcript-screen`), rather than a third
shell crate.** That project already has the Gradle module, the
`IrisView`/`MainActivity` Java, and the JNI registration; the only
thing a second screen needs on top is a different `AndroidAppState`,
the same axis `tabs_ui::build`/`transcript_ui::build` already vary
along on the winit side. `tabs-screen`/`transcript-screen` are
mutually exclusive and each pulls in only its own deps, so the plain
tabs build (I2/I4) is untouched.
- **Order of remaining work, updated 2026-09-05**: the two in-flight
pieces and I5's Android integration are all done; next is giving iris
its own frame-timing report so item 3 below can be decided by a number.
- **DECIDED by Iris, 2026-09-05: iris is the app's framework; Masonry was
the calibration.** Her words: "I think iris definitely makes more sense
based on the limitations we've found." The limitations: Masonry has no
touch scroll on Android (E2), no per-span rich text and no cross-row
selection on the pinned commit (E2), and its keyboard bridge is a TODO
(E1); iris carries the same screen under the Compose baseline on the
host GPU (p50 15.0 ms against Compose's 20.0 ms, RUST.md's I5 box). What
follows: the E-steps are closed as calibration, and the port proceeds
on iris — screens, the shell (E3/E5), and `client-core` underneath.
The item below is kept as the record of what she decided from.
- **Was DEFERRED — whether to commit to iris over Masonry for `ai-app`.**
Updated 2026-09-05 with the clean comparison the recommendation wanted:
same sandbox session content, same emulator, `EMU_GPU=software`, one
session. Headline numbers (RUST.md's I5 box, "Clean scroll comparison,
2026-09-05," has the full table and every caveat):
| app | build | frames | janky % | p50 | p90 | p99 | worst |
|---|---|---|---|---|---|---|---|
| Compose (in-app report) | debug | 1102 | 99.0% late | 33.8ms | 50.6ms | 79.5ms | -- |
| Compose (`dumpsys gfxinfo`) | debug | 1499 | 21.15% (95.66% legacy) | 32ms | 48ms | 150ms (p99) | -- |
| iris (`FrameReport`) | **release** | 299 | 94.65% | 79.1ms | 98.6ms | 117.8ms | 212.6ms |
| iris (`FrameReport`, repeat) | **release** | 233 | 94.42% | 109.3ms | 130.8ms | 147.1ms | 150.5ms |
**Not a clean apples-to-apples reading, stated plainly rather than
smoothed over**: iris had to be built **release** (debug `SIGSEGV`s on
this emulator's Vulkan loader, I4's finding) against Compose's mandated
**debug** build, so this asymmetry likely *understates* iris's gap
rather than the reverse; the three frame-time sources measure different
things (Compose's own phase accounting vs. Android's HWUI deadline-miss
definition vs. iris's redraw-start-to-present window, the last of which
`dumpsys gfxinfo` cannot see at all for iris's `SurfaceView`); and both
figures are emulator numbers under software rasterisation, which
Compose's *own* in-app report shows already costs 20-34ms/frame in
`swap`+`gpu` alone under this GPU mode, so a same-mode iris number well
above 16.7ms was expected going in for either app. A second pair under
`-gpu host` was not taken this pass. The earlier session's suspected
intermittent touch-delivery dropout was **not reproduced** this pass —
the zero-frame results this time traced to this pass's own script bug
(a `cd` that changed which emulator `ui-trace` targeted), not the
emulator; a CPU-load rise during the gesture was observed by a sampler
running throughout, but did not correlate with any failure, so the
original candidate is neither confirmed nor ruled out.
The choice in front of Iris, updated: decide now on the
structural-plus-functional case already made (iris works end-to-end
where Masonry's scroll gesture doesn't exist at all on Android) plus
this table — reading the two build profiles and three jank definitions
with the caveats above rather than as a single number — or ask for a
same-profile, same-GPU-mode rerun first. RUST.md's I5 box has the full
account.
**Updated 2026-09-05, the `-gpu host` pair taken.** Real GPU rendering
(`force-gles` -- the default Vulkan backend has no adapter at all under
plain host-GPU boot, confirmed by the exact `wgpu` error) reverses the
software-mode shape:
| app | build | GPU mode | frames | janky % | p50 | p90 | p99 | worst | cpu p50 | gpu-wait p50 |
|---|---|---|---|---|---|---|---|---|---|---|
| Compose (in-app report) | debug | host (virgl) | 1268 | 96.4% late | 20.0ms | 28.4ms | 37.7ms | -- | -- | -- |
| iris (`FrameReport`), **best of three, 2026-09-05** | release, `force-gles` | host (virgl) | 439 | 46.24% | 15.7ms | 23.3ms | 31.2ms | 57.4ms | 1.2ms | 13.2ms |
Under real GPU rendering iris's median frame is *faster* than
Compose's, not the 2-3x-slower shape the software-mode table shows. A
new split inside `FrameReport` (redraw-to-submit vs. submit-to-present,
commit `e2a1fad`) says why: iris's own CPU work per frame is a median
~1ms -- almost the entire frame is time spent handing the frame to the
driver, not in iris's layout/text/primitive code. This is consistent
with the earlier software-mode gap being mostly SwiftShader's CPU
rasterisation cost rather than an iris-specific slowness. **Still not
proof, and now closed as unanswerable rather than merely untaken**: a
same-mode software `force-gles` run to isolate the backend was retried
2026-09-05 after fixing the compute-limit crash the first attempt hit,
and hit a second, structural wall instead — SwiftShader's ES 3.0 GL
path has no storage-buffer capacity at all, and `shader.wgsl` reads
`var<storage>` buffers unconditionally, so reaching that path needs a
shader rewrite, not a limits fix (RUST.md's I5 box, "The three
remaining I5 verifications, closed 2026-09-05," item 2). The
intermittent touch-scroll dropout this pass also reproduced is
root-caused and fixed as of the same date (a missed `ACTION_DOWN` on a
row's padding/header left `DragArbiter` stuck in `Idle`); three clean
`iris-scroll.sh` runs post-fix each scrolled all 24/24 swipes, replacing
the single-attempt 62-frame reading this table used to carry. RUST.md's
I5 box, "Where iris's frame time goes, 2026-09-05, the `-gpu host`
pass," and "The three remaining I5 verifications, closed 2026-09-05,"
have the full account. The iris-vs-Masonry choice itself is still
Iris's to make.
## 2026-09-07: the enrolment link carries the CA, so an APK need not be built where its server runs
**Problem.** Every phone build pinned the CA of the machine that compiled
it -- the Compose app from `GeneratePinnedCert`, the iris app from
`build.rs` reading `$XDG_CONFIG_HOME/ai-app/certs/ca.pem`. That is fine
while the two are the same machine and impossible when they are not, which
is exactly the iris client's situation: cross-compiled in this VM,
delivered to a phone, run against `ai-server` on the host. Baking the
host/port/token as well made it worse -- a token in a built artifact.
**Decided: the CA rides in the enrolment link**, as `&ca=<base64url of the
DER>` (`wg_app_link::enroll::ca_param`), optional and per mint. The app
that opens the link pins what the link said, and an APK built anywhere
works against whatever server it is pointed at.
Two alternatives were worked out and rejected.
- **A CA *fingerprint* in the link, pinned at the TLS handshake.** The
smallest link (43 more characters) and the strongest shape, but `ureq`
3.4 exposes no hook for a custom `rustls` `ServerCertVerifier`: its
`TlsConfig` builds the `ClientConfig` itself, so this needs a hand-written
`Connector` on the `unversioned` API and `rustls` as a direct dependency
of `client-core`. A lot of machinery in the one crate that must stay
light.
- **A fingerprint in the link plus an unauthenticated `GET /ca.pem`.**
Small code, but it needs a first connection with verification disabled,
and it breaks a documented, tested posture -- `auth.rs`'s "gates every
route with zero unauthenticated endpoints", which is a load-bearing
decision rather than an implementation detail. Not something to change
silently for this.
**What it costs**, measured rather than guessed: on this project's P-256
CA the link goes from 89 bytes to 652, and `print_enrollment`'s terminal
QR from 45x23 to 93x47 characters. That is why the parameter is the
minter's choice per call: `ai-server` passes it (its iris client needs it),
`dev-updater` passes `None` (its app is built on the machine it talks to,
and its QR stays scannable in an 80-column terminal). The URI printed under
the QR is the fallback either way, and is the path Dev Updater's Enroll
button already uses -- it opens the link with `ACTION_VIEW`, so Android
offers whichever apps registered the scheme, which needed no change here.
The CA is a public certificate, so putting it in the QR leaks nothing the
token did not already: photographing the terminal still costs exactly the
token, which is rotatable.
**The log upload's destination is moot**, so it is not wired to this. On
the same day Iris decided Dev Updater will read an APK's runtime log from
an on-device ContentProvider instead, which removes `log_upload`,
`POST /client-log` and the `AI_APP_LOG_*` baking altogether -- so the
enrolment landed without touching any of them, for that change to delete
whole.
View File
File renamed without changes.
+1720
View File
File diff suppressed because it is too large. Load diff
+1280
View File
File diff suppressed because it is too large. Load diff
+219
View File
@@ -861,6 +861,58 @@ unspecified rather than getting them wrong:
conditions, so the remaining slack was accepted rather than chased
further.
## Density: `Len::dp`, resolved at `apply_rest` time (2026-09-06)
Iris asked for a third length kind beside `abs` (physical pixels) and
`rel`/`rest` (a fraction of the parent) — IRIS_TODO.md's "density-
independent length unit" — after the P0 phone pass found 16px text
drawing at roughly a third size on a real phone. The fix that shipped
first (RUST.md's P0 box) was a global stopgap: divide the whole window
into a "logical" coordinate space (physical ÷ `content_scale`) and let
the shader's NDC mapping stretch it back up onto the real framebuffer.
That fixed the *size* but not the *sharpness* — a glyph rasterised at the
small, pre-stretch size and then stretched onto more physical pixels than
it has texels for is blurry, which is exactly what Iris's next report
said.
**The fix**: `Len` gained a `dp` field, resolved against a `density: f32`
(physical pixels per dp) at the one place a `Len` becomes a `UiScalar`
(`Len::apply_rest`) — `abs + dp * density`. `density` lives on
`UiRenderState` (`set_density`/`density()`) and `Painter` (`density()`),
set once from `DisplayMetrics.density` in `android::view::new_peer`; the
desktop backend has no per-monitor density wired up yet and stays at
`1.0`. Every layout call site that used to call `.apply_rest()`/
`.to_uivec2()` now passes `painter.density()` (nine call sites — `Span`,
`Sized`, `MaxSize`, `Aligned`, `Scroll`, `List::place`, and
`UiRenderState::reposition` itself). This also meant the Android
boundary's global logical-space stopgap could come out entirely: window
size, touch coordinates and insets are physical pixels again, matching
`AndroidRenderer`'s own swapchain resolution, with `dp` doing the
per-length work the global divide used to do for everything at once.
**Text is the case that needed more than the `Len` plumbing.** A widget's
`font_size`/`line_height` are plain `f32`, not routed through `Len` at
all (there is no sensible `rel`/`rest` for a font size). `TextBuffer::
shape` now takes `density` directly and multiplies `font_size`/
`line_height` (and any span override) by it before handing them to
parley — so the size that reaches both the line-breaker and the
rasteriser (`TextData::place`, which reads back whatever `shape` set) is
the display's *physical* size, and the glyph atlas holds a bitmap at the
resolution it is actually shown at. `GlyphKey.size` already keys on the
resolved size, so a cache entry is naturally per-physical-size with no
further change. The one caller with no `Painter` to read density from
(`TextEditCtx::layout`, cursor movement and hit-testing) reads a second
copy kept directly on `TextData` (`TextData::density`) instead — an
accepted duplication rather than threading a `Painter` into every input
handler for one field, the same tradeoff `AndroidRenderer::content_scale`
already makes for the Diagnostics page.
**What did not change**: `rel`/`rest` are unaffected (already
resolution-independent, a fraction of the parent). `Span::gap` and
`Padding`'s four sides moved from bare `f32` to `Len` so `dp(...)` works
on them the same as any other size; a bare number is still `abs`,
physical pixels, unchanged.
## For IRIS.md
When this lands, copy this entry into `IRIS.md` (newest first):
@@ -895,3 +947,170 @@ When this lands, copy this entry into `IRIS.md` (newest first):
> `SizeCtx` and `Cache` are gone with it — see `LAYOUT.md` for the full
> design, the move-offset mechanism this shipped alongside, and the file
> list.
## Masks with a shape (decided 2026-09-07, built 2026-09-08)
Iris, on the code block's scrolling: "the code block scrolling currently
masks in an inner rectangle. Ideally masks should have a shape
associated with them, rounded rectangle being one of them, and/or
another widget you can select, so that the mask becomes the parent
container with rounded edges. Make sure alpha works properly with it,
eg. on the corners where alpha should be decreased / multiplied."
**What exists.** `Mask` in `shader.wgsl`/`data.rs` is two `UiSpan`s and
a `move_idx`; `fs_main` resolves it and does `color *= 0.0` outside the
rectangle -- a hard cut on a pixel boundary. `Masked` (`widget/mask.rs`)
sets the painter's mask to its own region. Separately, `draw_rounded_rect`
already produces an anti-aliased rounded edge from
`distance_from_rect(pos, center, corner, radius)` with a half-pixel
`smoothstep`, and the border variant multiplies a second coverage in.
**Design** (revised the same day on Iris's two corrections: hit-testing
applies the shape too, and a mask should reference a primitive rather
than carry a copy of its shape).
1. **A mask is a reference to a primitive already drawn, plus how to
use it.** `Mask { kind, idx, flags, parent }`: the primitive's
binding (`RECT`, `TEXTURE`, `GLYPH`) and slot, flags (today one:
*alpha only* -- take the primitive's coverage and ignore its colour,
which is the default and the only mode until a need for another
appears), and the enclosing mask's slot for nesting. The fragment
stage evaluates the referenced primitive *at the masked pixel* --
for a `Rect`, the same `draw_rounded_rect` coverage from the same
SDF; for a texture or glyph, the sampled alpha -- and does
`color.a *= coverage`. Nothing about the shape is copied: a rounded
container's corner and its children's clipped corner are the same
primitive's arithmetic, and a texture mask (an alpha image as the
clip) works with no new shader path.
What this needs from the data layout: evaluating a primitive at an
arbitrary pixel means its placement (its spans and `move_idx`, today
vertex attributes) has to be readable from a storage buffer in the
fragment stage. If it is not already there, put it there once, for
every primitive, rather than keeping a second copy for masks -- the
vertex stage can read the same buffer. Textures: the shader binds one
image at a time (see `masks_layout`'s comment on why an image's own
bind group must not name the masks buffer), so a texture mask is
limited to what the fragment can sample without a bind-group switch:
the atlas, and the primitive's own bound image when the masked
primitive is drawn in the same image's batch. Say so at the flag.
2. **Nested masks chain and multiply, like moves.** `parent` walks up
the chain, bounded like `resolve_move` (`MOVE_CHAIN_LIMIT`'s sibling;
debug-assert on overflow and print the chain); coverages multiply,
so a pixel inside two feathered corners is dimmed by both, which is
what a compositor does and what "alpha should be multiplied" asks.
3. **`.masked()` points the mask at the current widget's own
primitives.** `Masked` stops describing a region: it records which
primitive(s) the wrapping widget drew this frame (the painter knows
-- it just allocated the slots) and sets the mask to reference them.
So a rounded `Rect` widget's `.masked()` clips its children to
itself by pointing at the rect it already draws; an image widget's
`.masked()` clips to its alpha. No radius or shape argument exists to
fall out of sync. When a widget draws more than one primitive (a
bordered rect is one primitive; a card with a stripe is two), the
mask references the *first* and the doc says so; a widget that wants
another names it.
4. **Hit-testing applies the shape.** A press is inside a masked
subtree only if the mask's coverage at that point is above one half.
For a `Rect` that is the same rounded-rect SDF evaluated on the CPU
-- one function in the shared crate, with the WGSL a transliteration
of it and a test that compares the two at a grid of points
(`headless` renders to a buffer and reads back, or the Rust version
is checked against the values the shader produced once and recorded).
For a texture, the CPU needs the alpha: keep the alpha channel of an
image used as a mask readable on the CPU (it was uploaded from CPU
memory; keeping the alpha plane is a quarter of the image), and read
it at the point. A masked corner that cannot be tapped and a masked
corner that is not drawn are then the same corner.
**Rejected.** A stencil buffer (a second pass per mask level and no
anti-aliasing); the scissor rectangle (rectangles only, no alpha);
rendering a masked subtree to an offscreen texture and compositing
(a texture allocation per mask, every frame it scrolls, on the phone).
**Pass conditions.** A headless test draws a rounded container with a
masked child that overhangs all four sides and asserts the child's
coverage at a corner pixel equals the container's own coverage there
(same primitive evaluated, so exactly equal, not approximately); a
nested-mask test asserts the product at a pixel inside both feathers; a
texture-mask test clips a rect to an alpha image and asserts a
transparent texel masks fully; a hit-test asserts a press in a
container's clipped corner misses and one just inside the curve hits,
and that the CPU SDF and the shader agree at a grid of points; a
`run-headless.sh --phone` screenshot of a scrolled code block shows
rounded corners with no square pixels poking out at the top and bottom
of the scrolled content. Record the commands in RUST.md when it lands.
### What was built (2026-09-08), and where it differs
The commands and the screenshot are in docs/RUST.md's queue entry. Four
places the code is narrower than the design above, each deliberate:
- **No `kind` and no `flags` on `Mask`.** It is `{ primitive, parent }`.
The referenced instance already carries its own `binding`, so a copy
of it in the mask is a second thing to keep in step; *alpha only* is
the only mode there is, so there is nothing to select. Both are a
field away if a second mode appears.
- **A mask's shape must be a rect.** `Painter::set_mask_to` asserts it,
by name, rather than leaving the shader to read a `rects` entry that
is not there. A glyph would need a CPU-side alpha plane before the
hit test could agree with the shader, and a standalone image needs a
bind-group switch the fragment stage cannot make (`masks_layout`'s own
comment on why an image's bind group must not name the masks buffer).
So **the texture-mask pass condition is not met and no texture mask
exists** — the point of the reference design is that adding one is a
binding check and a sampled alpha, with no new shader path, and the
shader's `mask_coverage` already has the branch where it would go.
- **The shape is a primitive of its own, not always a drawn one.** A
plain `.masked()` writes an undrawn `RectPrimitive` at its region
(`Drawn::No`, `NOT_DRAWN`) and points the mask at that, so "clip to my
box" and "clip to that widget's rounded background" are one mechanism
and square-cornered clipping did not become a special case.
`.masked_by(shape)` draws `shape` behind the content — in its own
layer, the way `Stack` puts a background under its content — and
clips to the first primitive it drew.
- **The CPU/shader agreement is a GPU test**, `iris/tests/mask_sdf.rs`,
the only test in the workspace that needs an adapter. It lifts
`distance_from_rect` and `rounded_rect_coverage` out of
`iris_core::SHAPE_SHADER` *by name* and runs them in a compute pass,
so the thing under test is the shader itself rather than a copy of it
that would be edited alongside.
## What a widget's *offered* box may and may not be (2026-09-08)
Two rules that were each true in one place and missing from a sibling,
found together by Iris's 2026-09-08 phone report.
**Padding works in whatever container it is placed in, and is an inset or
an outset depending on how tight that container's region is.** Iris's
own words, 2026-09-08: "padding should work no matter what container a
widget is placed in, and acts as both inset and outset depending on how
tight the parent region is." `Pad` offers its child the region it was
handed, inset on each side, and reports `used + padding` — so given a
generous box it insets the child inside it, and given a box already the
size of the content it reports a larger size and the parent grows. What
this rules out is any container that offers a padded child a box and then
ignores what it reported, and any caller that reshapes its tree to avoid
a `Pad` (which `transcript-ui/src/tool.rs` did until 2026-09-08, at the
cost of a tool group's 4dp inset).
**A widget offered a box it does not fit is drawn again at the box its
own reported size implies, in the same frame.** Not next frame. The
temptation to defer is real — `List::place` offers a row its *cached*
height precisely so that an unchanged row hits `draw_inner`'s cheap
skip-or-move path, and `Scroll` sizes its child region from last frame's
content length for the same reason. But a `Rect` fills whatever region it
is given (`Size::REST`, and `rect.rs`'s `is_size_independent` doc says
why it must), and `.background(rect(..))` is the ordinary way to style
anything — so a one-frame-stale box is a background drawn at the wrong
size while the text inside it is already right. On screen that is a tool
card that looks closed while its text is there and open while it is not.
A `reposition` is not the fix and cannot be: it writes an offset, never a
size.
The cost is bounded and worth stating, because it is what makes the rule
safe to apply everywhere: the second draw happens only on the frame a
widget's own size actually changes, which is a frame that was already
redrawing it. A widget whose reported size is a function of the box it
was *offered* would disagree every frame and redraw every frame — which
is why `List` requires content-sized rows, and has since long before
this.
+147 -10
View File
@@ -147,12 +147,37 @@ turn.
Spawn: `claude -p --verbose --input-format stream-json --output-format
stream-json --permission-mode <mode>` in the chosen working directory, plus
`--model`. Wire-format notes are pinned against CLI 2.1.237 in
`session/claude.rs`'s module doc: permissions need the hidden
`--permission-prompt-tool stdio` flag, AskUserQuestion answers ride
`updatedInput.answers` keyed by question text, and `set_model`/`interrupt`
`--model` and, where one has been chosen, `--effort`. Wire-format notes are
pinned against CLI 2.1.237 in `session/claude.rs`'s module doc: permissions
need the hidden `--permission-prompt-tool stdio` flag, AskUserQuestion answers
ride `updatedInput.answers` keyed by question text, and `set_model`/`interrupt`
are control requests.
**The thinking level is settled at launch** (added 2026-09-04, because it is
the largest saving available on a long session: output is about an eighth of
what a session costs and thinking is the bulk of output, against the ~1.5% that
is prose). The CLI's only two setting control requests are `set_model` and
`set_permission_mode` -- checked against the 2.1.258 binary -- so there is no
way to ask a running process to think differently. `set_session_effort` is
therefore shaped like `set_session_cwd` rather than like `set_session_model`:
it records the level and **stops the process**, and the next message or Start
launches one that has it. It lives in the session settings dialog beside the
working directory for that reason, not on the session bar beside the model and
the mode, which do take effect mid-turn. `None` is a level in its own right --
the CLI's own default -- so the picker can return to it; a level this app named
as the default instead would be this app choosing one.
**What a new session starts at is `Config::default_effort`**, applied in
`spawn_session` rather than filled in by the spawn screen, so it holds for an
import and a bare API call as well. It is set by the spawn screen's own
picker, whose label says so: one control, where new sessions are made, rather
than a settings page for a single value. It is not on a provider, because
providers are discovered and the next rediscovery would erase it, and not on
the phone, because a second device would then spawn at a level nobody there
chose. `GET`/`POST /defaults` carry it, as a struct rather than a bare value
so the permission mode -- still hardcoded to `auto` on the spawn screen -- can
move there without a second route.
**`--resume` only ever runs when nothing else has that session open.** That
is the rule behind the import refusal, the single `ClaudeDriver::launch`
entry point, and the `Exited` correction below; two CLIs on one session file
@@ -169,10 +194,33 @@ deliberate and easy to undo by accident:
when the process restarts. That leaves the Claude driver as the odd one
out rather than this one — the CLI's memory is a cache in front of the same
transcript. Resolve any inconsistency in this direction.
- **A llama session on an ssh host is refused.** The model is reached over
HTTP and forwarding that port is not built, so refusing beats silently
talking to the wrong machine. A transport is "run this" plus "reach this
port", and only the first half exists.
- **A llama session runs on whatever machine its setup names** (2026-09-04,
the last of phase 5). A transport is "run this" plus "reach this port", and
the second half is `Transport::reserve_port` — the port the server binds
*there* and the port that reaches it *here*, the same number locally —
carried by `Launch::reaching` onto the connection that already runs the
command. `llama-server` binds loopback on the far machine, so nothing is
served to its network. The far port is a guess from a range below the
ephemeral one, because no portable way to ask a machine for a free port
avoids racing the bind anyway; a collision is not silent, since the server
fails to bind and the readiness poll reports what its log said.
- **The model file lives on the machine that serves it** (2026-09-04). Each
setup names its own models directory (`SshConfig::models_dir`, default
`~/.local/share/ai-app/models` expanded *there*), and a spawn resolves the
key on that machine — one round trip answering "at /abs/path" or "missing",
so a model that is not there is refused at the spawn rather than becoming a
server that never becomes ready. The spawn screen offers
`GET /setups/{id}/models`, that machine's list, rather than `GET /models`,
which is this backend's downloads. Downloading *to* another machine is
deliberately not built: a multi-gigabyte transfer with no progress
anywhere, and the file gets there however anything else on that machine
did.
- **The readiness poll watches the process, not only the port.** A model that
will not load, a port already taken, a flag an older build does not know:
all exit within a second and none will ever answer `/health`, so waiting
out the 300s timeout turned the server's own account of the problem into
"gave up". The failure carries the tail of `llama-server.log`, which on a
remote session is the only copy anybody reading the phone can see.
### Models (2026-08-28)
@@ -203,6 +251,13 @@ deliberate and easy to undo by accident:
A driver says what to run; something above it turns that into a process.
Otherwise transport knowledge sits inside a translator whose job is a wire
format, and every future driver has to remember to do the same.
- **A forwarded launch gets a pty and every other one does not** (measured
2026-09-04). Killing the ssh client ends a CLI because it closes the stdin
that CLI is reading; `llama-server` never reads its stdin, so the same kill
left it running on the far machine with the model loaded — one orphan per
stopped session. With `-tt` the far side takes SIGHUP when the connection
goes. Its log then arrives through a line discipline, which nothing parses.
`-T` stays everywhere else, where a pty would rewrite the JSONL.
- **`command -v` follows ssh's non-login PATH**, which is narrower than an
interactive shell's, so a binary somewhere unusual is invisible to
discovery. Point `command` at an absolute path.
@@ -499,6 +554,25 @@ rate-limited bucket). Poll at ≥180 s, only while a Claude session exists or
the usage screen is open, and cache the last answer. It is undocumented, so
`usage.rs` treats every field as optional and degrades rather than erroring.
**Per provider, not per machine (2026-09-04).** A machine is not what is
metered; the provider a session runs is. One machine offers echo, the Claude
CLI and a local model side by side, and only the second spends anything — so
pairing a session with a snapshot by machine alone drew the CLI's five-hour
window under every echo session on it, a quota that session cannot spend. A
session now names its meter (`usageProvider`, from
`DriverKind::usage_provider`, which `usage::providers_for` reads too, so the
two lists cannot disagree) and `GET /usage` is matched on machine *and*
provider. `None` is a session that meters nothing, and the phone draws
nothing at all for it — not a zero, and not "unknown".
`DriverKind::Echo` names a meter of its own that exists only when a test has
asked for one: `/usage` in an echo session sets an invented answer
(`usage::Fixture`), and with none set there is no snapshot and no bar. That
is what makes those screens' states reachable — a number near the top, a
window between blocks with no reset time, a machine nobody logged into, one
that could not be reached — without spending real quota to arrange them,
which is why none of them had ever been looked at.
**Per machine, not per backend (2026-08-29).** The credential store that
matters is the one on the machine the session runs on, because that is the
account being billed — and in the layout this aims at, `ai-server` is on the
@@ -522,6 +596,68 @@ always running. So absent means **not running**, and only a timestamp that
arrives and cannot be parsed is unknown. `WindowEnd` in `ResetCountdown.kt`
is the one rule both readers go through.
### Auto-resume (2026-09-05)
**A session may pick itself back up when the account's usage limit lifts.**
Off unless somebody switched that session to it, because it spends quota the
moment quota exists and does so with nobody looking — that is not a thing a
default may decide. It sends one message, `continue` unless another was
typed, and then it is done; there is no retry loop around the conversation
itself.
**Running out of quota is a state, not an error.** `Event::LimitReached`
carries the dialect's reset time where it gave one, and recognising it
belongs to the driver — the Claude CLI ends the turn with `is_error` and
`Claude AI usage limit reached|1788546972`, and nothing above the driver
matches on a string. The transcript draws it as a divider, like a clear or a
compaction: what a reader scrolling back wants from it is why the
conversation stops at that line.
**The schedule is a plan to ask, never a plan to send.** Every reset time
available here is untrustworthy in the direction that matters: the dialect's
is written when the turn fails, and the endpoint's moves when the window
does. So the wait ends in a question to `usage.rs`, and only `ok` with no
window at 100% sends anything. A window still spent reschedules to *its own*
reset time — which is what makes a limit that lifts later than promised wait
longer, and one that lifts sooner resume sooner. A meter that cannot be
asked at all is a longer wait too, never a send: "we could not find out"
must not be able to produce the same action as "there is room".
Bounded, because something has to be: a day after the limit was hit the wait
stops and says so in the session's own transcript. A machine that can never
be asked would otherwise be retried for ever with nothing on screen saying
so.
The schedule is persisted on the session (`resume: Some(ScheduledResume)`),
not held in memory: a five-hour window routinely outlasts a backend restart,
and a wait forgotten across one is a session that silently never comes back.
`resume.rs` is the top layer — it holds the manager and the monitor and
neither holds it — which is what lets the decision be a pure function of a
snapshot and a clock. The pump reports limits downward on a broadcast, for
the reason `Shared` exists: the pump runs underneath the manager.
**Exercised with echo, never with a real account.** `/limit [minutes]` in an
echo session reports the same event a real driver does, and `/usage` sets
what the meter answers — deliberately two commands, because the two
disagreeing is the state the whole design is about. The loop was driven end
to end that way on 2026-09-05: the wait moved from the dialect's two minutes
to the meter's seven when the meter changed its mind, and the message went
out on the first check after the meter came back under the limit.
### Subagents (2026-09-05)
**A subagent is a second transcript owned by a session, in the same event
model, with no process and no controls of its own.** Full design and wire
shape in `SUBAGENTS.md`, kept separate because the app half is being built
against it in parallel and it is the shared contract between the two. The
one-paragraph reason: a session's Task-tool helpers already speak the common
event model on the parent's own stdout (each line carrying
`parent_tool_use_id`), so giving each one its own small transcript — same
file format, same paging routes, same SSE stream, reused by addressing rather
than by copying — costs a routing step in the translator and a registry
(`session/subagent.rs`) rather than a second session type with a driver, a
process and a config entry it does not need.
### HTTP surface
**`routes.rs`'s module doc comment is the table.** REST for actions, one SSE
@@ -800,8 +936,9 @@ Noticed and deliberately not fixed, so they are not re-found from scratch.
Phases 13 (the skeleton pipe, the full Claude driver, the usage screen) done
2026-08-24. Phase 4 (llama.cpp: model browsing, downloads, and `llama-server`
through its OpenAI-compatible endpoint) and phase 5 (ssh) done 2026-08-28.
The file explorer and the transcript cache followed in September. What is
through its OpenAI-compatible endpoint) and phase 5 (ssh) done 2026-08-28,
except for the remote `llama-server` and its port forward, which landed
2026-09-04. The file explorer and the transcript cache followed in September. What is
left is real-phone/WireGuard bring-up, which is operational rather than code.
Each phase ended runnable and verified against the real thing. The backend
+214
View File
@@ -0,0 +1,214 @@
# Review: iris changes since 0e46293
Scope: `git diff 0e46293..HEAD -- iris/ client-core/` (58 files, +5224/-226).
Read-only review; no source changed. Ordered likely-bug, then invariant
guards, then rules, then tests/docs.
## Likely bugs
1. **`iris/transcript-ui/src/lib.rs:152-160` (`RowDiff::Rebuild` arm of
`TranscriptScreen::apply`) never unregisters the rows it drops from
`Selection`, so a stale `WeakWidget<TextEdit>` outlives the widget it
points to and the next touch on *any* row panics.**
`Selection::rows: BTreeMap<RowKey, WeakWidget<TextEdit>>` documents its
own contract at `selection.rs:69-71`: "every addition here needs its
removal ... called when `List` evicts the row." The `ReplaceLast` arm
above it honours this (`lib.rs:143-145`, `self.selection.borrow_mut()
.unregister(old_key)` when the key changes). The `Rebuild` arm calls
`(self.list)(rsc).clear()` and rebuilds every row from `new_rows`, but
never touches `self.selection` — any key present in `old_rows` and
*absent* from `new_rows` (exactly what `group_tool_runs` regrouping two
separate tool-call rows into one produces — see `diff_tests::
a_tool_run_closing_and_joining_an_earlier_call_is_a_regroup_fallback`,
which tests the diff decision but not `apply` itself) is left in
`self.rows` pointing at a widget `List::clear()` just freed.
`TextEditable::edit` (`iris/src/widget/text/edit.rs:582-587`) resolves
that handle with `ui.widgets.get_mut(self).unwrap()` — an unconditional
panic on the freed slot. `Selection::begin` (`selection.rs:88-101`)
iterates *every* registered row (`w.edit(ui).deselect()`) on an
ordinary fresh press, so the crash fires on the next tap anywhere in
the transcript after a regroup, not only on a tap targeting the
orphaned row.
Fix: give `Selection` a way to reconcile against the row set that
survived a rebuild (e.g. `Selection::retain(&self, keys: &BTreeSet<RowKey>)`
removing everything else, called from the `Rebuild` arm before
rebuilding), or simplest — call `self.selection.borrow_mut()` cleared
the same way `List::clear()` clears the list, then let the rebuild's
`push_row` calls re-`register` everything as they already do.
## Guarded invariants missing
2. **`iris/src/widget/list.rs:751` (`List::place`) indexes/expects on
`slot` with no assertion that it exists.** `slot_widget` (`:563-575`)
panics via `.expect(...)` for a sentinel with no widget set, and does
an unchecked `&self.items[s as usize]` for a real index — a bare
"index out of bounds" with no context if `place` is ever reached with a
stale slot. Every current caller happens to derive `slot` from
`repair_anchor`/`prev_slot`/`next_slot`, which already check existence,
but that invariant is enforced by convention across three call sites,
not by the function that depends on it. Add
`debug_assert!(self.slot_exists(slot), "place() called with a slot that doesn't exist: {slot:?}");`
at the top of `place`.
3. **`iris/src/widget/list.rs:426` (`List::fling`) and `sense.rs`'s
`FlingCalculator::distance`/`duration`/`position_at` never check that
the incoming velocity is finite.** A `NaN`/`inf` velocity (a
`VelocityTracker::velocity()` divide-by-near-zero span, or a caller
passing a raw device value straight through) propagates through
`deceleration_for`'s `.ln()` silently — the fling either never settles
(`settled_on_schedule` compares against a `NaN` `duration()`, which is
always `false`) or jumps to `NaN` positions with nothing on screen
saying why. Add `debug_assert!(velocity_px_per_s.is_finite())` in
`List::fling` and `FlingCalculator::new`/`distance`.
4. **`iris/src/sense.rs:592-604` (`VelocityTracker::velocity`) has no
assertion that samples are chronological.** `add_sample` trusts its
caller's `Instant` ordering; a caller that samples out of order (a
restored/replayed gesture, a test) would silently produce a negative
`span` handled only by the `span <= 0.0 => 0.0` catch-all, masking the
bug that produced it rather than surfacing it. Add
`debug_assert!(self.samples.back().is_none_or(|&(last, _)| at >= last))`
in `add_sample`.
5. **`iris/core/src/render/frame_report.rs:247-252` (`mark_phase`) has no
assertion that phases are pushed in non-decreasing `start_index`
order.** `phase_stats`'s slicing (`:274`, `idx >= phase.start_index &&
idx < end_index`) silently produces an empty or nonsensical slice for
an out-of-order phase rather than surfacing the misuse — cheap to add
given `self.phases.last()` is already in scope:
`debug_assert!(self.phases.last().is_none_or(|p| self.total_frames >= p.start_index));`
## Rules
6. **Two mechanisms answer "what row selection points at, still valid?"**
`Selection` relies on callers remembering to `unregister` (finding 1);
`List` relies on callers deriving slots only from already-checked
sources (finding 2). Both are the same class of problem — a derived
handle that silently outlives what it points to — solved ad hoc twice
rather than once. Not asking for a shared abstraction here, but the two
should at minimum cross-reference each other's doc comment so the next
caller who adds a third handle-into-`List`-rows type (the code rules'
"a rule that governs a set belongs to the set") finds both existing
examples.
7. **`iris/android-app/src/bench_client.rs:224-225` (`battery_line`)
calls `.min().unwrap()`/`.max().unwrap()` on `samples` guarded three
lines above by `if samples.is_empty()`, which is fine — but the guard
and the two unwraps are two statements apart with a `let mean = ...`
in between reading the same slice; a future edit reordering those
lines loses the guard's protection silently.** Low severity (this is
the bench tool, not the app), but worth a one-line comment tying the
unwraps back to the guard, or restructuring as
`let (Some(min), Some(max)) = (samples.iter().min(), samples.iter().max())`
pattern so the empty case can't be separated from the check by a future
edit.
## Tests
8. **No test exercises `TranscriptScreen::apply`'s `Rebuild` arm through
`Selection`.** `lib.rs`'s `diff_tests` module (`:284-379`) tests only
the pure `diff_rows` decision function, never `apply` itself wired to a
real `Selection`; `selection.rs`'s own tests (`a_missed_press_start_
recovers_on_the_next_pressing_frame`, `unregister_forgets_the_row_and_
clears_a_matching_anchor`) never go through `apply`/`List::clear`
either. This is exactly the gap that let finding 1 through: the two
pieces (`apply`'s fallback, `Selection`'s registration contract) are
each tested in isolation and never together. Add: build a
`TranscriptScreen`, force a `RowDiff::Rebuild` (two adjacent tool-call
rows regrouping, per the existing `diff_tests` case), then call
`selected_text`/simulate a fresh press on a surviving row and assert no
panic.
9. **`iris/src/widget/list.rs`'s fling tests check total distance and the
start/end clamp but not the speed profile in between.**
`fling_moves_the_list_and_then_settles`/`fling_distance_is_positive_
toward_the_end` only assert the fling started, moved in the right
direction, and eventually stopped — none checks that
`tick_fling`'s per-tick delta is *monotonically decreasing* once past
the fling's peak (the property `fling_calculator_tests::position_at_
is_monotonic_and_clamped_past_the_end` already checks one level down,
for `FlingCalculator` alone, but never through `List::tick_fling`'s own
`scroll`/`anchor.offset` accumulation). A regression that made
`tick_fling` apply the *total* distance every tick instead of the
incremental one, for instance, would still pass both existing tests
(final position and direction are unaffected by how the interior ticks
split it up) while being wildly wrong every intermediate frame.
10. **`iris/src/widget/list.rs::replacing_the_last_row_stays_pinned_to_
the_bottom` and its sibling test `replace_back`'s effect on the
displayed row, never that the row it evicted is actually gone from
`heights`/`extents`.** Both tests assert the *new* row's position;
neither asserts `old.key` is absent from `list_ref.heights`/`extents`
after the replace (the "stale primitive" class finding 1 is a
production instance of). A cheap addition: assert
`!list_ref.heights.contains_key(&old.key)` after `replace_back` in the
existing test, since `old.key` is already returned to the test as
`evicted`... (`lib.rs` calls it that way; the `list.rs` test would need
to capture the key from `old` similarly.)
## Docs
No missing `IRIS.md` entry found for a *public* API change in this diff —
`List::fling`/`VelocityTracker`/`FlingCalculator`, `List::
anchor_position_display`, `FrameReport::mark_phase`/`phase_stats`/
`late_at_hz`, `UiRenderNode::new`'s `Result` change, `Len::dp`, and
`List::replace_back`/`clear`/`TranscriptScreen::apply` all have entries.
The `List::replace_back`/`clear`/`TranscriptScreen::apply` entry
(`docs/IRIS.md:526`) predates this review's finding 1 and does not mention
`Selection`'s registration contract at all — once finding 1 is fixed,
that entry should gain a line noting what the fix requires of a caller
that keeps its own row-keyed side table (the same shape `Selection` is),
so the next such table doesn't reproduce the same gap.
## Fixed, 2026-09-06
All ten findings addressed after the `DragGesture` merge (`selection.rs`
was rewritten by that merge, but finding 1's shape and location were
unchanged — `TranscriptScreen::apply`'s `Rebuild` arm, `iris/transcript-ui/
src/lib.rs`).
1. **Fixed.** `Selection::clear()` (`selection.rs`) drops `rows` and
`anchor`, called from `apply`'s `Rebuild` arm right before
`List::clear()``push_row` re-`register`s whatever survives as it
rebuilds each row, the "simplest" fix option the finding named.
2. **Fixed.** `debug_assert!(self.slot_exists(slot), ...)` at the top of
`List::place` (`iris/src/widget/list.rs`).
3. **Fixed.** `debug_assert!(velocity_px_per_s.is_finite())` in
`List::fling`, and `debug_assert!(velocity.is_finite())` in
`FlingCalculator::distance`/`duration` (`iris/src/sense.rs`).
`position_at` calls both, so it inherits the guard rather than needing
its own.
4. **Fixed.** `debug_assert!` on chronological sample order in
`VelocityTracker::add_sample` (`iris/src/sense.rs`).
5. **Fixed.** `debug_assert!` on non-decreasing `start_index` in
`FrameReport::mark_phase` (`iris/core/src/render/frame_report.rs`).
6. **Fixed (doc cross-reference only, as asked).** `Selection::register`'s
doc now points at `List::place`'s `slot_exists` assertion and vice
versa isn't needed since finding 2's fix already cites this file in
its own comment; both are grep-able on "docs/REVIEW-2026-09-06.md" and
on each other's type names.
7. **Fixed.** `bench_client.rs::battery_line` restructured to
`let (Some(min), Some(max)) = (samples.iter().min(), samples.iter().max())`,
so the empty-guard and the two lookups can no longer be separated by a
future edit.
8. **Fixed.** `transcript-ui`'s new `apply_tests::
a_row_dropped_by_a_regroup_does_not_outlive_itself_in_selection`
(`lib.rs`) builds a real `TranscriptScreen`, forces the same regroup
shape `diff_tests` already covers at the pure-diff level, calls `apply`,
and then `Selection::begin` on a surviving row — which panicked before
fix 1, resolving a `WeakWidget` `List::clear()` had just freed.
9. **Fixed.** `list.rs`'s new `tick_fling_applies_shrinking_incremental_
deltas` flings toward the end from `jump_to_start` and asserts each
tick's `extents[&0]` delta is no larger than the previous one — would
fail against a `tick_fling` that applied the total spline distance
every tick instead of the incremental slice, which the two pre-existing
fling tests cannot catch.
10. **Fixed.** `list.rs`'s new `replace_back_forgets_the_evicted_keys_own_
height` replaces row 4 with a row keyed `100` (the two existing
`replace_back` tests always reuse the same key, so neither actually
exercises the removal) and asserts `heights` no longer contains the
evicted key.
Docs: `docs/IRIS.md`'s 2026-09-05 `List::replace_back`/`clear`/
`TranscriptScreen::apply` entry now has a line on what the fix requires of
a caller with its own row-keyed side table, naming `Selection` as the
example and dating the fix.
Verification run alongside the rest of this pass's checks: `cargo fmt
--all`, `cargo clippy --workspace --all-targets`, `cargo test --workspace`
from `iris/` — see docs/RUST.md's plan box for the pass/fail and any
caveats from this same session.
+459
View File
@@ -0,0 +1,459 @@
# Review, 2026-09-07 — `ba2afba..origin/rustify`
Read-only review of the day's 24 commits: the glyph-atlas fix, the fling
spline and Lsq2 velocity estimator, keyboard/IME insets and `targetSdk`,
historical touch samples and the input clock, list culling / clamp /
anchor re-homing, nested masks and `draw_again`, the headless harness +
`transcript-fixture` + `rig-input`, desktop density, the release profile,
platform fonts + the Android monospace patch, and the client-core log ring
with `POST /client-log`.
**Verified while reviewing** (working tree, which also carries three other
agents' uncommitted edits — `iris/src/sense.rs`, `iris/core/src/ui/render_state.rs`,
`iris/src/lib.rs`, `iris/core/src/orientation/axis.rs`, and an untracked
`iris/src/diagnostics.rs`): `cargo fmt --check` clean in `iris/`,
`client-core/` and `server/`; `cargo clippy --all-targets` clean in `iris/`
and `client-core/`; `cargo test --lib -p iris` 101 passed, `cargo test -p
transcript-fixture` 10 passed. The `iris` doctest target fails to link
(`extern location for iris_core does not exist`) — a stale build artefact,
not a code fault, but worth knowing before trusting `cargo test -p iris`
as a whole.
The work is unusually well documented and the two "a test that compared
the code with itself" findings the authors made themselves are real and
were fixed correctly. What follows is what is left.
Counts: **5 defects, 7 risks, 3 tests that cannot fail in the bug's
direction, 7 rule findings, 2 nits.**
## Fix pass, 2026-09-07 evening
Every finding below carries a **Status** line. In summary: **13 fixed**
(D1, D4, D5, R1, R5, R7, T1, T2, T3 and four of the rule findings and both
nits), **6 moot or deferred** (D2, D3, R3, R4 and two rule findings, all
of them in the phone-logging route that `06b8a1f` deleted or in files the
devlog agent held open), and **2 not done on purpose** (R2, which waits on
docs/LAYOUT.md's mask redesign, and R6, which needs Iris's own phone).
The commits are `2ec0fee` (D4), `7e79ec1` (D5), `551c013` (R1), `e10582a`
(T1-T3), `ff1d6ea` (R5, R7) and `a6a100e` (the rename and the nits). Each
fix that the rig can express carries a test, and each of those was
confirmed by breaking its subject on purpose -- the break is recorded
beside the assertion, so the next reader does not have to re-derive it.
---
## Defects
### D1 — the app's own log ring is drowned by the same day's per-frame `debug!` lines, so the route built to get Iris's logs to her carries almost none of them
`iris/android-app/src/lib.rs:132` installs the ring at `LevelFilter::Debug`,
and `client-core/src/log_ring.rs:279` (`RingLogger::enabled`) returns
`true` unconditionally by design, so **every `log::debug!` in the process
lands in a 2000-line / 256 KiB ring**. In the same commit range that ring
became the only way a line reaches Iris, three ungated per-frame `debug!`
callsites are live:
- `iris/src/android/view.rs:446` and `:509` — two lines *per rendered frame*.
- `iris/src/widget/list.rs:576``iris fling tick:`, one line per fling tick.
- `iris/src/widget/text/mod.rs:81` — one per text shape (many per frame while rows compose).
**Failure scenario.** Iris flicks the transcript on a 120 Hz phone. That is
~240360 debug lines a second; the ring's 2000-line bound is exhausted in
**under ten seconds**, so by the time she presses `Copy report` every
`log::info!` about what she was actually investigating has been evicted.
The uploader makes it worse: it sends at most the ring per 10 s wake
(2000 lines ≈ 200 lines/s) against ~350 lines/s produced, so it also runs
permanently behind and pushes tens of KB/s of frame spam over the tunnel.
Note that another agent has already built the right mechanism — the
untracked `iris/src/diagnostics.rs` has `set_trace`/`trace_enabled`, a
default-off gate, and its module doc states this exact problem in as many
words. It gates `iris::input`/`iris::frame`; it does **not** gate the four
callsites above.
*Fix*: put `List::tick_fling`'s line and `view.rs`'s two `render():` lines
behind `iris::diagnostics::trace_enabled()` (the mechanism that already
exists for exactly this), and/or record into the ring at `Info` while
leaving `android_logger` at `Debug`.
**Status:** fixed in `992c472` (verified 2026-09-07: all four callsites, plus `sense.rs`'s drag-release samples line, now sit behind `iris::diagnostics::trace_enabled`, and `input_log_roundtrip` proves both directions).
### D2 — `POST /client-log` can make `ai-server` write an unbounded runtime log at an authenticated client's request
`server/src/routes.rs:1473` bounds the **line count** (500) and nothing
else. The route sits inside the router that applies
`DefaultBodyLimit::max(32 * 1024 * 1024)` at `server/src/routes.rs:179`
(raised for phone photos), so one request may carry 500 lines of ~64 KiB
each, and each is re-emitted verbatim into `tracing`. There is no
per-message cap on the server, no rate limit, and the runtime log
`ai-server` writes is the file Dev Updater tails and never rotates.
`MAX_MESSAGE_BYTES` (4096) exists only in the *client*
(`client-core/src/log_upload.rs:33`), i.e. the server trusts a value the
attacker controls.
**Failure scenario.** A buggy client (a `log::debug!` in a loop is enough —
see D1) or one holding a leaked bearer token posts 32 MiB every 10 s; the
host's disk fills and every other component's log goes with it.
*Fix*: give the route its own `DefaultBodyLimit` (the attachments route at
`:175` is the precedent for a per-route limit) and truncate each `message`
server-side to the same 4096 bytes rather than assuming the client did.
**Status:** moot -- `POST /client-log` was deleted with the whole upload route (`06b8a1f`), the app hands its log to Dev Updater through an on-device ContentProvider instead. Nothing to bound.
### D3 — lines the ring drops before the uploader sends them vanish with nothing saying so
`LogRing::since` (`client-core/src/log_ring.rs:169`) filters `seq >= cursor`
and silently returns fewer lines when eviction has passed the cursor;
`LogUploader::flush_once` (`:94`) then advances to whatever came back.
`dropped` is counted (`log_ring.rs:109`) and shown in the *local*
diagnostics pane, but it is never put in the upload body, and
`ClientLogBody` has no field for it.
**Failure scenario.** The tunnel is down for two minutes; the ring wraps.
When it comes back, the server log jumps from `#812` to `#5106` with no
line saying anything was lost. This is precisely the "unknown state
sharing a value with the empty state" UI_RULES asks to design first, and
the module doc for `dropped` claims it is "reported rather than inferred"
— it is, but only on the half of the path nobody is reading.
*Fix*: carry `dropped` (or `firstSeq`) in the batch and have `client_log`
emit one `warn!` when the sequence is not contiguous with the last batch
from that `source`.
**Status:** moot -- `client-core/src/log_upload.rs` was deleted with the route (`06b8a1f`). Whatever the ContentProvider does about eviction is that design's question, not this one's.
### D4 — the input clock anchors on the first event's *own* time, so that event's historical samples are dated before the anchor: the ordering assert fires, and release silently collapses them onto one instant
`iris/src/android/view.rs:628` takes the anchor as
`(Instant::now(), event.event_time_nanos())` from the first `MotionEvent`
the view ever sees, and `at()` computes
`anchor_at + (sample_time - anchor_nanos).max(0)`. Historical samples of
that same event are by definition **earlier** than its own `event_time`.
**Failure scenario.** The first event this view receives is an
`ACTION_MOVE` (the `DOWN` was delivered to another view, or the view was
attached mid-gesture). Its historical samples are, say, 12 ms before
`anchor_nanos`; `at()` clamps all of them to `anchor_at`, so the tracker
receives three samples with identical timestamps, the Lsq2 fit is
degenerate, and the flick reads 0 px/s. In a debug build the
`debug_assert!(ht >= previous)` at `:653` fires first — but `previous`
starts at `anchor_nanos` (`:651`), which is a value from a *different*
event, so that assert is also the wrong comparison for the first sample of
every later event.
*Fix*: anchor on the earliest sample of the first event
(`historical_event_time_nanos(0)` when `history_size() > 0`, else
`event_time`), and seed `previous` from the previous event's last sample
rather than from the anchor.
**Status:** fixed in `2ec0fee`. The arithmetic moved into `sense::PointerClock`, which anchors at `now - (event_time - oldest_sample)` and carries the last sample seen *across* events, so the ordering assert compares against the previous event's last sample rather than the anchor. It lives in `sense` because `iris::android` is `cfg`'d out everywhere but the device: `sense_tests.rs`'s `the_first_events_batched_samples_are_dated_apart` reports `[0ns, 0ns, 0ns]` against the old anchoring.
### D5 — the "before" velocity quoted in four places is not what the reference script prints
`iris/benches/velocity_reference.py`, run today, prints **12250 px/s** for
`flick-120hz.touch`'s average and **12500 px/s** for "press and one move
frame". Four places say 11750 for both:
- `docs/RUST.md:900` (`flick-120hz.touch | 11750 px/s`)
- `docs/RUST.md:905` (`press + one move frame | 11750 px/s`)
- `docs/IRIS_TODO.md:1026`
- `iris/transcript-fixture/tests/phone_screen.rs:55`
`iris/src/sense.rs:1406` has the correct 12250, so the two halves of the
same change disagree. The file that carries the wrong number is the one
that says "every number below is printed by `velocity_reference.py` … do
not 'fix' one by running the Rust and copying what it said". One of the
two rows also being 11750 for a completely different sample set is the
tell.
*Fix*: replace 11750 with the script's own 12250 / 12500 in those four
places, or say which run produced 11750.
**Status:** fixed in `7e79ec1`. All four places now say 12250 / 12500, the 1.30x ratio becomes 1.24x, and RUST.md records where 11750 half came from (196 px over a 16.68 ms **60 Hz** frame rather than the recording's own 16 ms -- which explains the flick row and not the other one, so that one was copied).
---
## Risks
### R1 — every new invariant guard is a `debug_assert!`, and the phone runs release
The five guards added today —
`iris/src/widget/list.rs:1156` (a `List` must be inside a `.masked()`),
`:1218` (`extents` holds only on-screen rows),
`iris/src/android/view.rs:653` (historical sample ordering),
`iris/src/sense.rs:1076` (`poly_fit_least_squares` sample count), and
`iris/core/src/ui/painter.rs`'s doubled-`set_mask` check — are all
`debug_assert!`. `docs/RUST.md` records that the bench APK **must** be
installed as `release` on the emulator (the debug `libmain.so` is 325 MB
and will not install) and Iris's phone gets release too. So none of these
can fire on any build anybody actually runs; in release a `List` drawn
without a mask silently paints over its surroundings again — the exact
fault e922b73 was written to fix.
*Fix*: for the two that are cheap and once-per-draw (`is_masked`, the
extents check), consider a plain `assert!` or a one-shot `log::error!`, so
the guard survives into the build the defect was found in.
**Status:** fixed in `551c013`. `is_masked`, the `extents` check, `set_mask`'s doubled-call check, `Painter::glyphs`'s atlas generation and `List::fling`'s finiteness are `assert!`/`assert_eq!` now; `List::place`'s slot precondition, `poly_fit_least_squares`'s two, and `PointerClock::sample`'s ordering stay `debug_assert!` and say in a comment why. The layer-1 suites pass in `--release` as well as debug, which is what says the promoted ones do not fire on a real replayed flick.
### R2 — a straddling row is now invisible above the list and still tappable through the header
Masks are applied in the fragment shader
(`iris/core/src/render/shader.wgsl:203`); the CPU hit path
(`UiRenderState::resolved_region`, `iris/core/src/ui/render_state.rs:709`)
does not consult `masks` at all. Before today the top of a straddling row
was drawn over the header *and* hit-testable there; now it is clipped away
but still hit-testable, which is worse — a tap on "Run benchmark" can land
on an invisible link in the row behind it. `docs/LAYOUT.md:1012` ("Hit-
testing applies the shape") is design, not code.
*Fix*: until LAYOUT.md's mask redesign lands, intersect a widget's hit
region with its mask chain in `resolved_region`; the chain walk already
exists on the GPU side.
**Status:** not done, deliberately -- docs/LAYOUT.md's mask redesign ("masks reference a drawn primitive instead of copying a shape", `1121d7c`) is where hit-testing gets the shape, and intersecting a chain in `resolved_region` now would be a second mechanism to unpick. Pointer left here rather than a fix.
### R3 — three copies of one wire contract, none of them linked
`client-core/src/log_upload.rs:28` (`MAX_LINES_PER_BATCH = 500`) and
`server/src/routes.rs:1418` (`CLIENT_LOG_MAX_LINES = 500`) must agree, in
different crates, with only a comment saying so; the body itself is built
by hand with `serde_json::json!` on one side and parsed by a
`#[serde(deny_unknown_fields)]` struct on the other. This project already
has the mechanism for exactly this — `event-model`, a crate both `server`
and `client-core` depend on precisely so "the app hand-mirroring it" stops
happening (`server/Cargo.toml:16` says so).
**Failure scenario.** Somebody raises the client's batch to 1000. Every
upload now returns 400, the uploader retries the *same* batch from the same
cursor forever, and the only sign is one line in a diagnostics pane on a
phone.
*Fix*: move `ClientLogLine`/`ClientLogBody` and the batch constant into a
shared crate.
**Status:** moot -- both copies went with the route (`06b8a1f`). If a client/server contract comes back, `event-model` is still the answer.
### R4 — `build.rs` bakes in a CA it never asks Cargo to watch, and the bench build now has no rebuild trigger at all
`emit_log_config` (`iris/android-app/build.rs:92`) calls `read_pinned_ca()`
but emits only `rerun-if-env-changed` for `AI_APP_LOG_HOST/_PORT/_TOKEN`
no `rerun-if-changed` for the CA *file*, and (because the bench build
returns at `:65`, before the transcript path's declarations) no
`rerun-if-env-changed=AI_APP_CA`/`XDG_CONFIG_HOME` either. Emitting any
`rerun-if-*` directive turns off Cargo's default "rerun when anything in
the package changes" heuristic, so the bench build lost the only trigger it
had.
**Failure scenario.** `~/.config/ai-app` is wiped (AGENTS.md calls this the
one-way door), `ai-server` mints a new CA, the APK is rebuilt — and
`build.rs` does not re-run, so the APK still pins the dead CA and every
upload fails with a TLS error nobody can attribute.
*Fix*: `println!("cargo:rerun-if-changed={}", ca_path.display())` inside
`read_pinned_ca`, and move the `AI_APP_CA`/`XDG_CONFIG_HOME` declarations
above the bench early-return.
**Status:** moot -- `iris/android-app/build.rs` was deleted (`06b8a1f`/`d8562d9`): the destination comes from the enrolment link now, so nothing is baked in at build time and there is nothing for Cargo to watch.
### R5 — desktop density is read once and never updated
`iris/src/default/mod.rs:254` reads `content_scale(window)` at startup and
sets it on both `rsc.ui.text.density` and `render`. `WindowEvent::
ScaleFactorChanged` is not handled, and `UiRenderer::resize` deliberately
no longer consults `scale_factor`. Dragging the window to a monitor with a
different scale leaves every `dp(...)` and every rasterised glyph at the
old density — the same class of disagreement the commit removed elsewhere.
It is invisible here (every display on this machine is 1.0), which is why
it needs writing down.
**Status:** fixed in `ff1d6ea`. `WindowEvent::ScaleFactorChanged` re-reads `content_scale` -- through that function, so `IRIS_SCALE` still pins `--phone`'s density instead of following the monitor -- and `UiRenderState::set_density` marks the tree for a full redraw when the value actually changes, since `Text::shape` keys its cache on `(attrs, width, density)`.
### R6 — removing the bundled fonts removed the guard for a fault that was found on the phone, and the check was run on the desktop
`iris/core/src/primitive/text.rs`'s `register_bundled_fonts` existed
because "bold spans on a real phone rendered as blank gaps of the correct
advance width" — the deleted doc says so. Its removal is Iris's own call
and is recorded properly in `docs/DECISIONS.md`, but the verification
recorded there is "checked with CJK + emoji **on desktop**", which is the
half that cannot fail: the fault was Android's font enumeration resolving
a weight/style. `iris/transcript-ui/src/tool.rs:110`'s comment is honest
that `CLOSED_MARK`/`OPEN_MARK`/`UP_MARK` (U+25B8/BE/B4) are now "a bet"
that the platform monospace face has them — which is UI_RULES' "don't rely
on characters the platform might not have", stated and then accepted.
*Fix*: before the next phone build, look at a bold run and the three
chevrons on Iris's device specifically; the emulator's font set is not
evidence for hers.
**Status:** not done here -- it is a *look at it on Iris's phone* item, and no build in this VM is evidence about her device's font set. Carried forward as the review said: before the next phone build, look at a bold run and at `CLOSED_MARK`/`OPEN_MARK`/`UP_MARK` (U+25B8/BE/B4) on her device specifically.
### R7 — the least-squares fit clamps a degenerate norm instead of detecting it
`iris/src/sense.rs:1105`: `1.0 / dot(...).sqrt().max(1e-6)`. Compose's
`polyFitLeastSquares` treats `norm < 1e-6` as "vectors are linearly
dependent, no solution" and bails; clamping instead produces a `q` row of
zeros, a zero on `r`'s diagonal, and a `0/0` that the `is_finite` check at
`:1059` happens to catch. It works, but it works by accident and the escape
is not the one the source it is transcribed from takes.
**Status:** fixed in `ff1d6ea`. `poly_fit_least_squares` returns `Option` and bails at `DEGENERATE_NORM` (Compose's `0.000001f`) instead of clamping; `velocity()` answers 0 on `None`. `a_fit_through_linearly_dependent_points_has_no_solution` reports `Some([NaN, NaN, NaN])` with the clamp back in place.
---
## Tests that cannot fail in the direction the bug would go
### T1 — `iris/transcript-fixture/tests/phone_screen.rs:64` computes the expected fling duration with the calculator under test, and asserts it one-sidedly
`let expected = FlingCalculator::new(PHONE_SCALE).duration(velocity);` then
`assert!(ran_for <= expected + 2 frames)`. This is the same
"calculator compared with itself" shape the fling-spline commit
(73f956f) identified and fixed elsewhere, and the direction it can fail in
is "the fling ran too long" — never "the fling stopped dead", which is
literally Iris's reported symptom. The companion
`assert_ne!(before, after)` passes on one pixel of travel. A fling that
settles on the first tick passes this test.
*Fix*: add a lower bound from `velocity_reference.py`'s number (a fling at
-15250 px/s at density 2.55 must run ≥ ~1.4 s and travel ≥ ~6000 px), not
from `FlingCalculator`.
**Status:** fixed in `e10582a`. Both bounds come from `fling_spline_reference.py`, which gained this case's own line (`density=2.55 v=15250.0: distance=11057.424px duration=2.0716s`), and travel is measured in pixels from a row's own on-screen extent (10527px measured). Scaling `tick_fling`'s elapsed by 1000 reports "stopped after 8ms"; scaling its delta by 0.01 reports "travelled 111px".
### T2 — `top_edge.rs:150` checks a row *count* on the leg where the culling bug appeared, and the box only on the other leg
`rows_that_have_left_the_viewport_are_not_drawn` asserts `rows.len() <= 24`
on the outbound leg and the per-row `inside the box` predicate only on the
return leg. The doc explains why (an unmeasured row must be drawn to be
measured), which is correct — but it means the test's name is only true of
half of it, and a regression that draws 20 rows in the wrong *place* on the
outbound leg passes.
**Status:** fixed in `e10582a`. The first leg still cannot assert the box (an unmeasured row has to be drawn to be measured), so there is a third leg -- back again, every height known. Widening `intersects_viewport` downwards passes all 40 forward steps and fails at "back 6".
### T3 — `top_edge.rs:116` checks that a mask exists and where it is, not that it reaches anything
`the_list_is_clipped_to_its_own_box` asserts `active.mask != MaskIdx::NONE`
and that the mask's region lies within the list's box. It never checks the
row primitives actually reference that mask, so a broken `Mask::parent`
chain — the thing d507ae4 introduced — would leave this green while a code
fence inside a row drew unclipped again.
*Fix*: assert that a row primitive's mask chain contains the list's mask
slot.
**Status:** fixed in `e10582a`. It walks every row primitive's mask chain and requires the list's own slot on it, and rejects a chain that loops. Forcing `Painter::set_mask`'s `parent` to `NONE` fails it with "clips to [Id(1)], a chain that never reaches the list's own mask Id(0)".
---
## Rules
- **`iris/src/widget/list.rs:576` is a second mechanism for per-frame
instrumentation.** `iris::diagnostics::trace_enabled` exists for exactly
"a default-off `debug!` in a hot path" and this line does not use it.
(Cause of D1; the gate is in the untracked `diagnostics.rs`, so at the
reviewed commit the line is simply ungated.)
- **`server/src/routes.rs:1518` (`client_log_time`) duplicates
`client-core/src/log_ring.rs:76` (`clock_time`)** — the same arithmetic
written twice in two crates, with a comment noting they must agree. Same
shared-crate answer as R3.
- **`client-core/src/log_ring.rs:301`'s doc claims more than the code
delivers**: "the caller is named in the error so it is findable" —
`log::SetLoggerError` names nobody. `iris/android-app/src/app_log.rs:44`
repeats the claim.
- **Stale comment: `iris/src/android/view.rs:624`** cites
`VelocityTracker::add_sample`'s debug assert; the method was renamed to
`add_position` in the same commit range.
- **`MOVE_CHAIN_LIMIT` now bounds two different chains** (move offsets and
masks) under a name that says one, in both
`iris/core/src/ui/render_state.rs:63` and `shader.wgsl:97`. The shader's
comment already calls it "the bound on the parent walk"; the constant
should say that too, or masks should get their own.
- **`iris/src/sense.rs:1434`'s stated negative control is not reproducible
as written.** "Reverting `velocity` to `total / span` fails exactly this
one, the flick recording, and `phone_screen.rs`" — but `samples` now
holds *positions*, so `total / span` over them gives 2750 for the steady
drag too, and the commit message for the same change says "exactly seven
tests". Two numbers for one experiment.
- **`iris/android-app/src/bench_client.rs:393`'s `ime_visible` is right and
its sibling one line up is not.** `set_bottom_inset(rsc,
insets.bottom.max(insets.ime_bottom))` still infers "make room" from a
`max`, so during the slide-in the composer is padded by the system-bar
inset while `ime_visible` already says the keyboard is up. Harmless
today; it is the same conflation the comment beside it warns about.
**Status of the rule findings, 2026-09-07 evening.**
- `list.rs:576`'s ungated per-frame line -- **fixed in `992c472`** with
the rest of D1.
- `routes.rs:1518`'s `client_log_time` duplicating `log_ring.rs`'s
`clock_time` -- **moot**: the route was deleted (`06b8a1f`).
- `log_ring.rs:301`'s "the caller is named in the error" -- **deferred to
the devlog agent**; `client-core/src/log_ring.rs` is its file this pass,
and `app_log.rs` no longer repeats the claim.
- `view.rs:624`'s stale `VelocityTracker::add_sample` -- **fixed in
`2ec0fee`**; the paragraph was rewritten for the anchoring change and
now names `PointerClock` rather than a method that no longer exists.
- `MOVE_CHAIN_LIMIT` naming two chains -- **fixed in `a6a100e`**: renamed
to `PARENT_CHAIN_LIMIT` in `render_state.rs` and `shader.wgsl` at once
(it had no other users), with the doc naming both chains it governs.
- `sense.rs:1434`'s unreproducible negative control -- **fixed in
`7e79ec1`**. Rerun with `velocity` reverted to `(newest - oldest) /
span`: seven fail in `-p iris` (the flick recording, the accelerating
flick, the horizon, the stopped finger, the minimum sample count, both
`drag_gesture` flick tests) plus `phone_screen.rs`'s flick. RUST.md's
"exactly seven" was right; the doc comment's "exactly this one, the
flick recording, and `phone_screen.rs`" was not, and now says the same
thing RUST.md does.
- `bench_client.rs:393`'s `set_bottom_inset(.., max(..))` -- **deferred to
the devlog agent**; `iris/android-app/**` was open under it this pass.
## Nits
- `iris/src/sense.rs:798` computes `self.velocity.velocity()` twice on a
release when `info` logging is on (once for the outcome, once for the
log line) — a full Lsq2 fit each.
- `iris/transcript-ui/src/selection.rs:303` calls `ui.ui_mut().animate(id)`
even when `fling()` bailed (`|v| <= 1.0`, or no anchor). Harmless — the
first `tick` unregisters — but it registers an animation that is known
not to exist.
---
**Status of the nits, both fixed in `a6a100e`.** `DragGesture`'s release
computes `velocity()` once into a local both the outcome and the
`iris drag release:` line read. `selection.rs`'s `animate(id)` is behind
`is_scrolling()`, which is the same answer `List::fling` itself reached --
and `phone_screen.rs`'s recorded flick still flings, which is the half
that says the guard did not turn a working release off.
## Commits reviewed
```
7e4e26a iris: resolve fontique's Android monospace generic family ourselves
84a13e8 iris: a fling starts at Compose's velocity, which is a curve fit and not an average
452c442 docs/RUST.md: queue -- logging landed; iris app enrolment ...
238057a docs: the phone-logging decision, how to use it, and two build-apk traps
896c93a iris: drop bundled Noto Sans, match Compose's platform-font fonts
690161e docs: the transcript's edges were three faults, and what the rig found
e922b73 iris: a transcript row is drawn if it overlaps the viewport, and clipped to it
d507ae4 iris-core: masks nest instead of aborting, and a widget can ask to be drawn again
9ed01e2 docs: phone report 2026-09-07 later -- overscroll, low initial fling velocity ...
5be9f1b iris-android-app: keep the app's own log, put it in Copy report, upload it
977bdb9 client-core: the app's own log ring, and POST /client-log to get it off a phone
9cd1263 docs/RUST.md: queue -- APK size done, the embedded-fonts question left for Iris
42af780 iris android-app: strip+LTO+cgu1+opt-level=s halve libmain.so, no feature trim needed
4274b8b Merge remote-tracking branch 'origin/rustify' into worktree-agent-ace98b0bdaf33ffff
73f956f iris: the fling curve was the identity function, and the keyboard was a targetSdk
038f6a3 docs: the test rig's layers 1 and 2, with their commands and their limits
1121d7c docs/LAYOUT.md: masks reference a drawn primitive instead of copying a shape ...
232de0e iris: a phone-shaped desktop window, driven by the same touch recordings
e430880 docs: phone report 2026-09-07, rows at the transcript's top edge culled early ...
a999bd1 docs: masks with a shape (LAYOUT.md, decided 2026-09-07) and the orchestrator queue
6840edf iris-android-app: the bench's fixture half comes from transcript-fixture
3332201 iris: a headless in-process harness, and the bench fixture as a shared crate
7f4ea7e docs/TODO.md: Compose app crash from Iris's phone log export, reversed AnnotatedString range
591128e AGENTS.md: the phone app and the planned desktop app share widgets and styling
```
+8396
View File
File diff suppressed because it is too large. Load diff
View File
File renamed without changes.
+11
View File
@@ -33,3 +33,14 @@ one in place when it turns out to need a decision.
that would work today, for Claude sessions, and it is the option that was
not chosen.
## From Iris's phone log export, 2026-09-07 (Compose app)
- [ ] **Crash on 2026-09-03 11:40, `IllegalArgumentException: Reversed
range is not supported`** at `ToolInput.kt:200` (`highlighted`, inside
`ToolInputView` -> `RawBlock` -> `ToolCard`). An `AnnotatedString`
range was built with end before start while highlighting a tool
input. Found in the per-package system log she exported; the tool
input that triggered it is not in the log. Reproduce by fuzzing
`highlighted` with inputs whose token boundaries collapse, and guard
the range construction.
File renamed without changes.
+73
View File
@@ -0,0 +1,73 @@
# Compose bench report from Iris's phone, 2026-09-06
The Compose half of P0 (RUST.md), run by Iris on her own phone and pasted
back verbatim. The iris half's report goes beside it in this directory
when it exists. Her caveat, worth keeping with the numbers: "I don't think
this is entirely fair because the UI for iris is more minimal" -- the
Compose screen also draws the usage bar, the status row and tool cards,
which the iris bench screen does not yet. Her impression of the iris build
before its first-touch bug: "it already feels very smooth so far".
What to read first: the phone runs at 120 Hz, so the budget is 8.3 ms;
`late` is measured against that. Compose's tail is the streaming phase --
`markdown reparsed while streaming: 396, 8.5ms mean, 25.8ms worst` and
`record: one block: 398, 6.3ms mean, 19.9ms worst` -- which is exactly the
path iris's `TranscriptScreen::apply` (replace the last row only) is meant
to beat. Process CPU over the run is 20.9 s of a 38.5 s run; peak RSS
587 MB; battery current mean 419 mA.
```
ai-app render report
device: Pixel 9 Pro XL (Google), Android 17
build: release
transcript:
43 events, 41 rows, 93 units loaded
viewport 1333px, 2 units visible
on screen: the list's own 0px, AssistantMsg 24520px
0 tool calls and 0 groups open
frames:
1613 frames over 38.5s at 120Hz (8.3ms budget)
late: 742 (46.0%)
total p50 7.7ms p90 29.2ms p99 41.1ms
waited p50 0.5ms p90 12.9ms p99 27.2ms
input p50 0.0ms p90 0.0ms p99 0.0ms
anim p50 1.1ms p90 5.8ms p99 9.5ms
layout p50 0.0ms p90 0.1ms p99 0.2ms
draw p50 0.4ms p90 15.9ms p99 27.9ms
sync p50 0.1ms p90 0.5ms p99 1.0ms
issue p50 1.1ms p90 1.7ms p99 3.0ms
swap p50 0.4ms p90 0.5ms p99 0.7ms
gpu p50 1.8ms p90 2.1ms p99 6.6ms
where the draw phase went:
draw phase 3.83ms per frame, of which:
the transcript: 0.25ms (measure 0.15, place 0.10, record 0.00)
everything else: 3.58ms (93%)
work since this was last copied:
draw: the whole transcript: 12, 0.2ms total, 0.0ms mean, 0.0ms worst
grouped tool runs: 398, 10.3ms total, 0.0ms mean, 0.1ms worst
markdown cut into pieces: 1, 0.0ms total, 0.0ms mean, 0.0ms worst
markdown parsed while composing: 7, 3.0ms total, 0.4ms mean, 0.5ms worst
markdown ready: 46
markdown reparsed while streaming: 396, 3384.6ms total, 8.5ms mean, 25.8ms worst
markdown warmed: 1, 1.4ms total, 1.4ms mean, 1.4ms worst
measure: the whole transcript: 957, 248.7ms total, 0.3ms mean, 15.7ms worst
message composed: 403
message cut into parts: 1, 0.1ms total, 0.1ms mean, 0.1ms worst
place: the whole transcript: 1319, 162.4ms total, 0.1ms mean, 1.3ms worst
record: one block: 398, 2498.7ms total, 6.3ms mean, 19.9ms worst
session screen recomposed: 413
status row recomposed: 1
unit composed: 538
units flattened: 399, 44.3ms total, 0.1ms mean, 0.5ms worst
usage bar recomposed: 413
bench:
scroll: 6 cycles (24 swipes), streamed 400/400 fixture events
process CPU time over this run: 20907ms
peak RSS: 587356kB
battery current: mean -418509µA over 39 samples (min -1988281, max -107812)
```
+93
View File
@@ -0,0 +1,93 @@
# Compose bench v2 report from Iris's phone, 2026-09-06
Bench v2 (fling / stream / type / keyboard, RUST.md's P0 box) on the
Compose `bench` build, run by Iris on her Pixel 9 Pro XL, verbatim. Note
the display was at **60 Hz** for this run (16.7 ms budget) where the v1
run was at 120 Hz -- the phone's adaptive refresh rate decides, and
`late` is judged against whichever it was, so compare a run with a run at
the same rate. The iris v2 report goes beside this when it exists.
What it says: fling, type and keyboard are all essentially clean on
Compose (0.1%, 0.9% and 0% late; fling p50 5.5 ms, p99 11.6 ms). The
whole tail is the streaming phase again -- 41.9% late, p99 42.5 ms,
driven by `markdown reparsed while streaming` (8.6 ms mean, 30.3 ms
worst) and `record: one block` (6.3 ms mean, 25.5 ms worst). Process CPU
69.6 s over the 125.5 s run; peak RSS 577 MB; battery current mean
571 mA over 126 samples.
```
ai-app render report
device: Pixel 9 Pro XL (Google), Android 17
build: release
transcript:
108 events, 26 rows, 58 units loaded
viewport 1531px, 2 units visible
on screen: the list's own 0px, AssistantMsg 24520px
0 tool calls and 0 groups open
per phase:
fling: 3278 frames over 32.7s
late: 4 (0.1%)
total p50 5.5ms p90 8.7ms p99 11.6ms
worst 49.0ms
stream: 1041 frames over 21.3s
late: 436 (41.9%)
total p50 13.4ms p90 31.7ms p99 42.5ms
worst 52.5ms
type: 2446 frames over 61.5s
late: 23 (0.9%)
total p50 7.3ms p90 13.2ms p99 16.5ms
worst 38.9ms
keyboard: 358 frames over 10.0s
late: 0 (0.0%)
total p50 6.3ms p90 8.6ms p99 11.1ms
worst 12.0ms
frames:
7122 frames over 125.5s at 60Hz (16.7ms budget)
late: 463 (6.5%)
total p50 6.0ms p90 13.8ms p99 34.0ms
waited p50 0.5ms p90 1.1ms p99 19.9ms
input p50 0.0ms p90 0.0ms p99 0.0ms
anim p50 0.7ms p90 4.5ms p99 7.6ms
layout p50 0.1ms p90 0.1ms p99 0.2ms
draw p50 0.7ms p90 2.9ms p99 21.9ms
sync p50 0.1ms p90 0.2ms p99 0.6ms
issue p50 1.4ms p90 2.4ms p99 3.2ms
swap p50 0.4ms p90 0.8ms p99 1.2ms
gpu p50 1.5ms p90 2.1ms p99 6.6ms
where the draw phase went:
draw phase 1.74ms per frame, of which:
the transcript: 0.24ms (measure 0.10, place 0.14, record 0.00)
everything else: 1.51ms (86%)
work since this was last copied:
draw: the whole transcript: 280, 1.9ms total, 0.0ms mean, 0.0ms worst
grouped tool runs: 407, 18.6ms total, 0.0ms mean, 0.2ms worst
markdown cut into pieces: 40, 0.4ms total, 0.0ms mean, 0.0ms worst
markdown parsed while composing: 2, 1.6ms total, 0.8ms mean, 1.1ms worst
markdown ready: 323
markdown reparsed while streaming: 395, 3406.9ms total, 8.6ms mean, 30.3ms worst
markdown warmed: 40, 38.2ms total, 1.0ms mean, 4.3ms worst
measure: the whole transcript: 1978, 683.9ms total, 0.3ms mean, 15.9ms worst
message composed: 397
message cut into parts: 40, 2.9ms total, 0.1ms mean, 0.2ms worst
place: the whole transcript: 4308, 996.0ms total, 0.2ms mean, 2.4ms worst
record: one block: 394, 2472.7ms total, 6.3ms mean, 25.5ms worst
session screen recomposed: 1630
status row recomposed: 1
transcript page from server: 10
unit composed: 927
units flattened: 408, 85.5ms total, 0.2ms mean, 1.8ms worst
bench:
fling: 8 flings out + 8 back at 12000px/s, travel start=idx=0/off=0px outward=idx=188/off=182px end=idx=0/off=0px
scroll: 6 cycles (24 swipes, legacy tween), streamed 400/400 fixture events
type: 600 characters inserted then deleted, one per 50ms
keyboard: shown 5/5, hidden 5/5 (confirmed via isImeVisible)
process CPU time over this run: 69564ms
peak RSS: 577452kB
battery current: mean -571483µA over 126 samples (min -2361718, max -99218)
```
@@ -0,0 +1,31 @@
# iris bench report from Iris's phone, 2026-09-06, before the phone fixes
Build 46246ea (Vulkan, bench v1: 24-swipe scroll loop then 400 streamed
events), run by Iris on her Pixel 9 Pro XL before the first-touch wipe,
the missing bold faces, the density scale and the status-bar inset were
fixed -- so the rows were drawn at roughly a third of their intended size
and the run may have included frames after the wipe. Preliminary, kept
because it is the first iris number from real hardware. Compare with
`compose-phone-2026-09-06.md`, taken on the same phone with the same
fixture and gesture loop.
Reading it: the phone is 120 Hz (8.3 ms budget). `janky%` here counts
frames over 16.7 ms, so it is not Compose's `late` (over 8.3 ms). Like for
like: iris p50 6.2 ms vs Compose 7.7 ms; p90 32.0 vs 29.2; p99 42.1 vs
41.1. `cpu_p50=4.7ms` is iris's own per-frame CPU work on the phone,
against 0.2-0.4 ms on the emulator's x86 cores. Process CPU 15.6 s vs
20.9 s, but over a shorter run (692 frames vs 1613 -- iris only renders on
change and had no fling settle time), so per-second CPU is not directly
comparable; peak RSS 365 MB vs 587 MB. Battery current mean 563 mA vs
419 mA is the one figure that reads worse, and it is the least
comparable: 22 samples vs 39, over runs of different length and different
idle share. Bench v2's per-phase accounting is what makes these comparable.
```
iris bench report
frames=692 janky%=32.37 p50=6.2ms p90=32.0ms p99=42.1ms worst=52.6ms (measures redraw-start to after present() is called, not GPU/compositor completion) cpu_p50=4.7ms gpu_wait_p50=1.3ms (redraw-start-to-submit vs. submit-to-after-present)
scroll: 6 cycles (24 swipes), streamed 400/400 fixture events
process CPU time over this run: 15554ms
peak RSS: 365328kB
battery current: mean -563493µA over 22 samples (min -1807812, max -132812)
```
+66
View File
@@ -0,0 +1,66 @@
# iris bench v2 report from Iris's phone, 2026-09-06
Build 2e3f4ad (bench v2, fling physics, keyboard-wipe fix, dp unit), run
by Iris on her Pixel 9 Pro XL, verbatim. The display was at **120 Hz**
(8.3 ms budget) where `compose-phone-v2-2026-09-06.md` ran at 60 Hz, so
compare the millisecond percentiles, not `late`.
Side by side (Compose 60 Hz / iris 120 Hz, p50 / p90 / p99 ms): fling
5.5/8.7/11.6 vs 3.8/6.9/12.6; stream 13.4/31.7/42.5 vs 18.2/35.8/43.1;
type 7.3/13.2/16.5 vs 7.2/9.2/11.2; keyboard: iris could not show the IME
(phase invalid). Process CPU 69.6 s over 125 s vs 40.6 s over 150 s; peak
RSS 577 MB vs 379 MB; battery current mean 571 mA vs 452 mA.
Iris's observations on the same run: "the scrolling is not similar at
all. It does not fling for me yet [with a finger], and the test also seems
to give it a constant velocity and abruptly stop it at some point. Also
unsure what's going on in that image with the compaction" -- her
screenshot shows the `Compacted: 180000 -> 20000 tokens.` row drawn twice
overlapping, and once more below the composer bar: primitives of a
replaced/removed row surviving in the GPU buffers, the same shape as the
header drawn twice after a keyboard resize.
**Root-caused and fixed 2026-09-06** (commit `76b1f99`): the diagnosis in
that sentence was right and the location was not -- `UiRenderState::
draw_inner` read the `needs_redraw` mark without consuming it and skipped
the branch that frees a redrawn widget's old primitives. docs/RUST.md's
"Stale primitives, the phone's half" box has the full account, the guard
(`orphaned_primitives`, `debug_assert`ed every frame) and the emulator run
that exercises it.
```
iris bench report
per phase:
fling: 1783 frames over 53.2s
late: 104 (5.8%)
total p50 3.8ms p90 6.9ms p99 12.6ms
worst 29.1ms
stream: 401 frames over 21.3s
late: 306 (76.3%)
total p50 18.2ms p90 35.8ms p99 43.1ms
worst 43.8ms
type: 1202 frames over 65.7s
late: 309 (25.7%)
total p50 7.2ms p90 9.2ms p99 11.2ms
worst 15.3ms
keyboard: 9 frames over 9.7s
late: 9 (100.0%)
total p50 12.0ms p90 12.9ms p99 12.9ms
worst 12.9ms
frames:
3395 frames over 149.9s at 120Hz (8.3ms budget)
late: 728 (21.4%)
total p50 5.0ms p90 10.9ms p99 36.6ms
worst 43.8ms
cpu_p50 2.0ms gpu_wait_p50 2.6ms
bench:
fling: 8 flings out + 8 back at 12000px/s, travel start=idx=651/off=1217px outward=idx=651/off=101536px end=idx=651/off=1022px
scroll: 6 cycles (24 swipes, legacy tween), streamed 400/400 fixture events
type: 600 characters inserted then deleted, one per 50ms
keyboard: could not be shown (5 attempts, 0 confirmed visible)
process CPU time over this run: 40603ms
peak RSS: 379156kB
battery current: mean -452353µA over 149 samples (min -1753125, max -204687)
```
Binary file not shown.

After

Width:  |  Height:  |  Size: 110 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 198 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 72 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 294 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 14 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 98 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 114 KiB

+32
View File
@@ -150,6 +150,19 @@ pub enum Event {
ToolEnd {
id: String,
output: String,
/// Whether the tool reported that the call *failed*, from the
/// CLI's own `is_error` on the `tool_result`.
///
/// Added 2026-09-06 with the tool-call cards (RUST.md's P1b),
/// because without it a result is the only thing a card has and a
/// failed call is drawn as confidently as a successful one -- the
/// missing state, not a wrong one. `#[serde(default)]` so a
/// transcript written before this field, or a peer on an older
/// build, reads back as "not reported to have failed" rather than
/// failing to parse; that is the same claim the field's absence
/// used to make implicitly.
#[serde(default)]
is_error: bool,
},
/// An image the session produced or was sent, saved under the session
/// dir and referenced by id; the phone fetches it by URL.
@@ -305,6 +318,25 @@ pub enum Event {
/// it, which is why this is written down rather than left to be inferred
/// from a second example that does not exist.
Cleared,
/// The account behind this session has no quota left, so the turn stopped
/// without finishing.
///
/// Its own event rather than an [`Event::Error`] carrying the dialect's
/// sentence, because two things act on it that cannot read English: the
/// transcript draws it as a state the session is in rather than as a
/// failure of something it did, and `crate::resume` schedules the message
/// that picks the work back up. Recognising it belongs to the driver, which
/// is the only layer that knows its dialect's wording -- above here nothing
/// matches on strings.
///
/// `resets_at` is epoch seconds, and `None` is a real state: the dialect
/// said the limit was hit without saying when it lifts. Nothing here
/// invents one -- what the wait is actually decided against is the usage
/// endpoint, and this is the hint that starts the waiting.
LimitReached {
#[serde(default, skip_serializing_if = "Option::is_none")]
resets_at: Option<f64>,
},
Error {
message: String,
},
+704 -550
View File
File diff suppressed because it is too large. Load diff
+72 -11
View File
@@ -15,6 +15,11 @@ wgpu = { workspace = true }
image = { workspace = true }
accesskit = { workspace = true }
tokio = { workspace = true, features = ["sync", "rt", "rt-multi-thread"] }
# For diagnostics visible through android_logger (or whatever logger the
# app crate installs) -- this crate never installs one itself. Not in the
# android-only block below any more: the lines that matter most are in
# shared widget code, which the host backend compiles too.
log = "0.4.34"
# winit everywhere except Android; android-view (below) is what stands in
# for it there. Both backends live in this crate (see `src/android/mod.rs`'s
@@ -53,9 +58,25 @@ accesskit_android = "0.8.0"
# for `android/insets.rs`'s own id -> state map -- the same reason
# android-view's own `PEER_MAP` carries one.
send_wrapper = "0.6.0"
# For diagnostics visible through android_logger, wherever the app crate
# installs it -- this crate never installs a logger itself.
log = "0.4.28"
[features]
# RUST.md's I5 "Where iris's frame time goes" diagnosis: pins the
# `wgpu::Instance` to `Backends::GL` instead of `Backends::PRIMARY`, so one
# build can be measured on either backend. A compile-time feature rather
# than an env var because nothing on this machine can hand an env var to an
# already-launched Android process (there is no `am start` environment and
# no system-property reader here to add one).
#
# **Not needed to get GLES in the emulator**, whatever the history here
# says: the emulator's guest has no hardware Vulkan at all, so an ordinary
# build's runtime fallback lands on GLES by itself (docs/RUST.md, "What the
# emulator gives a GPU app"). Keeping the emulator on the same binary the
# phone runs is the point. What this feature is still for is forcing GLES
# on a machine that *does* have Vulkan -- the desktop -- which is why
# `default/render.rs` reads it too:
# ./run-headless.sh transcript --shot /tmp/x.png -- -p transcript-ui \
# --features iris/force-gles
force-gles = []
[dev-dependencies]
tokio = { workspace = true, features = ["sync", "rt", "rt-multi-thread", "time"] }
@@ -63,6 +84,9 @@ tokio = { workspace = true, features = ["sync", "rt", "rt-multi-thread", "time"]
# package is fine -- cargo excludes dev-dependencies from the graph used
# to build the library itself, so this only matters for `--examples`.
tabs-ui = { path = "tabs-ui" }
# `tests/mask_sdf.rs` only: the grid it hands the GPU and the coverages it
# reads back. wgpu and pollster are ordinary dependencies already.
bytemuck = { workspace = true }
# Plain Instant-timed binaries, not criterion -- see benches/message_list.rs's
# header for why. `harness = false` opts out of the unstable `#[bench]`
@@ -73,25 +97,62 @@ name = "message_list"
harness = false
[workspace]
members = ["core", "macro", "tabs-ui", "transcript-ui", "desktop-app"]
members = [
"core",
"macro",
"tabs-ui",
"transcript-ui",
"transcript-fixture",
"rig-input",
"desktop-app",
]
# android-app pulls in android-view, which needs the NDK sysroot to link
# -- excluded so `cargo build --workspace --all-targets` on the host stays
# buildable. Cross-compile it from its own directory (its own single-crate
# workspace, since it has no `[workspace]` table of its own and this
# exclusion stops it inheriting this one): `cd android-app && cargo ndk
# -t x86_64 -P 26 build`.
# -t x86_64 -P 29 build`.
exclude = ["android-app"]
[workspace.package]
version = "0.1.0"
edition = "2024"
# Debug info is the reason a `cargo test --workspace` here was taking half
# an hour, and it is worth the paragraph. Measured 2026-09-08: with rustc's
# default `debug = true`, linking this workspace's test binaries wrote
# **~54 GB** (one single test binary's linker wrote 16.9 GB) and left an
# **88 GB** `target/`. Eight test binaries each statically link the whole
# wgpu + naga + winit + parley graph, and at the default every one of them
# gets a full copy of that graph's DWARF written into it. On a btrfs at 83%
# full the linkers then sat in `handle_reserve_ticket` -- uninterruptible,
# waiting on space reservation -- at about 20 MB/s between them, which is
# what "the tests are slow" actually was. Not CPU: the machine was 87% idle
# throughout.
#
# `line-tables-only` keeps what is actually read from a backtrace -- the
# file and line of every frame, which is what a panicking test prints and
# what gdb needs to name the frames of a segfault. What it gives up is
# inspecting variables in a debugger; when that is wanted, ask for it on
# the command line for that one run rather than paying for it on every
# build:
#
# RUSTFLAGS="-C debuginfo=2" cargo test -p iris --test whatever
[profile.dev]
debug = "line-tables-only"
# The tests are what this is really for; `cargo test` uses `dev` for
# dependencies and `test` for the test targets themselves, so setting only
# `dev` leaves the eight big binaries at the default.
[profile.test]
debug = "line-tables-only"
[workspace.dependencies]
pollster = "0.4.0"
winit = "0.30.12"
wgpu = "28.0.0"
bytemuck = "1.23.1"
image = "0.25.6"
pollster = "1.0.1"
winit = "0.30.13"
wgpu = "30.0.1"
bytemuck = "1.25.2"
image = "0.25.10"
parley = "0.11.1"
swash = "0.2.10"
fxhash = "0.2.1"
@@ -99,7 +160,7 @@ arboard = "3.6.1"
accesskit = "0.25.0"
iris-core = { path = "core" }
iris-macro = { path = "macro" }
tokio = "1.49.0"
tokio = "1.53.1"
# Current stable as of 2026-09-05 (`cargo search`) -- I5's markdown block
# model, the same crate E2's uncommitted `e2-transcript` experiment used for
# the identical job (RUST.md), rather than reimplementing a CommonMark
+526 -148
View File
File diff suppressed because it is too large. Load diff
+64 -3
View File
@@ -17,13 +17,74 @@ crate-type = ["cdylib"]
[dependencies]
iris = { path = "../" }
tabs-ui = { path = "../tabs-ui" }
android-view = { git = "https://github.com/rust-mobile/android-view.git", rev = "bec6c62a96cef8239b0fd7fedeef9b184d02e3a1" }
android_logger = "0.15.0"
log = "0.4.28"
android_logger = "0.15.1"
log = "0.4.34"
# `tabs-screen` (default, I2/I4's demo) and `transcript-screen` (I5's
# Android integration) are mutually exclusive -- one `ActiveClient` type is
# compiled in, never both (`lib.rs`'s doc comment) -- so both sets of deps
# are optional and each screen's feature pulls in only its own. Without
# this, building `--features transcript-screen` alone (default features
# still on) left `tabs-ui` linked but never referenced under that cfg,
# which Cargo's `unused_dependencies` lint (on by default) correctly flags.
tabs-ui = { path = "../tabs-ui", optional = true }
transcript-ui = { path = "../transcript-ui", optional = true }
# P0's bench build only: the fixture and the folded screen both bench
# clients open, shared with the headless harness and the desktop window
# (docs/RUST.md's "Three test layers").
transcript-fixture = { path = "../transcript-fixture", optional = true }
client-core = { path = "../../client-core", optional = true }
event-model = { path = "../../event-model", optional = true }
serde_json = { version = "1", features = ["float_roundtrip"], optional = true }
# P0's bench build only (docs/RUST.md): `getrusage(RUSAGE_SELF)` for
# process CPU time, matching `libc::getrusage`'s mention in that box over
# parsing `/proc/self/stat` by hand and assuming `USER_HZ`. Already in the
# workspace's own dependency tree transitively (`iris/Cargo.lock`, pinned
# at 0.2.179) -- this makes it a direct dependency at the same version
# rather than a second, possibly-drifting resolution.
libc = { version = "0.2.189", optional = true }
# P0's bench build only: the scroll animation and the streaming phase are
# both a sequence of `sleep`s inside the async task `rsc.spawn_task` already
# runs on iris's own tokio runtime (`iris/src/task.rs`'s `Tasks::init`), and
# the battery sampler is a second, concurrent task on that same runtime
# (`tokio::spawn`) -- so this crate needs `tokio` directly rather than only
# through `iris`. `rt`+`time` only: no I/O, no macros, nothing this crate
# doesn't call. Version matches the one `iris`'s own dependency tree already
# resolves to (`iris/Cargo.lock`), so there is one copy of the runtime, not
# two.
tokio = { version = "1.53.1", features = ["rt", "time"], optional = true }
[features]
default = ["tabs-screen"]
tabs-screen = ["dep:tabs-ui"]
transcript-screen = ["dep:transcript-ui", "dep:client-core", "dep:event-model", "dep:serde_json"]
# RUST.md's I5 "Where iris's frame time goes": forces the GLES backend
# instead of SwiftShader's software Vulkan. See `iris/Cargo.toml`'s own doc
# on the feature this forwards to.
force-gles = ["iris/force-gles"]
# P0's iris half (docs/RUST.md, docs/AGENTS.md's "The rigs"): the same
# checked-in fixture, scroll loop and streaming phase the Compose `bench`
# build type drives, run here against `transcript-ui`'s real screen with no
# server. Depends on `transcript-screen` for `transcript-ui`/`client-core`/
# `event-model` -- `lib.rs`'s `ActiveClient` selection gives this feature
# priority over `transcript-screen`'s own `TranscriptClient` when both are
# listed, which is how this crate's build command names both explicitly.
bench = ["transcript-screen", "dep:transcript-fixture", "dep:libc", "dep:tokio"]
[profile.release]
panic = "abort"
# Measured 2026-09-07 (docs/RUST.md's "APK size" subsection): together these
# take libmain.so from 18,546,488 to 11,193,608 bytes (-39.7%) and the APK
# from 20,678,956 to 13,326,076 bytes (-35.5%), arm64-v8a release. `strip`
# also works around AGP's own stripReleaseDebugSymbols failing silently on
# this .so ("packaging them as they are"). `opt-level = "s"` over `"z"`:
# `z` measured another ~800 KB smaller but was not checked against iris's
# own frame-time bench, so it is not worth the unmeasured risk -- see the
# doc for the number and the follow-up this leaves.
strip = true
lto = "fat"
codegen-units = 1
opt-level = "s"
[profile.dev]
panic = "abort"
+61 -2
View File
@@ -13,15 +13,74 @@ android {
defaultConfig {
applicationId = "dev.iris.android.demo"
minSdk = 26
targetSdk = 34
// 29, not 26: `iris::android::view`'s touch handler dates each
// sample with `MotionEvent.getEventTimeNanos` and
// `getHistoricalEventTimeNanos`, both API 29, and a missing JNI
// method there is a hard crash on the first touch rather than a
// degraded fling. Raised deliberately rather than guarded at
// runtime: nothing this app is built for runs below 29, and an
// untested fallback path is its own defect. `build-apk.sh`'s
// `cargo ndk -P` is kept at the same number.
minSdk = 29
// 37, matching `compileSdk` and the Compose app in `app/` -- which
// is the one part of this that is measured rather than reasoned:
// that app targets 37 and its keyboard does push the transcript up
// on Iris's phone, and this one targeted 34 and does not
// (2026-09-07). The emulator here is API 36 and the push-up works
// there at either target, so the target is the only difference the
// two devices do not share.
//
// The mechanism, stated as the reading it is: below targetSdk 35
// a window keeps the legacy behaviour, where `adjustResize` shrinks
// the window for the IME and `getInsets(ime()).bottom` therefore
// measures the overlap with an already-shrunk window -- zero, with
// nothing left to push up. `MainActivity`'s
// `setDecorFitsSystemWindows(false)` opts out of that, and on API
// 36 it still takes; Android 16 deprecated it and Android 17 is
// where it appears not to. At 35+ edge-to-edge is not opt-in, so
// the app is handed the real overlap without relying on a
// deprecated call. If the phone still reports `ime_bottom=0` with
// a nonzero `dispatches` in the Diagnostics pane, this reading was
// wrong and the `WindowInsetsAnimation.Callback` in
// `MainActivity` is the other half to look at.
targetSdk = 37
versionCode = 1
versionName = "1.0"
}
// A release build must be signed, and the key is per machine rather than per repo -- same
// reasoning and the same key as `app/build-apk.sh` (the Compose app): it is what a phone
// recognises the app by, and a secret never lives in a checkout (the mount is shared with an
// untrusted VM). `build-apk.sh` generates this key once and points at it through the
// environment; without it a release build here is unsigned, which is fine for everything
// except installing.
def keystore = System.getenv("AI_APP_KEYSTORE")
signingConfigs {
if (keystore != null) {
release {
storeFile = file(keystore)
storePassword = System.getenv("AI_APP_KEYSTORE_PASSWORD")
keyAlias = "ai-app"
keyPassword = storePassword
}
}
}
buildTypes {
debug {
}
// P0's iris half (docs/RUST.md's P0 box): the build a phone actually runs. The `.so`
// itself is built separately with `cargo ndk --release --features "transcript-screen
// force-gles bench"` straight into src/main/jniLibs/ (this crate's own Cargo.toml) --
// Gradle here only packages and signs whatever is already there, the same division as the
// debug/tabs-screen build this project started with. `applicationIdSuffix` keeps it
// installable beside a debug build of the tabs demo rather than replacing it.
release {
applicationIdSuffix ".bench"
if (keystore != null) {
signingConfig = signingConfigs.release
}
}
}
compileOptions {
@@ -1,6 +1,15 @@
<?xml version="1.0" encoding="utf-8"?>
<manifest xmlns:android="http://schemas.android.com/apk/res/android">
<!-- Only needed by the transcript-screen feature (RUST.md's I5),
which talks to a real ai-server; the plain tabs demo (I2/I4) makes
no network call and never noticed this was missing. Absent,
UreqTransport::new's connect failed with EPERM (Operation not
permitted), not the ECONNREFUSED/ENETUNREACH a firewall or a dead
server would give: a seccomp-level socket denial reads nothing
like a network problem, which is what made it worth a comment. -->
<uses-permission android:name="android.permission.INTERNET" />
<application
android:allowBackup="true"
android:label="iris android-view demo"
@@ -15,8 +24,45 @@
<category android:name="android.intent.category.LAUNCHER" />
</intent-filter>
<!-- The enrollment link Dev Updater's Enroll button opens
(what `ai-server` mints), the same one the Compose app
in `app/` registers: which app answers it is the phone
owner's choice at the moment of the tap, and both being
offered is the intended behaviour rather than a clash.
BROWSABLE so a link tapped in another app reaches here,
and `android:host` so this app is not offered for every
aiapp:// URI a future route invents. -->
<intent-filter>
<action android:name="android.intent.action.VIEW" />
<category android:name="android.intent.category.DEFAULT" />
<category android:name="android.intent.category.BROWSABLE" />
<data android:scheme="aiapp" android:host="enroll" />
</intent-filter>
<meta-data android:name="android.app.lib_name" android:value="main" />
</activity>
<!-- This app's own recent log, for Dev Updater to read on the
phone. Iris runs these builds with no adb, and Android
forbids one app reading another's logcat, so this is the
only way a log::info! here reaches her. The shape is Dev
Updater's contract (its README.md, "An app's own log"), not
something invented for this app.
The authority carries ${applicationId}, so the bench package
and the ordinary one each get their own and neither can read
the other's log. Exported, because the whole point is
another app reading it, and guarded by a permission Dev
Updater declares at protectionLevel="normal" (a signature
permission is not available: the two apps are signed with
different locally generated keys). Read-only: insert,
update and delete throw. -->
<provider
android:name=".DevLogProvider"
android:authorities="${applicationId}.devlog"
android:exported="true"
android:readPermission="dev.updater.permission.READ_DEVLOG" />
</application>
</manifest>
@@ -0,0 +1,193 @@
package dev.iris.android.demo;
import android.content.ContentProvider;
import android.content.ContentValues;
import android.content.UriMatcher;
import android.database.Cursor;
import android.database.MatrixCursor;
import android.net.Uri;
/**
* This app's own recent log, exposed on the device.
*
* Iris runs these builds on a phone with no {@code adb}, and Android
* forbids one app reading another's {@code logcat} -- so nothing outside
* this process can recover what it wrote. The process already keeps a
* bounded copy of its log (Rust: {@code client_core::log_ring}); this
* hands it to Dev Updater, which is on the same phone, so it needs no
* tunnel, no token and no second enrolment.
*
* <p>The shape is <em>Dev Updater's contract</em>, not something invented
* here -- see that project's {@code README.md}, "An app's own log". Any
* app it delivers can implement the same and get the same Runtime tab.
* Two paths:
*
* <ul>
* <li>{@code lines?since=<seq>} -- every held line with a sequence at or
* after {@code since}, oldest first.
* <li>{@code status} -- one row: how many lines are held, how many the
* ring's own bound has dropped, and the newest sequence ({@code -1}
* for a log nothing has been written to, which is also how a reader
* notices this process restarted).
* </ul>
*
* <p>Read-only: there is nothing here for anyone else to change, so the
* three writing methods throw rather than silently doing nothing.
*
* <p>The authority is {@code <applicationId>.devlog}, filled in from
* Gradle so the bench build and the ordinary one each get their own and
* neither can read the other's. Read access is guarded by
* {@code dev.updater.permission.READ_DEVLOG}, declared in the manifest.
*
* <p>No {@code notifyChange}: the ring is filled by a {@code log::Log}
* backend on whatever thread logged, and giving that a way to reach a
* provider would mean plumbing a callback through {@code client-core} for
* every platform. Dev Updater polls while its tab is open, which its
* contract says it does precisely so implementing this stays cheap.
*/
public final class DevLogProvider extends ContentProvider {
static {
// The provider is created before any activity, so it cannot rely
// on MainActivity's own load. Loading twice is a no-op.
System.loadLibrary("main");
}
/** Matches {@link #nativeLinesSince}'s flat answer. Both sides say it once. */
private static final int FIELDS_PER_LINE = 5;
private static final String[] LINE_COLUMNS = {"seq", "t_ms", "level", "target", "message"};
private static final String[] STATUS_COLUMNS = {"held", "dropped", "newest_seq"};
private static final int LINES = 1;
private static final int STATUS = 2;
private UriMatcher matcher;
/** Every held line, {@link #FIELDS_PER_LINE} strings each, oldest first. */
private static native String[] nativeLinesSince(long since);
/** Three strings: held, dropped, newest sequence. */
private static native String[] nativeStatus();
/**
* Tells the Rust side which authority this build registered under, so
* the diagnostics pane can name somewhere a reader can actually query
* -- and so "declared but never created" is a state it can say. Only
* the provider knows it was instantiated; Android creates one lazily.
*
* <p>The files directory goes with it because <em>this is usually the
* only thing running</em>: after the app has died, Dev Updater's query
* starts the process for the provider alone, with no activity, so
* {@code MainActivity.nativeSetFilesDir} is never called and the line
* the panic hook left on disk is never replayed into the ring. That is
* exactly the run whose log somebody wants.
*/
private static native void nativeReady(String authority, String filesDir);
@Override
public boolean onCreate() {
// The authority is not a constant here: it is derived from this
// build's applicationId, so the bench package and the ordinary one
// do not share one. Read back from the manifest rather than
// recomposed, so there is one answer to what it is.
String authority = getContext().getPackageName() + ".devlog";
matcher = new UriMatcher(UriMatcher.NO_MATCH);
matcher.addURI(authority, "lines", LINES);
matcher.addURI(authority, "status", STATUS);
nativeReady(authority, getContext().getFilesDir().getAbsolutePath());
return true;
}
@Override
public Cursor query(
Uri uri,
String[] projection,
String selection,
String[] selectionArgs,
String sortOrder) {
switch (matcher.match(uri)) {
case LINES:
return lines(sinceOf(uri));
case STATUS:
return status();
default:
// Null rather than an exception: an unknown path is a
// reader asking for something this app does not have, and
// the contract's own answer for that is no cursor.
return null;
}
}
/**
* {@code ?since=} as a number, or 0 for a reader starting from the
* beginning. A value that is not a number is treated as 0 rather than
* refused -- what a caller wants from a malformed cursor is the log,
* not a stack trace about the query string.
*/
private static long sinceOf(Uri uri) {
String since = uri.getQueryParameter("since");
if (since == null) {
return 0;
}
try {
return Long.parseLong(since);
} catch (NumberFormatException ignored) {
return 0;
}
}
private static Cursor lines(long since) {
String[] fields = nativeLinesSince(since);
if (fields == null) {
return null;
}
MatrixCursor cursor = new MatrixCursor(LINE_COLUMNS, fields.length / FIELDS_PER_LINE);
for (int at = 0; at + FIELDS_PER_LINE <= fields.length; at += FIELDS_PER_LINE) {
cursor.addRow(
new Object[] {
Long.parseLong(fields[at]),
Long.parseLong(fields[at + 1]),
fields[at + 2],
fields[at + 3],
fields[at + 4],
});
}
return cursor;
}
private static Cursor status() {
String[] fields = nativeStatus();
if (fields == null || fields.length != STATUS_COLUMNS.length) {
return null;
}
MatrixCursor cursor = new MatrixCursor(STATUS_COLUMNS, 1);
cursor.addRow(
new Object[] {
Long.parseLong(fields[0]), Long.parseLong(fields[1]), Long.parseLong(fields[2]),
});
return cursor;
}
@Override
public String getType(Uri uri) {
// A MIME type is for something meant to be handed to another app
// as data; these rows are read by one reader that knows the
// columns. Saying nothing is the honest answer, not a gap.
return null;
}
@Override
public Uri insert(Uri uri, ContentValues values) {
throw new UnsupportedOperationException("this app's log is read-only");
}
@Override
public int update(Uri uri, ContentValues values, String selection, String[] selectionArgs) {
throw new UnsupportedOperationException("this app's log is read-only");
}
@Override
public int delete(Uri uri, String selection, String[] selectionArgs) {
throw new UnsupportedOperationException("this app's log is read-only");
}
}
@@ -1,6 +1,17 @@
package dev.iris.android.demo;
import android.app.Activity;
import android.content.ClipData;
import android.content.ClipboardManager;
import android.content.Context;
import android.view.Gravity;
import android.view.View;
import android.view.ViewGroup;
import android.widget.Button;
import android.widget.FrameLayout;
import android.widget.LinearLayout;
import android.widget.ScrollView;
import android.widget.TextView;
import org.linebender.android.rustview.RustView;
@@ -16,7 +27,7 @@ public final class IrisView extends RustView {
protected native long newViewPeer(Context context);
native void applyWindowInsetsNative(
long peer, int left, int top, int right, int bottom, int imeBottom);
long peer, int left, int top, int right, int bottom, int imeBottom, int imeVisible);
native void unregisterInsetsNative(long peer);
@@ -24,8 +35,9 @@ public final class IrisView extends RustView {
super(context);
}
void applyWindowInsets(int left, int top, int right, int bottom, int imeBottom) {
applyWindowInsetsNative(mViewPeer, left, top, right, bottom, imeBottom);
void applyWindowInsets(
int left, int top, int right, int bottom, int imeBottom, int imeVisible) {
applyWindowInsetsNative(mViewPeer, left, top, right, bottom, imeBottom, imeVisible);
}
@Override
@@ -33,4 +45,112 @@ public final class IrisView extends RustView {
unregisterInsetsNative(mViewPeer);
super.onDetachedFromWindow();
}
/**
* Called from the Rust side (iris/src/android/view.rs's
* `show_renderer_error`) when `AndroidRenderer::new` fails instead of
* drawing -- an ordinary instance method rather than a `native` one,
* since this call is Rust reaching into Java rather than the other
* direction. Replaces the whole activity content with plain,
* selectable, scrollable text rather than leaving the last frame (or a
* blank surface) on screen with no way to report what happened:
* UI_RULES.md's "a failure is reported where it happened, and says
* what to do next." No dialog and no styling beyond what is needed to
* read and copy the text -- this path exists for exactly the crash it
* replaces, so it must not depend on anything that could itself fail
* to render.
*/
void showRendererError(String report) {
Context context = getContext();
if (!(context instanceof Activity)) {
return;
}
Activity activity = (Activity) context;
TextView text = new TextView(activity);
text.setText(report);
text.setTextIsSelectable(true);
text.setGravity(Gravity.TOP | Gravity.START);
int pad = (int) (16 * activity.getResources().getDisplayMetrics().density);
text.setPadding(pad, pad, pad, pad);
ScrollView scroll = new ScrollView(activity);
scroll.addView(text);
activity.setContentView(scroll);
}
private static final String DIAGNOSTICS_OVERLAY_TAG = "iris-diagnostics-overlay";
/**
* The bench build's keyboard diagnostics capture
* (`bench_client.rs`'s `on_insets_changed` /
* `capture_keyboard_diagnostics`, via `bench_jni.rs`'s
* `PlatformHandle::show_diagnostics_overlay`): unlike
* `showRendererError` above, this adds a panel *over* this view
* (`MainActivity`'s `FrameLayout` still holds `IrisView` underneath,
* running) rather than replacing the activity's content, and gives it
* a Copy button and a Close that removes the panel -- so it draws
* (and can be read) whether or not iris itself is still putting
* anything on screen, without abandoning the session that produced
* it. Runs on the UI thread regardless of which thread calls it,
* since the call comes from a background task (a delayed capture
* after the keyboard opens), and touching the view tree off the UI
* thread is undefined.
*/
void showDiagnosticsOverlay(String report) {
Context context = getContext();
if (!(context instanceof Activity)) {
return;
}
Activity activity = (Activity) context;
activity.runOnUiThread(() -> {
ViewGroup parent = (ViewGroup) getParent();
if (parent == null) {
return;
}
View existing = parent.findViewWithTag(DIAGNOSTICS_OVERLAY_TAG);
if (existing != null) {
parent.removeView(existing);
}
float density = activity.getResources().getDisplayMetrics().density;
int pad = (int) (16 * density);
LinearLayout overlay = new LinearLayout(activity);
overlay.setTag(DIAGNOSTICS_OVERLAY_TAG);
overlay.setOrientation(LinearLayout.VERTICAL);
overlay.setBackgroundColor(0xEE000000);
overlay.setPadding(pad, pad, pad, pad);
TextView text = new TextView(activity);
text.setText(report);
text.setTextIsSelectable(true);
text.setTextColor(0xFFFFFFFF);
ScrollView scroll = new ScrollView(activity);
scroll.addView(text);
overlay.addView(scroll, new LinearLayout.LayoutParams(
LinearLayout.LayoutParams.MATCH_PARENT, 0, 1f));
LinearLayout buttonRow = new LinearLayout(activity);
buttonRow.setOrientation(LinearLayout.HORIZONTAL);
buttonRow.setPadding(0, pad, 0, 0);
Button copy = new Button(activity);
copy.setText("Copy");
copy.setOnClickListener(v -> {
ClipboardManager clipboard =
(ClipboardManager) activity.getSystemService(Context.CLIPBOARD_SERVICE);
if (clipboard != null) {
clipboard.setPrimaryClip(ClipData.newPlainText("iris diagnostics", report));
}
});
Button close = new Button(activity);
close.setText("Close");
close.setOnClickListener(v -> parent.removeView(overlay));
buttonRow.addView(copy);
buttonRow.addView(close);
overlay.addView(buttonRow);
parent.addView(overlay, new FrameLayout.LayoutParams(
FrameLayout.LayoutParams.MATCH_PARENT, FrameLayout.LayoutParams.MATCH_PARENT));
});
}
}
@@ -1,10 +1,14 @@
package dev.iris.android.demo;
import android.app.Activity;
import android.content.Intent;
import android.net.Uri;
import android.os.Build;
import android.os.Bundle;
import android.view.WindowInsets;
import android.view.WindowInsetsAnimation;
import android.widget.FrameLayout;
import java.util.List;
/**
* The android-view backend's demo activity (RUST.md's I2): one IrisView
@@ -18,9 +22,28 @@ public final class MainActivity extends Activity {
System.loadLibrary("main");
}
/**
* The app's private directory, where the Rust side keeps its enrollment
* (`src/enrollment.rs`). Handed over before the view is built, because
* the client the view creates reads the enrollment as it starts.
*/
private static native void nativeSetFilesDir(String path);
/**
* One `aiapp://enroll?host=&port=&token=&ca=` link, as Dev Updater's
* Enroll button opens it. Parsed and stored on the Rust side, which is
* where the enrollment lives for the desktop app too -- nothing about
* the link's format is known here.
*/
private static native void nativeEnroll(String uri);
@Override
public void onCreate(Bundle state) {
super.onCreate(state);
// Before the view: creating it starts the Rust client, which asks
// straight away which server it is enrolled with.
nativeSetFilesDir(getFilesDir().getAbsolutePath());
handleEnrollmentIntent(getIntent());
IrisView view = new IrisView(this);
view.setLayoutParams(new FrameLayout.LayoutParams(
FrameLayout.LayoutParams.MATCH_PARENT, FrameLayout.LayoutParams.MATCH_PARENT));
@@ -31,17 +54,136 @@ public final class MainActivity extends Activity {
setContentView(layout);
view.requestFocus();
// RUST.md's P0 box, defect 4 ("keyboard: could not be shown"):
// `logcat` showed the platform's own IME open/resize happening
// while `setOnApplyWindowInsetsListener` fired only once, at
// attach, and never again for a pure keyboard toggle -- a plain
// (non-edge-to-edge) window is only guaranteed that one initial
// dispatch; `adjustResize` handling the IME entirely by resizing
// the window is not itself a trigger for a fresh one. Opting into
// edge-to-edge (a platform call, API 30+, no new dependency) is
// what makes the system redeliver insets on every change,
// including the ones this activity actually cares about --
// `getSystemWindowInset*` below is unaffected by this (it has
// always reported the raw system-bar/IME overlap regardless of
// who consumes it), so the on-screen bars and the padding Rust
// already derives from those four numbers are unchanged; only the
// callback's firing became reliable.
if (Build.VERSION.SDK_INT >= Build.VERSION_CODES.R) {
getWindow().setDecorFitsSystemWindows(false);
}
// **The keyboard's height arrives twice, over two different
// paths, and the phone needs the second one** (Iris, 2026-09-07:
// the emulator pushed the composer up and her Pixel did not).
// `setOnApplyWindowInsetsListener` is the platform's *settled*
// answer; `WindowInsetsAnimation.Callback` is the running one, and
// an IME that animates in delivers every intermediate height
// through the callback with the static dispatch arriving only at
// the ends -- on some devices only at `onEnd`. Registering both
// means neither device depends on the other's timing, and it is
// also what makes the push-up *animate* with the keyboard rather
// than jump when it lands.
//
// The two do not disagree, because they are the same call with the
// same numbers read out of whichever `WindowInsets` is current.
// `DISPATCH_MODE_CONTINUE_ON_SUBTREE` so this view consuming
// nothing keeps the ordinary dispatch running underneath.
// `onEnd` re-reads the root's insets rather than trusting the last
// `onProgress`: an animation interrupted mid-flight never delivers
// its final frame, which is exactly the fault the Compose app hit
// (AGENTS.md, "the composer can get stuck floating above the
// bottom of the screen").
if (Build.VERSION.SDK_INT >= Build.VERSION_CODES.R) {
view.setWindowInsetsAnimationCallback(new WindowInsetsAnimation.Callback(
WindowInsetsAnimation.Callback.DISPATCH_MODE_CONTINUE_ON_SUBTREE) {
@Override
public WindowInsets onProgress(
WindowInsets insets, List<WindowInsetsAnimation> running) {
sendInsets(view, insets);
return insets;
}
@Override
public void onEnd(WindowInsetsAnimation animation) {
WindowInsets settled = view.getRootWindowInsets();
if (settled != null) {
sendInsets(view, settled);
}
}
});
}
view.setOnApplyWindowInsetsListener((v, insets) -> {
sendInsets((IrisView) v, insets);
return insets;
});
}
/**
* A link that arrives while the activity is already up. `singleTop` is
* not set, so this is the resumed case only -- the fresh-launch case
* goes through `onCreate`'s `getIntent`. `setIntent` so a later
* `getIntent` reports the one actually being acted on rather than the
* one this activity started with.
*/
@Override
protected void onNewIntent(Intent intent) {
super.onNewIntent(intent);
setIntent(intent);
handleEnrollmentIntent(intent);
}
/**
* Hands a VIEW intent's URI to the Rust side, which decides whether it
* is an enrollment link -- the scheme is checked here only so a launch
* intent (which carries no data) costs nothing.
*/
private static void handleEnrollmentIntent(Intent intent) {
if (intent == null) {
return;
}
Uri data = intent.getData();
if (data != null) {
nativeEnroll(data.toString());
}
}
/** Read one `WindowInsets` and hand it to the Rust side. The only
* place that reads these fields, so the static dispatch and the
* animation callback above cannot come to report different things. */
private static void sendInsets(IrisView view, WindowInsets insets) {
int left = insets.getSystemWindowInsetLeft();
int top = insets.getSystemWindowInsetTop();
int right = insets.getSystemWindowInsetRight();
int bottom = insets.getSystemWindowInsetBottom();
// **Two separate answers, because they are separate questions**
// (Iris's phone, 2026-09-06: "message box does not push up the
// scroll area"). `isVisible(ime())` says whether the keyboard is
// up; `getInsets(ime()).bottom` says how tall it is. An earlier
// pass sent the boolean *as* the height (0 or 1) because under
// plain `adjustResize` the window shrinks to make room and the ime
// inset therefore measures a zero overlap by construction -- true
// then, and no longer true now that this is an edge-to-edge window
// (`targetSdk` 35+, plus the `setDecorFitsSystemWindows` call
// above for the devices below that), which is exactly the case
// where the system stops resizing and hands the app the real
// overlap instead. Sending 1 for it left the Rust side padding the
// composer by one physical pixel, so the keyboard covered the bar
// and the transcript alike.
//
// The visibility is still sent in its own right rather than
// inferred from `height > 0`: the two disagree during the
// keyboard's slide-in and -out (visible, height still climbing),
// and "is the IME up" drives the bench's own state machine
// (`bench_client.rs`'s `ime_state`) where a half-open frame
// reading as "closed" is a miscount.
int imeBottom = 0;
int imeVisible = 0;
if (Build.VERSION.SDK_INT >= Build.VERSION_CODES.R) {
imeBottom = insets.getInsets(WindowInsets.Type.ime()).bottom;
imeVisible = insets.isVisible(WindowInsets.Type.ime()) ? 1 : 0;
}
((IrisView) v).applyWindowInsets(left, top, right, bottom, imeBottom);
return insets;
});
view.applyWindowInsets(left, top, right, bottom, imeBottom, imeVisible);
}
}
+117
View File
@@ -0,0 +1,117 @@
#!/bin/sh
# Builds iris-android-app end to end: the cdylib (cargo ndk, straight into
# app/src/main/jniLibs/) then the APK (Gradle). Written to stop re-typing
# the same incantation by hand every time (ANDROID_HOME/NDK exports, the
# cargo ndk invocation, the keystore env for a release build, apksigner/
# aapt2 verification) -- see docs/RUST.md's P0 box. Same shape as `app/
# build-apk.sh` (the Compose app's own build script) and `app/
# iris-scroll.sh` (no coordinates, set -eu, exit 0 on success).
#
# Usage: ./build-apk.sh [debug|release] [--abi arm64-v8a|x86_64] [--features "a b c"]
# debug/release default to debug (matches this-machine-android's "the
# emulator stays on debug" rule -- pass `release` explicitly for a phone
# build). --abi defaults to arm64-v8a (a phone/real device); pass
# x86_64 for this checkout's own AVD. --features defaults to
# "transcript-screen bench" -- deliberately *without* `force-gles`, and
# nothing should add it back for the emulator's sake.
#
# **The emulator does not need a GLES build, because it has no hardware
# Vulkan to be steered away from** (docs/RUST.md, "What the emulator
# gives a GPU app", 2026-09-08): its guest's only Vulkan is SwiftShader
# in software, its GLES is the host's real GPU through virgl, and iris's
# own runtime fallback -- `Backends::PRIMARY`, no adapter, rebuild on
# `Backends::GL` -- takes an ordinary build there by itself. So the
# emulator and the phone run the *same binary* and differ only in what
# that binary finds, which is the whole point: a build flag that changed
# the backend would mean the thing measured here is not the thing
# shipped.
#
# `force-gles` (`iris/Cargo.toml`'s own doc) pins the backend at compile
# time for a backend-isolation measurement (RUST.md's I5, "Where iris's
# frame time goes"), and the desktop is the better place to run it now
# (`run-headless.sh ... --features iris/force-gles`). It was never meant
# to reach a real device, but this script's old default put it in every
# arm64 build regardless, so the P0 bench APK delivered to Iris's phone
# forced GLES there too -- the named hypothesis in RUST.md's P0 box
# ("iris bench crash on the phone, 2026-09-06"). Never pass it for a
# build meant for a phone.
set -eu
cd "$(dirname "$0")"
BUILD_TYPE="debug"
ABI="arm64-v8a"
FEATURES="transcript-screen bench"
case "${1:-}" in
debug|release) BUILD_TYPE="$1"; shift ;;
esac
while [ $# -gt 0 ]; do
case "$1" in
--abi) ABI="$2"; shift 2 ;;
--features) FEATURES="$2"; shift 2 ;;
*) echo "build-apk.sh: unknown argument: $1" >&2; exit 1 ;;
esac
done
SDK_ROOT="$HOME/Android/Sdk"
export ANDROID_HOME="$SDK_ROOT"
export ANDROID_SDK_ROOT="$SDK_ROOT"
NDK_DIR=$(ls -d "$SDK_ROOT"/ndk/*/ 2>/dev/null | sort -V | tail -1)
if [ -z "$NDK_DIR" ]; then
echo "build-apk.sh: no NDK found under $SDK_ROOT/ndk" >&2
exit 1
fi
export ANDROID_NDK_HOME="$NDK_DIR"
# Only the ABI asked for goes into the APK. cargo ndk adds its output beside
# whatever earlier builds left here, and Gradle packages every directory it
# finds -- a debug x86_64 emulator build left behind made an arm64 "release"
# 339 MB on 2026-09-06.
rm -rf app/src/main/jniLibs
# ...and Gradle's own copy of them, which `rm -rf jniLibs` does not reach.
# `mergeReleaseNativeLibs` is *up to date* against its cached inputs, so a
# build that switches ABI packages the previous ABI: an `--abi x86_64`
# release APK containing `lib/arm64-v8a/libmain.so` installed fine and
# aborted at startup with `Could not get adapter!: NotFound {
# active_backends: VULKAN }` under libndk_translation -- which reads
# exactly like the phone's own Vulkan problem and is nothing of the kind.
# Scoped to the merge task's directory rather than all of `app/build`, so
# an ABI change costs the native merge and not the whole Gradle build.
rm -rf app/build/intermediates/merged_native_libs \
app/build/intermediates/stripped_native_libs \
app/build/intermediates/merged_jni_libs
echo "build-apk.sh: cargo ndk -t $ABI build ${BUILD_TYPE:+(${BUILD_TYPE})} --features \"$FEATURES\""
if [ "$BUILD_TYPE" = "release" ]; then
cargo ndk -t "$ABI" -P 29 -o app/src/main/jniLibs/ build --release --features "$FEATURES"
else
cargo ndk -t "$ABI" -P 29 -o app/src/main/jniLibs/ build --features "$FEATURES"
fi
GRADLE_TASK="assembleDebug"
APK_DIR="app/build/outputs/apk/debug"
APK_NAME="app-debug.apk"
if [ "$BUILD_TYPE" = "release" ]; then
GRADLE_TASK="assembleRelease"
APK_DIR="app/build/outputs/apk/release"
APK_NAME="app-release.apk"
# Same key `app/build-apk.sh` (the Compose app) generates once under
# ~/.config/ai-app/release.jks -- see AGENTS.md's "Checking your work".
export AI_APP_KEYSTORE="$HOME/.config/ai-app/release.jks"
if [ ! -f "$AI_APP_KEYSTORE" ]; then
echo "build-apk.sh: no release key at $AI_APP_KEYSTORE -- run app/build-apk.sh once first" >&2
exit 1
fi
export AI_APP_KEYSTORE_PASSWORD
AI_APP_KEYSTORE_PASSWORD=$(cat "$AI_APP_KEYSTORE.password")
fi
gradle ":app:$GRADLE_TASK" --console=plain
APK_PATH="$(pwd)/$APK_DIR/$APK_NAME"
BUILD_TOOLS=$(ls -d "$SDK_ROOT"/build-tools/*/ | sort -V | tail -1)
echo "--- aapt2 dump badging ---"
"${BUILD_TOOLS}aapt2" dump badging "$APK_PATH" | head -5
if [ "$BUILD_TYPE" = "release" ]; then
echo "--- apksigner verify ---"
"${BUILD_TOOLS}apksigner" verify --print-certs "$APK_PATH"
fi
echo "$APK_PATH"
+89
View File
@@ -0,0 +1,89 @@
#!/bin/sh
# Installs and runs the iris `bench` build on this checkout's own emulator
# (per this-machine-android's per-checkout-AVD rule; `emu serial` picks it)
# and prints the report -- the iris half of `app/transcript-bench.sh`'s
# job. No coordinates: the button is found by its accessibility label
# through `ui-trace`, per AGENTS.md's "Driving the UI".
#
# Usage: ./run-bench.sh [--apk PATH]
# Defaults to this checkout's own release APK
# (app/build/outputs/apk/release/app-release.apk) if it exists, else the
# debug one -- build one first with ./build-apk.sh.
set -eu
cd "$(dirname "$0")"
APK=""
while [ $# -gt 0 ]; do
case "$1" in
--apk) APK="$2"; shift 2 ;;
*) echo "run-bench.sh: unknown argument: $1" >&2; exit 1 ;;
esac
done
if [ -z "$APK" ]; then
if [ -f app/build/outputs/apk/release/app-release.apk ]; then
APK=app/build/outputs/apk/release/app-release.apk
else
APK=app/build/outputs/apk/debug/app-debug.apk
fi
fi
if [ ! -f "$APK" ]; then
echo "run-bench.sh: no APK at $APK -- run ./build-apk.sh first" >&2
exit 1
fi
SERIAL=$(emu serial)
PKG=$(aapt2 dump badging "$APK" 2>/dev/null | sed -n "s/^package: name='\\([^']*\\)'.*/\\1/p")
if [ -z "$PKG" ]; then
BUILD_TOOLS=$(ls -d "$HOME"/Android/Sdk/build-tools/*/ | sort -V | tail -1)
PKG=$("${BUILD_TOOLS}aapt2" dump badging "$APK" | sed -n "s/^package: name='\\([^']*\\)'.*/\\1/p")
fi
echo "run-bench.sh: installing $APK ($PKG) on $SERIAL"
adb -s "$SERIAL" install -r "$APK" >/dev/null
adb -s "$SERIAL" shell am force-stop "$PKG"
adb -s "$SERIAL" logcat -c
adb -s "$SERIAL" shell am start -n "$PKG/dev.iris.android.demo.MainActivity" >/dev/null
ui-trace record -s "$SERIAL" -d 3000 --do "tap 'Run benchmark'" -o /tmp/run-bench-tap.txt >/dev/null
# Which adapter drew, before any number is printed. The emulator is a GLES
# machine -- its guest has no hardware Vulkan (docs/RUST.md, "What the
# emulator gives a GPU app") -- so iris's runtime fallback lands on `Gl`,
# and `Gl (... virgl ...)` is the host's real GPU while `Gl (...
# SwiftShader ...)` is the CPU. Those two produce frame times an order of
# magnitude apart and are otherwise indistinguishable in this report, so
# the line is printed rather than left in logcat for somebody to think of.
ADAPTER=$(adb -s "$SERIAL" logcat -d -s iris-android-app:I 2>/dev/null \
| sed -n 's/.*\(iris renderer: .*\)/\1/p' | tail -1)
if [ -n "$ADAPTER" ]; then
echo "run-bench.sh: $ADAPTER"
else
echo "run-bench.sh: no 'iris renderer:' line in logcat -- cannot say what drew this run" >&2
fi
# Poll for the report line rather than a fixed sleep -- the run itself is
# a fixed script (RUST.md's "Benchmark v2": 16 flings, a 20s streaming
# phase, ~61s of typing, 10s of keyboard toggles, roughly 2.5 minutes end
# to end) but device speed varies. 260s cap rather than v1's 90s -- v2 is
# a longer script than v1's swipe-loop-only run.
# The report's own first line, not the bare "iris bench report:" prefix:
# `copy_report` logs that prefix too ("nothing to copy -- run the benchmark
# first", which the app emits at startup), so polling for the prefix
# returned instantly and the script printed a report that was never run.
REPORT_LINE="iris bench report: iris bench report"
i=0
while [ "$i" -lt 260 ]; do
LINE=$(adb -s "$SERIAL" logcat -d -s iris-android-app:I 2>/dev/null | grep "$REPORT_LINE" || true)
if [ -n "$LINE" ]; then
break
fi
i=$((i + 1))
sleep 1
done
if [ -z "$LINE" ]; then
echo "run-bench.sh: no report after 260s -- check logcat by hand" >&2
exit 1
fi
# -A 60 rather than v1's -A 6 -- v2's report has a per-phase block (four
# phases, four lines each) on top of the frames/bench sections v1 had.
adb -s "$SERIAL" logcat -d -s iris-android-app:I | grep -A 60 "$REPORT_LINE"
+186
View File
@@ -0,0 +1,186 @@
//! The platform half of this app's logging: what
//! `client_core::log_ring` needs that only Android can supply, which is
//! `android_logger` as the logger to forward to and nothing else.
//!
//! Everything general -- the ring, its bounds, the `log::Log` backend --
//! is in `client-core`, shared with the desktop app (AGENTS.md's sharing
//! rule).
//!
//! **Why an app carries its own log at all**: Iris tests these builds on a
//! GrapheneOS phone with no `adb`, and Android forbids one app reading
//! another's `logcat`. Nothing outside this process can recover what it
//! wrote, so the process keeps a copy -- and hands it to Dev Updater on
//! the same phone through `devlog`'s `ContentProvider`. See
//! `docs/DECISIONS.md`, 2026-09-07.
use client_core::log_ring::{self, LogRing};
/// Installs the ring in front of `android_logger`, so `logcat` still sees
/// exactly what it saw before and the ring sees it too.
///
/// Called once, from `JNI_OnLoad`. A second call is refused by `log`
/// itself; the message says which caller, since two initialisation paths
/// is a programmer error rather than something to recover from.
pub fn install(max_level: log::LevelFilter) {
let inner = android_logger::AndroidLogger::new(
android_logger::Config::default()
.with_max_level(max_level)
.with_tag("iris-android-app"),
);
if log_ring::install_process_logger(
Box::new(inner),
max_level,
iris::diagnostics::trace_enabled,
)
.is_err()
{
// Not a panic: a logger already installed means logging works,
// just without the ring, and taking the app down over a
// diagnostic would be worse than the diagnostic being missing.
// The line goes through whatever logger did win.
log::warn!("iris app log: a logger was already installed, so there is no ring");
}
install_panic_hook();
}
/// The process's ring -- what `Copy report` appends, what the diagnostics
/// pane counts, and what `devlog`'s provider hands to Dev Updater.
pub fn ring() -> &'static LogRing {
log_ring::process_ring()
}
/// Only the bench build has a diagnostics pane to put this in; the
/// transcript build's screen is the app's own and has no room for a
/// readout. Gated rather than left dead so the build stays warning-clean.
#[cfg(feature = "bench")]
/// Two lines for the diagnostics pane: how much of this app's log is held,
/// and where it can be read from.
///
/// The second names the provider's authority rather than saying "logging
/// is on", so a screenshot of this pane is enough to tell whether the
/// contract is live and which package's log it is -- the bench build and
/// the ordinary one have different ones.
pub fn diagnostics_line() -> String {
let where_to_read = match crate::devlog::authority() {
Some(authority) => format!("devlog provider: content://{authority}"),
// Not "off": Android creates a provider lazily, so this is what
// "nobody has asked for it yet" looks like, and it is a different
// thing from a build that does not have one.
None => "devlog provider: declared, not created yet".to_string(),
};
format!("{}\n{where_to_read}", ring().summary())
}
/// Where the panic hook leaves its report, under the app's private
/// directory. Read back and dropped by [`set_crash_dir`] on the next
/// start.
const CRASH_FILE: &str = "last-panic.txt";
/// How many of the dying run's own log lines the panic hook saves with
/// the panic, and [`set_crash_dir`] replays.
///
/// The panic's message and location say *what* broke; these say what the
/// app was doing on the way there, which is the half that is otherwise
/// unrecoverable -- the ring is memory only, so an abort takes every line
/// before the panic with it. Bounded rather than the whole ring because
/// this is written by a hook on a process that is about to die, and
/// because the replay pushes each line into the new run's ring, where an
/// unbounded paste would evict the run that is actually being watched.
const CRASH_CONTEXT_LINES: usize = 80;
/// The target the replayed context lines carry, so a reader can tell a
/// line from the run that died from one this run wrote. They keep their
/// original timestamp and level inside the text, which is why the level
/// they are re-pushed at is not meaningful and the target has to be.
const PREVIOUS_RUN_TARGET: &str = "previous_run";
static CRASH_PATH: std::sync::OnceLock<std::path::PathBuf> = std::sync::OnceLock::new();
/// Installs a `log`-level panic hook, so a panic's message and location
/// reach the ring and `logcat` rather than only the tombstone.
///
/// **Why this is needed at all**: these builds are `panic = "abort"`
/// (`Cargo.toml`), and the default hook writes to `stderr` plus
/// `android_set_abort_message` -- the crash report. Iris runs these on a
/// phone with no `adb`, so the crash report is exactly the surface she
/// cannot read, and an `assert!` that fired said nothing anywhere she
/// could see it. Routing it through `log::error!` puts it in front of
/// `android_logger` *and* in the ring `devlog`'s provider hands to Dev
/// Updater.
///
/// The ring is memory only, so after an abort the process that holds it
/// is gone -- hence the file half. [`set_crash_dir`] replays it.
fn install_panic_hook() {
let previous = std::panic::take_hook();
std::panic::set_hook(Box::new(move |info| {
let where_at = match info.location() {
Some(at) => format!("{}:{}:{}", at.file(), at.line(), at.column()),
None => "an unknown location".to_string(),
};
// `info`'s own `Display` repeats the location and a newline;
// the payload alone keeps this to the one line the ring wants.
let message = info.payload_as_str().unwrap_or("Box<dyn Any>");
let line = format!("iris panic at {where_at}: {message}");
log::error!("{line}");
if let Some(path) = CRASH_PATH.get() {
// The panic line first, then what the app was doing before
// it: one file, split again on that first newline by
// `set_crash_dir`.
let context = ring()
.try_tail_text(CRASH_CONTEXT_LINES)
// Said rather than left empty, so "the ring was locked as
// we died" cannot be read as "nothing had been logged".
.unwrap_or_else(|| {
"(the log ring was locked as this run died; no context)".to_string()
});
// Best effort by design: a panic is already the failure, and
// failing to record it must not become a second one.
let _ = std::fs::write(path, format!("{line}\n{context}"));
}
previous(info);
}));
}
/// Tells the panic hook where to leave its report, and replays the report
/// a previous run left there into the ring before deleting it.
///
/// Called from **both** `MainActivity.nativeSetFilesDir` and
/// `DevLogProvider.nativeReady` -- whichever of the two runs first in
/// this process, since after a crash Dev Updater's query starts the
/// process for the provider alone and no activity ever runs. Safe to call
/// twice: the file is gone after the first, so the second finds nothing
/// and says nothing. The panic itself is replayed at `error` level and
/// says it is from the previous run, so a crash loop shows the reason it
/// is looping in the Runtime tab of the run that is still up.
pub fn set_crash_dir(dir: &std::path::Path) {
let path = dir.join(CRASH_FILE);
if let Ok(previous) = std::fs::read_to_string(&path) {
// Delete before replaying rather than after: a replay that itself
// panicked would otherwise leave the file to be replayed again on
// every start, and a crash loop nothing can get out of is worse
// than one report lost.
let _ = std::fs::remove_file(&path);
replay_crash(&previous);
}
let _ = CRASH_PATH.set(path);
}
/// Puts a previous run's report back in the ring: its context lines in
/// the order they happened, then the panic itself.
///
/// Chronological, so the Runtime tab reads as one story -- the lines that
/// led to the crash, then the crash, then this run. The context goes in
/// through `LogRing::push` rather than through `log::info!` so it is not
/// stamped with this run's clock: each line already carries the time and
/// level it was written at, and [`PREVIOUS_RUN_TARGET`] is what says
/// whose run it was.
fn replay_crash(report: &str) {
let (panic_line, context) = report.split_once('\n').unwrap_or((report, ""));
for line in context.lines().filter(|line| !line.is_empty()) {
ring().push(log::Level::Info, PREVIOUS_RUN_TARGET, line.to_string());
}
log::error!(
"iris app log: the previous run died -- {}",
panic_line.trim()
);
}
File diff suppressed because it is too large. Load diff
+252
View File
@@ -0,0 +1,252 @@
//! JNI calls the `bench` feature needs that go through the shell's own
//! Java side rather than anything `iris`/`android-view` already wraps:
//! `BatteryManager.getIntProperty(BATTERY_PROPERTY_CURRENT_NOW)` for the
//! per-second battery sample, `ClipboardManager.setPrimaryClip` for the
//! "Copy report" control (P0's iris half, docs/RUST.md), and -- added for
//! RUST.md's "Benchmark v2" -- `Display.getRefreshRate()` for the phase
//! report's real late-frame budget and `InputMethodManager.
//! showSoftInput`/`hideSoftInputFromWindow` for the keyboard phase. None
//! of these are part of `android_view::context`'s own `Context`/
//! `Resources` wrappers (that file's own `// TODO: more methods?`), so
//! this calls them directly rather than growing that crate's wrapper for
//! calls this crate alone needs.
//!
//! Holds its own `JavaVM` + `GlobalRef` to the view (handed in through
//! [`iris::android::AndroidAppState::platform_ready`]) so it can attach
//! whichever thread calls it -- the battery sampler runs on a background
//! tokio task, not the UI thread the rest of `IrisViewPeer`'s JNI calls
//! run on. `JavaVM::attach_current_thread` is safe to call from a thread
//! already attached (the `jni` crate detects it and does not double
//! attach), so no caller here needs to know or care which thread it is.
use android_view::jni::{
JNIEnv, JavaVM,
objects::{GlobalRef, JObject, JValue},
};
/// `android.os.BatteryManager.BATTERY_PROPERTY_CURRENT_NOW` -- not exposed
/// as a constant anywhere reachable without the Android SDK jar, so named
/// here with its source rather than left as a bare `2`.
const BATTERY_PROPERTY_CURRENT_NOW: i32 = 2;
pub struct PlatformHandle {
vm: JavaVM,
view: GlobalRef,
}
impl PlatformHandle {
pub fn new(vm: JavaVM, view: GlobalRef) -> Self {
Self { vm, view }
}
fn context<'e>(&self, env: &mut JNIEnv<'e>) -> Option<JObject<'e>> {
env.call_method(
self.view.as_obj(),
"getContext",
"()Landroid/content/Context;",
&[],
)
.ok()?
.l()
.ok()
}
fn system_service<'e>(
&self,
env: &mut JNIEnv<'e>,
context: &JObject<'e>,
name: &str,
) -> Option<JObject<'e>> {
let jname = env.new_string(name).ok()?;
env.call_method(
context,
"getSystemService",
"(Ljava/lang/String;)Ljava/lang/Object;",
&[JValue::Object(jname.as_ref())],
)
.ok()?
.l()
.ok()
}
/// One sample of `BATTERY_PROPERTY_CURRENT_NOW`, in microamps. `None`
/// on any JNI failure, on a device with no `BatteryManager` service,
/// or when the platform itself answers "not supported" -- `0` or
/// `Integer.MIN_VALUE` are both documented SDK answers for that, and
/// both would read as a real (and wrong) measurement if folded into an
/// average rather than named apart. UI_RULES.md: never present an
/// inferred value as a measured one.
pub fn battery_current_ua(&self) -> Option<i32> {
let mut guard = self.vm.attach_current_thread().ok()?;
let env: &mut JNIEnv = &mut guard;
let context = self.context(env)?;
let battery_manager = self.system_service(env, &context, "batterymanager")?;
let value = env
.call_method(
&battery_manager,
"getIntProperty",
"(I)I",
&[JValue::Int(BATTERY_PROPERTY_CURRENT_NOW)],
)
.ok()?
.i()
.ok()?;
if value == 0 || value == i32::MIN {
None
} else {
Some(value)
}
}
/// Puts `text` on the system clipboard through `ClipboardManager` --
/// `true` only if the whole JNI chain (service lookup, `ClipData`,
/// `setPrimaryClip`) succeeded.
pub fn copy_to_clipboard(&self, label: &str, text: &str) -> bool {
self.try_copy_to_clipboard(label, text).is_some()
}
fn try_copy_to_clipboard(&self, label: &str, text: &str) -> Option<()> {
let mut guard = self.vm.attach_current_thread().ok()?;
let env: &mut JNIEnv = &mut guard;
let context = self.context(env)?;
let clipboard = self.system_service(env, &context, "clipboard")?;
let jlabel = env.new_string(label).ok()?;
let jtext = env.new_string(text).ok()?;
let clip = env
.call_static_method(
"android/content/ClipData",
"newPlainText",
"(Ljava/lang/CharSequence;Ljava/lang/CharSequence;)Landroid/content/ClipData;",
&[
JValue::Object(jlabel.as_ref()),
JValue::Object(jtext.as_ref()),
],
)
.ok()?
.l()
.ok()?;
env.call_method(
&clipboard,
"setPrimaryClip",
"(Landroid/content/ClipData;)V",
&[JValue::Object(&clip)],
)
.ok()?;
Some(())
}
/// The display's own refresh rate in Hz (`View::getDisplay()` ->
/// `Display::getRefreshRate()`), for RUST.md's "Benchmark v2": late
/// frames are judged against *this* device's real budget, not an
/// assumed 60Hz -- a 90Hz or 120Hz phone would otherwise call frames
/// "late" that met their own faster deadline. `None` if the view is
/// not yet attached to a window (`getDisplay` returns `null`) or the
/// platform reports a non-positive rate, which is not a real answer
/// either.
pub fn refresh_rate_hz(&self) -> Option<f32> {
let mut guard = self.vm.attach_current_thread().ok()?;
let env: &mut JNIEnv = &mut guard;
let display = env
.call_method(
self.view.as_obj(),
"getDisplay",
"()Landroid/view/Display;",
&[],
)
.ok()?
.l()
.ok()?;
if display.is_null() {
return None;
}
let rate = env
.call_method(&display, "getRefreshRate", "()F", &[])
.ok()?
.f()
.ok()?;
if rate > 0.0 { Some(rate) } else { None }
}
/// `InputMethodManager.showSoftInput(view, 0)` -- the keyboard phase's
/// own show, called directly rather than through the focus-driven
/// `pending_show_keyboard` path `android/view.rs` uses for a real tap,
/// since RUST.md's "Benchmark v2" spec asks for this "through the
/// shell's InputMethodManager" independent of focus state. `true` only
/// if the platform itself reports the request succeeded -- whether the
/// IME actually became visible is confirmed separately, from
/// `on_insets_changed`, per UI_RULES.md ("never present an inferred
/// value as a measured one").
pub fn show_ime(&self) -> bool {
self.try_toggle_ime(true).unwrap_or(false)
}
/// `InputMethodManager.hideSoftInputFromWindow(windowToken, 0)`.
pub fn hide_ime(&self) -> bool {
self.try_toggle_ime(false).unwrap_or(false)
}
fn try_toggle_ime(&self, show: bool) -> Option<bool> {
let mut guard = self.vm.attach_current_thread().ok()?;
let env: &mut JNIEnv = &mut guard;
let context = self.context(env)?;
let imm = self.system_service(env, &context, "input_method")?;
if show {
env.call_method(
&imm,
"showSoftInput",
"(Landroid/view/View;I)Z",
&[JValue::Object(self.view.as_obj()), JValue::Int(0)],
)
.ok()?
.z()
.ok()
} else {
let token = env
.call_method(
self.view.as_obj(),
"getWindowToken",
"()Landroid/os/IBinder;",
&[],
)
.ok()?
.l()
.ok()?;
env.call_method(
&imm,
"hideSoftInputFromWindow",
"(Landroid/os/IBinder;I)Z",
&[JValue::Object(&token), JValue::Int(0)],
)
.ok()?
.z()
.ok()
}
}
/// Shows `report` in the shell's plain-view diagnostics overlay
/// (`IrisView.showDiagnosticsOverlay`) -- a real `TextView` plus Copy
/// and Close controls, added over whatever iris itself is drawing
/// rather than replacing it (unlike `android::view::show_renderer_error`,
/// which exists for the case the renderer can never recover from and
/// intentionally never returns). Called from a background task after
/// the keyboard-open delay (`bench_client.rs`'s `on_insets_changed`),
/// so the Java side hops onto the UI thread itself before touching the
/// view tree -- see that method's own comment.
pub fn show_diagnostics_overlay(&self, report: &str) -> bool {
self.try_show_diagnostics_overlay(report).is_some()
}
fn try_show_diagnostics_overlay(&self, report: &str) -> Option<()> {
let mut guard = self.vm.attach_current_thread().ok()?;
let env: &mut JNIEnv = &mut guard;
let jreport = env.new_string(report).ok()?;
env.call_method(
self.view.as_obj(),
"showDiagnosticsOverlay",
"(Ljava/lang/String;)V",
&[JValue::Object(jreport.as_ref())],
)
.ok()?;
Some(())
}
}
+211
View File
@@ -0,0 +1,211 @@
//! The JNI half of `DevLogProvider`: reading this process's own log ring
//! for a `ContentProvider` that Dev Updater queries.
//!
//! **Why**: Iris runs these builds on a phone with no `adb`, and Android
//! forbids one app reading another's `logcat`, so nothing outside this
//! process can recover what it wrote. The app already keeps a bounded copy
//! (`client_core::log_ring`); this is how the copy leaves the process. Dev
//! Updater is on the same phone, so handing it over needs no tunnel, no
//! token and no second enrolment -- and it is Dev Updater's own contract
//! rather than something invented here, so any app it delivers can do the
//! same (its `README.md`, "An app's own log").
//!
//! **Everything general stays in `client-core`** (AGENTS.md's sharing
//! rule). What is here is only what Android forces: the JNI boundary and
//! the Java class on the other side of it.
//!
//! Both entry points answer a **flat `String[]`** rather than a row of
//! typed columns. That is the whole of the JNI, and it is one array type
//! instead of three interleaved ones for a payload the provider is about
//! to hand back over binder as a `MatrixCursor` anyway; `DevLogProvider`
//! parses the two numeric fields. Kept flat rather than nested for the
//! same reason -- an array of arrays is four more JNI calls per line.
use android_view::jni::JNIEnv;
use android_view::jni::objects::{JClass, JObject, JString};
use android_view::jni::sys::{jlong, jobjectArray};
use std::sync::OnceLock;
/// How many `String`s each log line occupies in the flat answer:
/// `seq`, `t_ms`, `level`, `target`, `message`, in that order. The Java
/// side has the same constant, and the two are the one place the shape is
/// written down on each side.
///
/// Gated with its one reader: the tabs demo links no `client-core` and so
/// has no ring to lay out, and an ungated constant is a warning in that
/// build (`iris-android-app` without `transcript-screen`).
#[cfg(feature = "transcript-screen")]
const FIELDS_PER_LINE: usize = 5;
/// The authority the provider registered itself under, once it has been
/// created. `None` until then, which is a state worth being able to say:
/// a provider Android never instantiated and one that is answering look
/// the same from inside this process otherwise.
static AUTHORITY: OnceLock<String> = OnceLock::new();
/// Where this app's log can be read from, for the diagnostics pane.
///
/// The provider's own answer rather than one composed from the package
/// name here: what makes the line worth showing is that it names an
/// authority somebody can actually query, and only the provider knows it
/// registered.
#[cfg(feature = "bench")]
pub fn authority() -> Option<&'static str> {
AUTHORITY.get().map(String::as_str)
}
/// `DevLogProvider.nativeReady` -- the provider announcing the authority
/// it registered under and the app's private directory, from its own
/// `onCreate`.
///
/// The directory is taken here as well as in
/// `MainActivity.nativeSetFilesDir` because **the provider is often the
/// only thing running**: once the app has died, Dev Updater's query
/// starts the process for the provider alone, so no activity ever runs
/// and the panic hook's file would never be replayed into the ring. That
/// is precisely the run whose log is being asked for. Whichever of the
/// two arrives first does the replay; `set_crash_dir` deletes the file,
/// so the second finds nothing and says nothing.
///
/// # Safety
/// Called by the JVM with the arguments its `native` declaration names.
#[unsafe(no_mangle)]
pub extern "system" fn Java_dev_iris_android_demo_DevLogProvider_nativeReady(
mut env: JNIEnv,
_class: JClass,
authority: JString,
files_dir: JString,
) {
// Before the authority line, so the previous run's death is above the
// line announcing this one rather than buried under it.
#[cfg(feature = "transcript-screen")]
if let Some(dir) = string_arg(&mut env, &files_dir) {
crate::app_log::set_crash_dir(std::path::Path::new(&dir));
}
#[cfg(not(feature = "transcript-screen"))]
let _ = &files_dir;
let Some(authority) = string_arg(&mut env, &authority) else {
return;
};
log::info!("iris devlog: serving this app's log at content://{authority}");
let _ = AUTHORITY.set(authority);
}
/// One `String` argument, or `None` for a null or unreadable one.
fn string_arg(env: &mut JNIEnv, value: &JString) -> Option<String> {
if value.is_null() {
return None;
}
env.get_string(value).ok().map(Into::into)
}
/// `DevLogProvider.nativeStatus` -- `held`, `dropped`, `newest_seq`, as
/// three strings.
///
/// `newest_seq` is `-1` for a ring nothing has been written to, which is
/// what tells a reader holding a cursor that this process **restarted**:
/// the ring is in memory, so a new process starts again at zero and a
/// stale cursor would otherwise skip everything silently.
///
/// Exported by name rather than registered, matching this crate's other
/// activity-side natives: the mangled name is the whole of what a class
/// this app owns needs.
///
/// # Safety
/// Called by the JVM with the arguments its `native` declaration names.
#[unsafe(no_mangle)]
pub extern "system" fn Java_dev_iris_android_demo_DevLogProvider_nativeStatus(
mut env: JNIEnv,
_class: JClass,
) -> jobjectArray {
string_array(&mut env, &status_fields())
}
/// `DevLogProvider.nativeLinesSince` -- every held line with a sequence at
/// or after `since`, oldest first, [`FIELDS_PER_LINE`] strings each.
///
/// Inclusive of `since` because [`client_core::log_ring::LogRing::since`]
/// is, and one definition of the cursor is what keeps the app's own
/// uploaded report and this provider describing the same lines.
///
/// # Safety
/// Called by the JVM with the arguments its `native` declaration names.
#[unsafe(no_mangle)]
pub extern "system" fn Java_dev_iris_android_demo_DevLogProvider_nativeLinesSince(
mut env: JNIEnv,
_class: JClass,
since: jlong,
) -> jobjectArray {
// A negative cursor is a caller asking for everything, not an error to
// take the app down over: the provider is a diagnostic.
string_array(&mut env, &line_fields(since.max(0) as u64))
}
/// The three status numbers, as the provider's row.
#[cfg(feature = "transcript-screen")]
fn status_fields() -> Vec<String> {
let ring = client_core::log_ring::process_ring();
vec![
ring.len().to_string(),
ring.dropped().to_string(),
ring.newest_seq().map_or(-1, |seq| seq as i64).to_string(),
]
}
/// The tabs demo links no `client-core` and keeps no ring, so it holds
/// nothing and has never dropped anything -- which is the truth, not a
/// stand-in. The natives are still exported there, because a `native`
/// method Java declares and the library does not is an
/// `UnsatisfiedLinkError` the moment the class loads.
#[cfg(not(feature = "transcript-screen"))]
fn status_fields() -> Vec<String> {
vec!["0".to_string(), "0".to_string(), "-1".to_string()]
}
#[cfg(feature = "transcript-screen")]
fn line_fields(since: u64) -> Vec<String> {
let (lines, _next) = client_core::log_ring::process_ring().since(since);
let mut fields = Vec::with_capacity(lines.len() * FIELDS_PER_LINE);
for line in lines {
fields.push(line.seq.to_string());
fields.push(line.at_ms.to_string());
fields.push(line.level.to_string());
fields.push(line.target);
fields.push(line.message);
}
fields
}
#[cfg(not(feature = "transcript-screen"))]
fn line_fields(_since: u64) -> Vec<String> {
Vec::new()
}
/// A Java `String[]` of those, or a null array if the JVM refused one.
///
/// Null rather than a panic across the JNI boundary: `DevLogProvider`
/// reads it as "the provider could not answer" and returns no cursor,
/// which Dev Updater already draws as a distinct state. Taking the app
/// down to report that its diagnostic is unavailable would be worse than
/// the diagnostic being unavailable.
fn string_array(env: &mut JNIEnv, fields: &[String]) -> jobjectArray {
let null = std::ptr::null_mut();
let Ok(class) = env.find_class("java/lang/String") else {
return null;
};
let Ok(array) = env.new_object_array(fields.len() as i32, class, JObject::null()) else {
return null;
};
for (index, field) in fields.iter().enumerate() {
let Ok(value) = env.new_string(field) else {
return null;
};
if env
.set_object_array_element(&array, index as i32, value)
.is_err()
{
return null;
}
}
array.into_raw()
}
+133
View File
@@ -0,0 +1,133 @@
//! Which `ai-server` this app talks to, and how it was told.
//!
//! The parsing, the file and its owner-only mode are
//! `client_core::config` (`EnrolledServer`/`EnrollmentStore`), shared with
//! the desktop app. What is genuinely this platform's, and all that is
//! here, is the intent plumbing: Android hands an `aiapp://enroll?...`
//! link to `MainActivity`, which passes it and the app's private files
//! directory across JNI (see `lib.rs`'s two exported functions).
//!
//! **Why the app is told at runtime rather than at build time.** The APK
//! is cross-compiled in a VM and run against the server on the host, whose
//! CA and token are not this machine's -- so nothing about the destination
//! can be baked in, and no token or CA may sit in a repo or a delivered
//! artifact either way. The CA arrives with the link (`ca` parameter,
//! `wg_app_link::enroll::ca_param`), which is what makes an APK built
//! anywhere able to pin the server it is pointed at.
//!
//! The files directory is process-wide state, which this project otherwise
//! avoids: it arrives from the activity, and `AndroidAppState::new` -- the
//! first thing that wants the enrollment -- has no parameter it could come
//! in through. Same shape, and the same reason, as
//! `client_core::log_ring`'s process ring.
#[cfg(not(feature = "bench"))]
use client_core::api::UreqTransport;
use client_core::config::{EnrolledServer, EnrollmentStore};
use std::path::PathBuf;
use std::sync::OnceLock;
/// `Context.getFilesDir()`, handed over by `MainActivity` before it builds
/// the view. Set once per process; a second call with a different path is
/// a programmer error rather than something to recover from, and a second
/// call with the same one is what a re-created activity does.
static FILES_DIR: OnceLock<PathBuf> = OnceLock::new();
pub fn set_files_dir(dir: PathBuf) {
if let Err(existing) = FILES_DIR.set(dir.clone()) {
assert_eq!(
existing, dir,
"the app's files directory was set twice with different paths"
);
}
}
/// `None` before `MainActivity` has handed the directory over -- which is
/// **not** the same as "not enrolled", and is why [`status`] has a state
/// for it (UI_RULES: design the unknown state first).
fn store() -> Option<EnrollmentStore> {
FILES_DIR.get().map(EnrollmentStore::new)
}
/// What this app has been told, or why it has not been.
pub enum Status {
Enrolled(EnrolledServer),
/// Nothing has been enrolled yet: the ordinary first-run state.
NotEnrolled,
/// The question could not be answered -- the activity never handed a
/// files directory over, or the file is there and unreadable. Kept
/// apart from `NotEnrolled` because the two want different actions
/// from whoever is looking.
Unknown(String),
}
pub fn status() -> Status {
let Some(store) = store() else {
return Status::Unknown("the activity never handed over a files directory".to_string());
};
match store.load() {
Ok(Some(server)) => Status::Enrolled(server),
Ok(None) => Status::NotEnrolled,
Err(error) => Status::Unknown(error.to_string()),
}
}
/// One line for the diagnostics pane. The three states read differently on
/// purpose: "not enrolled" says what to do about it, and "couldn't tell"
/// must not be mistaken for it.
///
/// Only the bench build has a pane to put this in -- same gate, and the
/// same reason, as `app_log::diagnostics_line`. The transcript build says
/// the same things where they matter to it, in the message
/// [`transport`]'s error becomes on screen.
#[cfg(feature = "bench")]
pub fn status_line() -> String {
match status() {
Status::Enrolled(server) => format!("enrolled: {}:{}", server.host, server.port),
Status::NotEnrolled => "not enrolled -- open the enrol link from Dev Updater".to_string(),
Status::Unknown(why) => format!("enrolment unreadable: {why}"),
}
}
/// Parses an `aiapp://enroll?...` link and saves it, replacing whatever
/// was enrolled before -- opening a link is how somebody says "this server
/// now", including after the old one's token was rotated.
///
/// The returned `Err` is the message for a person: this is called from a
/// tap on a link, and a link that did nothing with nothing said is the
/// failure the UI rules are most insistent about.
pub fn apply_link(uri: &str) -> Result<EnrolledServer, String> {
let server = EnrolledServer::parse_link(uri)?;
let store = store().ok_or("the app has no files directory to save an enrollment in")?;
store
.save(&server)
.map_err(|error| format!("couldn't save the enrollment: {error}"))?;
Ok(server)
}
/// A transport for the enrolled server, pinning the CA the link carried.
///
/// Gated to the same builds as `transcript_client`, its only caller: the
/// bench build opens a checked-in fixture and reaches no server, so
/// compiling this into it would be a warning about dead code that is
/// dead on purpose.
///
/// Every failure here is a sentence a screen can show, because there is
/// nowhere else for it to go: this app has no `logcat` on the phone it is
/// built for.
#[cfg(not(feature = "bench"))]
pub fn transport() -> Result<UreqTransport, String> {
let server = match status() {
Status::Enrolled(server) => server,
Status::NotEnrolled => {
return Err("Not enrolled yet -- open the enrol link from Dev Updater.".to_string());
}
Status::Unknown(why) => return Err(format!("Couldn't read the enrollment: {why}")),
};
let ca_pem = server.ca_pem.as_ref().ok_or(
"The enrollment link carried no CA, so there is nothing to pin. \
Enrol again with a link minted by this server.",
)?;
UreqTransport::new(server.base_url(), &server.token, ca_pem.as_bytes())
.map_err(|error| error.message)
}
+151 -3
View File
@@ -1,5 +1,5 @@
//! The android-view demo app: iris's `tabs` widget tree (`tabs_ui::build`,
//! shared with the winit example) running through
//! The android-view demo app: by default, iris's `tabs` widget tree
//! (`tabs_ui::build`, shared with the winit example) running through
//! `iris::android`'s `ViewPeer`. This is RUST.md's I2 pass condition made
//! concrete -- there is no UI here beyond what `tabs-ui` already draws.
//!
@@ -9,28 +9,81 @@
//! wrapping `iris::android::new_peer`'s generic function in a concrete
//! `extern "system" fn`, since `register_view_class` wants a plain
//! function pointer.
//!
//! **`transcript-screen` feature (RUST.md's I5 Android integration):** with
//! `--features transcript-screen`, `new_view_peer` instantiates
//! `transcript_client::TranscriptClient` instead of the tabs `Client`
//! below, against a real `ai-server` (see that module's doc). Chosen over a
//! third shell crate: this one already has the Gradle project, the
//! `IrisView`/`MainActivity` Java, and the JNI registration I2 built and
//! measured against, and the only thing a transcript screen needs on top
//! is a different `AndroidAppState` -- the same axis `tabs_ui::build` vs.
//! `transcript_ui::build` already varies along on the winit side (compare
//! `iris/examples/tabs.rs` and `iris/transcript-ui/examples/transcript.rs`).
//! A build picks one screen or the other, never both, so `Client` and
//! `TranscriptClient` are cfg-gated apart rather than switched at runtime --
//! there is no in-app navigation to switch *to* on either side yet.
//!
//! **`bench` feature (P0's iris half, docs/RUST.md):** a third
//! `AndroidAppState`, `bench_client::BenchClient`, on the same axis --
//! `transcript_ui::build_tree` again, this time against the checked-in
//! fixture (`app/bench-fixture/assets/transcript.jsonl`) instead of a real
//! server, with a "Run benchmark" control that drives the same scroll loop
//! and streaming phase the Compose `bench` build type's `BenchRun.kt`
//! does. `bench` depends on `transcript-screen` (Cargo.toml) for
//! `transcript-ui`/`client-core`/`event-model`, so both features end up
//! enabled together -- `ActiveClient` below gives `bench` priority in that
//! case, the same way `transcript-screen` already takes priority over the
//! default `tabs-screen`.
use android_view::{
Context, View,
jni::{
JNIEnv, JavaVM,
objects::{JClass, JString},
sys::{JNI_VERSION_1_6, JavaVM as RawJavaVM, jint, jlong},
},
register_view_class,
};
#[cfg(not(feature = "transcript-screen"))]
use iris::android::{AndroidAppState, AndroidRsc, AndroidUiState, HasAndroidUiState};
#[cfg(not(feature = "transcript-screen"))]
use iris::prelude::*;
use log::LevelFilter;
use std::ffi::c_void;
/// The app's own log ring and its upload -- only where `client-core` is
/// linked, which is every build that has a server to send to. The plain
/// tabs demo keeps `android_logger` alone, as it always had.
#[cfg(feature = "transcript-screen")]
mod app_log;
#[cfg(feature = "bench")]
mod bench_client;
#[cfg(feature = "bench")]
mod bench_jni;
/// This app's log ring, handed to Dev Updater on the phone through a
/// `ContentProvider`. Declared in every build for the reason the module
/// gives: the Java class is in the manifest either way, and a `native`
/// method the library does not export fails the class load.
mod devlog;
/// Which server this app talks to, told to it at runtime by an
/// `aiapp://enroll` link. Only where `client-core` is linked -- the plain
/// tabs demo makes no network call and has nothing to enrol against.
#[cfg(feature = "transcript-screen")]
mod enrollment;
#[cfg(all(feature = "transcript-screen", not(feature = "bench")))]
mod transcript_client;
/// The app's `View` subclass, matching the Java side's package --
/// `app/src/main/java/dev/iris/android/demo/IrisView.java`.
const VIEW_CLASS: &str = "dev/iris/android/demo/IrisView";
#[cfg(not(feature = "transcript-screen"))]
pub struct Client {
ui_state: AndroidUiState,
}
#[cfg(not(feature = "transcript-screen"))]
impl HasAndroidUiState for Client {
fn android_state(&self) -> &AndroidUiState {
&self.ui_state
@@ -40,6 +93,7 @@ impl HasAndroidUiState for Client {
}
}
#[cfg(not(feature = "transcript-screen"))]
impl AndroidAppState for Client {
fn new(mut ui_state: AndroidUiState, rsc: &mut AndroidRsc<Self>) -> Self {
// `widgets.info` is the winit example's frame-debug readout, kept
@@ -61,12 +115,19 @@ impl AndroidAppState for Client {
}
}
#[cfg(not(feature = "transcript-screen"))]
type ActiveClient = Client;
#[cfg(all(feature = "transcript-screen", not(feature = "bench")))]
type ActiveClient = transcript_client::TranscriptClient;
#[cfg(feature = "bench")]
type ActiveClient = bench_client::BenchClient;
extern "system" fn new_view_peer<'local>(
env: JNIEnv<'local>,
view: View<'local>,
context: Context<'local>,
) -> jlong {
iris::android::new_peer::<Client>(env, view, context)
iris::android::new_peer::<ActiveClient>(env, view, context)
}
/// # Safety
@@ -74,6 +135,13 @@ extern "system" fn new_view_peer<'local>(
/// mirrors android-view's own demo, which carries the same comment.
#[unsafe(no_mangle)]
pub unsafe extern "system" fn JNI_OnLoad(vm: *mut RawJavaVM, _: *mut c_void) -> jint {
// The ring in front of `android_logger` where there is one (see
// `app_log`), and `android_logger` alone otherwise. Both install the
// same tag and level, so `logcat` cannot tell the two builds apart --
// the ring only adds a second reader.
#[cfg(feature = "transcript-screen")]
app_log::install(LevelFilter::Debug);
#[cfg(not(feature = "transcript-screen"))]
android_logger::init_once(
android_logger::Config::default()
.with_max_level(LevelFilter::Debug)
@@ -85,3 +153,83 @@ pub unsafe extern "system" fn JNI_OnLoad(vm: *mut RawJavaVM, _: *mut c_void) ->
iris::android::register_native_methods(&mut env, VIEW_CLASS);
JNI_VERSION_1_6
}
/// `MainActivity.nativeSetFilesDir` -- the app's private directory, handed
/// over before the view exists because that is where the enrollment is
/// read from and written to (`enrollment`'s module doc).
///
/// Exported by name rather than registered through `RegisterNatives`: the
/// view's methods are registered because `android-view` owns that class
/// and hands out one function pointer, whereas these two are this app's
/// own activity and the mangled name is the whole of what is needed.
///
/// Declared in every build, including the tabs demo that has no
/// `client-core` to store anything -- a `native` method Java declares and
/// the library does not export is an `UnsatisfiedLinkError` when the class
/// loads, which would take down a build that merely shares the activity.
///
/// # Safety
/// Called by the JVM with the arguments its `native` declaration names.
#[unsafe(no_mangle)]
pub extern "system" fn Java_dev_iris_android_demo_MainActivity_nativeSetFilesDir(
mut env: JNIEnv,
_class: JClass,
dir: JString,
) {
let Some(dir) = jstring(&mut env, dir) else {
return;
};
#[cfg(feature = "transcript-screen")]
{
app_log::set_crash_dir(std::path::Path::new(&dir));
enrollment::set_files_dir(std::path::PathBuf::from(&dir));
}
log::debug!("iris app: files directory is {dir}");
}
/// `MainActivity.nativeEnroll` -- one `aiapp://enroll?...` link, from the
/// VIEW intent that started or resumed the activity.
///
/// Logged either way rather than answered: the activity has nothing to do
/// with the result, and where the enrollment shows up is the diagnostics
/// pane (`enrollment::status_line`), which reads the stored answer rather
/// than being told it.
///
/// # Safety
/// Called by the JVM with the arguments its `native` declaration names.
#[unsafe(no_mangle)]
pub extern "system" fn Java_dev_iris_android_demo_MainActivity_nativeEnroll(
mut env: JNIEnv,
_class: JClass,
uri: JString,
) {
let Some(uri) = jstring(&mut env, uri) else {
return;
};
#[cfg(feature = "transcript-screen")]
match enrollment::apply_link(&uri) {
// Never the token: `wg-app-link`'s enroll module forbids logging
// it, and this line would otherwise be the one place it leaked.
Ok(server) => log::info!("iris app: enrolled with {}:{}", server.host, server.port),
Err(error) => log::warn!("iris app: that enrolment link was refused -- {error}"),
}
#[cfg(not(feature = "transcript-screen"))]
log::warn!("iris app: {uri} arrived, but this build has no server to enrol with");
}
/// A `JString` as a Rust `String`, or `None` for a null or non-UTF-8 one --
/// neither is worth taking the app down for, and both are logged where
/// they happen.
fn jstring(env: &mut JNIEnv, value: JString) -> Option<String> {
if value.is_null() {
log::warn!("iris app: the activity passed a null string across JNI");
return None;
}
match env.get_string(&value) {
Ok(value) => Some(value.into()),
Err(error) => {
log::warn!("iris app: couldn't read a string from the activity -- {error}");
None
}
}
}
+386
View File
@@ -0,0 +1,386 @@
//! RUST.md's I5 Android integration: `transcript-ui`'s screen filling the
//! whole window on android-view, against a real `ai-server` through
//! `client-core` -- the missing half `iris-android-app` (I2) only had for
//! `tabs-ui` until now. Behind the `transcript-screen` Cargo feature so the
//! plain build (`cargo ndk build`, no `--features`) stays exactly the tabs
//! demo I2/I4 already measured against.
//!
//! **Deliberate simplification, recorded rather than left to be
//! rediscovered (RUST.md's I5 box has the full account)**: there is no
//! session list here -- the first session `ApiClient::fetch_sessions`
//! returns is opened automatically, since there is nothing to tap to get
//! there, which is what `transcript-bench.sh` and `ui-trace` need to land
//! straight on the screen under test.
//!
//! Which server it opens it against is no longer baked in: it is the
//! enrollment an `aiapp://enroll` link left behind (`crate::enrollment`,
//! and `desktop-app`'s identical `--link`), because an APK
//! cross-compiled here cannot pin the CA of a server on the host.
//!
//! **Reuses `iris/desktop-app`'s `app.rs` shape almost exactly** --
//! `fold_event`/`group_tool_runs`/`fold_page`/`raw_seq` from
//! `client_core::transcript_fold`, a `generation` counter guarding against
//! a stale background response. What differs is only the redraw
//! mechanism: android-view has no `winit::EventLoopProxy`, so this uses
//! `iris::task::Tasks::redraw_handle` (new, added alongside this box) to
//! request a frame after each `TaskCtx::update` instead of relying on
//! `Tasks::spawn`'s single end-of-future redraw -- see that method's own
//! doc for why.
//!
//! **Streaming no longer costs a full rebuild** (fixed after the P0 gate
//! showed why it mattered -- 20 events/second means 20 rebuilds/second of
//! a ~3,200-row transcript otherwise): `apply_event` calls
//! `transcript_ui::TranscriptScreen::apply` with the item list before and
//! after `fold_event`, which updates only the row(s) that actually
//! changed (almost always just the one open assistant message) instead of
//! refolding and rebuilding every row. `rebuild_transcript` still runs
//! the whole widget tree once, for the opening page and for `apply`'s own
//! rare regroup fallback.
use client_core::api::{ApiClient, UreqTransport};
use client_core::event_stream::{StreamItem, follow_session_events};
use client_core::transcript_fold::{TranscriptItem, fold_event, fold_page, group_tool_runs};
use event_model::SeqEvent;
use iris::android::{AndroidAppState, AndroidRsc, AndroidUiState, HasAndroidUiState};
use iris::prelude::*;
use std::sync::Arc;
use std::sync::atomic::{AtomicU64, Ordering};
pub struct TranscriptClient {
ui_state: AndroidUiState,
/// The screen's own content -- everything under the fixed
/// [`frame_report_controls`] bar, which is built once (`new`, below)
/// and never touched by `show_message`/`rebuild_transcript`'s own
/// `set` calls the way `desktop-app`'s `transcript_ptr` isn't touched
/// by rebuilding the session list beside it.
content: WeakWidget<WidgetPtr>,
screen: Option<transcript_ui::TranscriptScreen>,
/// The folded transcript as of the last rebuild -- kept here (not
/// re-derived) for the same reason `desktop-app`'s `Client::items`
/// exists: a live `StreamEvent` only carries one new wire event, and
/// `fold_event` needs everything folded so far to fold it in.
items: Vec<TranscriptItem>,
/// The session currently open -- `None` only before the first fetch
/// resolves. Read back by `apply_event`'s rebuild, which has no session
/// id of its own (a live `SeqEvent` doesn't carry one).
session_id: Option<String>,
/// Bumped every time a new session load starts; a background response
/// checks it before touching state, so a slow reply for a session this
/// screen has moved on from can't overwrite what replaced it. There is
/// only ever one session here (no list to switch away to), but the
/// guard still matters for the *first* fetch racing a `stop`/`start`.
generation: Arc<AtomicU64>,
}
impl HasAndroidUiState for TranscriptClient {
fn android_state(&self) -> &AndroidUiState {
&self.ui_state
}
fn android_state_mut(&mut self) -> &mut AndroidUiState {
&mut self.ui_state
}
}
/// Builds one `UreqTransport` from the stored enrollment. Called twice per
/// session load, same as `desktop-app`'s `build_transport` closure --
/// `ApiClient` and the live-stream follow each need their own, since
/// `UreqTransport` holds its own `ureq::Agent`.
///
/// Read afresh each time rather than held: opening a new enrolment link
/// while the app is running is how somebody points it at another server,
/// and a cached transport would keep talking to the old one.
fn build_transport() -> Result<UreqTransport, String> {
crate::enrollment::transport()
}
fn placeholder<Rsc: HasEvents>(rsc: &mut Rsc, message: &str) -> StrongWidget {
wtext(message.to_string())
.color(Color::WHITE)
.wrap(true)
.pad(16)
.add_strong(rsc)
.any()
}
/// The two named controls RUST.md's I5 box ("Measurements taken" (b))
/// drives by name over `ui-trace`, e.g. `ui-trace record --do "tap 'Frame
/// report'"`. `dumpsys gfxinfo` cannot see this screen's own GPU-drawn
/// frames at all -- this is the screen's own equivalent of the Compose
/// app's "Copy render timings" control, logged rather than clipboarded
/// (no clipboard wiring exists here) under this crate's own fixed
/// `android_logger` tag (`iris-android-app`, `lib.rs`'s `JNI_OnLoad`),
/// grep-able on the fixed string `"iris frame report"` the way
/// `transcript-bench.sh` greps `"ai-app render report"`.
fn frame_report_controls(rsc: &mut AndroidRsc<TranscriptClient>) -> WeakWidget {
type Rsc = AndroidRsc<TranscriptClient>;
let report_rect = rect(Color::rgb(50, 50, 60))
.on(
CursorSense::click(),
|ctx: EventIdCtx<'_, Rsc, _, _>, _rsc: &mut Rsc| match ctx
.state
.android_state()
.frame_report
.report()
{
Some(stats) => log::info!("iris frame report: {stats}"),
None => log::info!(
"iris frame report: no frames recorded -- scroll first, then press this"
),
},
)
.label("Frame report");
let report = (
report_rect,
wtext("Frame report").size(18).text_align(Align::CENTER),
)
.stack()
.pad(8)
.add(rsc);
let reset_rect = rect(Color::rgb(70, 40, 40))
.on(
CursorSense::click(),
|ctx: EventIdCtx<'_, Rsc, _, _>, _rsc: &mut Rsc| {
ctx.state.android_state_mut().frame_report.reset();
log::info!("iris frame report: reset");
},
)
.label("Reset frame report");
let reset = (
reset_rect,
wtext("Reset").size(18).text_align(Align::CENTER),
)
.stack()
.pad(8)
.add(rsc);
(report, reset).span(Dir::RIGHT).height(56).add(rsc)
}
impl AndroidAppState for TranscriptClient {
fn new(mut ui_state: AndroidUiState, rsc: &mut AndroidRsc<Self>) -> Self {
let content = WidgetPtr::new().add(rsc);
let loading = placeholder(rsc, "Loading sessions...");
content(rsc).set(loading);
let tree = (frame_report_controls(rsc), content.height(rest(1)))
.span(Dir::DOWN)
.add_strong(rsc)
.any();
ui_state.set_root(tree);
let mut client = Self {
ui_state,
content,
screen: None,
items: Vec::new(),
session_id: None,
generation: Arc::new(AtomicU64::new(0)),
};
client.spawn_fetch_sessions(rsc);
client
}
fn back_pressed(&mut self, _rsc: &mut AndroidRsc<Self>, _render: &mut UiRenderState) -> bool {
// No screen stack of its own -- same "let the activity finish"
// answer `iris-android-app`'s tabs `Client` already gives.
false
}
}
impl TranscriptClient {
fn show_message(&mut self, rsc: &mut AndroidRsc<Self>, message: &str) {
let widget = placeholder(rsc, message);
(self.content)(rsc).set(widget);
self.screen = None;
}
fn spawn_fetch_sessions(&mut self, rsc: &mut AndroidRsc<Self>) {
let redraw = rsc.tasks.redraw_handle();
let my_generation = self.generation.load(Ordering::SeqCst);
let generation = self.generation.clone();
rsc.spawn_task(async move |mut ctx| {
let outcome = match build_transport() {
Ok(transport) => ApiClient::new(transport)
.fetch_sessions()
.map_err(|e| e.to_string()),
Err(e) => Err(format!("couldn't set up TLS: {e}")),
};
ctx.update(move |state: &mut TranscriptClient, rsc| {
if generation.load(Ordering::SeqCst) != my_generation {
return;
}
match outcome {
Ok(sessions) => match sessions.into_iter().next() {
Some(session) => state.select_session(rsc, session.id),
None => state.show_message(rsc, "No sessions on the sandbox server."),
},
Err(message) => {
state.show_message(rsc, &format!("Couldn't list sessions: {message}"))
}
}
});
redraw.request_redraw();
});
}
/// Loads the opening page, then follows the live SSE stream for the
/// rest of this session's life -- `desktop-app`'s `select_session`
/// almost verbatim, with `Proxy::send_event` replaced by `ctx.update` +
/// `redraw.request_redraw()` (see this module's doc).
fn select_session(&mut self, rsc: &mut AndroidRsc<Self>, session_id: String) {
let my_generation = self.generation.fetch_add(1, Ordering::SeqCst) + 1;
self.items.clear();
self.session_id = Some(session_id.clone());
self.show_message(rsc, "Loading transcript...");
let redraw = rsc.tasks.redraw_handle();
let live_generation = self.generation.clone();
rsc.spawn_task(async move |mut ctx| {
let transports =
build_transport().and_then(|rest| build_transport().map(|stream| (rest, stream)));
let (rest, stream_transport) = match transports {
Ok(pair) => pair,
Err(e) => {
let message = format!("couldn't set up TLS: {e}");
ctx.update(move |state: &mut TranscriptClient, rsc| {
if live_generation.load(Ordering::SeqCst) == my_generation {
state.show_message(rsc, &message);
}
});
redraw.request_redraw();
return;
}
};
let api = ApiClient::new(rest);
// The most recent 200 events, coalesced -- the same page size
// `desktop-app` uses; RUST.md's I3/history-paging work is what
// a real scrollback would reuse (out of scope here, same as
// E4).
let page: Result<Vec<serde_json::Value>, String> = api
.fetch_transcript_page(&session_id, None, 200, true)
.map_err(|e| e.to_string());
// The wire `seq` of the last line, not a folded item's `seq()`
// -- see `client_core::transcript_fold::raw_seq`'s doc for why
// resuming from the latter re-delivers deltas already folded
// into an in-progress reply.
let after = page
.as_ref()
.ok()
.and_then(|values| values.last())
.and_then(client_core::transcript_fold::raw_seq)
.unwrap_or(0);
let result = page.and_then(|values| fold_page(&values));
{
let live_generation = live_generation.clone();
ctx.update(move |state: &mut TranscriptClient, rsc| {
if live_generation.load(Ordering::SeqCst) != my_generation {
return;
}
match result {
Ok(items) => {
state.items = items;
state.rebuild_transcript(rsc);
}
Err(message) => {
state.show_message(rsc, &format!("Couldn't load transcript: {message}"))
}
}
});
}
redraw.request_redraw();
if live_generation.load(Ordering::SeqCst) != my_generation {
return;
}
// The outer closure here is an `FnMut` -- `follow_session_events`
// calls it once per line -- so it captures `live_generation` by
// move and re-clones it for each inner `ctx.update` closure
// rather than moving a shared `stop`-style helper into itself:
// a value moved out of an `FnMut`'s captures on one call leaves
// nothing there for the next.
let _ =
follow_session_events(
&stream_transport,
&session_id,
after,
move |item| match item {
StreamItem::Open | StreamItem::Reset => {
live_generation.load(Ordering::SeqCst) == my_generation
}
StreamItem::Event { event, .. } => {
if live_generation.load(Ordering::SeqCst) != my_generation {
return false;
}
let live_generation = live_generation.clone();
ctx.update(move |state: &mut TranscriptClient, rsc| {
if live_generation.load(Ordering::SeqCst) != my_generation {
return;
}
state.apply_event(rsc, &event);
});
redraw.request_redraw();
true
}
},
);
});
}
/// Rebuilds the whole widget tree from `self.items` -- same tradeoff as
/// `desktop-app`'s `rebuild_transcript` (this module's doc comment).
/// Reads `self.session_id` rather than taking one, since every caller
/// (the opening page, and every live event) already has it set there.
fn rebuild_transcript(&mut self, rsc: &mut AndroidRsc<Self>) {
let in_progress = self
.screen
.as_ref()
.map(|screen| screen.composer.field.edit(rsc).text.text().to_string())
.filter(|t| !t.is_empty());
let rows = group_tool_runs(&self.items);
let (screen, tree) = transcript_ui::build_tree(rsc, rows);
if let Some(text) = in_progress {
screen.composer.field.edit(rsc).set(&text);
}
if let Some(session_id) = self.session_id.clone() {
let field = screen.composer.field;
rsc.register_event(field, Submit, move |ctx, rsc| {
let text = field.edit(rsc).take();
let text = text.trim().to_string();
if !text.is_empty() {
ctx.state.send_message(session_id.clone(), text);
}
});
}
(self.content)(rsc).set(tree);
self.screen = Some(screen);
}
fn apply_event(&mut self, rsc: &mut AndroidRsc<Self>, event: &SeqEvent) {
let old_items = self.items.clone();
self.items = fold_event(&self.items, event);
match &self.screen {
// The common path: update only the row(s) that actually
// changed instead of refolding and rebuilding all ~3,200 of
// them per event (RUST.md's P0 streaming-phase fix).
Some(screen) => screen.apply(rsc, &old_items, &self.items),
// No screen yet (the opening page hasn't landed) -- build one
// the ordinary way once it has.
None => self.rebuild_transcript(rsc),
}
}
fn send_message(&mut self, session_id: String, text: String) {
std::thread::spawn(move || {
if let Ok(transport) = build_transport() {
let api = ApiClient::new(transport);
let _ = api.send_message(&session_id, &text, &[]);
}
});
}
}
+156
View File
@@ -0,0 +1,156 @@
#!/usr/bin/env python3
"""AOSP's fling spline, transcribed independently of the Rust port.
This exists so the numbers in `sense.rs`'s `the_spline_matches_aosps_own_table`
and `a_flick_decelerates_the_way_aosp_says_it_does` are not the Rust code
grading its own homework. Every test iris's fling had before 2026-09-07
compared the curve with itself -- monotonic, signed, integrates to the closed
form -- and all of them passed while `distance_fraction(t)` was returning
exactly `t` (see `android_fling_spline`'s doc comment). Numbers checked into a
test have to come from somewhere else, and this is the somewhere else.
Transcribed by hand from, and only from:
* frameworks/base `core/java/android/widget/OverScroller.java`,
`SplineOverScroller`'s static initialiser, `getSplineDeceleration`,
`getSplineFlingDistance`, `getSplineFlingDuration` and `update`.
* androidx.compose.animation:animation:1.12.0 `SplineBasedDecay.kt`
(`computeSplineInfo`, `AndroidFlingSpline.flingPosition`) and
`FlingCalculator.kt` (`computeDeceleration`, `flingDistance`,
`flingDuration`, `FlingInfo.position`/`velocity`). The two agree line for
line, which is why iris ports one curve rather than two.
Run it with no arguments; it prints the table entries and the (velocity,
density, t) points the Rust tests assert on.
"""
NB_SAMPLES = 100
INFLEXION = 0.35
START_TENSION = 0.5
END_TENSION = 1.0
P1 = START_TENSION * INFLEXION
P2 = 1.0 - END_TENSION * (1.0 - INFLEXION)
# ViewConfiguration.getScrollFriction(), and SplineOverScroller's own
# "look and feel tuning" constant -- a different number in a different place
# of the same formula, which is the pair iris got the wrong way round once.
SCROLL_FRICTION = 0.015
TUNING = 0.84
GRAVITY_EARTH = 9.80665
INCHES_PER_METER = 39.37
import math
DECELERATION_RATE = math.log(0.78) / math.log(0.9)
def spline_positions():
"""SPLINE_POSITION: distance fraction at each of 101 even time steps."""
position = [0.0] * (NB_SAMPLES + 1)
x_min = 0.0
for i in range(NB_SAMPLES):
alpha = i / NB_SAMPLES
x_max = 1.0
while True:
x = x_min + (x_max - x_min) / 2.0
coef = 3.0 * x * (1.0 - x)
# Solved on the P1/P2 curve...
tx = coef * ((1.0 - x) * P1 + x * P2) + x * x * x
if abs(tx - alpha) < 1e-5:
break
if tx > alpha:
x_max = x
else:
x_min = x
# ...and sampled on the tension curve.
position[i] = coef * ((1.0 - x) * START_TENSION + x * END_TENSION) + x * x * x
position[NB_SAMPLES] = 1.0
return position
POSITION = spline_positions()
def fling_sample(t):
"""(distance fraction, velocity fraction) at time fraction `t`."""
t = min(max(t, 0.0), 1.0)
index = int(t * NB_SAMPLES)
if index >= NB_SAMPLES:
return 1.0, 0.0
t_inf = index / NB_SAMPLES
t_sup = (index + 1) / NB_SAMPLES
velocity_coef = (POSITION[index + 1] - POSITION[index]) / (t_sup - t_inf)
return POSITION[index] + (t - t_inf) * velocity_coef, velocity_coef
def physical_coefficient(density):
return GRAVITY_EARTH * INCHES_PER_METER * density * 160.0 * TUNING
def deceleration(velocity, density):
return math.log(
INFLEXION * abs(velocity) / (SCROLL_FRICTION * physical_coefficient(density))
)
def fling_distance(velocity, density):
l = deceleration(velocity, density)
return (
SCROLL_FRICTION
* physical_coefficient(density)
* math.exp(DECELERATION_RATE / (DECELERATION_RATE - 1.0) * l)
)
def fling_duration_s(velocity, density):
l = deceleration(velocity, density)
return math.exp(l / (DECELERATION_RATE - 1.0))
def position_at(velocity, density, t_seconds):
d = fling_duration_s(velocity, density)
return fling_distance(velocity, density) * fling_sample(t_seconds / d)[0]
def velocity_at(velocity, density, t_seconds):
d = fling_duration_s(velocity, density)
return fling_sample(t_seconds / d)[1] * fling_distance(velocity, density) / d
if __name__ == "__main__":
print("SPLINE_POSITION at a few indices (index: value)")
for i in (0, 1, 10, 25, 50, 75, 99, 100):
print(f" {i:3}: {POSITION[i]:.6f}")
print()
print("distance/velocity fraction at time fractions")
for t in (0.0, 0.1, 0.25, 0.5, 0.75, 0.9, 1.0):
d, v = fling_sample(t)
print(f" t={t:<5} distance={d:.6f} velocity={v:.6f}")
print()
# 2.55 is Iris's Pixel 9 Pro XL (docs/bench/iris-phone-v2-2026-09-06.md);
# 2.75 is this checkout's emulator.
for density in (2.55, 2.75):
# 15250 is `transcript-fixture/touch/flick-120hz.touch`'s own
# release velocity (velocity_reference.py), so `phone_screen.rs`
# can bound the fling it produces from *here* rather than from the
# `FlingCalculator` under test (docs/REVIEW-2026-09-07.md's T1).
for velocity in (5000.0, 11064.0, 15250.0):
dur = fling_duration_s(velocity, density)
print(
f"density={density} v={velocity}: "
f"distance={fling_distance(velocity, density):.3f}px "
f"duration={dur:.4f}s"
)
# Deliberately not round fractions. The velocity coefficient is
# piecewise *constant* across each of the 100 samples, so it
# steps at t = k/100 and a test asserting on 0.75 is asserting
# on which side of a discontinuity the last float landed --
# which is genuinely different between Python and Rust and says
# nothing about the curve.
for frac in (0.125, 0.335, 0.505, 0.755):
t = frac * dur
print(
f" t={frac:>4} of duration ({t:.4f}s): "
f"pos={position_at(velocity, density, t):.3f}px "
f"vel={velocity_at(velocity, density, t):.3f}px/s"
)
+5 -5
View File
@@ -142,7 +142,7 @@ fn bench_first_frame(n: usize) {
let start = Instant::now();
render.update(&root, &mut rsc);
let elapsed = start.elapsed();
let (draws, rewrites, moves) = render.take_counters();
let (draws, rewrites, moves, _shapes) = render.take_counters();
report(
&format!("(a) first frame, N={n}"),
elapsed,
@@ -177,7 +177,7 @@ fn bench_scroll(n: usize, ticks: usize) {
let start = Instant::now();
render.update(&root, &mut rsc);
total += start.elapsed();
let (draws, rewrites, moves) = render.take_counters();
let (draws, rewrites, moves, _shapes) = render.take_counters();
total_draws += draws;
total_rewrites += rewrites;
total_moves += moves;
@@ -245,7 +245,7 @@ fn bench_input_grows(n: usize, lines: usize) {
let start = Instant::now();
render.update(&root, &mut rsc);
total += start.elapsed();
let (draws, rewrites, moves) = render.take_counters();
let (draws, rewrites, moves, _shapes) = render.take_counters();
total_draws += draws;
total_rewrites += rewrites;
total_moves += moves;
@@ -302,7 +302,7 @@ fn bench_insert_above_anchor(n: usize, inserts: usize) {
let start = Instant::now();
render.update(&root, &mut rsc);
total += start.elapsed();
let (draws, rewrites, moves) = render.take_counters();
let (draws, rewrites, moves, _shapes) = render.take_counters();
total_draws += draws;
total_rewrites += rewrites;
total_moves += moves;
@@ -384,7 +384,7 @@ fn bench_expand_holds_edge(n: usize, growths: usize) {
let start = Instant::now();
render.update(&root, &mut rsc);
total += start.elapsed();
let (draws, rewrites, moves) = render.take_counters();
let (draws, rewrites, moves, _shapes) = render.take_counters();
total_draws += draws;
total_rewrites += rewrites;
total_moves += moves;
Loaded 100 of 222 files, more files were not shown because too many files have changed in this diff. Show more