Files
ai-app/docs/DECISIONS.md
T
irisandClaude Opus 5 6d5a231f5c iris is the framework alone; the app is one crate in app-rust/
Iris: "the organization of the rust rewrite is a mess right now... there
shouldn't be anything related to the app inside of iris. Iris is supposed
to be the UI framework alone." And, on the crate count: "I'm confused why
the app only code needs more than one crate though."

Nine cargo workspaces become three, and the port's project code -- which
sat in five places, four of them inside the framework -- becomes one crate,
`ai-app`, in `app-rust/`:

  client-core                -> app-rust/src/client
  iris/transcript-ui         -> app-rust/src/ui
  iris/transcript-fixture    -> app-rust/src/ui/fixture.rs + tests/ + touch/
  iris/desktop-app           -> app-rust/src/desktop + src/bin_desktop.rs
  iris/android-app           -> app-rust/src/android + android-project/
  android-shell              -> app-rust/src/shell

iris/ keeps core, macro, the iris crate, tabs-ui and rig-input, and now
mentions no session, transcript, setup or server anywhere.

Only two of the old splits had a reason that survived reading. event-model
stays a crate at the repo root because server/ depends on it too, so a
crate is what makes the backend and the app agree by construction. The two
Android .so names looked like a hard constraint -- a package produces one
library artifact -- until P2 turned out to already plan merging those two
Android apps into one; both faces now come out of libai_app.so, picked
apart by features so `--no-default-features --features shell` keeps wgpu,
parley and iris out of the Compose app's APK. docs/RUST.md's "One app
crate" has the rest, including what each remaining feature is for.

DECISIONS.md and SUBAGENTS.md move into docs/ with everything else.

Verified: ./run-tests.sh and `cd iris && cargo test` green, clippy and fmt
clean in all five workspaces, `cargo ndk -t x86_64` links libai_app.so,
build-apk.sh produces an APK that installs and launches on this checkout's
emulator (Gl ... virgl, as expected), and the phone-sized headless
screenshot renders the transcript unchanged.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-08 23:36:38 -04:00

51 KiB

Decisions taken for Iris to review

Short list of design choices made by the design agent without asking, so they can be judged and reversed later. Detail lives in RUST.md (and IRIS.md for iris API changes); this file is only the summary. Newest first. Items marked DEFERRED are ones the agent chose not to decide alone.

2026-09-08 (the port is one crate, and iris is framework-only)

Asked for by Iris directly, so the shape rather than the fact is what is open to review. app-rust/ is one crate, ai-app, with client, ui, desktop, android and shell as modules; iris/ holds only core, macro, iris, tabs-ui and rig-input. docs/RUST.md's "One app crate" has the table and the reason each old split did or did not survive.

Three calls made inside that, none of which she named:

  1. event-model stayed at the repo root rather than moving into app-rust/ -- her choice when asked, since server/ depends on it too.
  2. The Android application id and Java package were left alone (dev.iris.android.demo, label "iris android-view demo"), though both now name the wrong thing. Renaming them makes the next install side-by-side rather than an upgrade on her phone, and changes the DevLogProvider authority Dev Updater reads. Say the word and it is a small change.
  3. The bench fixture is behind a fixture feature that is on by default, with build-apk.sh passing --no-default-features so an APK carries the 1.9 MB fixture only when it asked for bench. The alternative -- default off -- would have made cargo test silently skip the six harness suites, which is the worse failure.

2026-09-08 (last: scrolling moves out of the list)

Agreed with Iris in the exchange that followed, so most of this is her call rather than mine. IRIS.md has the account. What I decided along the way, and would flag for reversal:

  • A third Widget method, scroll_offset, beyond the two we agreed. apply_scroll's remainder is exact only when the wall was already in view, and a lazy span usually cannot see its wall until it has walked there -- so Scroll reads the child's accumulated movement after the placing draw instead of adding remainders up, which would drift.
  • One scroll-delta convention, the finger's. The two widgets had opposite ones under the same name; LazySpan::scroll is now private and the single negation lives in its apply_scroll. Call sites that passed -dy/-v pass them straight through, and one fixture's expected velocity flipped sign with its magnitude unchanged.
  • The transcript builds its Scroll by hand rather than through .scrollable_to_end(), because that helper registers a finger drag and Selection is already the arbiter for those frames -- two DragGestures seeing one gesture is what DragGesture's own doc rules out. The wheel is registered identically; only the drag differs.
  • DEFERRED: the pin stays in each widget. Iris asked for amt and the at-end control to live in Scroll; amt does, the pin does not, because applying a pin happens when a row is appended -- between frames, with no painter -- so moving it needs a fourth Widget method or a parameter on apply_scroll. Nothing external edits a pin today. docs/IRIS_TODO.md carries it.

2026-09-08 (later still: the list's overscroll clamp, in frame)

Finishes the item the previous entry deferred. IRIS.md has the account.

  • List lays out a second time within the frame when its walk lands off the end of the content, instead of writing the correction to the anchor and asking for another frame. The extra walk is paid only on an overscrolled frame, and it is mostly O(1) moves.
  • Painter::draw_again is removed, List having been its only caller -- so the framework no longer offers a way to ask for a corrective frame at all.
  • List::place's top-known and bottom-known cases are one path (Placement::edges), which is the "write the logic once" rule applied to two symmetric directions rather than a behaviour change.

2026-09-08 (later: a scroll area measures and places in one frame)

From Iris's phone report about the composer's padding while typing newlines, and the rule she stated when she read the first fix: layout is a pure function of the state, nothing self-heals, and two draws to place something happen in the same frame. IRIS.md's entry has the account.

  • Scroll::draw draws its child twice -- once at last frame's length to measure it, once at the measured length to place it -- instead of placing against the stale length and leaving a wrong frame on screen. The second draw is free unless the content's length changed.
  • An end-anchored Scroll is at its end on its first drawn frame, a consequence of the above. Two layout tests now build their area with at_end: false, which is what they meant: they scroll down from the top.
  • List::clamp_to_content's next-frame correction is left in place and written down in docs/IRIS_TODO.md instead of fixed here, because List::place is a larger piece of machinery and deserves its own before/after on the phone.

2026-09-08 (every crate to its latest version, wgpu 28 -> 30)

At Iris's request. RUST.md's "Every crate to its latest version" box has the full list and the migration.

  • wgpu 30 taken now rather than pinned at 28. Two majors of API change, all mechanical (instance descriptor, optional bind-group and vertex-buffer slots, Queue::present, a CurrentSurfaceTexture enum), and one that would have been a startup abort on a device rather than a compile error: naga now demands @interpolate(flat) on the shader's integer varyings. Verified on both backends before this was called done, since a renderer that compiles proves nothing.
  • The desktop instance now carries winit's display handle. wgpu 30 asks for it when a GLES surface will be presented on Wayland, which is what this machine's Vulkan-to-GLES fallback produces. Android passes none: its surface comes from a NativeWindow.
  • syn 2 -> 3, pollster 0.4 -> 1.0 with no source change in iris/macro or anywhere else.

2026-09-08 evening (the fling is shared; a cancel is not a release)

From Iris's four-item phone report; RUST.md's "2026-09-08 (evening)" box has the reasoning and the tests, IRIS.md the summary.

  • A Flinger that does not know which way the content moves. Every scroll area flings now, on either axis, as Iris asked -- and the physics is one type shared by List and Scroll rather than a copy each. The choice worth reviewing is the seam: Flinger owns the curve and the clock, and the caller owns the sign convention and where the content ends. Rejected: teaching Flinger a direction, which would have to be told to it -- and being told is the same thing as not knowing, with an extra field to get wrong.
  • A cancel is a first-class end to a gesture, not an early release. CursorState::cancelled is new state on the pointer sample, set by Android's ACTION_CANCEL and the harness's TouchAction::Cancel. Rejected: mapping a cancel to PressEnd and having each widget decide what to suppress, which is what shipped and is why leaving the app flung the transcript.
  • A DragGesture ignores a Cancel it caused. One gesture is driven by several widgets, so the widget that was pressed can be a "loser" on the frame its own gesture won. The test is whether the gesture's own capture id is the holder. This is what makes it safe for every widget driving a gesture to register the whole drag_senses() set, which is now the rule without exception.
  • List::place draws a resized row twice in one frame. The old comment accepted a one-frame lag by analogy with Scroll's content length. That analogy was wrong: a stale length only misplaces the next thing, while a stale box is drawn, because a background fills whatever box it is handed. The extra draw is bounded to frames where a row's height actually changed.

2026-09-08 (iris ships an icon font, and the drawn mark is deleted)

  • Directed by Iris. Her question on seeing widget::mark: "why does mark exist? The font should be working if it's working for compose and nerd fonts are bundled." It was not: the Compose app draws its icons from its own committed Nerd Fonts subset, while iris was setting the disclosure mark with bare Unicode geometric codepoints (U+25B8/25BE/25B4) out of whatever face the platform resolved -- an empty box on her phone, a dot on this VM. The 2026-09-07 entry below, which said "iris had no equivalent icon font to keep", is what left that gap: iris had no icon font because it had never had one, not because it needed none.
  • So iris now bundles the same kind of subset: iris/core/build-icon-font.sh writes iris/core/assets/fonts/nerd_icons.ttf (992 bytes, three Material Design glyphs today), iris::icon names the codepoints, and Family::Icons draws them. This does not reopen the platform-fonts decision: body and monospace text still come from the platform, and an icon is the opposite case -- a small, closed, known set of codepoints, which is exactly the division the Compose app already makes.
  • iris::widget::mark is deleted (added earlier the same day). It drew a correct triangle, but only a triangle, and every further icon would have been another bespoke rasteriser. An icon as text also takes the size, colour and baseline of the line it sits in for free.

2026-09-08 (the emulator is a GLES machine, and Vulkan is verified elsewhere)

  • Directed by Iris, carried out here: "make sure the setup uses GL for the android emulator and remove any vulkan requirements. That'll be tested through both the desktop version as well as my phone." So the emulator is settled as a GLES rig and nothing chases hardware Vulkan in it any more; the Vulkan path is covered by the desktop build and by her phone.
  • Nothing had to be forced to make that true. Measured in the guest the same day: the emulator has no hardware Vulkan at all (its only Vulkan is SwiftShader, in software) and its GLES is the host's real RX 7900 XT through virgl at ES 3.1. iris's existing runtime fallback -- Backends::PRIMARY, no adapter, rebuild on Backends::GL -- already lands there, verified end to end with an ordinary (no force-gles) debug APK.
  • The emulator and the phone therefore run the same binary, differing only in what that binary finds. That is deliberate and worth not undoing: a build flag that changed the backend would mean the thing measured on the emulator is not the thing shipped. force-gles stays, but only for pinning the backend on a machine that does have Vulkan (the desktop), and never for a phone build.
  • Every run now says which adapter drew it. The Android renderer logs the full adapter line at startup the way the desktop already did -- only the backend enum was logged before, which cannot separate Gl on the host's GPU from Gl on SwiftShader, or a phone's real Vulkan from a software one. run-bench.sh prints that line before any number.
  • No Vulkan requirement was found in iris to remove. device_limits() asks for nothing beyond wgpu's defaults (and zeroes the compute fields), neither backend requires a feature, and both probe rather than .expect() an adapter. What was removed was the documentation telling people to boot the emulator with SwiftShader Vulkan.

2026-09-07 (platform fonts, not bundled ones)

  • Iris's own decision, carried out as directed: removed the 3.6 MB of bundled Noto Sans/Noto Sans Mono TTFs from iris-core and load text from the platform's own font collection instead (fontique's system discovery, already on by default). Matches what the Compose app does -- it takes body text from FontFamily.Default and code text from FontFamily.Monospace, both platform-resolved, and ships no text font of its own. Rejected alternative (the one this pass had left open 2026-09-06): subsetting the bundled Noto Sans to Latin/common punctuation instead of removing it outright, which would have kept identical rendering across devices for a smaller (not zero) size cost; Iris chose to match Compose instead.
  • .so -3,748,136 bytes (11,193,608 -> 7,445,472), matching the original 3.6 MB estimate. Fallback still lands on the platform's own tofu for a codepoint no resolved face has (checked with CJK + emoji on desktop) rather than blank space, so the UI_RULES unknown-glyph rule still holds.
  • Gap found, then closed same day: this fontique version's Android backend never resolved the Monospace generic family at all (confirmed on this checkout's emulator, mono=None in the startup diagnostic) -- two pre-existing bugs in fontique's own fonts.xml parsing stacked (an ordering bug, and a <family name="monospace"> declaration whose <font> children the backend's parser never reads), not something this change introduced, but this change is what stopped masking it (the bundled mono font used to be registered ahead of the broken platform lookup, so it always won). Checked linebender/parley's main branch on GitHub: neither bug is fixed there, so there was no newer release to bump to. Fixed instead in iris-core itself (TextData::patch_android_monospace, Android-only): reads /system/etc/fonts.xml's own "monospace" declaration for the font filename it names, then registers whichever of fontique's actually- scanned families owns that file as the Monospace generic -- the same authority Compose's Typeface.MONOSPACE resolves through, without pinning an OEM-specific family name. Verified on this checkout's emulator: mono=Some("Droid Sans Mono"), and a screenshot showing the bench-fixture's code block and tool-card values in a visibly monospaced face beside sans body text; the desktop fontconfig backend is unaffected (still resolves monospace correctly, confirmed unchanged). docs/RUST.md's "Platform fonts (2026-09-07)" has the full account.

2026-09-07 (a phone log reaches Iris through Dev Updater's own tab)

Supersedes the "how a phone log reaches Iris" entry below, same day. Iris's call once the route was working: put it in Dev Updater properly rather than smuggling the lines through ai-server's log.

  • The app exposes its own log on the device, and Dev Updater reads it there. A ContentProvider at <applicationId>.devlog, one table of lines queried with ?since=<seq> so a poll is incremental, plus a status row (held, dropped, newest_seq). Dev Updater's phone app polls it while the component's Runtime tab is open and forwards what is new to its own build machine, into that APK component's runtime log -- so the same tab renders both kinds and the history outlives the phone. No tunnel, no token, no second enrolment: the two apps are on the same phone.

    It is a contract, not a feature for iris. Written down in dev-updater's README.md ("An app's own log"), so any app that server delivers gets the tab by implementing it; the Compose app in app/ can do the same later. That is the reason it beat the route below on its second look -- the earlier one only ever worked for the one project that had a server, and put a phone's lines under a different component than the one they came from.

  • Read access is protectionLevel="normal", and that is a real trade. signature is what this wants and is not available: Dev Updater and the apps it delivers are built on one machine but signed with different locally generated keys, so a signature permission would be held by nothing at all. What normal costs is that any app on that phone which requests dev.updater.permission.READ_DEVLOG by name can read another app's dev log. Accepted because these are development builds on a development phone and the alternative was no log; stated in the manifest beside the declaration and in dev-updater's README so it is not rediscovered as a surprise.

  • The provider polls rather than notifying. notifyChange was not implemented: the ring is filled by a log::Log backend on whatever thread logged, and giving that a route to a ContentProvider means plumbing a callback through client-core for every platform. Dev Updater's contract therefore says it polls (about a second, only while the tab is open), which is what keeps implementing the contract cheap -- a provider that does notify loses nothing.

  • What was deleted, so there is one mechanism: client-core's log_upload module, POST /client-log on ai-server, the AI_APP_LOG_HOST/_PORT/_TOKEN baking in iris/android-app/build.rs (which left that file with nothing to do, so it is gone too), and the uploader fields on both Android clients. Kept: the ring, RingLogger, install_process_logger, and the Diagnostics line counting what is held. The upload-status line there is now "devlog provider: content://" -- named from what the provider registered rather than composed from the package here, so a screenshot of that pane is evidence the contract is live and says which package's log it is.

2026-09-07 (how a phone log reaches Iris) -- superseded, see above

  • The app sends its own log to ai-server, and Dev Updater shows it as ai-server's runtime log. Iris has no adb/logcat on her phone, and Android forbids one app reading another's logcat, so the app has to carry its own copy and post it somewhere. POST /client-log on ai-server re-emits each line into that server's own tracing output; Dev Updater already runs ai-server as a Managed component, whose stdout its own service script redirects to a file and reports through GET /apps/{key}/components/{name}/logs?kind=runtime, which the phone app's log dialog already offers as a Runtime tab for a server component. So no change to Dev Updater at all -- one route on ai-server, and the client in client-core.

    Rejected: posting to Dev Updater's own server (the first candidate, and what the entry above went on to build -- the estimate below was right about the work and wrong about it being too much). It would need a new authenticated write route on a TLS surface whose module doc says every route on it "is, or decides, the bytes that get handed to REQUEST_INSTALL_PACKAGES next"; a per-app device-log store; a change to component_logs so an APK component can have a runtime log; a change to the phone app's hasBothKinds = component.kind == "server" gate and to what hasRuntimeLogs means on the wire; and -- the real cost -- a second enrollment for the iris app, since it has no CA or token for Dev Updater and Dev Updater mints tokens per device by QR. Five changes across two repos against one route, for the same line landing in the same viewer.

    Rejected: a share intent from a debug button (a log file in the app's external files dir, shared by hand). It works today and needs no server, but every line costs Iris a manual export and a message, which is the round trip through a person this was meant to remove. It is still the fallback when the tunnel is down, and GrapheneOS's own per-app log export already covers the crash case (that is how the ToolInput.highlighted crash was reported).

  • The ring is in client-core, not in the Android crate. A bounded in-memory ring (2000 lines or 256 KiB, whichever bites first) behind a log::Log backend that forwards to whichever logger the platform already installed, so logcat and a desktop terminal see exactly what they saw before. The platform supplies only its own logger and its destination. Copy report appends the ring to what goes on the clipboard, and flushes the uploader first.

  • The destination is baked in at build time, from the build machine's own files (AI_APP_LOG_HOST/_PORT/_TOKEN plus the pinned CA) -- gone; the provider above replaced it. What is worth keeping from it is the reason it went: an APK good only for the server that built it cannot be built in this VM for Iris's phone, which is the case that mattered. all three or none, never two. The same trust boundary the transcript config and the Compose APK's CA already use: nothing secret is committed, and an APK is good for the server that built it. A build told nothing still keeps its ring and still copies it; the diagnostics pane says which of "not tried yet", "failing -- " and "no server configured" it is, because otherwise all three look like silence.

2026-09-06 (how a tool call looks, P1b)

  • A card that never got a result says "no result", in yellow, and it is a state Compose cannot say. A call that finished having printed nothing and a call whose turn was interrupted before anything came back both leave an empty output. Compose draws both as an ordinary finished call, which reads as a fact somebody established. There are five states now, each with a word and a colour: nothing at all for a call that worked, "running" (grey), "your turn" (peach, Compose's own wording and colour), "failed" (red), "no result" (yellow).

  • A failed call is drawn as failed, which needed a field on the wire. is_error is on the CLI's tool_result and was being dropped; the server now carries it to the phone. Reversible, but the alternative is a card that says a call succeeded because it cannot tell.

  • A group's cards do not each carry their own surface. Compose gives each card a fill and squares the corners where it faces a neighbour, so a run reads as one object broken into parts. iris has no per-corner radius, and -- more to the point -- a group built the way Compose builds it hit a framework layout defect that drew every card's text a card below its own box. So a group is one surface with its cards on it, separated by a small gap, and the 4dp inset Compose holds them off the edge by is gone. Worth revisiting once the layout defect is fixed (docs/IRIS_TODO.md).

  • A long tool output is capped at 80 lines or 4 kB with a "Show all N lines". Compose draws the whole thing, and gets away with it because its Text inside a LazyColumn lays out lazily; here the output is one text widget and shaping a hundred kilobytes of it costs what the file editor's 32 kB limit was measured against. If iris's text gets cheaper, this is the number to move.

  • A card's command is clipped, not pannable, and its summary line is clipped rather than ellipsised. Both are framework gaps rather than choices (scrollable_on on a non-editable text draws nothing; there is no overflow ellipsis), and both are worse than Compose today. Named here because they are visible.

2026-09-06 (how a markdown block looks, P1a)

  • A table is drawn as padded monospace columns, not as a grid. Your call to reverse. Compose draws a real grid: cells on a tint, each column with a 136dp floor, scrolling sideways when there are too many. iris has no grid widget, and building one would be a widget per markdown feature -- which is the thing the block model exists to avoid. In a monospace face a character count is a pixel width, so padding each cell to its column's width is alignment, the widths are still measured from the cells, and a table that is too wide pans sideways through the same mechanism a code fence already uses. The header is bold with a rule under it, and a long cell wraps inside its column (capped at 28 characters, which is what fits three columns across a phone). What it trades: no cell borders, and a table looks like code rather than like a table. If you want the grid, it is a new widget and it is a day's work.
  • Three block frames, and only three. A heading, paragraph and list are plain text with spans; a fence and a table are a rounded panel that does not wrap; a quote is a bar with the text padded past it. Everything else markdown says is expressed in span styles, which cost no widgets and no layout nodes. So a new markdown feature is a span, not a widget.
  • A list's marker is part of the text, so a wrapped item's second line returns to the left margin. Compose keeps it indented by giving the marker its own column. Doing the same here needs per-line indent in iris's text attributes; it is written down rather than done, because the list items in a real reply are usually one line.
  • A link opens on a tap and not on the end of a drag. A press that panned the transcript past a link, or that held long enough to start a selection, does not follow it -- decided by the same gesture machine that decides pan-versus-select, so there is one rule rather than two that can disagree.

2026-09-06 (composer scroll and the streaming block model)

  • A streamed message becomes a column of per-block widgets. Decided by the design agent; recorded here because it is the shape of every message on screen. A transcript row is one TextEdit today, so a streamed delta re-shapes the entire message through parley on every event -- the stream phase is the one place iris is behind Compose on your phone (p50 18.2ms vs 13.4ms). A row becomes a column of one widget per markdown block (paragraph, heading, fence, list, table) and a delta replaces only the last block, keeping every earlier block's layout. Rejected: splitting parley's layout at block boundaries inside one text widget (couples iris's text widget to markdown structure, and parley has no incremental API), and caching shaped runs per paragraph inside TextEdit (a second cache with its own invalidation beside the glyph cache). Chosen because P1's markdown block model is needed anyway, so the split happens once, in client-core, and iris stays a text renderer. Status: designed, not built -- this pass spent its budget on the composer's three layout defects; docs/RUST.md has the design and the pass conditions.
  • The composer's overflowing text now scrolls on a finger, capped at six lines and clipped to the bar. Reverses the "still does not scroll" item below.
  • A widget may not report a dp length (see IRIS.md). A rule for widget authors, enforced by a debug_assert!; nothing changes for app code.

2026-09-06 (stale-primitives and touch-scroll pass)

  • A vertical drag inside a focused composer now scrolls rather than selects. Android's own EditText does this -- a vertical drag scrolls the field, and only a long press starts a selection -- so the platform decided it. What it costs: you can no longer drag straight down inside the composer to select several lines of what you typed; use a long press and then drag, or drag sideways. Say if that trade is wrong for you.
  • Scroll gets a finger pan but no fling. List flings; a scroll area does not, because it has no per-frame tick to animate one and the areas it wraps are at most a screenful (Android does not fling a six-line text box either). Easy to add later if a scroll area ever wraps something long.
  • The composer still does not scroll its overflowed text, though the mechanism it needs is now in place. Wrapping the field in .scrollable() was tried and reverted the same day: Scroll measures its content and container against the window, so inside the MaxSize that caps the composer at six lines the two are in different spaces and the field pans itself entirely out of the bar (measured on the emulator with 474 characters in it -- the bar collapsed to its padding). Fixing that means Scroll measuring against its own offered box, which is a change to a widget the transcript and the bench shell both use, so it is its own piece of work rather than a rider on this one.

2026-09-06 (defect pass)

  • The keyboard-open diagnostics overlay is gone; the capture only logs now. It was added when on_insets_changed was not firing at all and there was no way to get a report off the phone. It fires reliably since the activity went edge-to-edge -- and what that looks like in use is a full-screen report covering the app every time the keyboard opens, with its own Copy/Close buttons sitting underneath the keyboard, so it cannot be dismissed (reproduced on the emulator this pass: two tap 'CLOSE' runs left it up). An interruption for something nobody asked for, over the app you are trying to type into. The named Diagnostics button still shows the same text on demand, and the new iris surface:/iris insets: log lines carry the lifecycle a logcat pull needs. Reversible: capture_keyboard_diagnostics is still the one place this is decided, and PlatformHandle::show_diagnostics_overlay is still there.

  • The bench shell's report pane is sized to its report, not to a share of the window. It held .height(rest(1)) beside the transcript's rest(2), so an empty TextEdit reserved a third of every screen -- which is what Iris's "the app does not start with keyboard spacing correct" screenshot was showing, with the composer two thirds down and black below it. It is .max_height(dp(260)) now and sits above the transcript rather than under the composer, where it was eating the navigation-bar clearance. Cost: a filled report is clipped at 260dp rather than scrolling (a Scroll there drew itself off the top of the screen, since Scroll pins to the end of its content and reports its content's full length to the parent -- worth fixing in Scroll, not worked around here). "Copy report" and logcat still have the whole thing.

2026-09-05

  • iris no longer asks every device for compute-shader limits it never uses. adapter.request_device (both iris/src/android/render.rs and iris/src/default/render.rs) used Limits::default() plus an override for max_buffer_size, and Limits::default() unconditionally requests desktop-tier compute limits (max_compute_workgroups_per_dimension: 65535, per wgpu_types) even though nothing in iris/iris-core creates a ComputePipeline or writes a @compute shader stage — confirmed by grepping the whole tree, not assumed. That crashed request_device outright on the Android emulator's software GL path (EMU_GPU=software, --features force-gles): SwiftShader's GL reports itself as OpenGL ES 3.0, which has no compute shaders at all, so the adapter's real limit is 0 against the unconditional request for 65535 — RUST.md's "Software mode ... crashes for a third, different reason," 2026-09-05, earlier today. The same would happen on any real GLES-3.0-only Android device, not just the emulator. Fixed by a new iris_core::device_limits() (iris/core/src/render/mod.rs), shared by both platform backends so the two requests cannot drift, that zeros the six max_compute_* fields explicitly rather than switching to a downlevel Limits preset — Limits::downlevel_webgl2_defaults() was considered and rejected: it also zeros max_storage_buffers_per_shader_stage, and shader.wgsl's vertex stage reads four var<storage> buffers (rects, glyphs, masks, move_offsets), so that preset would trade the compute crash for a bind-group-layout one on the same downlevel hardware this is meant to support. No capability check or fallback path was needed since nothing is being disabled — the request is simply narrowed to what the pipeline actually uses. rigs/gpu-probe's own mirrored limits (it is deliberately its own crate, not a workspace member, so it cannot call device_limits() directly) were updated to match, and confirm IRIS DEVICE: ok against this VM's own Vulkan and GL adapters. Not verified this pass: the specific SwiftShader-ES-3.0 crash this fixes, on-device — the EMU_GPU=software cold boot this needs would have force-restarted this checkout's emulator while another session was actively running its own app on it (com.example.aiapp had window focus at the time), so it was left for a pass when the emulator is free rather than disrupting that session. Everything reachable without the emulator is clean: cargo fmt/clippy --workspace --all-targets/test --workspace, cargo ndk build/clippy for iris-android-app with force-gles, and gpu-probe against this VM's own Vulkan and GL(ES 3.2, which still has compute and so would not have reproduced the crash even before this fix — not a substitute for the real ES-3.0 test).

  • P0's Compose half is built and smoke-tested on the emulator — the bench build type, the shared app/bench-fixture/ transcript, and an in-process fake backend (BenchFixture.kt/BenchNetwork.kt) that answers TranscriptSource/EventStream from an in-memory event log instead of a real server, so the fold and paging under test are the real ones. Full account, the smoke run's report, and what is deliberately left (the iris half, the real on-phone runs) are in RUST.md's P0 box. Not a decision to review so much as the gate itself now being runnable — flagged here because it is the first half of something Iris explicitly asked to see before P1.

  • P0's iris half is also built and smoke-tested on the emulator, 2026-09-05. A new bench Cargo feature on iris-android-app, on top of transcript-screen: the same checked-in fixture (include_str!, no asset pipeline needed), the same 24-swipe scroll loop animated through List::scroll and the same 400-event/20s streaming phase through fold_event, "Run benchmark"/"Copy report" as named accessible controls, and the same three added report fields (process CPU time, peak RSS, battery current) via direct JNI calls (bench_jni.rs::PlatformHandle) since android_view has no BatteryManager/ClipboardManager wrapper of its own. One small public API addition to get there: AndroidAppState::platform_ready (IRIS.md), a default-no-op lifecycle hook handing an implementor a JavaVM + GlobalRef it can call Java through from any thread. Packaged with a new release build type on iris-android-app's own Gradle project (there was previously only debug), signed with the same key app/build-apk.sh generates. Smoke run and the full report are in RUST.md's P0 box; not attempted this pass: the real on-phone runs and Iris's pass/fail call, which is the actual gate.

  • The intermittent touch-scroll dropout is root-caused and fixed: a missed ACTION_DOWN hit-test, not the previously-suspected coalesced first ACTION_MOVE. Diagnosed by temporary logcat tracing of every touch event, DragArbiter state transition and Selection::drag dispatch (removed once confirmed), reproduced on this checkout's own emulator against a real sandbox session. The trace showed the actual mechanism: a gesture's ACTION_DOWN lands wherever the finger actually is, which is not guaranteed to fall inside the same row-local sensor region a later ACTION_MOVE in the same gesture lands in (a row's own padding/gap, or its non-selectable sender-name header, is pointer-transparent to iris::sense::CursorSense). When that happens, the widget that ends up handling the gesture never saw PressStart, so DragArbiter sits in Idle — which answers every subsequent frame with Undecided and has no way to tell "no press is happening" from "a press is happening but I missed its start," so it never recovers on its own for the rest of that gesture. One real trace showed exactly this: touch Down/Move/Up all delivered correctly, but zero PressStart reaching the arbiter, state=Idle unchanged from first frame to last. Fixed at the call site that has the context to recover (iris::transcript_ui::selection::Selection::drag, iris/transcript-ui/src/selection.rs): a new DragArbiter::is_idle() (iris/src/sense.rs) lets it notice a Pressing frame arriving with the arbiter still Idle — which can only mean a missed PressStart, since a Pressing sense requires the button to genuinely be down — and start the press there instead of where it was missed. Three new unit tests in sense.rs's drag_arbiter_tests and one in transcript-ui's selection::tests (the latter fails on the code before this fix). Commit follows. Not the same failure the earlier pass's DECISIONS.md DEFERRED item speculated about (a coalesced first ACTION_MOVE skipping slop detection) — that hypothesis is now ruled out; the arbiter's own slop/long-press logic was never wrong. RUST.md's I5 box, "Touch-scroll dropout root-caused, 2026-09-05" has the full trace.

  • P0, a phone benchmark gate before any porting, asked for by Iris 2026-09-05: "before P1 I'd like to see benchmarks & also maybe stress test on my own phone ... If it doesn't match compose reasonably well then I don't think I'd wanna continue." Design (RUST.md's P0 box has the detail): the same embedded synthetic fixture in both apps with no server needed; the same scripted scroll loop then a streaming phase, run programmatically since the phone has no usable system tracing and no agent can drive it; the same report from both (frames, janky %, p50/p90/ p99, process CPU time, peak RSS, battery current where readable) with a copy button; the iris app under its own id and the Compose one as a new bench build type with an id suffix, so neither replaces her production install; two arm64 APKs plus instructions delivered under ~/host/bench/. The gate is hers: iris within a reasonable margin of Compose release on p50, p99 and CPU time, no crashes, no visible stutter. If it fails, the port stops.

  • The rest of the port is one UI crate, iris/app-ui, grown out of iris/transcript-ui rather than started beside it. It holds a Screen enum plus a back stack — the Rust equivalent of AppRoot.kt's when — and iris/desktop-app/iris/android-app become thin entry points over it. Chosen over a fresh crate because transcript-ui already has the right generic shape (Rsc: HasEvents + Rsc::State: FocusHost) and the client-core/event-model path dependencies every later screen needs, so growing it in place is the smaller diff. Platform-only code (notification service, share target, QR scanner, Keystore token, deep-link enrolment) stays in the E3/E5 Java shell (android-shell/ + app/shellApp) rather than moving into this crate, since none of it is a screen. The Android APK is built by cargo xtask apk (E5), merging the app-ui cdylib into the E3 shell so there is one app rather than a demo shell plus a service shell. app/androidApp (the Compose app) stays untouched and is the baseline every step is measured against, until parity is reached (P7 decides the switch, and is itself a load-bearing decision left to Iris). Order is by risk to the daily-use path: session screen first (P1, where every hard behaviour already lives), then the shell merge and a real phone install (P2), then root tabs (P3), the explorer (P4), settings/enrolment (P5), desktop parity (P6), and the cutover itself (P7). Full plan: RUST.md's "The port, in order (decided 2026-09-05)".

  • iris gets its own measured frame report, rather than waiting on a dumpsys/gfxinfo answer that cannot see a SurfaceView's GPU-drawn frames. iris_core::FrameReport (iris/core/src/render/frame_report.rs) times each frame's wall clock from the same point render()'s redraw starts to just after queue.submit + present() — the span Compose's own render report and gfxinfo both count — into a fixed 4096-entry ring (no allocation per frame; report() is the only place that allocates, and only on a button tap). The report gives total frames, janky % over the same 16.7ms budget gfxinfo uses, P50/P90/P99 and the worst, plus a reset. Exposed the way the Compose app's copy-button report already is: two named controls ("Frame report", "Reset frame report") on the transcript screen, tappable by accessibility name via ui-trace, logging under this crate's fixed android_logger tag (iris-android-app) so a script can grep "iris frame report" the way transcript-bench.sh greps "ai-app render report". The report's own Display line says plainly that it measures up to the present() call returning, not GPU/compositor completion — wgpu's present() is not fenced against either, so presenting that span as "time to reach the screen" would be a measured-looking number that is actually inferred, which the standing UI rule forbids.

  • ui-trace gains a hold-then-drag gesture, additive, in emulator-tools. Neither of its two existing actions can produce "hold stationary for LONG_PRESS, then move without lifting" — tap has no hold and swipe X1 Y1 X2 Y2 MS interpolates motion across its whole duration from t=0. A new action presses, waits, then moves to a second point and releases as one continuous touch (raw sendevent/MotionEvent injection, extending whatever mechanism the existing swipe already uses), so DragArbiter's pan-vs-select rule (iris/src/sense.rs, already covered by 8 unit tests against a synthetic clock) can finally be driven on a real device instead of only in a test harness.

  • Touch drag on a transcript row follows Android's own rule: a vertical drag pans the list immediately; a stationary press held 500 ms starts a text selection which further dragging extends; a horizontal drag while something is already selected extends that selection without the wait. One DragArbiter per list decides it (iris/src/sense.rs). Chosen over a "text layer always wins" or "list always wins" rule because either loses one of the two gestures a reader expects.

  • E4's desktop shape is a new iris/desktop-app crate: a winit window holding transcript-ui's screen beside a session list, talking to a real ai-server through client-core. It enrols by pasting the same aiapp://enroll?… link a phone scans (client-core::config::EnrolledServer) and keeps it owner-only under $XDG_CONFIG_HOME/ai-app-desktop/. The pinned CA is a path given on the command line, not baked in. Chosen so the phone and desktop share one enrolment format and no second one is invented.

  • I5's Android integration extends iris-android-app (I2's shell) behind a Cargo feature (transcript-screen), rather than a third shell crate. That project already has the Gradle module, the IrisView/MainActivity Java, and the JNI registration; the only thing a second screen needs on top is a different AndroidAppState, the same axis tabs_ui::build/transcript_ui::build already vary along on the winit side. tabs-screen/transcript-screen are mutually exclusive and each pulls in only its own deps, so the plain tabs build (I2/I4) is untouched.

  • Order of remaining work, updated 2026-09-05: the two in-flight pieces and I5's Android integration are all done; next is giving iris its own frame-timing report so item 3 below can be decided by a number.

  • DECIDED by Iris, 2026-09-05: iris is the app's framework; Masonry was the calibration. Her words: "I think iris definitely makes more sense based on the limitations we've found." The limitations: Masonry has no touch scroll on Android (E2), no per-span rich text and no cross-row selection on the pinned commit (E2), and its keyboard bridge is a TODO (E1); iris carries the same screen under the Compose baseline on the host GPU (p50 15.0 ms against Compose's 20.0 ms, RUST.md's I5 box). What follows: the E-steps are closed as calibration, and the port proceeds on iris — screens, the shell (E3/E5), and client-core underneath. The item below is kept as the record of what she decided from.

  • Was DEFERRED — whether to commit to iris over Masonry for ai-app. Updated 2026-09-05 with the clean comparison the recommendation wanted: same sandbox session content, same emulator, EMU_GPU=software, one session. Headline numbers (RUST.md's I5 box, "Clean scroll comparison, 2026-09-05," has the full table and every caveat):

    app build frames janky % p50 p90 p99 worst
    Compose (in-app report) debug 1102 99.0% late 33.8ms 50.6ms 79.5ms --
    Compose (dumpsys gfxinfo) debug 1499 21.15% (95.66% legacy) 32ms 48ms 150ms (p99) --
    iris (FrameReport) release 299 94.65% 79.1ms 98.6ms 117.8ms 212.6ms
    iris (FrameReport, repeat) release 233 94.42% 109.3ms 130.8ms 147.1ms 150.5ms

    Not a clean apples-to-apples reading, stated plainly rather than smoothed over: iris had to be built release (debug SIGSEGVs on this emulator's Vulkan loader, I4's finding) against Compose's mandated debug build, so this asymmetry likely understates iris's gap rather than the reverse; the three frame-time sources measure different things (Compose's own phase accounting vs. Android's HWUI deadline-miss definition vs. iris's redraw-start-to-present window, the last of which dumpsys gfxinfo cannot see at all for iris's SurfaceView); and both figures are emulator numbers under software rasterisation, which Compose's own in-app report shows already costs 20-34ms/frame in swap+gpu alone under this GPU mode, so a same-mode iris number well above 16.7ms was expected going in for either app. A second pair under -gpu host was not taken this pass. The earlier session's suspected intermittent touch-delivery dropout was not reproduced this pass — the zero-frame results this time traced to this pass's own script bug (a cd that changed which emulator ui-trace targeted), not the emulator; a CPU-load rise during the gesture was observed by a sampler running throughout, but did not correlate with any failure, so the original candidate is neither confirmed nor ruled out. The choice in front of Iris, updated: decide now on the structural-plus-functional case already made (iris works end-to-end where Masonry's scroll gesture doesn't exist at all on Android) plus this table — reading the two build profiles and three jank definitions with the caveats above rather than as a single number — or ask for a same-profile, same-GPU-mode rerun first. RUST.md's I5 box has the full account.

    Updated 2026-09-05, the -gpu host pair taken. Real GPU rendering (force-gles -- the default Vulkan backend has no adapter at all under plain host-GPU boot, confirmed by the exact wgpu error) reverses the software-mode shape:

    app build GPU mode frames janky % p50 p90 p99 worst cpu p50 gpu-wait p50
    Compose (in-app report) debug host (virgl) 1268 96.4% late 20.0ms 28.4ms 37.7ms -- -- --
    iris (FrameReport), best of three, 2026-09-05 release, force-gles host (virgl) 439 46.24% 15.7ms 23.3ms 31.2ms 57.4ms 1.2ms 13.2ms

    Under real GPU rendering iris's median frame is faster than Compose's, not the 2-3x-slower shape the software-mode table shows. A new split inside FrameReport (redraw-to-submit vs. submit-to-present, commit e2a1fad) says why: iris's own CPU work per frame is a median ~1ms -- almost the entire frame is time spent handing the frame to the driver, not in iris's layout/text/primitive code. This is consistent with the earlier software-mode gap being mostly SwiftShader's CPU rasterisation cost rather than an iris-specific slowness. Still not proof, and now closed as unanswerable rather than merely untaken: a same-mode software force-gles run to isolate the backend was retried 2026-09-05 after fixing the compute-limit crash the first attempt hit, and hit a second, structural wall instead — SwiftShader's ES 3.0 GL path has no storage-buffer capacity at all, and shader.wgsl reads var<storage> buffers unconditionally, so reaching that path needs a shader rewrite, not a limits fix (RUST.md's I5 box, "The three remaining I5 verifications, closed 2026-09-05," item 2). The intermittent touch-scroll dropout this pass also reproduced is root-caused and fixed as of the same date (a missed ACTION_DOWN on a row's padding/header left DragArbiter stuck in Idle); three clean iris-scroll.sh runs post-fix each scrolled all 24/24 swipes, replacing the single-attempt 62-frame reading this table used to carry. RUST.md's I5 box, "Where iris's frame time goes, 2026-09-05, the -gpu host pass," and "The three remaining I5 verifications, closed 2026-09-05," have the full account. The iris-vs-Masonry choice itself is still Iris's to make.

Problem. Every phone build pinned the CA of the machine that compiled it -- the Compose app from GeneratePinnedCert, the iris app from build.rs reading $XDG_CONFIG_HOME/ai-app/certs/ca.pem. That is fine while the two are the same machine and impossible when they are not, which is exactly the iris client's situation: cross-compiled in this VM, delivered to a phone, run against ai-server on the host. Baking the host/port/token as well made it worse -- a token in a built artifact.

Decided: the CA rides in the enrolment link, as &ca=<base64url of the DER> (wg_app_link::enroll::ca_param), optional and per mint. The app that opens the link pins what the link said, and an APK built anywhere works against whatever server it is pointed at.

Two alternatives were worked out and rejected.

  • A CA fingerprint in the link, pinned at the TLS handshake. The smallest link (43 more characters) and the strongest shape, but ureq 3.4 exposes no hook for a custom rustls ServerCertVerifier: its TlsConfig builds the ClientConfig itself, so this needs a hand-written Connector on the unversioned API and rustls as a direct dependency of client-core. A lot of machinery in the one crate that must stay light.
  • A fingerprint in the link plus an unauthenticated GET /ca.pem. Small code, but it needs a first connection with verification disabled, and it breaks a documented, tested posture -- auth.rs's "gates every route with zero unauthenticated endpoints", which is a load-bearing decision rather than an implementation detail. Not something to change silently for this.

What it costs, measured rather than guessed: on this project's P-256 CA the link goes from 89 bytes to 652, and print_enrollment's terminal QR from 45x23 to 93x47 characters. That is why the parameter is the minter's choice per call: ai-server passes it (its iris client needs it), dev-updater passes None (its app is built on the machine it talks to, and its QR stays scannable in an 80-column terminal). The URI printed under the QR is the fallback either way, and is the path Dev Updater's Enroll button already uses -- it opens the link with ACTION_VIEW, so Android offers whichever apps registered the scheme, which needed no change here.

The CA is a public certificate, so putting it in the QR leaks nothing the token did not already: photographing the terminal still costs exactly the token, which is rotatable.

The log upload's destination is moot, so it is not wired to this. On the same day Iris decided Dev Updater will read an APK's runtime log from an on-device ContentProvider instead, which removes log_upload, POST /client-log and the AI_APP_LOG_* baking altogether -- so the enrolment landed without touching any of them, for that change to delete whole.