docs: close I5's three blocked verifications (24/24-swipe, backend isolation, cold-boot bench)
Ran three clean iris-scroll.sh passes on a cold -gpu host boot (all 24/24 swipes confirmed scrolling via clustered render() timestamps, not inferred from frame count) and retook the host-GPU table's iris row as a best-of-three. EMU_GPU=software + force-gles still cannot produce a GLES number on this hardware -- after the earlier compute-limit crash was fixed, device creation now aborts on max_storage_buffer_binding_size instead (SwiftShader ES 3.0 has no SSBOs, and shader.wgsl reads four var<storage> buffers unconditionally), so the SwiftShader-Vulkan-vs-GLES question is closed as structurally unanswerable rather than answered. A fresh cold-boot run-bench.sh reading for P0's bench build is in line with the earlier warm-AVD readings, closing that box's own caveat too. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
1 parent
c07d544aeb
commit
a27fbdb029
2 files changed
+140
-72
No files matched your search
+20
-14
@@ -264,24 +264,30 @@ marked **DEFERRED** are ones the agent chose not to decide alone.
|
||||
| app | build | GPU mode | frames | janky % | p50 | p90 | p99 | worst | cpu p50 | gpu-wait p50 |
|
||||
|---|---|---|---|---|---|---|---|---|---|---|
|
||||
| Compose (in-app report) | debug | host (virgl) | 1268 | 96.4% late | 20.0ms | 28.4ms | 37.7ms | -- | -- | -- |
|
||||
| iris (`FrameReport`) | **release**, `force-gles` | host (virgl) | 62 | 41.94% | 15.0ms | 21.8ms | 37.1ms | 37.1ms | 0.2ms | 12.9ms |
|
||||
| iris (`FrameReport`), **best of three, 2026-09-05** | release, `force-gles` | host (virgl) | 439 | 46.24% | 15.7ms | 23.3ms | 31.2ms | 57.4ms | 1.2ms | 13.2ms |
|
||||
|
||||
Under real GPU rendering iris's median frame is *faster* than
|
||||
Compose's, not the 2-3x-slower shape the software-mode table shows. A
|
||||
new split inside `FrameReport` (redraw-to-submit vs. submit-to-present,
|
||||
commit `e2a1fad`) says why: iris's own CPU work per frame is a median
|
||||
0.2ms -- almost the entire frame is time spent handing the frame to the
|
||||
~1ms -- almost the entire frame is time spent handing the frame to the
|
||||
driver, not in iris's layout/text/primitive code. This is consistent
|
||||
with the earlier software-mode gap being mostly SwiftShader's CPU
|
||||
rasterisation cost rather than an iris-specific slowness, but is not
|
||||
proof of it: a same-mode software `force-gles` run to isolate the
|
||||
backend crashed for an unrelated reason (SwiftShader's GL path reports
|
||||
itself as OpenGL ES 3.0, which has no compute shaders, and iris's device
|
||||
request assumes them unconditionally) — real scope to fix, not done
|
||||
here — and the two apps' frame populations still differ in kind the same
|
||||
way the software-mode caveats describe. A real intermittent touch-
|
||||
scroll dropout was also reproduced this pass (six consecutive swipes
|
||||
produced zero redraws while taps kept working; an identical retry then
|
||||
succeeded) and is not explained. RUST.md's I5 box, "Where iris's frame
|
||||
time goes, 2026-09-05, the `-gpu host` pass," has the full account. The
|
||||
iris-vs-Masonry choice itself is still Iris's to make.
|
||||
rasterisation cost rather than an iris-specific slowness. **Still not
|
||||
proof, and now closed as unanswerable rather than merely untaken**: a
|
||||
same-mode software `force-gles` run to isolate the backend was retried
|
||||
2026-09-05 after fixing the compute-limit crash the first attempt hit,
|
||||
and hit a second, structural wall instead — SwiftShader's ES 3.0 GL
|
||||
path has no storage-buffer capacity at all, and `shader.wgsl` reads
|
||||
`var<storage>` buffers unconditionally, so reaching that path needs a
|
||||
shader rewrite, not a limits fix (RUST.md's I5 box, "The three
|
||||
remaining I5 verifications, closed 2026-09-05," item 2). The
|
||||
intermittent touch-scroll dropout this pass also reproduced is
|
||||
root-caused and fixed as of the same date (a missed `ACTION_DOWN` on a
|
||||
row's padding/header left `DragArbiter` stuck in `Idle`); three clean
|
||||
`iris-scroll.sh` runs post-fix each scrolled all 24/24 swipes, replacing
|
||||
the single-attempt 62-frame reading this table used to carry. RUST.md's
|
||||
I5 box, "Where iris's frame time goes, 2026-09-05, the `-gpu host`
|
||||
pass," and "The three remaining I5 verifications, closed 2026-09-05,"
|
||||
have the full account. The iris-vs-Masonry choice itself is still
|
||||
Iris's to make.
|
||||
+120
-58
@@ -42,14 +42,29 @@ session spending an afternoon on them again.
|
||||
replace the "refold + rebuild the whole ~3,200-row tree per event" path
|
||||
in all three clients. Worst/p99 frame time in the streaming phase
|
||||
dropped roughly 3x on this checkout's emulator (see the box for the
|
||||
exact numbers and their caveats). Blocked and not done this same pass:
|
||||
the three `iris-scroll.sh` runs, the host-GPU `FrameReport` retake, and
|
||||
the `EMU_GPU=software` cold boot the dropout-fix pass left open --
|
||||
`app/ui-sandbox.sh`'s `ai-server` currently fails to build
|
||||
(`event_model::Event::LimitReached` missing, unrelated to this pass's
|
||||
diff). Two new scripts, `iris/android-app/build-apk.sh` and
|
||||
`iris/android-app/run-bench.sh`, now do the build/install/tap/read-report
|
||||
cycle that used to be typed out by hand each time.
|
||||
exact numbers and their caveats). Two new scripts,
|
||||
`iris/android-app/build-apk.sh` and `iris/android-app/run-bench.sh`, now
|
||||
do the build/install/tap/read-report cycle that used to be typed out by
|
||||
hand each time.
|
||||
- **The three items the dropout-fix pass left open are all closed,
|
||||
2026-09-05** (the `ai-server` build break -- `event_model::
|
||||
Event::LimitReached` -- was already fixed on `rustify` by the time this
|
||||
pass started). Three clean `-gpu host` cold-boot `iris-scroll.sh` runs
|
||||
all scrolled 24/24 swipes (checked directly via clustered `render():`
|
||||
timestamps, not inferred from frame count), and the host-GPU table's
|
||||
iris row is now a best-of-three. `EMU_GPU=software` + `force-gles`
|
||||
still cannot produce a GLES number on this hardware, now for a third,
|
||||
structural reason found this pass: SwiftShader's ES 3.0 GL path reports
|
||||
zero storage-buffer capacity, and `shader.wgsl` reads four
|
||||
`var<storage>` buffers unconditionally -- reaching that path needs a
|
||||
shader rewrite, not a limits fix, so the SwiftShader-Vulkan-vs-GLES
|
||||
question is closed as unanswerable on this hardware rather than
|
||||
answered with a number. A fresh cold-boot `run-bench.sh` reading for
|
||||
P0's bench build (`frames=690 janky%=62.03 p50=19.6ms worst=62.5ms`) is
|
||||
in line with or better than the P0 box's three warm-AVD readings, so
|
||||
that box's "needs a clean cold boot" caveat is resolved too. Full
|
||||
account in I5's box, "The three remaining I5 verifications, closed
|
||||
2026-09-05."
|
||||
- **The intermittent touch-scroll dropout is root-caused and fixed,
|
||||
2026-09-05.** Not the previously-suspected coalesced first
|
||||
`ACTION_MOVE` (ruled out) -- a gesture's `ACTION_DOWN` can land on a
|
||||
@@ -63,16 +78,12 @@ session spending an afternoon on them again.
|
||||
tests (three in `iris/src/sense.rs`'s `drag_arbiter_tests`, one in
|
||||
`transcript-ui`'s `selection::tests`, the latter failing on the
|
||||
pre-fix code). See this box's own "Touch-scroll dropout root-caused,
|
||||
2026-09-05" subsection for the trace and what could and could not be
|
||||
re-verified this pass -- **this checkout's emulator turned out to be
|
||||
concurrently in use by another session's P0 benchmark work partway
|
||||
through verification** (its sandbox server was restarted, wiping this
|
||||
pass's test session, and its Compose `bench` app took window focus),
|
||||
so the "run iris-scroll.sh three times cleanly" and "re-take the
|
||||
host-GPU FrameReport row" pass conditions could not be completed
|
||||
end-to-end this pass. The fix itself is verified by direct, targeted
|
||||
logcat traces taken before that interference began, not by the
|
||||
aggregate script.
|
||||
2026-09-05" subsection for the trace. **The aggregate verification an
|
||||
earlier pass could not complete (peer-emulator interference) is now
|
||||
done, 2026-09-05**: three separate cold-`-gpu host`-boot
|
||||
`iris-scroll.sh` runs each scrolled all 24/24 swipes, confirmed by
|
||||
clustered `render():` timestamps rather than frame count alone -- see
|
||||
this same subsection's "Update, 2026-09-05" paragraph.
|
||||
- **iris no longer requests compute-shader limits it never uses,
|
||||
2026-09-05.** `adapter.request_device`'s `Limits::default()` asks for
|
||||
desktop-tier compute limits unconditionally even though nothing in
|
||||
@@ -82,12 +93,15 @@ session spending an afternoon on them again.
|
||||
ES 3.0, no compute at all). New shared `iris_core::device_limits()`
|
||||
zeros exactly the six compute fields; `rigs/gpu-probe`'s own mirrored
|
||||
limits were updated and confirm `IRIS DEVICE: ok` on this VM's own
|
||||
Vulkan and GL adapters. **The specific SwiftShader-ES-3.0 crash this
|
||||
fixes was not re-verified on-device this pass** -- the cold boot needed
|
||||
would have force-restarted this checkout's emulator while another
|
||||
session had its own app focused on it, so it was left rather than
|
||||
disrupted. See this box's "Fixed, 2026-09-05, later the same day"
|
||||
subsection (under the software-mode crash it fixes) and `DECISIONS.md`.
|
||||
Vulkan and GL adapters. **Verified on-device 2026-09-05**: a cold
|
||||
`EMU_GPU=software` boot no longer aborts on the compute-limit request
|
||||
this fix targeted -- adapter selection now succeeds -- but device
|
||||
creation still aborts, on a different, unfixed limit
|
||||
(`max_storage_buffer_binding_size`, SwiftShader ES 3.0 has no SSBOs
|
||||
either); see this box's "The three remaining I5 verifications, closed
|
||||
2026-09-05" subsection, item 2. See this box's "Fixed, 2026-09-05,
|
||||
later the same day" subsection (under the software-mode crash it fixes)
|
||||
and `DECISIONS.md`.
|
||||
- **Decided 2026-09-05: iris over Masonry**, by Iris, from the host-GPU
|
||||
numbers in I5's box and E1/E2's findings. See the Recommendation's item
|
||||
3 and `DECISIONS.md`. Next: the remaining screens and the app on iris —
|
||||
@@ -3152,7 +3166,7 @@ silently on real hardware.
|
||||
| app | build | GPU mode | frames | janky % | p50 | p90 | p99 | worst | cpu p50 | gpu-wait p50 |
|
||||
|---|---|---|---|---|---|---|---|---|---|---|
|
||||
| Compose (in-app report) | debug | host (virgl) | 1268 | 96.4% late | 20.0ms | 28.4ms | 37.7ms | -- | -- | -- |
|
||||
| iris (`FrameReport`) | **release**, `force-gles` | host (virgl) | 62 | 41.94% | 15.0ms | 21.8ms | 37.1ms | 37.1ms | 0.2ms | 12.9ms |
|
||||
| iris (`FrameReport`), **best of three, 2026-09-05** | release, `force-gles` | host (virgl) | 439 | 46.24% | 15.7ms | 23.3ms | 31.2ms | 57.4ms | 1.2ms | 13.2ms |
|
||||
|
||||
**Under real GPU rendering, iris's median frame is faster than
|
||||
Compose's, not 2-3x slower** -- the opposite shape from the
|
||||
@@ -3196,7 +3210,27 @@ silently on real hardware.
|
||||
Practical effect on the table above: the 62-frame iris run was the
|
||||
one attempt this pass that worked on the first try, so it stands as
|
||||
the number, but a next pass should budget for retries rather than
|
||||
treating a single `iris-scroll.sh` invocation as reliable.
|
||||
treating a single `iris-scroll.sh` invocation as reliable. **Update,
|
||||
2026-09-05, the dropout fix (`e692429`, this box's own "Touch-scroll
|
||||
dropout root-caused" subsection) verified against this table**: with
|
||||
the fix in, three separate `iris-scroll.sh` invocations against a
|
||||
fresh cold `-gpu host` boot each scrolled all 24/24 swipes, no
|
||||
retries needed --
|
||||
frames=450 janky%=48.44 p50=16.3ms p90=22.7ms p99=26.2ms worst=29.7ms cpu_p50=1.0ms gpu_wait_p50=13.1ms
|
||||
frames=462 janky%=50.22 p50=16.9ms p90=25.0ms p99=45.4ms worst=59.0ms cpu_p50=1.3ms gpu_wait_p50=13.9ms
|
||||
frames=439 janky%=46.24 p50=15.7ms p90=23.3ms p99=31.2ms worst=57.4ms cpu_p50=1.2ms gpu_wait_p50=13.2ms
|
||||
Per-swipe coverage was checked directly, not inferred from the frame
|
||||
count alone: a continuous `adb logcat -v time -s iris-android-app:D`
|
||||
capture started before each of the second and third runs (the first
|
||||
run's capture was taken with `logcat -d` after the fact and lost
|
||||
earlier lines to the ring buffer, so it is corroborating rather than
|
||||
direct) shows `render():` timestamps clustering into exactly 24
|
||||
groups per run, one per swipe, each with 22-42 render calls and no
|
||||
gap over 0.3s inside a cluster -- i.e. every one of the 24 swipes
|
||||
produced real redraw activity in all three runs, closing this table's
|
||||
open verification. The third run (lowest janky% and p50) is the
|
||||
"best of three" row now in the table above, replacing the earlier
|
||||
single-attempt 62-frame reading.
|
||||
|
||||
**Redundant-work check (no optimising, as instructed), host GPU,
|
||||
`force-gles`.** (1) **Idle redraw: zero**, confirmed fresh this pass
|
||||
@@ -3799,38 +3833,66 @@ device.
|
||||
number**: two fresh cold boots of the same AVD, one per build,
|
||||
`run-bench.sh` on each, nothing else running.
|
||||
|
||||
**The three remaining I5 verifications the dropout-root-cause pass
|
||||
left open (this box's "Touch-scroll dropout root-caused" subsection)
|
||||
-- attempted this pass, blocked, not silently dropped.**
|
||||
`app/iris-scroll.sh` needs `dev.iris.android.demo`'s plain
|
||||
`transcript-screen` (non-`bench`) debug build, which needs a live
|
||||
sandbox server (`app/ui-sandbox.sh`) to bake in at build time
|
||||
(`build.rs`'s `AI_APP_TRANSCRIPT_HOST`/etc, skipped only under
|
||||
`bench`). `./ui-sandbox.sh start` fails to build **the Rust
|
||||
`ai-server` itself**, unrelated to anything in this box's diff:
|
||||
`error[E0599]: no variant named 'LimitReached' found for enum
|
||||
'event_model::Event'` in `server/src/session/mod.rs:2633`,
|
||||
`server/src/session/claude/translate.rs:307`, and
|
||||
`server/src/session/echo.rs:389` -- `event_model` and `server` have
|
||||
drifted out of sync on this branch, most likely from concurrent
|
||||
work elsewhere on `rustify` (this pass touched nothing under
|
||||
`server/` or `event-model/`, confirmed by `git status`). Since
|
||||
fixing that is a separate, unrelated repair (and risks colliding
|
||||
with whatever pass is mid-edit there), this pass did not attempt
|
||||
it. Consequently, not done this pass: the three `iris-scroll.sh`
|
||||
runs, the host-GPU `FrameReport` table row retake, and the
|
||||
`EMU_GPU=software` + `force-gles` cold-boot `FrameReport` (the
|
||||
backend-isolation question) -- none of the three need the broken
|
||||
server directly, but the first two need the same debug build the
|
||||
server outage blocks, and by the time that was found there was not
|
||||
enough of this pass left to justify a fresh `EMU_GPU=software` cold
|
||||
boot (several minutes) for the third alone without also covering the
|
||||
other two on the same session. A future pass: fix or wait out the
|
||||
`server`/`event_model` drift, rebuild `dev.iris.android.demo` with
|
||||
plain `transcript-screen` via `./build-apk.sh debug --abi x86_64
|
||||
--features "transcript-screen force-gles"`, then run
|
||||
`app/iris-scroll.sh` three times and retake the host-GPU row, and
|
||||
separately cold-boot with `EMU_GPU=software` for the third.
|
||||
**The three remaining I5 verifications, closed 2026-09-05.** The
|
||||
`server`/`event_model` drift the previous pass hit (`no variant
|
||||
named 'LimitReached'`) was already fixed upstream on `rustify` by
|
||||
the time this pass started (commit `c07d544`, "carry main's
|
||||
`LimitReached` event") -- `cargo build --release` under `server/`
|
||||
is clean, `./ui-sandbox.sh start` builds and runs. All three items
|
||||
this left open:
|
||||
|
||||
1. **Three clean `iris-scroll.sh` runs against a cold `-gpu host`
|
||||
boot, all 24/24 swipes scrolling in every run** -- results and
|
||||
the per-swipe verification method are recorded just above, in
|
||||
this same "Touch-scroll dropout root-caused" subsection's
|
||||
"Update, 2026-09-05" paragraph, and the host-GPU table above now
|
||||
carries the best-of-three row.
|
||||
2. **`EMU_GPU=software` + `force-gles`, cold boot -- still cannot
|
||||
isolate SwiftShader-Vulkan from SwiftShader-GL, now for a third
|
||||
and structural reason.** Same release build
|
||||
(`transcript-screen force-gles`), fresh cold boot under
|
||||
`EMU_GPU=software`. The earlier compute-limit abort this same
|
||||
box's "Fixed, 2026-09-05, later the same day" paragraph resolved
|
||||
(`iris_core::device_limits()` zeroing `max_compute_*`) no longer
|
||||
fires -- adapter selection now succeeds and picks up
|
||||
SwiftShader's ES 3.0 GL path -- but device creation aborts on a
|
||||
*different* limit immediately after:
|
||||
Abort message: 'Could not get device!: RequestDeviceError { inner: Core(LimitsExceeded(
|
||||
FailedLimit { name: "max_storage_buffer_binding_size", requested: 134217728, allowed: 0 } )) }'
|
||||
i.e. SwiftShader's ES 3.0 reports zero storage-buffer capacity at
|
||||
all -- SSBOs are an ES 3.1+ feature, the same generation gap the
|
||||
compute-limit failure came from, and exactly the trap this box's
|
||||
own "Fixed" paragraph flagged when it rejected
|
||||
`Limits::downlevel_webgl2_defaults()` for zeroing
|
||||
`max_storage_buffers_per_shader_stage` while `shader.wgsl`'s
|
||||
vertex stage reads four `var<storage>` buffers unconditionally.
|
||||
**Closing the open question, one sentence**: whether
|
||||
SwiftShader-Vulkan or GLES-in-general explains the ~80-150ms
|
||||
software-mode numbers cannot be answered on this hardware at
|
||||
all, because `shader.wgsl`'s storage-buffer reads make a GLES
|
||||
path on downlevel (ES 3.0) SwiftShader structurally unreachable
|
||||
rather than merely unmeasured -- reaching it is a shader rewrite
|
||||
(moving those reads off `var<storage>`), which is real scope, not
|
||||
a measurement task, and was not attempted here.
|
||||
3. **P0's bench build, cold-boot `run-bench.sh` number**: same
|
||||
`-gpu host` cold boot as item 1 (re-cold-booted after the
|
||||
`EMU_GPU=software` boot above), `./build-apk.sh release --abi
|
||||
x86_64 --features "transcript-screen force-gles bench"`,
|
||||
`./run-bench.sh`:
|
||||
frames=690 janky%=62.03 p50=19.6ms p90=42.4ms p99=56.5ms worst=62.5ms cpu_p50=4.4ms gpu_wait_p50=12.6ms
|
||||
scroll: 6 cycles (24 swipes), streamed 400/400 fixture events
|
||||
process CPU time over this run: 15394ms
|
||||
peak RSS: 164348kB
|
||||
battery current: mean 900000µA over 21 samples (min 900000, max 900000, the emulator's fixed mocked-charger reading, not a real battery -- see P0's own box)
|
||||
Against the P0 box's own three same-warm-AVD readings
|
||||
(`frames=690/691/691`, `p50` 18.9-28.3ms, `worst` 101-130ms),
|
||||
this cold-boot run's `worst` (62.5ms) and `p99` (56.5ms) are
|
||||
*lower* than any of the three warm-AVD runs, and its `p50`
|
||||
(19.6ms) sits inside their range -- so the P0 box's caveat that
|
||||
the warm-AVD numbers might be inflated by AVD staleness does not
|
||||
hold up under a fresh cold boot; if anything this run is cleaner.
|
||||
`cpu_p50` (4.4ms) is within the 3.5-7.4ms noise band the P0 box
|
||||
already flagged as run-to-run noisy on a shared warm AVD.
|
||||
|
||||
**Redelivered, 2026-09-05.** `./build-apk.sh release --abi
|
||||
arm64-v8a` (arm64-only jniLibs; an earlier step in this same pass
|
||||
|
||||
Reference in new issue
Block a user