From 048738bb9dc907c3b37a575ee7b7210dd4bccb17 Mon Sep 17 00:00:00 2001 From: iris-ai <4+iris-ai@noreply.localhost> Date: Sun, 20 Sep 2026 17:32:55 -0400 Subject: [PATCH] Record PR 19 performance sweep and request invalidation fix --- docs/HANDOFF.md | 28 +++++++++++++- docs/LAYOUT.md | 22 +++++++++-- docs/LAYOUT_LOG.md | 92 ++++++++++++++++++++++++++++++++++++++++++++++ 3 files changed, 137 insertions(+), 5 deletions(-) diff --git a/docs/HANDOFF.md b/docs/HANDOFF.md index ff86754..40ee3c4 100644 --- a/docs/HANDOFF.md +++ b/docs/HANDOFF.md @@ -4,9 +4,35 @@ Where the work in flight stands. The settled layout design, the vocabulary and the measurement method are in `docs/LAYOUT.md`; what the review of #19 found is in `docs/LAYOUT_LOG.md`. +## Performance sweep is submitted on PR #19 + +`layout/one-ask` is at `0e838e9` in `/home/bob/repos/iris-deferred`. The fix +keeps dependencies only for size requests the allocator uses, and borrows +`Widgets::size_rules` rather than cloning both axes at every lookup. Used +hints retain their reader dependency; adding a cap after layout is checked for +both hinted and measured shares. Ordinary text measurement still tracks its +own dependencies. The app's Iris pin is untouched. + +Repeated release instruction counts reproduce the reported +5.9% resize cost +and find a larger defect: an unchanged paragraph repaint invalidated its span, +turning 0.376B instructions into 3.010B in a 40-paragraph, 2,000-frame fixture. +The fix is 0.370B. Text resize is now +0.99% against `de1eb7e`, the width sweep ++1.23%, and edits/scroll are within 0.3%. `docs/LAYOUT_LOG.md` has the full +performance sweep, including deep-tree regressions, RSS, upstream-base +comparison, and GPU timestamp measurements. This is not a universal speedup. + +Passed: format, workspace Clippy with and without diagnostics, all 197 ordinary +and 201 diagnostic workspace tests, zero steady allocation fixtures, 400 +depth-5 / 1,000 depth-6 / 2,000 depth-4 trees, 400 depth-5 expression trees, +400 relative-bound trees, GPU recording and move-chain rigs, and identical +cold geometry for all 34,986 bounded boxes against `8780b40`. Code and comments +were reviewed separately; the final test/comment refinements passed their +focused checks. Artifacts are under `/tmp/pr19-perf/`. + ## Deferred comparisons are submitted on PR #19 -The `layout/one-ask` tip is `8780b40`, in `/home/bob/repos/iris-deferred`. +The deferred-comparison implementation landed at `8780b40`, before the +performance correction above. `/home/bob/repos/iris` remains on `layout/bounds`; that branch's `de1eb7e` was fast-forwarded into #19 before this work. The app's Iris pin is untouched. diff --git a/docs/LAYOUT.md b/docs/LAYOUT.md index 64abc97..4ff66c6 100644 --- a/docs/LAYOUT.md +++ b/docs/LAYOUT.md @@ -66,10 +66,24 @@ The allocation rig forces every widget to redraw across 100 resize frames: both the plain and clamped eight-row fixtures allocate zero times after warm-up. This is a measured property of those stable visible trees, not a guarantee for arbitrary widgets or visibility changes. An eight-level known nested span paints -its leaf once. The 40-row, 1,000-frame text resize rig costs 9,891,085,554 retired -instructions versus 9,341,015,256 at `de1eb7e` (5.89% more), with the same printed -geometry. Request dependency tracking has a cost even when text still needs -measurement. The 400-tree unbounded cold dump preserves all 34,986 boxes. +its leaf once. Discovery retains dependencies only for requests it actually +uses; discarded discovery leaves dependency tracking to the measured draw. +Used hints still register request readers so a newly added cap reallocates the +parent even when the child's reported share is unchanged. + +The performance sweep of 2026-09-20 found and removed an eightfold increase in +the unchanged-paragraph repaint fixture caused by discarded discovery +invalidating its span. The 40-row, 1,000-frame text resize rig now costs 9.432B +retired instructions against 9.340B at `de1eb7e` (+0.99%), down from 9.890B +at `8780b40` (+5.90%). Text draw, line-break and glyph-placement counts agree; +the remaining cost is layout bookkeeping. The width sweep is +1.23%, and +paragraph edits and scrolling are within 0.3% in the measured fixtures. +These are medians of nine process totals, not GPU or phone frame times. +Deep generated trees still have workload-dependent wins and regressions; +`docs/LAYOUT_LOG.md` records the matrix, storage costs and measurement limits. +The 400-tree unbounded cold dump preserves all 34,986 boxes against `de1eb7e`; +the performance correction also preserves all 34,986 bounded boxes against +`8780b40`. The app's submodule pin is unchanged. This design lives on Iris PR #19. diff --git a/docs/LAYOUT_LOG.md b/docs/LAYOUT_LOG.md index 0bdc6e2..a9dbd3f 100644 --- a/docs/LAYOUT_LOG.md +++ b/docs/LAYOUT_LOG.md @@ -6,6 +6,98 @@ nothing here is rediscovered. Each entry says who found it and when. it (settled design, the measurement method) belongs in `docs/LAYOUT.md`, and the current plan is in `docs/HANDOFF.md`. +## Performance sweep of #19 (2026-09-20) + +`0e838e9`, by Codex. + +The 5.89% text-resize increase at `8780b40` was real, but was not extra +shaping: `revision_cost` at both `de1eb7e` and `8780b40` drew 168 widgets, +rendered seven texts and broke seven layouts per resize frame. Sampling the +latter put most CPU work in relocating retained primitives. This is a CPU +layout/recording fixture, not a GPU text-rendering benchmark. + +**Discarded discovery invalidated things it did not decide.** Both +`Painter::size_request` and `measured_request` kept dependencies and pinned +the rel base even when they discarded the discovered answer and used the +ordinary measured size. Repainting one unchanged paragraph consequently +redrew its enclosing span and remapped the other 39 paragraphs. Discovery now +keeps those dependencies only when its answer is used; the measured path +already records the size it reads. Used hints still register request readers: +adding a cap to a previously unbounded hinted share must reach its allocator. +The same edit is checked for an unhinted share too. `Widgets::size_rules` now +borrows its rules instead of cloning both axes at every lookup; callers that +need an owned snapshot clone explicitly. + +All instruction measurements below are medians of nine executable runs with +`perf stat -e instructions:u`, release builds without diagnostics. Every set +kept all nine readings within 2% of its median. Totals include initialization +and the cold frame. Resize uses 40 rows and 1,000 frames; the other phases use +40 paragraphs in a scroll and 2,000 frames. Edit alternates a short suffix, +rather than growing the workload every frame. `SWEEP=1` cycles 256 widths; +the default toggles two widths. Billion retired instructions: + +| Workload | Before deferred requests (`de1eb7e`) | Incoming (`8780b40`) | After this fix | After vs before deferred | +| --- | ---: | ---: | ---: | ---: | +| Text resize, two widths | 9.340 | 9.890 | 9.432 | +0.99% | +| Text resize, width sweep | 7.044 | 7.478 | 7.131 | +1.23% | +| Unchanged paragraph repaint | 0.376 | 3.010 | 0.370 | -1.58% | +| Paragraph edit | 3.485 | 3.607 | 3.492 | +0.22% | +| Scroll | 0.428 | 0.430 | 0.428 | +0.10% | +| Idle | 0.306 | 0.306 | 0.306 | +0.04% | + +The repaint regression was eightfold in these process totals; diagnostics +now show one text draw and no span draw. The remaining roughly 1% resize +cost is discovery around content that still needs measurement. It is not +necessary to weaken wrapping, overflow or bounds correctness to recover the +regression. Differences below about half a percent should not be ranked: +rebuilding changes instruction counts too. + +The full PR is substantially cheaper than its upstream base `ca2b4b2` in +these fixtures: that base takes 124.072B instructions for the two-width resize, +13.516B for repaint, 14.155B for edit and 13.517B for scroll. Its geometry also +differs (paragraphs run past the 900-pixel output), so this is a whole-version +comparison, not equal-work evidence for a single optimization. The fixture +was identical apart from spelling the fixed width `Len::abs(40.0)` there and +omitting unavailable diagnostic calls. + +**The generated trees are not all faster.** At depth 8, 1,000 frames, seeds +1 and 13, the incoming branch's bounded resize workloads take 0.561B and +4.297B instructions; the fix takes 0.539B and 4.016B. The pre-deferred branch +uses 0.359B and 1.898B, but does different work: seed 13 draws 67 widgets per +resize there and 170 with deferred requests, and even the active cold tree +changes. Those are not equal-work comparisons. Removing intrinsic bounds +with `IRIS_UNBOUNDED=1` isolates the ordinary path: seed 1 resize is 0.359B +before vs 0.363B after, and seed 13 is 1.897B vs 1.296B. Unbounded seed 13's +multi-widget repaint still rises from 0.980B to 1.208B. A consumed request's +conservative invalidation remains a cost; eliminating it would require +checking whether the request changed before invalidating its allocator. +This sweep does not establish a universal speedup or make that cost inevitable. + +**Storage and limits.** The forced-redraw plain and clamped allocation rigs +still allocate zero times after warm-up. A single 2,000-row stress run +(10,001 widgets, all retained) settles at 133,400 KiB RSS before deferred +requests, 141,912 KiB incoming and 138,820 KiB after this fix. These process +RSS snapshots show the retained bookkeeping cost, not a per-widget allocation +size or a leak. The spare glyph store still copies text when keeping a prior +width; its 128-entry bound does not bound bytes. The allocator visits expression +nodes at each crossing; many distinct caps can require quadratic work. Neither +case should be mistaken for the small fixed-width text-update workload above. + +**GPU work is a separate tradeoff.** The first attempt detected RADV on the +RX 7900 XT but could not allocate its 64 MiB device buffer (`No space left on +device`). After allocation became available, both rigs ran successfully. +Across three timestamp runs of `chain_cost` (200,000 two-pixel rectangles), +depth 1 took 48–78 µs, depth 4 took 76–79 µs, depth 16 took 145–150 µs, and +depth 64 took 430–443 µs. The first run's depth-2 result was anomalous; the two +repeats agreed on approximately 78 µs for depths 1–4. Chain depth counts opted-in +movable regions, not every widget. Deep chains have a real GPU cost, but 200,000 +small quads through 64 movable ancestors is a stress case, not a measurement +of a normal paragraph or a phone. The CPU recording-only `draw_cost` rig also +passed; its wall times under concurrent load are not used as regression evidence. +No renderer or shader behavior changed in this performance fix. + +Verification is recorded in `HANDOFF.md`. The app's Iris pin is unchanged. + ## One ask, the root's included (2026-09-20) `0d03267`. `root_layout` read the root's declared lengths against the window