Withdraw the 3x, and record what a move actually costs

The performance section reported a cost the fixture had invented. `Branch`
picks which subtree to draw from a measured length, so seed 1 draws 88
widgets at `5ed9e87` and 115 at `bd6de71` -- the ms table was comparing
three and a half times the work. Measured on one tree, fixed point costs
about a fifth to a quarter.

Also corrects where the cost is. Cycles say the `i64` division is 0.00% of
`apply_scalar` and the multiply 1.5%; it is `saturating_add`, five
instructions with no `i32` vector form. Cycle counts are steady here to
0.23% -- it is wall time that varies 2x, not the counters -- and IPC is
unchanged across fixed point, so this is instruction count and not stalls.

Records the four wins taken (-7.4% cycles, -12% instructions), the four
things tried that made it worse, and the one lever left that needs a
decision rather than a measurement.

Also: the `Len`/`LayoutLen` split and `has_exact_size`; that `Scroll`
reports `LEFTOVER` rather than its content's size, which this said wrongly
in one place and rightly in another; and that the clipping `debug_assert`
holds mask-setters only, which is why a `LazySpan` does not trip it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
iris-aiandClaude Opus 5 committed 2026-09-16 14:08:00 -04:00
1 parent 388a6a060d
commit 9ed90505b3
1 file changed
+94 -67
+94 -67
View File
@@ -8,7 +8,7 @@ log.
Canonical Iris `main` is **`ca2b4b2`** (#17, the headless rig). **#18 Canonical Iris `main` is **`ca2b4b2`** (#17, the headless rig). **#18
`split/18-position-chain`** is open in `/home/bob/repos/iris-pr18`; its local `split/18-position-chain`** is open in `/home/bob/repos/iris-pr18`; its local
head is **`4f5e27c`**, sixty-three commits. Built-in alignment is complete there; head is **`394d514`**, seventy commits. Built-in alignment is complete there;
see "Built-in alignment" below for the retained-layout details. No PR reviews see "Built-in alignment" below for the retained-layout details. No PR reviews
were present when checked on 2026-09-15. were present when checked on 2026-09-15.
@@ -37,6 +37,17 @@ length is reported and any relative or `leftover` child makes it report
`Holds` ranges propagate through final-box drawing, so a resize repositions `Holds` ranges propagate through final-box drawing, so a resize repositions
them without redrawing when their own contracts permit it. them without redrawing when their own contracts permit it.
`a8898aa` gives a length with no share in it its own type again, which is the
separation canonical `main` had and this branch had collapsed: `Len` is pixels
plus a fraction of a box -- a position being the length from the box's start,
which is why a span is two of them -- and `LayoutLen` is a `Len` plus a claim
only a container dividing its room can answer. What was `UiScalar` is `Len`,
what was `Len` is `LayoutLen`, and `From<Len> for LayoutLen` is the one-way
step. `490918b` renames `Painter::ruled` to `has_exact_size`, over
`SizeRule::exact`: "is there a rule here" and "is my report moot" are the same
question only while `Exact` is the only rule, and `Min`/`Max`/`Clamp` are
queued.
`71c9c39` replaces the public `Painter::place` distinction with an opt-in `71c9c39` replaces the public `Painter::place` distinction with an opt-in
widget property. `.region_node()` gives a widget one independently movable widget property. `.region_node()` gives a widget one independently movable
retained region; `Widgets::set_region_node` can change that choice at runtime retained region; `Widgets::set_region_node` can change that choice at runtime
@@ -78,7 +89,7 @@ rewrite:
`core/src/ui/holds.rs`, the retained tests, and the generated cold-layout oracle `core/src/ui/holds.rs`, the retained tests, and the generated cold-layout oracle
pin those rules. Seeds 10 and 86 are now in the ordinary generated set. pin those rules. Seeds 10 and 86 are now in the ordinary generated set.
## Verification at `4f5e27c` ## Verification at `394d514`
- `cargo fmt --all --check` - `cargo fmt --all --check`
- `cargo build --workspace --all-features` - `cargo build --workspace --all-features`
@@ -92,7 +103,8 @@ pin those rules. Seeds 10 and 86 are now in the ordinary generated set.
(159,024 widgets per case, largest tree 587) at `5ed9e87`, before fixed (159,024 widgets per case, largest tree 587) at `5ed9e87`, before fixed
point; re-run that before quoting it again. point; re-run that before quoting it again.
- `tabs`, `view`, `minimal`, `text` and `random` render byte-identical at - `tabs`, `view`, `minimal`, `text` and `random` render byte-identical at
1920x1200 across the whole fixed-point sequence. 1920x1200 across the whole fixed-point sequence, and across this session's
renames and arithmetic changes on top of it.
- `tests/cases/drift.rs` passed its 20,000-move exactness check in release - `tests/cases/drift.rs` passed its 20,000-move exactness check in release
mode. mode.
- The seeded `random` example, live-resized from 1920x1200 to 1280x800, is - The seeded `random` example, live-resized from 1920x1200 to 1280x800, is
@@ -127,69 +139,78 @@ choice and the incompatible visual effect of making `Full` the default.
### Performance ### Performance
Measured 2026-09-16 on the release `layout_diagnostics` fixture, seed 1, **The 3x this section used to report was a fixture artifact, withdrawn
depth 8, 500 frames, as median milliseconds a frame. `5ed9e87` is the commit 2026-09-16.** `random.rs`'s `Branch` picks which of two subtrees to draw by
before fixed point; `39e4ca2` is fixed point complete. comparing a measured pixel length with a threshold, so the fixture's shape
moves with the thing being measured. Seed 1 at depth 8 draws 88 widgets and
writes 2,298 primitives a frame at `5ed9e87`, and 115 and 8,209 at `bd6de71`
-- three and a half times the work behind a number read as three and a half
times the cost. `394d514` gives the rig `Edits::fixed_branches`; the oracle
keeps measured branches, which is the whole point of them.
| phase | `5ed9e87` | `39e4ca2` | `4f5e27c` | **Check the work counters before comparing two commits' times.** The rig
| --- | ---: | ---: | ---: | prints drawn widgets, widget draws and primitive writes. An undrawn
| many | 0.179 ms | 0.544 | 0.274 | `leftover` child still moves them and no flag removes that, so the trees are
| resize | 0.020 | 0.035 | 0.033 | near rather than equal -- and `9d8415d` changed the generator itself, since
| scroll | 0.011 | 0.030 | 0.019 | an X span's `OrthoSize::Full` became a size rule.
| repaint | 0.012 | 0.031 | 0.021 |
**Fixed point cost 3x, and two thirds of that was not the grid.** The Measured on the fixed-shape fixture, seed 1, depth 8, 500 frames of `many`:
counters said eight more "placed by redrawing" a frame in `scroll`, all of
them "reuse: another layer": a retained drawing belongs to the layer it was
made on, and `Stack` measured the child that sizes it on its own layer before
drawing it again on the child layer. `97cc8b3` measures on the layer the
child ends up on, which puts the scroll phase's counters back exactly where
they were -- 4 widget draws, 12 draw requests.
What is left is per-operation cost, not more work: the `many` phase does 125 | | drawn widgets | primitive writes | instructions | cycles |
widget draws against 121 before, and takes 1.5x as long. Three roundings and | --- | ---: | ---: | ---: | ---: |
divisions were taken out after measuring (`11c55bc`, `4f5e27c`, and halving | `5ed9e87`, before fixed point | 100 | 4,272 | 1,761M | ~699M |
the divisions in `Holds::through`); `RegionRemap::apply_span` is still a | `4cbb242` | 97 | 3,951 | 2,093M | ~819M |
fifth of the phase, and the rest is spread thin. Measure with `perf stat -e
instructions:u` rather than the clock, which varies 2x here.
**Why the grid is slower than floats, from the disassembly at `4f5e27c`.** So the grid costs something like a fifth to a quarter, on a head that draws
`RegionRemap::apply_span` -- two scalars, every branch -- is 352 slightly less. Normalised by primitive writes it is about 1.27x.
instructions, of which two are 64-bit `idiv`, six `imul`, and about sixty
are jumps and `cmov`s over twenty-eight compares. The same arithmetic in
`f32` is a couple of dozen SSE instructions. The difference is in three
places, and none of them is the grid being a worse idea:
- **A float's rounding is free and a fixed-point one is code.** `Fixed::mul` **Cycles are measurable on this machine and instructions alone are not
widens to `i64`, multiplies, adds a half step, shifts back, and saturates enough.** Three runs of one binary varied 0.23%. It is *wall time* that
into `i32` -- roughly eight instructions and a sign branch where `mulss` varies 2x here, which is what the machine notes say; the counters are
is one instruction with the rounding in hardware. steady. IPC is 2.58 before fixed point and 2.59 after, so the cost is
- **Integer division is the slowest instruction on the core**, is not instruction count at unchanged throughput -- not stalls, not branch misses
pipelined, and has no vector form; `divss` is pipelined and vectorises. (0.69%), not the division.
`div_round` also needs the remainder and its sign. This is why removing
divisions is where the wins have been, and why the remaining ones are
worth structural effort rather than micro-optimisation.
- **Saturation costs what floats get from infinities**, and it breaks the
pairing. Adding four `i32` lanes with saturation is a dozen SSE ops
(`pcmpgtd`/`paddd`/`pandn`/`psrad`/`pxor`/`por`) against one `addps`; and
the widening multiply cannot stay in a vector register at all, since there
is no packed 64-bit `imul` here, so lanes that an `f32` `Vec2` did at once
serialise.
"Fixed point is faster" is a rule from machines without an FPU. On this one **What the cycle profile says the cost is**, against what the previous
the grid buys exactness, and exactness is what the warm-against-cold oracle session recorded from the disassembly. In `AxisRemap::apply_scalar` the
demands -- so the cost is the price of the property, not a defect to chase `i64` division is **0.00% of cycles** and the multiply 1.5%: the time is in
back to 1.0x. `saturating_add`, which is five instructions and has no `i32` vector form.
The division is rare because most moves are translations.
One thing tried and reverted, recorded so it is not tried again: short- Recovered this session, each measured before and after on one commit with
circuiting `apply_scalar` where the fraction is nought or one. Those are not byte-identical renders proving the layout unchanged -- on the old fixture,
the common cases, and the comparisons cost 17% more than the divisions they 689.6M cycles and 1,886M instructions down to 638.9M and 1,657M:
saved.
- `cb1bba4` works a move out once per subtree rather than per scalar: -3.8%
instructions, -0.8% cycles. The gap between those two is the finding.
- `1940e85` translates a whole region at once, since that is what a
translation is: -4% instructions, -3.4% cycles.
- `d75a1e2` and `4cbb242` skip multiplies that cannot change anything --
composing through a full box, and `lerp` where the ends are equal:
together -2.7% instructions, -1.8% cycles.
Tried and rejected, with numbers, so they are not tried again:
- A float reciprocal for the division: **+6% cycles**. The conversions cost
more than the `idiv` that is not on the hot path anyway.
- Branchless `shift_round`: **+6.7% cycles**. The sign of a product predicts
well, and the positive path is two instructions.
- Removing the per-child hash lookup in `remap_subtree` with `mem::take`:
**0.0%**. `FxHashMap` on a hot line is not the cost.
- Short-circuiting `apply_scalar` where the fraction is nought or one: +17%.
(From the previous session.)
**The one lever left that is worth a decision:** `Fixed`'s `saturating_add`
and `saturating_sub`, measured at **8.1% of instructions and 2.7% of
cycles**. Wrapping instead is one instruction where saturating is five. That
trades away "a clamped coordinate keeps the ordering a wrapped one inverts",
which is a posture rather than a micro-optimisation, so it needs Bryan.
Everything else is spread thin.
Earlier, against #18's own history: the retained rewrite took `many` from Earlier, against #18's own history: the retained rewrite took `many` from
25.17M instructions a frame at `691e3eb` to 6.14M at `29c7881`, and `resize` 25.17M instructions a frame at `691e3eb` to 6.14M at `29c7881`, and `resize`
from 16.08M to 8.39M. That fixture has since changed; do not compare across from 16.08M to 8.39M. That fixture has since changed twice; do not compare
it. across it.
## Retained-layout invariants ## Retained-layout invariants
@@ -200,9 +221,16 @@ it.
-- `Painter::child_layer_at` addresses one -- or it pays two draws a frame -- `Painter::child_layer_at` addresses one -- or it pays two draws a frame
forever and keeps whichever the second ask left. forever and keeps whichever the second ask left.
- A widget that clips its contents to its box reports its box: `Scroll` and - A widget that clips its contents to its box reports its box: `Scroll` and
`Masked` both report `LEFTOVER`, and a `debug_assert` in `draw_at` holds `Masked` both report `LEFTOVER`, and a `debug_assert` holds any widget that
any widget that set a mask to it. Overflowing is otherwise ordinary and a **set a mask this draw** to it. Overflowing is otherwise ordinary and a text
text too tall for its box says so. too tall for its box says so -- which is why the assertion is narrowed to
mask-setters rather than to every widget.
`Masked` is the only one that sets a mask; `Scroll` reports its box because
a drawing placed in a longer box is *scaled* by recomposition, not because
it masks. So a `LazySpan` is unaffected: it clips by not drawing, and may
report its whole content length, which is what the `Scroll` above it
measures. One that set its own mask would have to report its box and hand
the content length up another way.
- A span is as long across itself as its longest child, unless a rule beside - A span is as long across itself as its longest child, unless a rule beside
it says how long it is -- and then it does not read its children there at it says how long it is -- and then it does not read its children there at
all, since the answer is not wanted and reading one is what makes its size all, since the answer is not wanted and reading one is what makes its size
@@ -261,9 +289,10 @@ it.
where there is none. The box a parent hands back and the sum of what the where there is none. The box a parent hands back and the sum of what the
children asked for are counts of the same step, so the boundary needs no children asked for are counts of the same step, so the boundary needs no
margin and the validity interval is split exactly at it. margin and the validity interval is split exactly at it.
- `Scroll` reports its content's first measured size. Its drawing can survive - `Scroll` reports `Size::LEFTOVER`; what it takes from the first box it asked
container-length changes only over the interval in which clamping and its about is its *content's* length, which it clamps and holds by rather than
current offset do not change. reports. Its drawing can survive container-length changes only over the
interval in which clamping and its current offset do not change.
- **A move that keeps a box's length is a translation, and an offset is - **A move that keeps a box's length is a translation, and an offset is
exact on the grid.** A box that also changed length has to re-express each exact on the grid.** A box that also changed length has to re-express each
part as a fraction of the new one, and that division and multiplication part as a fraction of the new one, and that division and multiplication
@@ -562,11 +591,9 @@ Queued from this work, in order:
one offer per child. `min(report, cap)` is also not a `Len`: `Len` is a sum one offer per child. `min(report, cap)` is also not a `Len`: `Len` is a sum
of the three parts, and the smaller of two of them is not. of the three parts, and the smaller of two of them is not.
So a cap takes pixels and a fraction and no share, which is `UiScalar`'s So a cap takes pixels and a fraction and no share, which is what `Len` is
shape (`core/src/orientation/pos.rs`) -- either that type or a `Len` since `a8898aa`. Still awaiting Bryan: whether a `Max` narrows the box the
without `leftover`. Awaiting Bryan: whether to split the type, and whether child draws in, or only what the parent reports for it.
a `Max` narrows the box the child draws in or only what the parent reports
for it.
Other queued work, in dependency order: Other queued work, in dependency order: