Files
iris/core
iris-aiandClaude Opus 5 1940e85c70 Move a whole box at once, since that is what a move does
Profiling a move by cycles rather than by instructions says the cost is not
where the last session recorded it. In `apply_scalar` the `i64` division is
**0.00%** of cycles and the multiply 1.5%: the time is in `saturating_add`,
which is five instructions and no vector form for an `i32`, and a box that
only moved does eight of them. Asking for them one scalar at a time, each
behind a match on which kind of move this is, gives the compiler four short
sequences where it had four adds in a row to pair up.

So a translation is now asked for once for the whole region -- which is what
a translation is -- and the match happens once above it rather than per
scalar. `many` over 500 frames: 684M cycles to 660M, and 1,815,666,327
instructions to 1,742,553,104.

Cycle counts are worth trusting here, which is the other thing to keep: three
runs of one binary varied 0.23%. It is wall time that varies 2x on this
machine, not the counters, and instructions alone cannot see a stall.

Checked: fmt, clippy, 105 tests, all five shrinker cases at 300 seeds, and
`tabs`, `text` and `random` byte-identical at 1920x1200.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-16 13:50:58 -04:00
..