Files
iris/tests/revision_cost.rs
T
iris-aiandClaude Opus 5 e5f8b6b244 Shape a text once per width, not once per ask
A container measures a child by drawing it in a box it may not keep, so
one layout asks a text for a dozen widths and comes back to widths it
has already had -- the hottest text in the depth-8 tree draws 32 times.
Each ask re-ran the shaper, because the two caches in front of it held
one entry each and a trial width alternating with a final width evicts
the answer about to be wanted again. `perf record` put 63% of a resize
frame in text and 0.9% in `draw_inner`.

So keep more than one: a bounded store of shapings on `TextData`, keyed
by the text, the attrs and the width, holding the parley layout and the
glyphs placed from it. Bounding the store rather than each buffer is
what keeps it a fixed cost -- +4 MB on a tree of 4,000 texts, which is
19 MB less than the code before #16 holds after the same resizes.

`TextBuffer` now holds the glyphs of the shaping it is drawn as, which
is where `TextView::tex` was. That leaves one place to invalidate rather
than two, so the `MutDetect` flags on a view's text and attrs have no
reader and go, along with the `buf.changed = true` after every edit.

On a 40-row tree of distinct random paragraphs, 500 resize frames:
124.2M instructions per frame before, 17.7M after, and 45.9M when the
width never repeats. The five reference renders and the resize render
are byte-identical, and the 100-seed sweep passes.

`tests/revision_cost.rs` is that tree, written in the API subset
`43ce8c7` shares so the same source measures the code this replaced.
Report the worst frame and p99 beside the median, since a stutter is
what somebody sees. Count glyph placements, and count a text render per
ask rather than per shaping, so the store cannot hide how many times a
layout drew the same text.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-14 18:49:45 -04:00

204 lines
5.9 KiB
Rust

//! What a resize frame costs and what it holds, on a tree the revision before
//! #16 also builds.
//!
//! Deliberately written in the API subset `43ce8c7` and this branch share, so
//! the same source can be dropped into an old worktree and measured there:
//! that is the only like-for-like comparison with the code the retained
//! layout replaced. The random tree cannot carry one, because the generator
//! itself changed with the work.
//!
//! ROWS=40 FRAMES=500 cargo test --release --test revision_cost \
//! -- --ignored --nocapture resize_cost
//! ROWS=2000 cargo test --release --test revision_cost \
//! -- --ignored --nocapture text_memory
//!
//! Wall time on this machine varies with CPU frequency; take the number from
//! `perf stat -e instructions:u` on the test binary directly.
use iris::harness::Harness;
use iris::prelude::*;
use std::time::Instant;
/// xorshift64, so one seed is one set of paragraphs on any machine.
struct Rng(u64);
impl Rng {
fn bits(&mut self) -> u64 {
self.0 ^= self.0 << 13;
self.0 ^= self.0 >> 7;
self.0 ^= self.0 << 17;
self.0
}
fn below(&mut self, n: usize) -> usize {
(self.bits() % n as u64) as usize
}
}
const WORDS: [&str; 24] = [
"wrapping",
"shapes",
"one",
"source",
"into",
"as",
"many",
"lines",
"as",
"the",
"box",
"leaves",
"room",
"for",
"paragraph",
"height",
"answer",
"setting",
"container",
"width",
"before",
"knows",
"measured",
"again",
];
/// A run of its own words, so nothing here is fast for two texts being the
/// same string.
fn words(rng: &mut Rng, least: usize, most: usize) -> String {
let words = least + rng.below(most - least);
let mut out = String::new();
for _ in 0..words {
if !out.is_empty() {
out.push(' ');
}
out.push_str(WORDS[rng.below(WORDS.len())]);
}
out
}
const OUTPUT: (f32, f32) = (900.0, 1200.0);
fn env<T: std::str::FromStr>(name: &str, fallback: T) -> T {
std::env::var(name)
.ok()
.and_then(|value| value.parse().ok())
.unwrap_or(fallback)
}
/// A row of a fixed-width rect beside a column of one wrapping and one
/// overflowing text: the shape that makes a container measure a child in a
/// box it will not keep.
fn build(h: &mut Harness, rows: usize) -> Vec<WidgetId> {
let mut rng = Rng(1);
let mut paragraphs = Vec::new();
let mut col = Span::empty(Dir::DOWN);
for _ in 0..rows {
let mut row = Span::empty(Dir::RIGHT);
row.push(
rect(Color::RED)
.width(Len::abs(40.0))
.add_strong(&mut h.rsc),
);
let mut body = Span::empty(Dir::DOWN);
let para = wtext(words(&mut rng, 12, 52))
.size(16)
.wrap(true)
.add_strong(&mut h.rsc);
paragraphs.push(para.id());
body.push(para);
body.push(
// Short, or its unwrapped width decides the row and the
// paragraph beside it never wraps.
wtext(words(&mut rng, 2, 6))
.size(16)
.wrap(false)
.add_strong(&mut h.rsc),
);
row.push(body.add_strong(&mut h.rsc));
col.push(row.add_strong(&mut h.rsc));
}
let root = col.add(&mut h.rsc);
h.set_root(root);
paragraphs
}
#[test]
#[ignore = "measurement, not a check"]
fn resize_cost() {
let rows = env("ROWS", 40_usize);
let frames = env("FRAMES", 500_usize);
let mut h = Harness::new(OUTPUT);
let paragraphs = build(&mut h, rows);
// What it cost is only half the comparison: the old code is cheaper
// partly because it wraps at the container's whole width rather than the
// part left beside the rect, and draws past the edge of the output.
println!("output width {}", OUTPUT.0);
for (at, id) in paragraphs.iter().enumerate().take(3) {
println!("paragraph {at}: {:?}", h.region(id));
}
// Two widths in turn is the friendly case for anything that remembers an
// answer, so `SWEEP=1` never repeats one -- a drag rather than a toggle.
let sweep = env("SWEEP", 0_usize) != 0;
let mut elapsed = Vec::with_capacity(frames);
for frame in 0..frames {
let narrower = match sweep {
true => (frame % 256) as f32,
false => ((frame + 1) % 2) as f32 * 8.0,
};
h.resize((OUTPUT.0 - narrower, OUTPUT.1));
let start = Instant::now();
h.frame();
elapsed.push(start.elapsed().as_secs_f64() * 1000.0);
}
elapsed.sort_by(|a, b| a.partial_cmp(b).unwrap());
println!(
"resize: {frames} frames, min {:.3} ms, median {:.3} ms, p99 {:.3} ms, \
max {:.3} ms, total {:.1} ms",
elapsed[0],
elapsed[frames / 2],
elapsed[frames * 99 / 100],
elapsed[frames - 1],
elapsed.iter().sum::<f64>()
);
}
fn kb(field: &str) -> u64 {
std::fs::read_to_string("/proc/self/status")
.unwrap()
.lines()
.find(|line| line.starts_with(field))
.and_then(|line| line.split_whitespace().nth(1)?.parse().ok())
.unwrap()
}
fn report(label: &str) {
println!(
"{label:24} rss {:>7} kB peak {:>7} kB",
kb("VmRSS:"),
kb("VmHWM:")
);
}
/// Run this one on its own: the figures are the whole process's.
#[test]
#[ignore = "measurement, not a check"]
fn text_memory() {
let rows = env("ROWS", 2000_usize);
report("before");
let mut h = Harness::new(OUTPUT);
let paragraphs = build(&mut h, rows);
report("after cold frame");
for frame in 0..40 {
h.resize((OUTPUT.0 - ((frame + 1) % 2) as f32 * 8.0, OUTPUT.1));
h.frame();
}
report("after 40 resizes");
// Settled: the output holds still and one leaf repaints per frame.
for _ in 0..10 {
let _ = h.rsc.widgets_mut().get_dyn_mut(paragraphs[0]);
h.frame();
}
report("after settling");
}