`tests/chain_cost.rs` times the pass on the GPU with timestamp queries,
which this adapter supports, rather than by the clock. 200,000 two-pixel
instances at 1024x1024, so vertex work dominates, best of eight batches:
depth 1 77.9 us +0.0%
depth 2 78.1 us +0.2%
depth 4 79.0 us +1.3%
depth 8 81.8 us +5.0%
depth 16 111.1 us +42.6%
depth 32 159.6 us +104.9%
depth 64 250.3 us +221.3%
Free to about depth 8 and then roughly 3 us per level. Each step is a
storage load whose address is the previous load's result, so it is the
chaining that costs rather than the arithmetic at each level -- which
means the number would look the same for a slot carrying a whole region
instead of a delta.
That matters because every active widget owns a slot, so a primitive
resolves through its full depth in the widget tree, and LAYOUT.md notes
real trees have exceeded 16. At the couple of hundred primitives an
example draws it is nothing; a transcript's glyphs are tens of thousands
of primitives, which is the regime measured here.
No behaviour change. Recorded rather than acted on: keeping the chain
shallow means not giving every widget a slot, which is a design decision
of LAYOUT.md §2 and §6 and the owner's to make.