# How iris should render an unbounded number of images ## Status (2026-09-04) **Implemented**, on the `rustify` branch of `ai-app-2`, in `iris/core` and `iris/src/default/render.rs`. See "Implemented, 2026-09-04" at the bottom for what landed, what differs from the proposal below and why, and what was verified versus merely reasoned about. The short version: the binding array is gone, `request_device` asks for no features and no binding-array limits, and that is now proven on the emulator's software Vulkan (`rigs/gpu-probe`), not just read from the code. `RUST.md`'s blocking item is resolved. Iris (the person) asked whether iris's (the library's) approach to "draw however many images happen to be on screen" — relevant here because a transcript can hold an unbounded number of attached screenshots — actually works on mobile, her recollection being that it does not. Checked rather than assumed, on 2026-09-04, on the `rustify` branch of `ai-app-2`. This file is that investigation and the resulting recommendation, written for a second agent to review before anything in iris's render core changes — no code has been written against this yet. ## The problem Every texture iris ever creates — every `Image` widget (`iris/src/widget/image.rs`) and every glyph atlas page — gets a permanent slot in one array via `Textures::add` (`iris/core/src/primitive/texture.rs:65`). Both of iris's texture-sampling primitives (`TEXTURE` and `GLYPH`) read that array by index: `core/src/render/shader.wgsl:56` declares `var views: binding_array>`, sized by `UiLimits::default()` (`core/src/render/mod.rs:347`) at **100,000 textures, 1,000 samplers**. Getting a device to accept that layout needs three wgpu features — `TEXTURE_BINDING_ARRAY`, `SAMPLED_TEXTURE_AND_STORAGE_BUFFER_ARRAY_NON_UNIFORM_INDEXING`, `PARTIALLY_BOUND_BINDING_ARRAY` — which correspond to Vulkan's `VK_EXT_descriptor_indexing` ("bindless"), promoted to Vulkan core at 1.2. A transcript with an unbounded number of image attachments is exactly the case that grows this array without bound: each attachment becomes its own `Image` widget, which takes its own permanent array slot until dropped. ## What was measured **A new rig, `rigs/gpu-probe`**, asks a device for exactly iris's features and limits with no window and no APK — a plain executable pushed with `adb push` and run from `/data/local/tmp`. It has two parts: `wgpu::Adapter::request_device` with iris's exact `Features`/`Limits` (`src/main.rs`), and a raw Vulkan query bypassing wgpu entirely via `ash` (`src/vk.rs`), to tell "the driver doesn't have it" apart from "wgpu didn't detect it." - **On this VM's own GPU** (Vulkan via Venus onto an RX 7900 XT): `IRIS DEVICE: ok`. Not the case that matters — nobody's phone is a discrete desktop GPU — but it is why the design was never checked before now: it always worked in the one place it was tried. - **On the Android emulator's guest Vulkan**, both ICDs it ships (`vk_swiftshader_icd.json` and, cold-booted, `lvp_icd.json`/lavapipe): `request_device` **fails** — `Unsupported features were requested: TEXTURE_BINDING_ARRAY | SAMPLED_TEXTURE_AND_STORAGE_BUFFER_ARRAY_NON_UNIFORM_INDEXING | PARTIALLY_BOUND_BINDING_ARRAY`. The raw `ash` query on lavapipe shows the driver itself reporting all seven descriptor-indexing sub-features as `true` at device API version 1.3 — so wgpu-hal's own feature detection is being more conservative than the driver here, for a reason not chased further (a likely instance-version negotiation gap, since the extension only promoted to core at 1.2). That part is a wgpu-hal/emulator question, not the finding that matters, and is **not** why this design is rejected. **The finding that matters is about real phones, sourced rather than recalled:** - The **Android Vulkan Profile 2025** — Google and Khronos's current baseline, covering **80.1% of active Vulkan-capable Android devices** as of October 2025 ([developer.android.com/ndk/guides/graphics/android-vulkan-profile](https://developer.android.com/ndk/guides/graphics/android-vulkan-profile)) — does **not** require `VK_EXT_descriptor_indexing` or any descriptor- indexing feature. It requires `shaderSampledImageArrayDynamicIndexing` (indexing by a value uniform across the invocation — Vulkan 1.0 baseline, unrelated to bindless) and stops there; true of the 2021 and 2022 profiles as well. - Arm's own developer documentation states **"`VK_EXT_descriptor_indexing` is supported on all Valhall and 5th Gen GPUs"** ([developer.arm.com/mobile-graphics-and-gaming/vulkan-api-best-practices-on-arm-gpus](https://developer.arm.com/mobile-graphics-and-gaming/vulkan-api-best-practices-on-arm-gpus)) — Mali generations from roughly 2019 (Mali-G77) onward, with no claim made for Bifrost, Midgard or Utgard, which are still common in budget and older Android phones that are still in daily use. - A search engine's summarized claim of "1% support on Android" for this extension was checked against its cited source (an Arm blog post from 2021) and **was not actually there** — that number does not appear in any primary source found and should not be repeated. The baseline- profile finding above is the one with an attributable source; use it instead. So this is not a software-renderer artifact. A real, currently-shipping share of the Android fleet lacks the feature iris's texture pipeline asks for unconditionally, and neither the emulator's failure nor the current official hardware baseline gives any reason to expect that to change soon. ## What growth already costs today, before any redesign Checked directly in `core/src/render/mod.rs` and `core/src/render/texture.rs`, because "does this redesign make things worse" needs the current baseline first: - The `RenderPipeline` (`UiRenderNode::new`) is created **once** and never rebuilt for any reason related to texture count — its bind group *layouts* declare fixed slot counts (`limits.max_textures`, `limits.max_samplers`) up front and that never changes at runtime. Growth was never at risk of recreating the pipeline, in the current design or any redesign discussed below. - What **does** get rebuilt: `UiRenderNode::update` calls `self.textures.update(&mut ui.textures)`, and if that reports any change, rebuilds `self.rsc_group` — one `BindGroup` whose entries are `BindingResource::TextureViewArray(&tex_manager.views())`, collected fresh over **every currently-live texture**, plus the sampler array and the mask buffer. This happens on every texture `Push`, `Set`, or `Free` — an image added anywhere in the whole app rebuilds one shared structure referencing every other image too. - The one path already excluded from this, on purpose, is a `Patch` — writing into an existing texture's pixels without changing which textures exist. The code says why directly (`core/src/render/texture.rs`, in `GpuTextures::update`): *"A patch changes texture contents, not the binding array, so it must not report `changed` — rebuilding the bind group per glyph is the cost this exists to avoid."* This is exactly the mechanism I1 built for the glyph atlas: growing an existing atlas page costs a `write_texture` into a sub-rect, nothing else. So today, growth that stays inside an existing texture (glyphs added to an atlas page) is already free. Growth that adds a *new* texture — a new atlas page, or any standalone image — already rebuilds the one shared array regardless of how the array is populated, before any change discussed below. That existing cost is O(live texture count) in CPU work to collect the view list and in however expensive the driver finds a descriptor-set-sized-for-N-descriptors to be. ## Prior art, checked rather than assumed Two independent projects were checked to see whether "atlas for images" is actually how this is normally done, rather than a guess: - **egui_wgpu** (`crates/egui-wgpu/src/renderer.rs` in emilk/egui), the closest prior art to iris — an immediate-mode wgpu-backed UI library that ships on Android. It keeps a `HashMap` and gives **each texture its own ordinary `BindGroup`** — one texture, one sampler, no array, no descriptor indexing of any kind. Draw calls are batched by texture id and the bind group is switched between batches within the render pass. - **Vello** — the renderer Masonry (E1/E2's Linebender stack) draws through — hit the identical problem and wrote down why in their own roadmap document ([github.com/linebender/vello/blob/main/doc/roadmap_2023.md](https://github.com/linebender/vello/blob/main/doc/roadmap_2023.md)): *"The number of images that may appear in a scene is not bounded, which is not a good fit for the basic descriptor binding model... Until then, we'll do a workaround of having a single atlas image containing all the images in the scene."* Their reason is broader than Android — WebGPU 1.0 has no descriptor indexing at all — but it reaches the same conclusion for the same shape of problem: atlas, not a bigger bindless array. **This is also a live hazard, not a solved one.** Vello's own changelog (Sparse Strips v0.2.0) lists a fix titled *"WebGL image-atlas allocation and growth on Mali-G52 GPUs, avoiding application-not-responding errors"* — an actual ANR, from atlas growth, on an actual mid-range Android GPU, in the renderer Masonry is built on. The same release added `AtlasSpaceDiagnostics`/`AtlasLayerDiagnostics` (per-layer free-space, utilization, fragmentation) because growth needed instrumenting in production, not because it turned out to be free. ## Recommendation (not yet implemented) 1. **Small, plentiful textures** — glyphs (already done, I1), thumbnails, downscaled attachment previews, icons — go through a shared atlas, the same technique as `core/src/render/atlas.rs` generalized beyond glyphs. Adding one to an existing page is a `Patch`, already free per the section above. 2. **Large or one-off images** — a photo attachment opened at full resolution, anything that would fragment a shared page — get their **own ordinary, non-array bind group**, the egui_wgpu way. Creating one is O(1): it references only itself, and does not touch any other texture's binding, unlike today's shared array where every push rebuilds a structure listing everything. 3. **Opening a new atlas page** is the one case that still resembles today's rebuild — infrequent (bounded by how many *pages* are needed, not by how many images have ever been attached) but not free, and Vello's Mali-G52 fix says this specifically deserves care: it should never be allowed to block a frame, and it is worth having the equivalent of Vello's atlas diagnostics before trusting it under load. 4. **Net effect**: dropping `TEXTURE_BINDING_ARRAY`, `SAMPLED_TEXTURE_AND_STORAGE_BUFFER_ARRAY_NON_UNIFORM_INDEXING`, and `PARTIALLY_BOUND_BINDING_ARRAY` from iris's device request entirely. Every path above is plain Vulkan 1.0 / GLES-level texture sampling. This is also what fixes the emulator failure measured above, regardless of the unresolved wgpu-hal question: a device that never asks for the feature cannot be refused for lacking it. ## What this touches, and what is still open Implementing this reworks iris's rendering core: the shader's binding group layout (`shader.wgsl`), `Textures` and `GpuTextures` (`core/src/primitive/texture.rs`, `core/src/render/texture.rs`), both texture-sampling primitives, and `core/src/ui/painter.rs`'s draw-call batching (today one draw call can reference any texture by index; the per-texture-bind-group path needs draws grouped by which bind group they use). Nothing has been started. Open questions a reviewer should weigh in on: - **The size threshold** between "goes in an atlas page" and "gets its own bind group." Too low and ordinary attachment thumbnails end up as one-off bind groups, losing the batching benefit the atlas exists for; too high and a page fragments on a handful of medium images. - **Eviction policy** for atlas pages once the working set does not fit — today's `GlyphAtlas` never evicts, because a font's glyph set is small and bounded; images are not. An LRU at the page level, or at the individual-image level within a page, has not been designed. - **Whether iris should keep any binding array at all**, even a small fixed one (say, capped at a few dozen slots) for atlas pages themselves, or whether every atlas page should also be its own ordinary bind group like standalone images — the array's only remaining justification would be avoiding a bind-group-per-draw-call switch cost that has not been measured on this project's actual target hardware. - **How this interacts with I2/E2's virtualised list** (I3): a bottom-anchored transcript composes only visible rows, so the live texture set should already be bounded by what is on screen rather than by the whole conversation — worth confirming that invariant holds before relying on it to keep atlas/bind-group churn small. ## Review, 2026-09-04 A second pass over the file above against the code, done before anything is implemented. Iris's worry going in: a bind group per texture means a draw call per image, and she wants this as efficient as it can be. ### What checked out Every code reference above is accurate as of this commit: the 100,000 / 1,000 limits, the one-time pipeline, the `rsc_group` rebuild on every `Push`/`Set`/`Free`, and the `Patch` exclusion. The device request that asks for the three features is `iris/src/default/render.rs:96`, which the text above does not name. egui-wgpu and Vello are described correctly. ### The emulator refusal is a wgpu-hal gap, now located The file guessed "a likely instance-version negotiation gap." It is narrower than that and it is in wgpu-hal, not the emulator. wgpu-hal 28.0.0 (`src/vulkan/adapter.rs:1618`) only queries `PhysicalDeviceDescriptorIndexingFeaturesEXT` **when the device advertises the `VK_EXT_descriptor_indexing` extension string**. A Vulkan 1.2+ driver that has descriptor indexing as core need not list the extension, and lavapipe at 1.3 evidently does not, so wgpu never asks and reports the features absent, which is why `ash` sees seven `true`s and wgpu sees none. The properties query beside it (line 1486) correctly accepts `device_api_version >= 1.2 || extension`; the features query does not. wgpu-hal 30.0.1 in the local registry has the same asymmetry (lines 1872 and 2036). Worth an upstream issue, but not a reason to keep the design: on real phones the gate that matters is stricter still. **wgpu's `TEXTURE_BINDING_ARRAY` needs six sub-features, not one** (`adapter.rs:160-177`): non-uniform indexing *and* update-after-bind for sampled images, storage images and storage buffers, all together, because wgpu marks every array-bearing descriptor set update-after-bind. So Arm's "the extension is supported on Valhall" is necessary but not sufficient; a driver with sampled-image indexing and without storage-buffer update-after-bind is refused too. That widens the excluded set beyond what the Arm quote suggests and strengthens the conclusion. ### A live bug in the current code, found on the way `GpuTextures::update` (`core/src/render/texture.rs:33`) implements "a patch must not report changed" as `changed = false`, unconditionally, which also **cancels a `Push` earlier in the same batch**. That ordering is exactly what opening a new atlas page produces: `GlyphAtlas::allocate` pushes the page and `insert` patches it in the same frame, so the bind group is not rebuilt and the new page's view is not bound until some unrelated texture change happens to rebuild it. It is hidden today only because the masks path also sets `changed`. The fix is one line (`changed |= !matches!(update, Patch)` in spirit); it should go in with the redesign since that code is being replaced, and it is recorded here so it is not rediscovered. ### In-layer draw order is already undefined Relevant to any batching redesign: `Primitives::apply_free` (`core/src/render/primitive.rs:147`) uses `swap_remove`, so the instance order within a layer is permuted whenever anything is freed. Overlap order inside one layer is therefore not something the renderer promises today; ordering is done with layers. That means grouping a layer's draws by texture, or drawing a layer's images after its rects and glyphs, loses nothing that currently exists. It should be written down as an invariant when the redesign lands, because the new code will depend on it. ### On "a draw call per image" Two corrections to the worry. First, it is a draw per *distinct texture per layer*, not per image primitive: every glyph quad in a layer shares the atlas and stays one instanced draw, and a thumbnail atlas would do the same for previews. Second, the count is bounded by what is on screen, which I3's virtualised transcript already bounds, and a mobile GPU is not draw-call bound at tens of draws per frame; egui ships exactly this on Android. What does cost is per-frame *bind group creation* and per-frame *sorting*, and the current code already creates a `primitive_group` bind group every time a layer updates (`render/mod.rs:103`), so one more per new image is not a regression in kind. ### Recommended shape (proposal, for Iris to accept or change) Aimed at the fewest moving parts that need no feature beyond Vulkan 1.0: 1. **Atlas pages become layers of one `texture_2d_array`**, not separate textures. Every page is already `PAGE`x`PAGE` RGBA8, which is the one constraint an array texture imposes. A layer index is an ordinary sampling operand in WGSL and needs no indexing feature, so `GLYPH` (and any future atlased-image primitive) carries a layer instead of a `view_idx` and all of a layer's text stays **one draw**. This answers the open question above about keeping a small binding array: no. Cost of opening a page: recreate the array with one more layer and `copy_texture_to_texture` the old ones, GPU-side, no readback; grow with headroom (double) so it is rare. wgpu's default `max_texture_array_layers` is 256, at 4 MB each, so the cap is memory rather than the API. 2. **Every standalone image is its own texture with its own bind group**, and its instances live in a **separate per-layer instance list**, not the main one. Then the main instance buffer never contains an image, there is nothing to sort, no handle remapping beyond what `apply_free` already does, and each image is `draw(0..4, k..k+1)` with its bind group set first. Group 2's layout becomes `{atlas array, one image texture, sampler, masks}`; the main draw binds a 1x1 null image in the image slot, each image draw binds its own. One pipeline, one shader, one layout. 3. **No thumbnail atlas in the first version.** With images on their own textures, the threshold and eviction questions above disappear: an image is freed when the row that owns its `TextureHandle` scrolls out. Add an image atlas only if a measured screen shows enough small images to matter, which a transcript rarely does. 4. **Drop the three features and the two `max_binding_array_*` limits from `src/default/render.rs`**, and the `UiLimits` counts with them. 5. **Sampling is `NonFiltering` today** (`render/mod.rs:290,299`), so a downscaled attachment will alias. Either request a filtering sampler for the image slot or downscale on the CPU before upload; decide when the image widget is touched, not as part of this. What this costs against the file's original recommendation: `Textures` needs to know an image from a page (two kinds of handle, or a kind on `TextureHandle`), and `Primitives` gets a second instance list per layer. What it saves: the sort, the size threshold, the eviction policy, and any per-page bind group switch. ## Implemented, 2026-09-04 The shape above, built as proposed with one structural addition the proposal didn't need to spell out and one bug it predicted made moot rather than literally fixed. Files: `core/src/primitive/texture.rs` (`Textures`, `TextureHandle`), `core/src/render/texture.rs` (`GpuTextures`), `core/src/render/primitive.rs` (`Primitives`, `GlyphPrimitive`), `core/src/render/atlas.rs`, `core/src/ui/painter.rs`, `core/src/render/mod.rs` (`UiRenderNode`, `UiLimits` removed), `core/src/render/shader.wgsl`, `src/default/render.rs`, and `rigs/gpu-probe/src/main.rs`. **1. Atlas pages as array layers.** `GpuTextures` owns one `texture_2d_array` (`array_texture`/`array_view`), grown by doubling (`grow_array`): a new texture is created at twice the layer capacity, the old layers are copied across with `copy_texture_to_texture` (GPU-side, no readback), and every bind group that referenced the old view — the main one and every live standalone image's — is rebuilt, since the view's identity changed. `GlyphPrimitive` carries `layer: u32` instead of `view_idx`/`sampler_idx`; the layer number is assigned synchronously in `Textures::add_page` (a plain counter, `next_page_layer`), not by the renderer, because `GlyphAtlas::insert` needs it in the same call, before any GPU sync happens — the renderer only finds out later, when it processes the queued `Push`. **2. Standalone images, one bind group each.** `TextureKind` on `TextureHandle`/`Textures` distinguishes `Image` (a plain bind-group index, `slot`) from `Page { layer }`. `Primitives` gained a second per-layer list — `images: Vec`, tagged `IMAGE_BINDING` — separate from `instances` (rects and glyphs), written by `Painter::write_image` rather than through the generic `Primitive` trait, since an image has nowhere in `PrimitiveData` to put a per-instance entry once the bind group already picks the texture. `UiRenderNode::draw` draws a layer's `instance` buffer once as before, then walks `image_instance` one entry at a time, binding that texture's `BindGroup` (`GpuTextures::image_bind_group`) and issuing `draw(0..4, k..k+1)` per image. Group 2's layout is exactly the proposed `{atlas array, one image texture, sampler, masks}`; the main draw binds a 1x1 null view in the image slot. **The one addition beyond the proposal**: the masks storage buffer lives in every per-image bind group (group 2, binding 3), and `ArrBuf` recreates its buffer whenever the mask count changes size (`render/util/mod.rs`'s `ArrBuf::update` now returns whether it resized). A resize invalidates every bind group holding the old buffer, not just the main one, so `GpuTextures::update` takes a `masks_resized: bool` and calls `rebuild_image_bind_groups` when it's set, alongside the same rebuild the array-growth path already needed. This wasn't a design question the proposal had to answer (it treated bind-group construction as a given), but it's exactly the shape of trap layer growth already had, so it uses the same fix. **3. No thumbnail atlas.** Not built, as proposed. **4. Removed**: `TEXTURE_BINDING_ARRAY`, `PARTIALLY_BOUND_BINDING_ARRAY`, `SAMPLED_TEXTURE_AND_STORAGE_BUFFER_ARRAY_NON_UNIFORM_INDEXING` from `src/default/render.rs`'s `request_device`, and `UiLimits` (the type itself, not just its binding-array methods — once its two fields were gone there was nothing left in it, and `UiRenderNode::new` no longer takes a limits parameter). `binding_array` no longer appears anywhere in `shader.wgsl`. **5. Sampling** is still `NonFiltering`, unchanged, per the proposal's own note that this is a separate decision for whenever the image widget itself is touched. **The `changed = false` bug is structurally gone, not patched.** The old `GpuTextures::update` held one `changed: bool` that a `Patch` reset unconditionally, which could erase an earlier `Push` in the same batch (a new atlas page's `Push` immediately followed by `GlyphAtlas::insert`'s `Patch`, both queued before the renderer ever runs). The new `update` computes the rebuild signal by OR-ing each event's own answer (`rebuild_main |= self.push(...)`), and `Patch`'s arm simply never contributes to it — there is no shared mutable flag left for a `Patch` to stomp on. Documented at the call site (`core/src/render/texture.rs`, `GpuTextures::update`'s doc comment and the `Patch` match arm's comment) rather than fixed as a one-line diff, since the mechanism that could go wrong no longer exists. **In-layer draw order is an explicit invariant now, not just a fact about `swap_remove`.** `UiRenderNode::draw` draws every layer's images after its rects and glyphs, and `Primitives::apply_free`'s doc comment states directly that both of a layer's lists (`instances` and `images`) free with `swap_remove` and that nothing may assume adjacency survives a free — recorded there because `apply_free` is the one place a change to either list's ordering would have to be reconciled. **Verified:** - `cargo fmt --all -- --check`, `cargo build --workspace --all-targets`, `cargo clippy --all-targets`, `cargo test --workspace` all clean in `iris/`, on the pinned `nightly-2026-09-03` toolchain. 14 tests pass (unchanged from I1; nothing here is pure-logic enough to add a unit test to — it's all GPU resource wiring). - `iris/run-headless.sh minimal --shot /tmp/minimal.png` and `iris/run-headless.sh tabs --shot /tmp/tabs.png`: both render correctly on this VM's GPU (Venus) — `tabs`'s glyph-atlas text renders in every panel, confirming `GlyphPrimitive.layer` addresses the array correctly. - The standalone-image path specifically: a throwaway example (not committed) with an `image(...)` widget as part of the root, run the same way, rendered the image next to glyph-atlas text in one frame — confirming a live `BindGroup` built by `GpuTextures::create_image` and bound per-`draw()` call actually samples the right texture. `tabs`'s own "image span" tab exercises the same widget but needs a click to reach, which the headless compositor can't deliver (no seat devices, per I1's own note on this file) — the throwaway example is what stood in for it. - **Not separately stress-tested**: triggering a second atlas page (the `grow_array` doubling-and-copy path) under a real glyph load large enough to fill the first 1024x1024 page. The code path was reasoned through and matches the existing single-page write exactly except for the `z` origin and the extra copy, but nobody has watched a real second-page grow happen on screen. Worth doing before trusting this under a transcript with a large or unusual glyph set (many distinct fonts/sizes, or a font with an unusually large character set). - **The decisive check**, `rigs/gpu-probe` rewritten to request iris's new (empty) feature/limit set and run on this checkout's own emulator (`ai-app-2`, via `emu`), booted with `EMU_GPU=software` so the guest gets a real Vulkan device (SwiftShader) rather than the `-gpu host` default, which disables Vulkan in this VM entirely (`-feature -Vulkan`, because gfxstream can't pair Venus with the real GPU here — worth remembering, since the *default* `emu up` gives a device with **no** Vulkan adapter at all, which reads exactly like the old bindless failure if you don't know to ask for `EMU_GPU=software`): cd rigs/gpu-probe ANDROID_NDK_HOME=$HOME/Android/Sdk/ndk/29.0.14206865 \ cargo ndk -t arm64-v8a -P 26 build --release EMU_GPU=software emu up # from ~/repos/emulator-tools adb push target/aarch64-linux-android/release/gpu-probe /data/local/tmp/ adb shell chmod 755 /data/local/tmp/gpu-probe adb shell /data/local/tmp/gpu-probe Output: `adapters: 1 — Vulkan SwiftShader Device (Subzero) (Cpu)`, `features iris requires:` (none listed — the set is empty), `max_buffer_size … ok`, and **`IRIS DEVICE: ok`**. This is the fix measured working, on the exact rig that first measured it failing. Emulator stopped afterward (`emu down`); nothing was left running.