Files
ai-app/.claude/skills/ai-app-rigs/SKILL.md
T
iris-ai 8c323fc7a9 Serve a machine's models from one shared llama-server
A llama.cpp session had its own `llama-server`: two sessions on one model
held two copies of it in memory, a model change bought a load only that
session benefited from, and the process was a session's to end. A machine's
models are now served by one `llama-server` in **router mode** -- no `-m`,
a preset file naming models and their flags, a child server per model asked
for, and each request routed by its `model` field. So one server per model
with that model's own settings is what a machine runs, while this backend
has one process, one port and one record per machine to keep track of.

The record is the mechanism every other driver already uses, so a restart
adopts it; a session records the same pid in its own directory as
`Detail::Shared`, and `process::signal` refuses to signal one of those --
which is what keeps stopping, deleting or cleaning up after one session
from unloading a model every other session is using. Nothing stops a router
on its own. That is deliberate (a loaded model is minutes of disk) and it is
why the machines tab now has a card per provider that opens its own screen:
how each model is loaded, how many stay in memory, Unload, and Stop.

How a model is *loaded* therefore belongs to the model on its machine rather
than to a session -- context size, GPU layers, threads, slots, speculative
decoding -- written into the preset as llama-server's own argument names.
Saving them re-reads that file, which unloads the model; that is the change
taking effect, and the dialog says so before you save. What stays a
session's is everything that rides on a request, including which tools it
offers: the router hosts one set for the machine and the choice is a filter
applied here, so it costs no reload (2,181 tokens of prompt with all seven,
698 with none).

Verified end to end against the scratch backend and the emulator: two
sessions sharing one loaded model with one child process, a second session
joining it with a 26ms prefill, a backend restart adopting the router and
answering with the prompt cache intact, the same over ssh to this VM, a
model's settings reaching the running server, Unload, and Stop leaving every
session `exited` with no error line.
2026-09-19 17:37:31 -04:00

329 lines
20 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
---
name: ai-app-rigs
description: ai-app's test rigs, harness scripts and reference measurements - ui-sandbox.sh, debug-transcript.sh, transcript-bench.sh, stream-bench.sh, trace-draw.sh, the /usage fixture vocabulary, the fake CLI, the rule that no UI-driving script may tap a coordinate, how to test llama.cpp and ssh on this machine, how importing behaves, and the scroll/stream/explorer numbers not worth re-measuring. Read before running or writing a benchmark, driving the app's UI from a script, exercising the session lifecycle, testing a llama or remote session, or touching the import screen.
---
# ai-app: rigs, harnesses and measurements
Moved out of `AGENTS.md` on 2026-09-04 so it is read when it is relevant
rather than sent with every request in this repo -- it was 12 KB of the 35 KB
that file cost on every one. Unchanged in the move, and still the only copy.
## The rigs
Each exists because something was invisible without it.
- **`app/ui-sandbox.sh`** — a second `ai-server` with its own `$HOME`, config
and data directory, holding eight invented Claude Code transcripts and a
`claude` that is two lines of shell. **That isolation is the point**: the
import screen lists whatever is in `~/.claude/projects`, which in this VM is
real agent transcripts, so exercising *delete* against the ordinary server
deletes somebody's conversation and exercising *import* starts a real
`--resume` on the owner's account.
Its port and root derive from the checkout's name, so two checkouts'
sandboxes cannot reach each other, and its token is generated once into
`~/.config/ai-app/sandbox-token` and carried across restarts along with any
the enrolment flow appended — so the emulator app is enrolled **once** (the
start banner prints the command) and stays enrolled. It shares the real TLS
certificates, because the installed APK pins that CA.
Driving verbs, so none of this is re-derived per session:
`./ui-sandbox.sh spawn [title]` (an echo session, prints its id),
`./ui-sandbox.sh send SID text|@file`, and
`./ui-sandbox.sh api /path [curl args]`.
`./ui-sandbox.sh keep` restarts the server without wiping the sessions and
enrolment already there — for when the fixture under test was expensive to
build; plain `start` wipes them, which is right for the list-screen
fixtures and wrong for that.
It passes `--delay` by default, and `AI_SANDBOX_BIG_MB` puts one large
transcript among the small ones while `AI_SANDBOX_SPAWN_DELAY` makes the
fake CLI slow to start. Both exist because operations that finish in
milliseconds have states on the way that nothing can observe, and an
unobservable state is one where broken and working look identical.
It also builds a fixture tree at the sandbox home's `~/files` for the
explorer, holding the states otherwise only reachable by finding a real
machine in one: an empty directory, a name with a tab and one with an
apostrophe, a binary file, one over `FILE_LIMIT`, one `chmod 000`, a
symlink to a directory and a broken one, a source file per language, and
the three sizes the limits were measured against (`edit-32k.rs`,
`edit-128k.rs`, `big-source.rs`). Point a session at it with
`./ui-sandbox.sh api /sessions/<id>/cwd -X POST -H 'content-type: application/json' -d '{"cwd":"~/files"}'`.
The explorer's 409 is produced by editing the file on the machine
(`printf … > file`) between pressing the pencil and pressing save.
- **`app/debug-transcript.sh`** — a real conversation on the emulator. The
echo driver is the right rig for most things and the wrong one for anything
whose cost scales with what was actually written: a real reply is longer,
is real markdown, and carries tool calls whose input and output are
kilobytes. Two faults were invisible until a real transcript was loaded — a
page of history landing mid-fling threw the reader back to the newest end,
and parsing one real reply took 51ms against 4.6ms for a synthetic one.
`-b` takes the biggest conversation on the machine rather than the newest,
which is what a scrolling test wants; `--stop` takes it down.
It copies the transcript into `/tmp` and gives the server a `HOME` of its
own, so the import can only see the copy — importing spawns `claude
--resume`, and against the real file that is a second CLI writing to a
conversation somebody may still be in. **A transcript never goes in this
repository**: they hold whatever was said, read and written in that
session, and `~/repos` is shared with the host besides.
- **`/usage` in an echo session puts up an invented meter**, which is how the
rate-limit screens' states are reached without spending quota: `/usage 42`,
`/usage 95 20` (minutes left), `/usage 42 never` (the between-blocks window
with no reset time), `/usage 42 unreadable`, `/usage notloggedin`,
`/usage unreachable`, `/usage failed`, `/usage off`. The vocabulary is
`usage::Fixture`'s, since those are its states. With none set an echo
session meters nothing, which is the ordinary case and draws no bar.
- **A fake CLI exercises the process lifecycle without a token.** Point a
`claude_cli` provider's `command` at a script that ordinarily runs
`cat > /dev/null` and it behaves the way the lifecycle code cares about:
it holds the fifo open, records a real pid, writes nothing, and dies on a
signal. So adopt, stop, restart and start are all drivable without a real
`--resume` and without spending a turn on somebody's account. Reach for
this when what is under test is *whether a process is running*, and for
`debug-transcript.sh` when it is *what the transcript draws*. The sandbox's
version also handles `auth login`: it prints an inert Anthropic-shaped URL,
rejects any code except `sandbox-code`, and exits successfully for that one.
- **`/think [seconds]` in an echo session puts up a thinking card**, long
enough to watch it spin before it closes with the span it actually took.
The rest of the turn is the ordinary echo reply, so it is also the rig for
a block and a reply meeting.
- **`app/transcript-bench.sh`** is the standard scroll measurement: it opens
the first session (or `-k` keeps the current screen), scrolls a fixed
gesture loop, and prints the app's render report — the same one the in-app
copy button produces, whose `on screen:` line names what the viewport was
holding. Compare two runs with the same gestures; the emulator's absolute
frame times transfer nothing, the report's accounting does. Run it either
side of any change under `Markdown*.kt`, `Transcript*.kt` or
`SessionScreen.kt`'s list, and put the report in the commit. The numbers
that move first are the worst `record: one block`, the reparse mean while
streaming, and the draw phase's accounting line.
- **`app/stream-bench.sh [-k] FILE`** is that measurement for a reply still
arriving. It taps "Jump to latest" so the list is pinned to the newest end,
resets the report, sends FILE, waits for the transcript to stop growing,
and prints. Both of those are corrections to a first version that measured
nothing: a transcript parked further back never redraws while a reply
streams into it, and a session is idle at *both* ends of a turn, so polling
for idle answers before the turn has started.
- **`app/trace-draw.sh`** names what a scrolling frame spends inside the
framework, from `atrace` text output with no trace processor needed. It is
how the cost of a layout node per link was attributed to the framework
rather than guessed at.
### Driving the UI
**No script that drives this app's UI presses a coordinate.** Every control
is found by the name it already carries for assistive technology —
`ui-trace record --do "tap 'Session settings'"` — which resolves the label
against the screen at the moment of the gesture and fails the whole run when
it is not there. `app/bench-lib.sh` is what the bench scripts share for it. A
coordinate is a position measured once by hand, and anything that moves the
control makes the tap land on whatever now sits there — the bench then
reports a number that was never measured, which reads exactly like a result.
Both bench scripts pressed the render report at `tap 723 205` until that
button moved into the session settings dialog on 2026-09-03. The check that
none has crept back:
grep -n "tap [0-9]" app/*.sh
Swipes are still coordinates, deliberately: a gesture across a scrolling area
is a distance rather than a control.
**Two traps in the emulator bench loop**, each of which cost a run.
`adb shell pm clear` removes the enrolment and the notification permission
along with the saved anchors, so the next run measures a permission dialog —
re-enrol with the command `ui-sandbox.sh` prints, and
`pm grant … POST_NOTIFICATIONS`. And a saved scroll anchor is per session id,
so the only way two builds start a scroll from the same place is a *fresh
session for each*.
**The emulator is `~/repos/emulator-tools`' business, not this repo's.**
`emu up` creates and boots the AVD named after this checkout — whatever `emu
name` prints, never a name typed out here, since this file is the same in
every clone. `run-android.sh` is that plus a build and an install. The `adb`
on `PATH` after sourcing `android-env.sh` is that repo's wrapper, which fills
in `-s` from the same rule. Gradle does not go through it, so a Gradle init
script from `emulator-tools` runs `emu check` before `installDebug`,
`uninstallDebug` and `connectedAndroidTest` and fails rather than fanning out
to every attached device; when it refuses, say which device you mean at the
moment you use it — `ANDROID_SERIAL=$(emu serial) ./gradlew …`.
### Testing llama.cpp and ssh here
**Both are set up here** and need nothing typed. The prebuilt llama.cpp lives
outside the repo at `~/.local/opt/llama.cpp-vk` — a **Vulkan** build as of
2026-09-19, replacing the CPU one that was there before — and is symlinked as
both `~/.local/bin/llama-server` and `/usr/local/bin/llama-server`. The second
is what makes **discovery find it over ssh**: `~/.local/bin` is not on the
PATH a non-interactive ssh session gets. It resolves its own libraries through
`$ORIGIN`, so no `LD_LIBRARY_PATH` is needed.
Two models are downloaded under `~/.local/share/ai-app/models`:
- `unsloth/Qwen3-0.6B-GGUF/Qwen3-0.6B-Q8_0.gguf`, 639 MB, loads in ~4s. It
calls tools correctly and is the right rig for the driver's shape. Do not
judge *answers* by it — asked for the second line of a file it read from
line 2 and then named the third.
- `ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF/Qwen3.8-27B-GSQ-RCO-IQ3_S-mtp.gguf`,
12 GB, ~20s to load, and the only one here with a multi-token-prediction
head. It is the rig for anything about `loading` being a state of its own,
since 20s is long enough to send into.
**Do not test with a 2-bit quant**: the IQ2_XXS of the 0.6B produces fluent
nonsense, which reads exactly like a broken driver — `llama-cli` produces the
same from the file directly, which is how to tell the two apart in a hurry.
**The GPU is shared and llama-server dies loudly when it runs out.** A second
server loading a model while the 27B holds VRAM fails with `radv/amdgpu:
Failed to allocate a buffer` / `MESA: error: buffer allocation failed` and
exits mid-request. `-ngl 0` runs it on the 8 cores instead, which is the way
to test the driver while something else holds the card -- through the app, that
is the model's "Layers on the GPU" set to 0 in the machines tab's provider
view, and `--models-max` above 1 is how two models come to be loaded at once
in the first place.
**Testing tools and MCP without the app**: `llama-server --tools all` publishes
its built-in tools at `GET /tools` and runs one at `POST /tools` with
`{"tool": …, "params": …}` and an `x-tool-cwd` header — so a whole agent loop
is drivable with `curl` and no model at all. The Exa MCP server at
`https://mcp.exa.ai/mcp` answers **without an API key** and needs a
`User-Agent` header (Cloudflare answers 403 without one, which reads as a
refusal rather than a missing header).
There is no second machine, so **ssh this VM to itself**. That is set up
too: the key is `~/.config/ai-app/ssh-self` (its public half is in
`~/.ssh/authorized_keys`, labelled removable), and the real config carries a
machine called **"this vm over ssh"** — `bob@127.0.0.1` with that
`identityFile` plus
`options: ["StrictHostKeyChecking=no", "UserKnownHostsFile=/tmp/ai-app-known-hosts"]`
so it touches nothing real — offering `claude-cli` and `llama-cpp`. It is the
whole rig for "does a remote llama session work", since the far machine is
this one and the model file is the same file. For a throwaway machine of your
own, point a provider's `command` at something harmless like `/bin/echo`
rather than at `claude`: the transport is what is under test, the process
exiting immediately is the signal, and it costs no tokens. The remote login
shell here is **fish**; the
remote script and `ssh.rs`'s POSIX quoting happen to mean the same thing in
both, but that is luck rather than design, and a shell that is neither is the
thing to suspect first if a remote spawn ever mangles an argument.
## Importing
The import list reports each session's **size as well as its line count**,
because the two disagree in the way that matters: these transcripts embed
screenshots as base64, so one line can be a megabyte. On this machine a 69 MB
session has 3,427 lines and a 44 MB one has 6,792 — nothing about a line
count tells you what continuing a session will cost. Shown, not warned about;
importing a large session is a choice somebody is entitled to make.
**Never import a Claude Code session that is open in a terminal.** The app
refuses it — see PLAN.md for the incident that made that a refusal rather
than a warning.
**One Claude Code session id can name two files, and the listing offers it
once.** Resuming from a different working directory makes the CLI write a
second transcript with the same id under that directory's project folder — an
ordinary state of a machine, not corruption. Everything downstream addresses
a session by id, and the phone keyed its list on it, so two rows sharing one
**closed the app** on a Compose duplicate-key throw. `parse_listing` keeps
the copy with the most lines, because the other is usually a few-hundred-byte
stub and is often the *newer* of the two, so recency is the wrong key.
Deleting removes every copy rather than the first, or the row came back after
a delete that reported success. The phone's half is `uniqueItems`, which
every list keyed on a server-chosen id goes through: a repeat there must
never be able to close the app, whatever produced it.
**Deleting a session offers to take the machine's own transcript with it**
`DELETE /sessions/{id}?deleteForeign=true`, behind a switch in the
confirmation, and only where the driver keeps a record of its own
(`keepsOwnTranscript`, currently Claude Code or Codex). Off by default,
because leaving that copy is what makes an ordinary delete recoverable — and
the dialog's paragraph is rewritten when it is on rather than appended to,
since the sentence promising the conversation "should still be there to
import again" is exactly the one the switch makes false. The server deletes
the machine's copy *first*, so a machine it cannot reach leaves the session
where it was instead of half-deleted.
## Measurements worth not re-taking
- **`-np 1` is what makes the MTP draft head pay.** Taken 2026-09-19 on the
27B above, decode speed for a 300-token reply, from `llama-server`'s own
timings rather than the clock:
| flags | tok/s |
| --- | --- |
| plain, any `-np` | 41.5 |
| `--spec-type draft-mtp -np 1` | 61.4 |
| `--spec-type draft-mtp -np 2` (n-max 2) | 65.9 |
| `--spec-type draft-mtp`, default `-np` (4 slots) | 28 |
Draft acceptance is 0.530.73 in every case, so the head is working in all
of them: what changes is that speculating against a KV cache split four ways
is slower than not speculating. A model's preset gets `parallel = 1` unless
its settings say otherwise (the machines tab's provider view, since
2026-09-19), so this is recorded for whoever next sees MTP look broken or
next raises the slot count to answer two sessions at once. `--spec-draft-n-max 2` was
worth another 7% in a single sample and is deliberately *not* passed — one
sample on a virtualised GPU is not a number to hardcode.
- **Prompt processing is the expensive part of a llama turn here, and decode
speed falls only slowly with context.** Taken 2026-09-19 on a free GPU, the
27B with `--spec-type draft-mtp -np 1`, generating 160 tokens each time:
| context | decode | prefill of that prompt |
| --- | --- | --- |
| 88 | 43.4 tok/s (cold) | 21s |
| 1,569 | 55.5 tok/s | (model still warming) |
| 6,068 | 53.2 tok/s | 9.5s |
| 14,068 | 50.3 tok/s | 22s |
So a turn on a long conversation spends tens of seconds before the first
token, and that is what `SessionStatus::Reading` exists to say. The same
sweep on the 0.6B **on the CPU** falls much harder -- 30.1 tok/s at 44
tokens of context to 11.5 at 6,024 -- which is the shape somebody means by
"it gets slower as the conversation goes on". The figure the app draws is
`timings.predicted_per_second`, decode only, so prefill is never mixed into
it.
- **A busy GPU is a model that will not load at all**, not a slow one:
`radv/amdgpu: Failed to allocate a buffer` and `failed to load model` while
something else holds VRAM. A 0.6B that had been decoding at 149 tok/s ran at
16.7 in that window before its server died, so a tok/s figure taken while
the card is shared says nothing about the model.
- **Asking for the head when the file has none is fatal**, not ignored:
`context type MTP requested but model doesn't contain MTP layers` and the
server exits. Without the flag the same file logs `unused tensor
blk.N.nextn.* — ignoring` and runs normally, which is the state to look for
when MTP is silently not happening.
- **What the transcript screen costs to scroll.** Taken 2026-08-30 on the GPU
emulator against a real imported transcript with the server at
`--delay 120`. Settled and flinging fast, both into fresh history and back
through rows already drawn: **5.25.9% janky frames, 99th percentile
2932ms, 02 slow UI-thread frames.** The stock Settings app on the same
device is 3.3% and 38ms, so this is at the platform floor. The number that
is *not* at the floor is the first few seconds after opening a session,
where every row on the way is being composed for the first time; that is
inherent to a lazy list and it is why a measurement taken before the screen
settles reads three times worse. **Settle first, then reset `gfxinfo`.**
- **The reset path is not reachable by reopening a session.** Measured
2026-09-04 against a session streaming at 20 events a second: reopening one
with an anchor 1,800 events back connects **87119 events behind**, well
under `CATCH_UP_LIMIT`'s 200, because the restore is two requests — the
opening page, then one span covering the whole distance. To exercise the
reset at all you have to lower `CATCH_UP_LIMIT` in a throwaway build; at 5
the app takes the reset on a live connection, clears, refills and carries
on without reconnecting.
- **The session screen's stream survives backgrounding here** — 20 seconds at
the launcher while 415 events were produced brought no reconnect at all,
which is not what the comment above that loop expects, and is most likely
this emulator being headless rather than the phone's behaviour.
- **Reopening a cached session costs one request for one event** (the probe),
and scrolling the whole conversation back costs nothing more; a cold open
of the same 500-event session is two pages, 100 events. Measured
2026-09-04 on the emulator against the sandbox.
- **Reading is cheap and editing is not.** The viewer handles a 1 MiB,
28,000-line file because it draws one row per line; the editor is one
`BasicTextField`, which costs two seconds a frame at 128 kB and stops the
app at 1 MiB, so `EDIT_LIMIT` caps it at 32 kB with the reason said on
screen. If you make the editor faster, that number is what to move.
EXPLORER.md's "What the measurements said" has the rest.