Say which half of the wait a llama turn is in

A turn has two waits in front of the first token and they were one word.
`SessionStatus::Loading` was already the model coming off disk; this adds
`SessionStatus::Reading` for llama-server processing the prompt -- emitted when
the request goes out, cleared by the first thing the model says of any kind, so
it covers every generate in a tool loop rather than only the first.

Prefill is the expensive half on this machine: measured 9.5s for 6,068 tokens
and 22s for 14,068 on the 27B with the GPU to itself. Reported as `running`
that was indistinguishable from a model thinking, which is the thing the reader
is waiting for. The phone draws both with the working spinner and its own
words -- "loading model" and "reading prompt" -- and the session screen's
status row now spins for all three busy states instead of only `running`,
which is also how `loading` stops being a bare word with nothing moving.

Measured while checking the tok/s figure, and recorded in the rigs skill: the
27B holds 55.5 to 50.3 tok/s between 1.5k and 14k of context, so decode decays
gently, while the 0.6B on the CPU falls 30.1 to 11.5 over 6k. A shared GPU is a
different failure -- the model does not load at all.

Verified on the emulator against a real llama session: "loading model" while
the server started, then "reading prompt" with the spinner through prompt
processing, then the thinking card.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
iris-aiandClaude Opus 5 committed 2026-09-19 15:44:04 -04:00
1 parent bb5ac1a242
commit b660905098
8 files changed
+114 -6

No files matched your search

+8
View File
@@ -70,6 +70,14 @@ Module-by-module intent is in PLAN.md's "Backend layout".
`timings.predicted_per_second` off the same stream becomes `UsageDelta`'s
`tokensPerSecond`, which is the "149 tok/s" under a finished reply — nothing
else here measures one, so every other driver sends `None`.
**A turn's wait has two halves and says which** (2026-09-19):
`SessionStatus::Loading` is the model coming off disk and
`SessionStatus::Reading` is `llama-server` processing the prompt -- emitted
when the request goes out and cleared by the first thing the model says, of
any kind. Prefill is the expensive half here (~10s at 6k tokens, ~22s at
14k), and as `running` it looked exactly like thinking. The phone draws
both with the working spinner and its own words, "loading model" and
"reading prompt".
**Every one of those is a default rather than a constant** (2026-09-19):
`DriverKind::params` declares what a provider takes — key, label, shape,
and whether a change waits for a restart — and the phone renders whatever