Say which half of the wait a llama turn is in
A turn has two waits in front of the first token and they were one word. `SessionStatus::Loading` was already the model coming off disk; this adds `SessionStatus::Reading` for llama-server processing the prompt -- emitted when the request goes out, cleared by the first thing the model says of any kind, so it covers every generate in a tool loop rather than only the first. Prefill is the expensive half on this machine: measured 9.5s for 6,068 tokens and 22s for 14,068 on the 27B with the GPU to itself. Reported as `running` that was indistinguishable from a model thinking, which is the thing the reader is waiting for. The phone draws both with the working spinner and its own words -- "loading model" and "reading prompt" -- and the session screen's status row now spins for all three busy states instead of only `running`, which is also how `loading` stops being a bare word with nothing moving. Measured while checking the tok/s figure, and recorded in the rigs skill: the 27B holds 55.5 to 50.3 tok/s between 1.5k and 14k of context, so decode decays gently, while the 0.6B on the CPU falls 30.1 to 11.5 over 6k. A shared GPU is a different failure -- the model does not load at all. Verified on the emulator against a real llama session: "loading model" while the server started, then "reading prompt" with the spinner through prompt processing, then the thinking card. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
1 parent
bb5ac1a242
commit
b660905098
8 files changed
+114
-6
No files matched your search
@@ -155,7 +155,14 @@ seq N", so there is no separate history path to drift from the live one.
|
||||
- `Answered { id, answer }` — so a question card resolves on every connected
|
||||
device, not just the one that answered.
|
||||
- `Status { state }` — idle / running / awaiting-input / compacting /
|
||||
**waiting** / exited / unknown. `waiting` (2026-09-06) is the session's own
|
||||
**loading** / **reading** / **waiting** / exited / unknown. `loading` is a
|
||||
process that is up and cannot be spoken to yet (a model coming off disk);
|
||||
`reading` (2026-09-19) is the model holding the prompt and not yet
|
||||
answering, which on a long conversation is tens of seconds -- measured at
|
||||
9.5s for 6,068 tokens and 22s for 14,068 on the 27B here. Reported as
|
||||
`running` that was indistinguishable from a model thinking, which is the
|
||||
thing the reader is actually waiting for. Both are working states: nothing
|
||||
settles and nothing is invited. `waiting` (2026-09-06) is the session's own
|
||||
turn being over while work it started is not: a backgrounded subagent, or a
|
||||
command left running. Its own state because `idle` and it differ in *kind* —
|
||||
`idle` means the session is waiting for a person, and this means it is
|
||||
|
||||
Reference in new issue
Block a user