Say which half of the wait a llama turn is in

A turn has two waits in front of the first token and they were one word.
`SessionStatus::Loading` was already the model coming off disk; this adds
`SessionStatus::Reading` for llama-server processing the prompt -- emitted when
the request goes out, cleared by the first thing the model says of any kind, so
it covers every generate in a tool loop rather than only the first.

Prefill is the expensive half on this machine: measured 9.5s for 6,068 tokens
and 22s for 14,068 on the 27B with the GPU to itself. Reported as `running`
that was indistinguishable from a model thinking, which is the thing the reader
is waiting for. The phone draws both with the working spinner and its own
words -- "loading model" and "reading prompt" -- and the session screen's
status row now spins for all three busy states instead of only `running`,
which is also how `loading` stops being a bare word with nothing moving.

Measured while checking the tok/s figure, and recorded in the rigs skill: the
27B holds 55.5 to 50.3 tok/s between 1.5k and 14k of context, so decode decays
gently, while the 0.6B on the CPU falls 30.1 to 11.5 over 6k. A shared GPU is a
different failure -- the model does not load at all.

Verified on the emulator against a real llama session: "loading model" while
the server started, then "reading prompt" with the spinner through prompt
processing, then the thinking card.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
iris-aiandClaude Opus 5 committed 2026-09-19 15:44:04 -04:00
1 parent bb5ac1a242
commit b660905098
8 files changed
+114 -6

No files matched your search

+14
View File
@@ -658,6 +658,20 @@ pub enum SessionStatus {
/// state -- a driver that reports it is responsible for holding what it
/// is sent until it can deliver it -- but this is what says so on screen.
Loading,
/// The model is reading what it was given, and has not begun answering.
///
/// Its own state for the same reason [`SessionStatus::Loading`] is, one
/// level down: prompt processing is work the machine does before a turn
/// produces anything, and on a long conversation it is the part of the
/// wait somebody is looking at. Reported as `Running` it was
/// indistinguishable from a model thinking, which is what the reader is
/// actually waiting to see.
///
/// It is still a turn in progress -- nothing settles, nothing is invited
/// -- and it is not a measurement of the prompt: it says the request is
/// out and the model has said nothing yet, which for a server serving one
/// slot is what reading the prompt looks like.
Reading,
/// The session's own turn is over, but work it started is still going:
/// a backgrounded subagent, or a command left running.
///
+35
View File
@@ -1588,6 +1588,12 @@ fn generate(
map.insert(key.clone(), value.clone());
}
// Prompt processing starts the moment this is sent and nothing comes back
// until it is done, so this is where the wait somebody is watching begins.
// Cleared by the first thing the model says, whatever kind it is.
shared.emit(Event::Status {
state: SessionStatus::Reading,
});
let mut response = ureq::post(format!("{endpoint}/v1/chat/completions"))
.config()
// A turn can be long: a slow model on a long prompt, and the whole
@@ -1627,6 +1633,9 @@ fn generate(
// timestamps. `None` between blocks -- a turn can think, speak, call a
// tool and think again.
let mut thinking: Option<std::time::Instant> = None;
// Whether the model has produced anything at all yet; until it has, this
// turn is still reading its prompt.
let mut speaking = false;
// What the server says its own generation ran at. Read from it rather than
// divided out of the wall time here, which would count the request, the
// prompt processing and this loop's own scheduling as generation.
@@ -1672,6 +1681,12 @@ fn generate(
let Some(delta) = chunk.pointer("/choices/0/delta") else {
continue;
};
if !speaking && says_something(delta) {
speaking = true;
shared.emit(Event::Status {
state: SessionStatus::Running,
});
}
if let Some(fragment) = delta.get("reasoning_content").and_then(Value::as_str)
&& !fragment.is_empty()
{
@@ -1717,6 +1732,26 @@ fn generate(
Ok(Reply { text, calls })
}
/// Whether a delta carries anything the model produced, of any kind.
///
/// The first one of these is the end of prompt processing. The chunk that only
/// opens the message -- `{"role": "assistant", "content": null}` -- is not one,
/// which is why this asks what is *in* the delta rather than that one arrived.
fn says_something(delta: &Value) -> bool {
let said = |key| {
delta
.get(key)
.and_then(Value::as_str)
.is_some_and(|text| !text.is_empty())
};
// Not `is_some`: a dialect that sends the key as an explicit null on every
// chunk would end prompt processing on the one that opens the message.
let calling = delta
.get("tool_calls")
.is_some_and(|calls| !calls.is_null());
said("content") || said("reasoning_content") || calling
}
/// How long one completion may take before the turn is abandoned.
const GENERATE_TIMEOUT: std::time::Duration = std::time::Duration::from_secs(1800);