Report what a reply spent reading its prompt, and pin the clock right
`UsageDelta` gains `prefillMs`, llama-server's own `timings.prompt_ms`, so the footer under a finished reply is "read 9.5s · 50.3 tok/s · 3:00 PM". Prefill is the half of a turn that was invisible and is often the larger: measured on the 0.6B here, 1m 4s for the first turn after a model loads against 22ms for the next, whose prompt the server still had cached. The clock moves to the end of the line. Everything in front of it is a provider's own measurement, so a session on another provider has fewer of them or none, and a reader who has learned where the time is should not have to find it again because the model changed. The costs grow leftwards into the space instead, and a test asserts every shape of the line ends with the same thing. Verified on the emulator against a real llama session: three replies reading "read 1m 4s · 193 tok/s · 3:54 PM", "read 25ms · 308 tok/s · 3:54 PM" and "read 22ms · 194 tok/s · 3:54 PM", with the clock in one column. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
1 parent
b660905098
commit
369b8f7e52
14 files changed
+135
-42
No files matched your search
@@ -67,9 +67,11 @@ Module-by-module intent is in PLAN.md's "Backend layout".
|
||||
the span the *driver* measured, and the phone draws a card that spins while
|
||||
the block is open and says "Thought for 12.4s" once it is not. The reasoning
|
||||
is deliberately not part of the next prompt (`conversation` ignores it), and
|
||||
`timings.predicted_per_second` off the same stream becomes `UsageDelta`'s
|
||||
`tokensPerSecond`, which is the "149 tok/s" under a finished reply — nothing
|
||||
else here measures one, so every other driver sends `None`.
|
||||
`timings.predicted_per_second` and `timings.prompt_ms` off the same stream
|
||||
become `UsageDelta`'s `tokensPerSecond` and `prefillMs`, which is the
|
||||
"read 9.5s · 50.3 tok/s · 3:00 PM" under a finished reply — nothing else here
|
||||
measures either, so every other driver sends `None`, and the clock is last so
|
||||
that it does not move when a provider reports fewer of them.
|
||||
**A turn's wait has two halves and says which** (2026-09-19):
|
||||
`SessionStatus::Loading` is the model coming off disk and
|
||||
`SessionStatus::Reading` is `llama-server` processing the prompt -- emitted
|
||||
|
||||
Reference in new issue
Block a user