Commit Graph
21 Commits
Author SHA1 Message Date
iris-aiandClaude Opus 5 369b8f7e52 Report what a reply spent reading its prompt, and pin the clock right
`UsageDelta` gains `prefillMs`, llama-server's own `timings.prompt_ms`, so the
footer under a finished reply is "read 9.5s · 50.3 tok/s · 3:00 PM". Prefill is
the half of a turn that was invisible and is often the larger: measured on the
0.6B here, 1m 4s for the first turn after a model loads against 22ms for the
next, whose prompt the server still had cached.

The clock moves to the end of the line. Everything in front of it is a
provider's own measurement, so a session on another provider has fewer of them
or none, and a reader who has learned where the time is should not have to find
it again because the model changed. The costs grow leftwards into the space
instead, and a test asserts every shape of the line ends with the same thing.

Verified on the emulator against a real llama session: three replies reading
"read 1m 4s · 193 tok/s · 3:54 PM", "read 25ms · 308 tok/s · 3:54 PM" and
"read 22ms · 194 tok/s · 3:54 PM", with the clock in one column.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-19 15:57:56 -04:00
iris-aiandClaude Opus 5 bb5ac1a242 Draw a model's thinking, and what a reply cost to produce
A llama.cpp session's `reasoning_content` becomes `Event::Thinking` deltas
closed by an `Event::ThinkingDone` carrying the span the driver measured, and
the phone draws it as a card of its own: "Thinking" with the spinner a running
command has, then "Thought for 12.4s". Deliberately not a tool call, so a run
of calls cannot collapse the reasoning into "Called 6 tools"; the reasoning is
also kept out of the next prompt, which `conversation` already ignored.

`UsageDelta` gains `tokensPerSecond`, the provider's own figure or nothing --
llama.cpp reports `timings.predicted_per_second` and the coding CLIs report no
such thing -- and a finished reply carries a small line under it saying when it
was sent and, where there is one, how fast it came out: "3:00 PM · 149 tok/s".

The compact usage bar drops the provider's name for the window and puts its
length after the time left instead: "42% · 3h 20m left / 5h".

Three things that had to come with it: the transcript coalesces runs of
thinking deltas as it does reply deltas, so one block is one row of a page
rather than a page of its own; `joinPages` welds a block cut by a page boundary
(`healSplitThinking`), since the half with no ending spun for ever; and
`UsageDelta` now reaches the fold, which is what carries the rate to the reply.

Verified on the emulator against a real Qwen3-0.6B session and the echo rig's
new `/think [seconds]`: the spinner while it runs, "Thought for 1.4s" and
"2:54 PM · 149 tok/s" after, the reasoning on tapping the card, and the usage
bar reading "42% · 3h 19m left / 5h".

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-19 15:07:25 -04:00
iris-ai 827a30768c Count Codex background terminals 2026-09-16 01:07:40 -04:00
iris-ai a9cfea89e5 Show Codex background task counts 2026-09-15 23:26:37 -04:00
iris-ai 3d1b1e304d Render whole-file patches from change metadata 2026-09-15 16:18:18 -04:00
iris-ai 46831520e3 Recover Codex sessions with missing rollouts 2026-09-14 15:01:47 -04:00
iris fe25108c51 Count every Codex model request, not just the last
App-server sends `thread/tokenUsage/updated` once per *model request*, and a
Codex turn makes as many as it made tool calls. The translator held the last
one until `turn/completed`, so a turn's cost was reported as its final
request alone -- measured against the real rollout, 28,878 tokens for a turn
that spent 51,399 -- and the gap grows with how much work the turn did. The
context figure also stood still for the whole turn, which is exactly when it
is moving most.

Reported as each arrives instead: `tokens` now adds up to what the turn
spent, and the context figure climbs during the turn (28,921 -> 33,190 ->
35,978 on a two-file read here, matching Codex's own `last_token_usage`
exactly at every step).
2026-09-13 03:58:47 -04:00
iris 898e6b92d0 Clarify subagent coordination cards 2026-09-13 02:48:21 -04:00
iris cad0cbcfbe Keep subagent delivery out of assistant text 2026-09-13 01:05:35 -04:00
iris 83b113ef0f Support Codex subagent transcripts 2026-09-13 00:41:59 -04:00
iris 57e1cec09c Render Codex web searches as common tools 2026-09-11 02:31:21 -04:00
iris 3c19b5a9bb Fix Codex transcript convergence 2026-09-10 01:24:16 -04:00
iris f9c8f640ce Fix quoted Bash tool titles 2026-09-10 00:46:09 -04:00
iris cbdd8493ed Unwrap double-quoted Codex Bash commands 2026-09-09 22:17:14 -04:00
iris 26fe9895e7 Unwrap rendered Codex Bash commands 2026-09-09 22:11:52 -04:00
iris b00e89795e Parse Codex app-server patch payloads 2026-09-09 21:30:15 -04:00
iris 10ce1a216b Defer Codex patches until their diff arrives 2026-09-09 20:58:48 -04:00
iris b507656abd Normalize shell and patch tool cards 2026-09-09 20:30:52 -04:00
iris 4dc3e3d784 Fix Codex transcript streaming and images 2026-09-09 15:14:24 -04:00
iris 8c88a7e991 Use native Codex steering and transcript deletion 2026-09-09 12:19:11 -04:00
iris 6a0202b1b5 Add Codex JSON sessions and usage limits 2026-09-07 23:29:15 -04:00