Draw a model's thinking, and what a reply cost to produce

A llama.cpp session's `reasoning_content` becomes `Event::Thinking` deltas
closed by an `Event::ThinkingDone` carrying the span the driver measured, and
the phone draws it as a card of its own: "Thinking" with the spinner a running
command has, then "Thought for 12.4s". Deliberately not a tool call, so a run
of calls cannot collapse the reasoning into "Called 6 tools"; the reasoning is
also kept out of the next prompt, which `conversation` already ignored.

`UsageDelta` gains `tokensPerSecond`, the provider's own figure or nothing --
llama.cpp reports `timings.predicted_per_second` and the coding CLIs report no
such thing -- and a finished reply carries a small line under it saying when it
was sent and, where there is one, how fast it came out: "3:00 PM · 149 tok/s".

The compact usage bar drops the provider's name for the window and puts its
length after the time left instead: "42% · 3h 20m left / 5h".

Three things that had to come with it: the transcript coalesces runs of
thinking deltas as it does reply deltas, so one block is one row of a page
rather than a page of its own; `joinPages` welds a block cut by a page boundary
(`healSplitThinking`), since the half with no ending spun for ever; and
`UsageDelta` now reaches the fold, which is what carries the rate to the reply.

Verified on the emulator against a real Qwen3-0.6B session and the echo rig's
new `/think [seconds]`: the spinner while it runs, "Thought for 1.4s" and
"2:54 PM · 149 tok/s" after, the reasoning on tapping the card, and the usage
bar reading "42% · 3h 19m left / 5h".

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
iris-aiandClaude Opus 5 committed 2026-09-19 15:07:25 -04:00
1 parent 45f249ae91
commit bb5ac1a242
18 files changed
+900 -101

No files matched your search

@@ -130,6 +130,24 @@ sealed class TranscriptUnit {
get() = "u$seq:$ordinal"
}
/**
* The line under a finished reply: when it was sent, and how fast it was generated.
*
* A unit of its own rather than something drawn inside the last block, because a settled reply
* *is* its blocks -- there is no row left to hang it on, and the last block is a piece of
* markdown that knows nothing about the message it came from.
*/
data class ReplyFoot(
override val seq: Long,
override val ordinal: Int,
val ts: Double,
val tokensPerSecond: Double?,
override val gap: Dp,
) : TranscriptUnit() {
override val key: Any
get() = "f$seq"
}
/** One memory note of a settled reply; see [MemoryNote]. */
data class Memory(
override val seq: Long,
@@ -237,6 +255,18 @@ fun transcriptUnits(
}
}
}
// Unconditional, because being in this branch is what says the reply is over:
// [splitWanted] is settled-or-overtaken. The case to keep out is a message still
// arriving, whose "sent at" is not yet the one it ends up with, and that is drawn
// whole.
units +=
TranscriptUnit.ReplyFoot(
row.startSeq,
ordinal,
item.ts,
item.tokensPerSecond,
gap(FOOT_SPACING),
)
} else {
units += TranscriptUnit.Whole(row, rowGap)
}
@@ -280,6 +310,14 @@ fun unwarmedReplies(rows: List<TranscriptRow>, replies: ParsedReplies): List<Tra
* length has lines that wrap, so its bubble is at the full width already and the slices match it
* exactly. Below it, one item of at most a few screens is nothing the list minds composing.
*/
/**
* The room between a reply's last block and the line under it.
*
* Tighter than the gap between blocks: the footer belongs to the message above it, and at a block's
* spacing it reads as a row of its own floating between two replies.
*/
private val FOOT_SPACING: Dp = 2.dp
const val USER_SPLIT_CHARS = 4000
/**
@@ -375,6 +413,7 @@ private val TranscriptUnit?.kind: String
is TranscriptUnit.PeerBlock -> "peer block"
is TranscriptUnit.UserChunk -> "user slice"
is TranscriptUnit.Memory -> "memory note"
is TranscriptUnit.ReplyFoot -> "reply footer"
is TranscriptUnit.Whole ->
when (val row = row) {
is TranscriptRow.Tools -> "tool group"