A turn waits on three things that look nowhere at all: a permission question, a tool call with a minute to run, and the completion itself, which says nothing while the prompt is read -- tens of seconds on a long conversation. Setting a flag left the turn exactly where it was until whichever it was came back. The wait is now what ends, not the work. Each of those runs on a thread of its own and the interrupt answers the wait; the abandoned thread finishes into a channel nobody is reading. 43ms to end a turn in every state, measured against a real model -- including mid prompt-processing, which used to be a twenty-second wait. That makes cancellation a token per turn rather than a flag on the session: the abandoned thread wakes up some time later, and a flag the next turn had reset would let it write into a conversation it is no longer part of. The open thinking block moves to Shared for the same reason -- the thread that knows one is open is no longer the thread that ends the turn. Stop, meanwhile, did nothing at all to a llama session: it signals the session's recorded process and process::stop refuses a Shared one, which is the whole point of that record -- so the session sat at idle. A Shared record routes to the driver now, because what stopping means for a session that borrows the machine's process is the driver's to say. It ends the turn, says exited itself, and gives up its claim on the model; each live session claims the model it is on, and the model is unloaded when the last claim goes. A model another session is using stays where it is.
4.6 KiB
TODO
Working list from Iris, 2026-09-03. Remove an entry when it lands; annotate one in place when it turns out to need a decision.
App — transcript
-
Decide how running background tasks can be inspected. For now the session status shows only the provider-reported count; command details stay in their existing tool cards and must not become subagent cards.
-
Messages received from other agents are inconsistent — sometimes they appear, sometimes they don't. Needs a rig. Read the code rather than measured: a live Claude session only learns of a peer message from the
originobject on a turn'sresult(session/claude/translate.rs), which the CLI attaches to a turn the message started. So a message that arrives mid-turn, or a second one within one turn, has nowhere to be reported — while an imported session, which syncs from the CLI's own file, picks up every one of them. That would show exactly as "sometimes". Confirming it means driving a real stream-json session and sending it messages in both states.
Session settings
Asked for by Bryan on 2026-09-21, in one run while other work was in flight. The first two are one change; the rest can land separately.
-
A screen, not a modal. The settings dialog has outgrown one: it scrolls inside itself and covers the session it is about.
-
Two tabs on that screen, the way the main screen has three. One is the session's own settings; the other is the same provider screen reached from the machines tab (
ProviderScreen), for this session's provider -- two ways in, one screen, no second copy of the truth. -
Compact fields everywhere. The label goes above the box rather than floating inside it, and the padding around the value comes down. The value's own text size does not change: what is costing a row its height is the framing, not the text.
-
A system prompt per session. For llama.cpp it is a
systemmessage on each request, so it is one entry inDriverKind::paramsand no app change; whether the CLI drivers get one (--append-system-prompt) is a separate question. -
Stop unloads the model where nothing else is using it. Today
Driver::stopdeliberately leaves it in memory, because the server is the machine's and another session may be on the same model. The answer is a claim per live session on the router, and an unload when the last one goes -- not an unconditional unload. -
Autocompact belongs in session settings; empty disables it, which is the default. Iris chose "hand it to the driver" — only where a driver has auto-compaction of its own. That option was offered on a false premise and is not buildable yet. It named pi's
set_auto_compaction, but pi was never built as a driver here:session/llama.rstalks tollama-server's OpenAI-compatible endpoint directly, and itscompact()refuses outright. Claude Code's auto-compaction is the CLI's own and nothing in the stream-json control protocol this app uses configures it. So the setting would be stored, passed to a driver, refused by every one of them, and the field would never appear on any session. What is needed first is either a driver that can take it, or a different rule — the server watchingcontextTokensand running/compactitself is the one that would work today, for Claude sessions, and it is the option that was not chosen.
A session the server could not load
- A session whose transcript will not parse is skipped with nothing but a
log line, and from the phone it looks exactly like an idle unresponsive
one.
SessionManager::newcatches a failinglaunchand logs "couldn't relaunch session ", so the session has no pump and no driver: no status, no history, nothing sendable. That is what thetaskNoteincident (fd71d87) looked like from Bryan's phone, and why it needed a report from him rather than being visible in the app.Event::Unreadableremoves the cause that time, but not the class — an unreadableprocess.json, a provider edited away and an unreachable host all reach the same place. This is the "design the unknown state first" rule: a session the server could not load is not a session with nothing to say, and only the phone can show the difference. It needs a status the wire can carry for it — the failure with its reason, reported on the session itself — rather than the reader having to tell it apart from silence.