Give llama.cpp sessions tools, web search and a model picker

A llama session was a chat box: no tools, a fixed model, no permission
mode, and a model name drawn as the path the file sits at. It now runs the
agent loop itself, which is what the pieces below all hang off.

Tools are `llama-server`'s own (`--tools all`), which that server both
publishes and runs -- `GET /tools` for the definitions, `POST /tools` to
call one. Web search is Exa's MCP server, reached from this backend rather
than from the machine serving the model: that is what llama.cpp's own web
UI does, and it puts the search on the machine with a route out instead of
the one with the GPU. `llama-server`'s `--mcp-servers-json` can only spawn
local commands, so using it would have meant a Node bridge on every
machine that serves a model.

Driving the loop is what makes the permission gate ours. Two modes,
`manual` and `bypassPermissions`, which is what the mechanism has: the web
UI asks before every call and remembers the tools you say "always" to. The
allowances fold back out of the transcript's own answers, so they survive
a restart and a model change without being stored anywhere else.

Also here, because tools made each of them matter:

- **Loading is a state.** A 12 GB model takes twenty seconds to reach
  memory and refuses everything until it has; the session used to report
  `running` for that whole time, and a message sent meanwhile came back as
  an error. It is `loading` now, and the message waits.
- **The model can be changed.** A `llama-server` holds one model, so this
  stops it and starts another. The conversation survives because it was
  never in the server.
- **Models are named, not pathed.** `general.name` read out of the file
  itself -- over ssh too, in the round trip the spawn was already making.
  Where two models share a name the file name breaks the tie.
- **`-np 1`, and the MTP draft head where the file has one.** Measured on
  the 27B here: 41.5 tok/s plain, 61.4 with `--spec-type draft-mtp` at one
  slot, and 28 with it at four -- speculating against a split KV cache is
  worse than not speculating. The flag is conditional because asking for a
  head that is not there makes `llama-server` exit.
- **A refusal says what to do.** Tool results are thousands of tokens, so
  an overrun context is now ordinary; it was "http status: 400" and is now
  the server's own "exceeds the available context size, try increasing it".

`GET /machines/{id}/models` is gone: the provider models route answers the
same question, and two answers to one question is how a picker comes to
offer a model the spawn screen does not.

Verified end to end against real models: a tool call asked and allowed, an
Exa search, a shell command, a 27B loaded while a message waited on it, a
model switch mid-session, a second message queued behind a running turn,
and the whole of it again on a session running over ssh.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
iris-aiandClaude Opus 5 committed 2026-09-19 08:11:49 -04:00
1 parent 392cc5413d
commit ac476ab0c9
23 files changed
+3609 -1078

No files matched your search

+60 -12
View File
@@ -143,18 +143,42 @@ moment you use it — `ANDROID_SERIAL=$(emu serial) ./gradlew …`.
### Testing llama.cpp and ssh here ### Testing llama.cpp and ssh here
**Both are set up here as of 2026-09-04** and need nothing typed. The **Both are set up here** and need nothing typed. The prebuilt llama.cpp lives
prebuilt CPU llama.cpp lives outside the repo at `~/.local/opt/llama.cpp` outside the repo at `~/.local/opt/llama.cpp-vk` — a **Vulkan** build as of
(the 15 MB `ubuntu-x64` release asset) and is symlinked as 2026-09-19, replacing the CPU one that was there before — and is symlinked as
`/usr/local/bin/llama-server`, which is what makes **discovery find it over both `~/.local/bin/llama-server` and `/usr/local/bin/llama-server`. The second
ssh**: `~/.local/bin` is not on the PATH a non-interactive ssh session gets. is what makes **discovery find it over ssh**: `~/.local/bin` is not on the
It resolves its own libraries through `$ORIGIN`, so no `LD_LIBRARY_PATH` is PATH a non-interactive ssh session gets. It resolves its own libraries through
needed. One model is downloaded — `unsloth/Qwen3-0.6B-GGUF/Qwen3-0.6B-Q8_0.gguf`, `$ORIGIN`, so no `LD_LIBRARY_PATH` is needed.
639 MB under `~/.local/share/ai-app/models` — and answers at usable speed on
this VM's 8 cores. **Do not test with a 2-bit quant**: the Two models are downloaded under `~/.local/share/ai-app/models`:
IQ2_XXS of that model produces fluent nonsense, which reads exactly like a
broken driver — `llama-cli` produces the same from the file directly, which - `unsloth/Qwen3-0.6B-GGUF/Qwen3-0.6B-Q8_0.gguf`, 639 MB, loads in ~4s. It
is how to tell the two apart in a hurry. calls tools correctly and is the right rig for the driver's shape. Do not
judge *answers* by it — asked for the second line of a file it read from
line 2 and then named the third.
- `ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF/Qwen3.8-27B-GSQ-RCO-IQ3_S-mtp.gguf`,
12 GB, ~20s to load, and the only one here with a multi-token-prediction
head. It is the rig for anything about `loading` being a state of its own,
since 20s is long enough to send into.
**Do not test with a 2-bit quant**: the IQ2_XXS of the 0.6B produces fluent
nonsense, which reads exactly like a broken driver — `llama-cli` produces the
same from the file directly, which is how to tell the two apart in a hurry.
**The GPU is shared and llama-server dies loudly when it runs out.** A second
server loading a model while the 27B holds VRAM fails with `radv/amdgpu:
Failed to allocate a buffer` / `MESA: error: buffer allocation failed` and
exits mid-request. `-ngl 0` runs it on the 8 cores instead, which is the way
to test the driver while something else holds the card.
**Testing tools and MCP without the app**: `llama-server --tools all` publishes
its built-in tools at `GET /tools` and runs one at `POST /tools` with
`{"tool": …, "params": …}` and an `x-tool-cwd` header — so a whole agent loop
is drivable with `curl` and no model at all. The Exa MCP server at
`https://mcp.exa.ai/mcp` answers **without an API key** and needs a
`User-Agent` header (Cloudflare answers 403 without one, which reads as a
refusal rather than a missing header).
There is no second machine, so **ssh this VM to itself**. That is set up There is no second machine, so **ssh this VM to itself**. That is set up
too: the key is `~/.config/ai-app/ssh-self` (its public half is in too: the key is `~/.config/ai-app/ssh-self` (its public half is in
@@ -212,6 +236,30 @@ where it was instead of half-deleted.
## Measurements worth not re-taking ## Measurements worth not re-taking
- **`-np 1` is what makes the MTP draft head pay.** Taken 2026-09-19 on the
27B above, decode speed for a 300-token reply, from `llama-server`'s own
timings rather than the clock:
| flags | tok/s |
| --- | --- |
| plain, any `-np` | 41.5 |
| `--spec-type draft-mtp -np 1` | 61.4 |
| `--spec-type draft-mtp -np 2` (n-max 2) | 65.9 |
| `--spec-type draft-mtp`, default `-np` (4 slots) | 28 |
Draft acceptance is 0.530.73 in every case, so the head is working in all
of them: what changes is that speculating against a KV cache split four ways
is slower than not speculating. The driver passes `-np 1` always, so this is
recorded for whoever next sees MTP look broken. `--spec-draft-n-max 2` was
worth another 7% in a single sample and is deliberately *not* passed — one
sample on a virtualised GPU is not a number to hardcode.
- **Asking for the head when the file has none is fatal**, not ignored:
`context type MTP requested but model doesn't contain MTP layers` and the
server exits. Without the flag the same file logs `unused tensor
blk.N.nextn.* — ignoring` and runs normally, which is the state to look for
when MTP is silently not happening.
- **What the transcript screen costs to scroll.** Taken 2026-08-30 on the GPU - **What the transcript screen costs to scroll.** Taken 2026-08-30 on the GPU
emulator against a real imported transcript with the server at emulator against a real imported transcript with the server at
`--delay 120`. Settled and flinging fast, both into fresh history and back `--delay 120`. Settled and flinging fast, both into fresh history and back
+24
View File
@@ -47,6 +47,21 @@ Module-by-module intent is in PLAN.md's "Backend layout".
readiness poll watches the process as well as the port, since a model that readiness poll watches the process as well as the port, since a model that
will not load exits in a second and was being reported as "gave up after will not load exits in a second and was being reported as "gave up after
300s". See PLAN.md's "Transport" and "llama-server management". 300s". See PLAN.md's "Transport" and "llama-server management".
**A llama session has tools and runs the loop itself** (2026-09-19):
`--tools all` gives it `llama-server`'s built-in set, which that server also
*runs* (`GET /tools` for the definitions, `POST /tools` to call one), while
web search comes from an MCP server this backend connects to directly
(`session/llama/mcp.rs`, Exa preset in a discovered provider's
`mcpServers`). Driving the loop is what makes the permission gate ours:
`manual` asks before every call and remembers a tool you answer
"Always allow …" to, `bypassPermissions` never asks, and the allowances are
folded back out of the transcript. Three more things fall out of it and are
easy to get wrong again — a model change **reloads the server** rather than
being refused, since the conversation lives in the transcript rather than in
`llama-server`; `-np 1` is always passed, and it is what decides whether the
MTP draft head is a 50% speed-up or a 33% loss; and `--spec-type draft-mtp`
is conditional on the file actually having a head, because asking for one
that is not there makes `llama-server` **exit**.
Codex is one persistent `codex app-server --stdio` process per session; its Codex is one persistent `codex app-server --stdio` process per session; its
driver uses native turn steering and interruption, persists the protocol driver uses native turn steering and interruption, persists the protocol
state and thread id, and reads subscription limits through the same CLI state and thread id, and reads subscription limits through the same CLI
@@ -310,6 +325,15 @@ written, and the fold uses that same predicate to decide a reply is settled.
## Things that have bitten ## Things that have bitten
- **A llama session reports `loading`, and a message sent into it waits.**
Before 2026-09-19 the session showed `running` from the moment the process
started, so a minute of reading a model off disk was indistinguishable from
a minute of thinking -- and anything sent in that window came back as an
error, because `llama-server` refuses everything until the model is in
memory. `SessionStatus::Loading` is the state and `Shared::await_ready` is
the waiting. A driver that reports `Loading` owes the holding as well as the
word.
- **A transcript outlives the enum.** Removing `Event::TaskNote` hours after - **A transcript outlives the enum.** Removing `Event::TaskNote` hours after
adding it made every transcript that had recorded one unreadable, so adding it made every transcript that had recorded one unreadable, so
`launch` failed for those sessions and `SessionManager::new` skipped them — `launch` failed for those sessions and `SessionManager::new` skipped them —
+53 -2
View File
@@ -345,6 +345,55 @@ deliberate and easy to undo by accident:
out the 300s timeout turned the server's own account of the problem into out the 300s timeout turned the server's own account of the problem into
"gave up". The failure carries the tail of `llama-server.log`, which on a "gave up". The failure carries the tail of `llama-server.log`, which on a
remote session is the only copy anybody reading the phone can see. remote session is the only copy anybody reading the phone can see.
- **Loading is a state of its own** (2026-09-19, `SessionStatus::Loading`).
A multi-gigabyte model takes tens of seconds to reach memory and refuses
everything until it has, and the session used to report `running` for that
whole time — indistinguishable from a model thinking, with the added
detail that any message sent meanwhile came back as an error. It is now
`loading` on both screens, and a message sent into a load **waits** for it
rather than failing. The waiting is the driver's (`Shared::await_ready`, a
condvar on a three-state `Serving`), because "there is a process and it is
not ready" is a fact only a driver can have. The third state matters as
much as the first two: a model that will never load has to answer a waiting
message with what went wrong rather than holding it for ever.
- **The driver runs the agent loop, and therefore owns the permission gate**
(2026-09-19). `llama-server --tools all` *hosts* the built-in tools —
`GET /tools` is their definitions, `POST /tools` runs one — but it does not
drive a conversation: a completion comes back with tool calls in it and
stops. So the loop is here, which is what puts "may I run this?" somewhere
a phone can answer it. Two modes, `manual` and `bypassPermissions`, which
is what the mechanism actually has: llama.cpp's own web UI asks before
every call and remembers the tools you said "always" to, and a third mode
between them would have to invent a rule about which tools count as edits.
The allowances are folded out of the transcript's `Answered` events, like
everything else this driver remembers, which is why the answer carries the
tool's name in it.
- **Tools run where the model does; MCP runs here** (2026-09-19). The
built-in tools are the far machine's, for the same reason the model file
is — they act on that machine's disk. An MCP server is reached from *this*
backend instead (`session/llama/mcp.rs`), which is both what llama.cpp's
own web UI does (it connects to `https://mcp.exa.ai/mcp` from the browser)
and the right side to be on: a web search wants the machine with a route
out, not the machine with the GPU. `llama-server`'s own `--mcp-servers-json`
is deliberately not used — it can only spawn local commands, so a remote
server would mean a Node bridge on whichever machine serves the model.
- **A model change reloads the server rather than being refused** (2026-09-19).
A `llama-server` holds one model, so switching stops it and starts another;
the conversation survives because the conversation was never in the server.
What is lost is the prompt cache, which is exactly what the phone already
warns about before a switch.
- **One slot, and the draft head where the file has one** (measured
2026-09-19). `-np 1` always: a session is one conversation making one
request at a time, so the other three slots `llama-server` picks on its own
are context this session could have had. It is also what decides whether
multi-token prediction pays — on the 27B here, **41.5 tok/s** plain at any
slot count, **61.4** with `--spec-type draft-mtp` at one slot, and **28**
with the head at four. Speculating against a split KV cache is worse than
not speculating, and it reads exactly like the head being broken.
The flag is conditional because it must be: asked for on a model without a
head, `llama-server` exits. `crate::gguf::has_mtp_head` reads the answer out
of the file — on the machine that will serve it, in the round trip the spawn
was already making — and `params["speculative"] = "off"` is the way out.
### Models (2026-08-28) ### Models (2026-08-28)
@@ -1386,8 +1435,10 @@ verified by running it, matching dev-updater's posture.
it truncates old KV cache entries, which is silent forgetting with no it truncates old KV cache entries, which is silent forgetting with no
summary, and it corrupts the harness's view of what the model knows. Fine summary, and it corrupts the harness's view of what the model knows. Fine
as a server-side safety net; not memory management. as a server-side safety net; not memory management.
- **Remote llama-server** needs its port forwarded (`ssh -L`) and is not - **MCP servers are configured in `config.ron`, not from the phone**
built; such a session is refused rather than misdirected. (2026-09-19). `mcpServers` on a llama provider, with Exa preset on a newly
discovered one. A phone screen for them is the obvious next step and was
deliberately left out of the change that added them.
- **Claude sessions over ssh need the remote machine logged in to Claude.** - **Claude sessions over ssh need the remote machine logged in to Claude.**
Usage reporting reads each machine's own credentials, and the Machines tab Usage reporting reads each machine's own credentials, and the Machines tab
can run that machine's CLI login without requiring an interactive SSH shell. can run that machine's CLI login without requiring an interactive SSH shell.
@@ -1334,7 +1334,18 @@ fun deleteSession(settings: ServerSettings, sessionId: String, deleteForeign: Bo
// Browsing is proxied by the server rather than done here, because this app trusts exactly one // Browsing is proxied by the server rather than done here, because this app trusts exactly one
// certificate and has no general internet trust to spend on huggingface.co. // certificate and has no general internet trust to spend on huggingface.co.
data class LocalModel(val key: String, val repo: String, val file: String, val bytes: Long) data class LocalModel(
val key: String,
val repo: String,
val file: String,
val bytes: Long,
/**
* What the file itself says it is called, or null when it does not say. Not what to draw: see
* the server's `models::labels`, which needs the whole list to decide -- two quantisations of
* one model share a name.
*/
val name: String?,
)
/** /**
* A download in flight or finished. [total] is null when the server never said how big the file is * A download in flight or finished. [total] is null when the server never said how big the file is
@@ -1371,36 +1382,28 @@ private fun parseDownload(o: JSONObject) =
) )
/** /**
* The models on one machine, which is the list a llama.cpp session there can choose from. * One model a picker can offer.
* *
* Not [fetchModels], which is what the *backend* has downloaded. A session serves its model from * Two fields because for one provider they differ: a llama.cpp session names its model by the path
* the machine it runs on, so for a machine reached over ssh those are two different lists -- and * it lives at and reads it as the name its own metadata gives it. Every other provider's [id] is
* offering the backend's would name files that are not there, turning a choice that cannot work * already what a person calls it, and the server says so by repeating it -- which is what keeps
* into a session that fails when it tries to load one. * every picker here free of a branch on the session kind.
*/ */
fun fetchMachineModels(settings: ServerSettings, machineId: String): List<LocalModel> = data class OfferedModel(val id: String, val label: String)
requestFromServer(settings, "/machines/${machineId.urlEncoded()}/models") { connection ->
JSONArray(connection.inputStream.bufferedReader().readText()).mapObjects { m ->
LocalModel(
key = m.getString("key"),
repo = m.getString("repo"),
file = m.getString("file"),
bytes = m.getLong("bytes"),
)
}
}
/** The current model catalog for one CLI provider on the machine where it runs. */ /** The current model catalog for one provider on the machine where it runs. */
fun fetchProviderModels( fun fetchProviderModels(
settings: ServerSettings, settings: ServerSettings,
machineId: String, machineId: String,
provider: String, provider: String,
): List<String> = ): List<OfferedModel> =
requestFromServer( requestFromServer(
settings, settings,
"/machines/${machineId.urlEncoded()}/providers/${provider.urlEncoded()}/models", "/machines/${machineId.urlEncoded()}/providers/${provider.urlEncoded()}/models",
) { connection -> ) { connection ->
JSONArray(connection.inputStream.bufferedReader().readText()).strings() JSONArray(connection.inputStream.bufferedReader().readText()).mapObjects { m ->
OfferedModel(id = m.getString("id"), label = m.getString("label"))
}
} }
fun fetchModels(settings: ServerSettings): Models = fun fetchModels(settings: ServerSettings): Models =
@@ -1414,6 +1417,7 @@ fun fetchModels(settings: ServerSettings): Models =
repo = m.getString("repo"), repo = m.getString("repo"),
file = m.getString("file"), file = m.getString("file"),
bytes = m.getLong("bytes"), bytes = m.getLong("bytes"),
name = m.optString("name").ifEmpty { null },
) )
}, },
downloads = body.getJSONArray("downloads").mapObjects(::parseDownload), downloads = body.getJSONArray("downloads").mapObjects(::parseDownload),
@@ -325,7 +325,8 @@ fun parseSeqEvent(json: String): SeqEvent {
* first time the server grows a state, and the drift would be a reply that never splits or one * first time the server grows a state, and the drift would be a reply that never splits or one
* split mid-stream. * split mid-stream.
*/ */
fun sessionWorking(state: String): Boolean = state == "running" || state == "compacting" fun sessionWorking(state: String): Boolean =
state == "running" || state == "compacting" || state == "loading"
/** Whether the latest events still say this session needs an explicit provider login. */ /** Whether the latest events still say this session needs an explicit provider login. */
internal fun authenticationPromptAfter(open: Boolean, event: SessionEvent): Boolean = internal fun authenticationPromptAfter(open: Boolean, event: SessionEvent): Boolean =
@@ -21,12 +21,25 @@ const val DEFAULT_MODEL = "default"
* one model rather than one model from another. Anything that does not look like that is returned * one model rather than one model from another. Anything that does not look like that is returned
* untouched. * untouched.
* *
* A llama.cpp session's model is not an identifier at all -- it is `owner/repo/file.gguf`, where
* the file was downloaded from -- so what is kept is the file, which is the part that tells two
* models apart, and the extension goes with the directories. The model's *own* name is better still
* and is not derivable here: it is inside the file, and only the server has ever opened it. Where a
* screen has the server's answer it should prefer it; this is the floor under every screen that
* does not.
*
* A display decision, not a correction: the full name is what the session reports. * A display decision, not a correction: the full name is what the session reports.
*/ */
fun modelLabel(model: String?): String { fun modelLabel(model: String?): String {
val name = model?.takeIf { it.isNotBlank() } ?: return DEFAULT_MODEL val name = model?.takeIf { it.isNotBlank() } ?: return DEFAULT_MODEL
if (name.endsWith(GGUF)) {
return name.substringAfterLast('/').removeSuffix(GGUF)
}
return name.removePrefix("claude-").replace(DATED_SUFFIX, "") return name.removePrefix("claude-").replace(DATED_SUFFIX, "")
} }
/** A trailing `-YYYYMMDD`, which is how these identifiers carry their release date. */ /** A trailing `-YYYYMMDD`, which is how these identifiers carry their release date. */
private val DATED_SUFFIX = Regex("""-\d{8}$""") private val DATED_SUFFIX = Regex("""-\d{8}$""")
/** What every model a llama.cpp session can run is stored as. */
private const val GGUF = ".gguf"
@@ -302,7 +302,7 @@ fun SessionScreen(
var permissionMode by remember { mutableStateOf(summary.permissionMode ?: "auto") } var permissionMode by remember { mutableStateOf(summary.permissionMode ?: "auto") }
// The models this provider actually offers, asked of the server rather than listed here: a // The models this provider actually offers, asked of the server rather than listed here: a
// hardcoded list is a claim about a machine. // hardcoded list is a claim about a machine.
var offeredModels by remember { mutableStateOf<List<String>>(emptyList()) } var offeredModels by remember { mutableStateOf<List<OfferedModel>>(emptyList()) }
var offeredPermissionModes by remember { mutableStateOf<List<String>>(emptyList()) } var offeredPermissionModes by remember { mutableStateOf<List<String>>(emptyList()) }
val lifecycleOwner = LocalLifecycleOwner.current val lifecycleOwner = LocalLifecycleOwner.current
// The resume cursor, written from the stream's IO thread. // The resume cursor, written from the stream's IO thread.
@@ -1194,6 +1194,17 @@ fun SessionScreen(
} }
} }
/**
* What to call a model on screen.
*
* The provider's own answer where it has one, because only the server can have it: a llama
* model is identified by the path it lives at and named by what is written inside the file, and
* the phone has never opened that file. [modelLabel] is the fallback and the right one for the
* rest -- a coding CLI's identifier already is its name.
*/
fun label(id: String?): String =
offeredModels.firstOrNull { it.id == id }?.label ?: modelLabel(id)
// Only for the model picker, which a subagent does not have. // Only for the model picker, which a subagent does not have.
if (!isSubagent) { if (!isSubagent) {
LaunchedEffect(summary.machine, summary.provider) { LaunchedEffect(summary.machine, summary.provider) {
@@ -1889,8 +1900,8 @@ fun SessionScreen(
) { ) {
pendingModel?.let { chosen -> pendingModel?.let { chosen ->
ModelSwitchWarning( ModelSwitchWarning(
from = modelLabel(model), from = label(model),
to = modelLabel(chosen), to = label(chosen),
onDismiss = { pendingModel = null }, onDismiss = { pendingModel = null },
onConfirm = { onConfirm = {
pendingModel = null pendingModel = null
@@ -2013,22 +2024,28 @@ fun SessionScreen(
) { ) {
if (offeredModels.isNotEmpty()) { if (offeredModels.isNotEmpty()) {
PickerButton( PickerButton(
current = modelLabel(model), current = label(model),
// What the machine offers, plus the state a session is in when // What the machine offers, plus the state a session is in when
// it has chosen none of them. The button has always been able // it has chosen none of them. The button has always been able
// to // to
// say "default"; until this the list could not, so leaving it // say "default"; until this the list could not, so leaving it
// was a one-way trip. // was a one-way trip.
options = listOf(DEFAULT_MODEL) + offeredModels, options =
listOf(DEFAULT_MODEL) + offeredModels.map { it.label },
// Not set here. The button follows what the session reports it // Not set here. The button follows what the session reports it
// is set to, which arrives a moment later and is sometimes a // is set to, which arrives a moment later and is sometimes a
// different answer -- a name the CLI resolved, or no change at // different answer -- a name the CLI resolved, or no change at
// all on a provider whose model is fixed. Asked about first, // all on a provider whose model is fixed. Asked about first,
// unless there is nothing to lose by it -- see // unless there is nothing to lose by it -- see
// [ModelSwitchWarning]. // [ModelSwitchWarning].
onPick = { chosen -> onPick = { picked ->
// Back to the id, because that is what the server resolves
// and it is not always the word on the chip.
val chosen =
offeredModels.firstOrNull { it.label == picked }?.id
?: picked
if ( if (
modelLabel(chosen) == modelLabel(model) || label(chosen) == label(model) ||
!worthWarningAbout(status, contextTokens, items) !worthWarningAbout(status, contextTokens, items)
) { ) {
act { setSessionModel(settings, summary.id, chosen) } act { setSessionModel(settings, summary.id, chosen) }
@@ -22,6 +22,10 @@ fun sessionStatusWord(status: String, subagent: Boolean = false): String =
"idle" -> "idle" "idle" -> "idle"
"running" -> "running" "running" -> "running"
"compacting" -> "compacting" "compacting" -> "compacting"
// Not "running": a model coming off disk is not a model answering, and the difference is
// minutes. Said in its own word so a first message that waits is explained rather than
// looking like a session that has stopped responding. See `SessionStatus::Loading`.
"loading" -> "loading"
// Its own word, because the state it is easily mistaken for means the opposite: "idle" // Its own word, because the state it is easily mistaken for means the opposite: "idle"
// invites the reader to type something, and a waiting session is going to carry on without // invites the reader to type something, and a waiting session is going to carry on without
// them. See `SessionStatus::Waiting`. // them. See `SessionStatus::Waiting`.
@@ -53,6 +57,9 @@ fun sessionStatusColour(status: String): Color =
"awaitingInput" -> awaitingColor "awaitingInput" -> awaitingColor
"running" -> runningColor "running" -> runningColor
"compacting" -> commandColor "compacting" -> commandColor
// The same accent as the other states that are busy on their own account, because that is
// what this is: something is happening and nothing is wanted from the reader.
"loading" -> commandColor
"waiting" -> waitingColor "waiting" -> waitingColor
else -> MaterialTheme.colorScheme.onSurfaceVariant else -> MaterialTheme.colorScheme.onSurfaceVariant
} }
@@ -58,7 +58,7 @@ fun SpawnScreen(
var providerName by remember { mutableStateOf<String?>(null) } var providerName by remember { mutableStateOf<String?>(null) }
var title by remember { mutableStateOf("") } var title by remember { mutableStateOf("") }
var model by remember { mutableStateOf("") } var model by remember { mutableStateOf("") }
var providerModels by remember { mutableStateOf<List<String>>(emptyList()) } var providerModels by remember { mutableStateOf<List<OfferedModel>>(emptyList()) }
var providerModelsLoading by remember { mutableStateOf(false) } var providerModelsLoading by remember { mutableStateOf(false) }
var providerModelsError by remember { mutableStateOf<String?>(null) } var providerModelsError by remember { mutableStateOf<String?>(null) }
var cwd by remember { mutableStateOf("") } var cwd by remember { mutableStateOf("") }
@@ -73,12 +73,13 @@ fun SpawnScreen(
// Only the spawn's own failure. The fetch's lives in `options`: this one leaves a filled-in // Only the spawn's own failure. The fetch's lives in `options`: this one leaves a filled-in
// form worth keeping, and that one leaves nothing to fill in. // form worth keeping, and that one leaves nothing to fill in.
var spawnError by remember { mutableStateOf<String?>(null) } var spawnError by remember { mutableStateOf<String?>(null) }
// GGUFs on the chosen machine, for a llama provider to choose between. Kept separate from a
// coding CLI's provider catalog and refetched when the machine changes.
var models by remember { mutableStateOf<List<LocalModel>>(emptyList()) }
var modelKey by remember { mutableStateOf<String?>(null) }
var contextSize by remember { mutableStateOf("") } var contextSize by remember { mutableStateOf("") }
var temperature by remember { mutableStateOf("") } var temperature by remember { mutableStateOf("") }
// Whether a model that carries a multi-token-prediction head drafts with it. Left to the
// server by default, which turns it on exactly where the file has one -- see `SPECULATIVE` in
// the llama driver. Here so a machine where drafting does not pay has a way out that is not
// an edit to config.ron.
var speculative by remember { mutableStateOf(SPECULATIVE_AUTO) }
LaunchedEffect(Unit) { LaunchedEffect(Unit) {
// Separate from the machines fetch below and deliberately not fatal: failing to learn the // Separate from the machines fetch below and deliberately not fatal: failing to learn the
@@ -126,17 +127,6 @@ fun SpawnScreen(
is LoadState.Loaded -> state.value is LoadState.Loaded -> state.value
} }
val machine = machines.firstOrNull { it.name == machineName } val machine = machines.firstOrNull { it.name == machineName }
// Whichever machine is chosen now, asked again when that changes. The old machine's list
// is dropped first rather than left on screen: a file name from another machine looks
// exactly like one from this one.
LaunchedEffect(machine?.id) {
models = emptyList()
modelKey = null
val id = machine?.id ?: return@LaunchedEffect
models =
runCatching { withContext(Dispatchers.IO) { fetchMachineModels(settings, id) } }
.getOrDefault(emptyList())
}
val current = machine?.providers?.firstOrNull { it.name == providerName } val current = machine?.providers?.firstOrNull { it.name == providerName }
// Coding CLIs take a working directory, model, permission mode and thinking level. Keying // Coding CLIs take a working directory, model, permission mode and thinking level. Keying
// the extra fields on the kind rather than the provider name keeps a second installation // the extra fields on the kind rather than the provider name keeps a second installation
@@ -145,13 +135,26 @@ fun SpawnScreen(
val isCodex = current?.kind == "codex_cli" val isCodex = current?.kind == "codex_cli"
val isCodingCli = isClaude || isCodex val isCodingCli = isClaude || isCodex
val isLlama = current?.kind == "llama_cpp" val isLlama = current?.kind == "llama_cpp"
// Echo is the only kind with nothing to choose between.
val offersModels = isCodingCli || isLlama
// Where a session's tools act, which is the only thing a working directory decides.
val takesCwd = isCodingCli || isLlama
// Whichever machine and provider are chosen now, asked again when either changes. The
// previous answer is dropped first rather than left on screen: a model name from another
// machine looks exactly like one from this one.
LaunchedEffect(machine?.id, current?.name) { LaunchedEffect(machine?.id, current?.name) {
model = "" model = ""
providerModels = emptyList() providerModels = emptyList()
providerModelsError = null providerModelsError = null
permissionMode = current?.defaultPermissionMode.orEmpty() permissionMode = current?.defaultPermissionMode.orEmpty()
if (isCodingCli) { // Every kind that offers models at all, not only the coding CLIs: a llama provider
// answers with the GGUFs on the machine it runs on, through the same call. One
// question with one answer is what keeps the picker free of a branch on the kind.
if (machine == null || current == null || !offersModels) {
providerModelsLoading = false
return@LaunchedEffect
}
providerModelsLoading = true providerModelsLoading = true
try { try {
providerModels = providerModels =
@@ -163,9 +166,6 @@ fun SpawnScreen(
} finally { } finally {
providerModelsLoading = false providerModelsLoading = false
} }
} else {
providerModelsLoading = false
}
} }
// The machine first, because it decides what can be run at all. // The machine first, because it decides what can be run at all.
@@ -220,29 +220,70 @@ fun SpawnScreen(
modifier = Modifier.fillMaxWidth(), modifier = Modifier.fillMaxWidth(),
) )
if (isLlama) { if (offersModels) {
// A llama session names one of the models on the machine it will run on, so the when {
// choice is that list rather than free text -- a name that is not on that machine's providerModelsLoading ->
// disk is a session that cannot start.
if (models.isEmpty()) {
Text( Text(
"No models on ${machine.name}. The Models screen downloads " + "Loading model choices…",
style = MaterialTheme.typography.bodySmall,
color = MaterialTheme.colorScheme.onSurfaceVariant,
)
providerModelsError != null ->
Text(
"Model choices unavailable: $providerModelsError",
style = MaterialTheme.typography.bodySmall,
color = MaterialTheme.colorScheme.error,
)
// A llama session cannot start without one, so this says what to do about it
// rather than only that there is nothing -- the models it needs are on the
// machine that will serve them, which is not always this backend.
providerModels.isEmpty() && isLlama ->
Text(
"No models on ${machine?.name}. The Models screen downloads " +
"to the backend; another machine needs the file put there itself.", "to the backend; another machine needs the file put there itself.",
style = MaterialTheme.typography.bodyMedium, style = MaterialTheme.typography.bodyMedium,
color = MaterialTheme.colorScheme.onSurfaceVariant, color = MaterialTheme.colorScheme.onSurfaceVariant,
) )
} else { providerModels.isEmpty() ->
Text(
"This machine reported no selectable models.",
style = MaterialTheme.typography.bodySmall,
color = MaterialTheme.colorScheme.onSurfaceVariant,
)
else -> {
Spacer(Modifier.height(16.dp))
ChipGroup( ChipGroup(
label = "Model", label = "Model",
// The file, not the whole key: the repository is the same for every // The label, and the id is what is sent: for a llama model those differ,
// quantisation of a model, so the file name is what tells two of them apart. // since it is chosen by path and named by what is inside the file.
options = models.map { it.file }, options = providerModels.map { it.label },
selected = models.firstOrNull { it.key == modelKey }?.file, selected = providerModels.firstOrNull { it.id == model }?.label,
onSelect = { file -> modelKey = models.first { it.file == file }.key }, onSelect = { chosen ->
val id = providerModels.first { it.label == chosen }.id
// A llama session has to have one, so choosing the same chip twice
// must not clear it -- there is nothing to fall back to.
model = if (model == id && !isLlama) "" else id
},
) )
} }
}
Spacer(Modifier.height(16.dp)) Spacer(Modifier.height(16.dp))
}
if (isCodingCli) {
// Free text as well as the chips above: the catalog is a shortcut, and a CLI will
// take a name it did not list.
OutlinedTextField(
value = model,
onValueChange = { model = it },
label = { Text("Model (blank = the CLI's default)") },
singleLine = true,
modifier = Modifier.fillMaxWidth(),
)
Spacer(Modifier.height(16.dp))
}
if (isLlama) {
OutlinedTextField( OutlinedTextField(
value = contextSize, value = contextSize,
onValueChange = { contextSize = it }, onValueChange = { contextSize = it },
@@ -260,48 +301,21 @@ fun SpawnScreen(
modifier = Modifier.fillMaxWidth(), modifier = Modifier.fillMaxWidth(),
) )
Spacer(Modifier.height(16.dp)) Spacer(Modifier.height(16.dp))
}
if (isCodingCli) { // Said as what it is rather than as "MTP": the reader is choosing whether the session
when { // goes faster, and most models have nothing to turn on here at all.
providerModelsLoading ->
Text(
"Loading model choices…",
style = MaterialTheme.typography.bodySmall,
color = MaterialTheme.colorScheme.onSurfaceVariant,
)
providerModelsError != null ->
Text(
"Model choices unavailable: $providerModelsError",
style = MaterialTheme.typography.bodySmall,
color = MaterialTheme.colorScheme.error,
)
providerModels.isEmpty() ->
Text(
"This machine reported no selectable models.",
style = MaterialTheme.typography.bodySmall,
color = MaterialTheme.colorScheme.onSurfaceVariant,
)
else -> {
Spacer(Modifier.height(16.dp))
ChipGroup( ChipGroup(
label = "Model", label = "Speculative decoding (models that carry a draft head)",
options = providerModels, options = listOf(SPECULATIVE_AUTO, SPECULATIVE_OFF),
selected = model.ifEmpty { null }, selected = speculative,
onSelect = { chosen -> model = if (model == chosen) "" else chosen }, onSelect = { speculative = it },
)
}
}
Spacer(Modifier.height(8.dp))
OutlinedTextField(
value = model,
onValueChange = { model = it },
label = { Text("Model (blank = the CLI's default)") },
singleLine = true,
modifier = Modifier.fillMaxWidth(),
) )
Spacer(Modifier.height(16.dp)) Spacer(Modifier.height(16.dp))
}
// Every session whose tools act on files needs one, which is both kinds that have
// tools -- a llama session's built-in tools run in it exactly as a CLI's do.
if (takesCwd) {
OutlinedTextField( OutlinedTextField(
value = cwd, value = cwd,
onValueChange = { cwd = it }, onValueChange = { cwd = it },
@@ -311,7 +325,11 @@ fun SpawnScreen(
modifier = Modifier.fillMaxWidth(), modifier = Modifier.fillMaxWidth(),
) )
Spacer(Modifier.height(16.dp)) Spacer(Modifier.height(16.dp))
}
// Offered wherever the provider has modes, rather than where this screen believes it
// does: the server is what knows, and llama.cpp grew them without this line changing.
if (current != null && current.permissionModes.isNotEmpty()) {
ChipGroup( ChipGroup(
label = "Permissions", label = "Permissions",
options = current.permissionModes, options = current.permissionModes,
@@ -319,7 +337,9 @@ fun SpawnScreen(
onSelect = { permissionMode = it }, onSelect = { permissionMode = it },
) )
Spacer(Modifier.height(16.dp)) Spacer(Modifier.height(16.dp))
}
if (isCodingCli) {
// Says what it does to *later* spawns as well, because it does: the level chosen here // Says what it does to *later* spawns as well, because it does: the level chosen here
// is stored as the default, which is the whole way that default is set. A picker that // is stored as the default, which is the whole way that default is set. A picker that
// quietly changed a global would be the same control with the fact left out. // quietly changed a global would be the same control with the fact left out.
@@ -363,11 +383,9 @@ fun SpawnScreen(
machine = machine.id, machine = machine.id,
provider = chosen.name, provider = chosen.name,
title = title.trim(), title = title.trim(),
model = model = model.trim().takeIf { offersModels },
if (isLlama) modelKey cwd = cwd.trim().takeIf { takesCwd },
else model.trim().takeIf { isCodingCli }, permissionMode = permissionMode.takeIf { it.isNotEmpty() },
cwd = cwd.trim().takeIf { isCodingCli },
permissionMode = permissionMode.takeIf { isCodingCli },
effort = effort.takeIf { isCodingCli }, effort = effort.takeIf { isCodingCli },
// Sent only when set, so blank means "whatever llama.cpp does // Sent only when set, so blank means "whatever llama.cpp does
// by default" rather than a zero. // by default" rather than a zero.
@@ -382,6 +400,13 @@ fun SpawnScreen(
.trim() .trim()
.takeIf { it.isNotEmpty() } .takeIf { it.isNotEmpty() }
?.let { put("temperature", it) } ?.let { put("temperature", it) }
// Only the choice that changes anything: "auto"
// is the absence of the setting, not a value of
// it, so a session spawned without an opinion
// carries none.
if (speculative == SPECULATIVE_OFF) {
put("speculative", "off")
}
} }
}, },
) )
@@ -393,13 +418,23 @@ fun SpawnScreen(
} }
} }
}, },
enabled = !busy && current != null && !(isLlama && modelKey == null), // A llama session names the file to load, so there is nothing to spawn without one.
enabled = !busy && current != null && !(isLlama && model.isEmpty()),
) { ) {
Text(if (busy) "Spawning..." else "Spawn") Text(if (busy) "Spawning..." else "Spawn")
} }
} }
} }
/**
* Leave the draft head to the server, which uses one wherever the model file has one. Spelled the
* same as the absence of the `speculative` parameter, because that is what it means.
*/
private const val SPECULATIVE_AUTO = "auto"
/** The `speculative` parameter's only other value; see the llama driver's `SPECULATIVE`. */
private const val SPECULATIVE_OFF = "off"
/** /**
* A labeled row of choices that wraps onto as many lines as it needs. * A labeled row of choices that wraps onto as many lines as it needs.
* *
+37 -2
View File
@@ -92,6 +92,32 @@ pub struct ProviderConfig {
/// too; this is a shortcut list, not a restriction. /// too; this is a shortcut list, not a restriction.
#[serde(default, skip_serializing_if = "Vec::is_empty")] #[serde(default, skip_serializing_if = "Vec::is_empty")]
pub models: Vec<String>, pub models: Vec<String>,
/// MCP servers whose tools this provider's sessions can use, on top of
/// whatever the provider runs itself.
///
/// On the provider rather than the machine, because it is a statement
/// about what a session can do rather than about where it runs -- and
/// because only a driver that runs its own agent loop can use one. Today
/// that is llama.cpp; the coding CLIs have their own MCP configuration
/// and this would be a second, quieter answer to the same question.
#[serde(default, skip_serializing_if = "Vec::is_empty")]
pub mcp_servers: Vec<McpServerConfig>,
}
/// An MCP server reached over HTTP.
///
/// A URL and nothing else: this backend connects to remote servers rather than
/// spawning local ones, so there is no command, no arguments and no
/// environment to configure. See `session::llama::mcp` for why that is the
/// shape -- in short, it is what llama.cpp's own web UI does, and it keeps the
/// tools on the machine with a route out rather than the machine with the GPU.
#[derive(Debug, Clone, Serialize, Deserialize)]
#[serde(rename_all = "camelCase")]
pub struct McpServerConfig {
/// Prefixes every tool this server offers, so two servers with a `search`
/// are two tools. Also what a failure to connect is named by.
pub name: String,
pub url: String,
} }
impl ProviderConfig { impl ProviderConfig {
@@ -278,7 +304,11 @@ impl DriverKind {
match self { match self {
Self::ClaudeCli => &["manual", "acceptEdits", "auto", "bypassPermissions", "plan"], Self::ClaudeCli => &["manual", "acceptEdits", "auto", "bypassPermissions", "plan"],
Self::CodexCli => &["workspace-write", "read-only", "danger-full-access"], Self::CodexCli => &["workspace-write", "read-only", "danger-full-access"],
Self::Echo | Self::LlamaCpp => &[], // Named by the driver that enforces them rather than repeated
// here: this list and the one the gate matches on being two
// literals is how a mode comes to be offered and then refused.
Self::LlamaCpp => crate::session::llama::PERMISSION_MODES,
Self::Echo => &[],
} }
} }
@@ -287,7 +317,8 @@ impl DriverKind {
match self { match self {
Self::ClaudeCli => Some("auto"), Self::ClaudeCli => Some("auto"),
Self::CodexCli => Some("workspace-write"), Self::CodexCli => Some("workspace-write"),
Self::Echo | Self::LlamaCpp => None, Self::LlamaCpp => Some(crate::session::llama::DEFAULT_PERMISSION_MODE),
Self::Echo => None,
} }
} }
} }
@@ -499,6 +530,7 @@ impl Config {
kind: DriverKind::Echo, kind: DriverKind::Echo,
command: None, command: None,
models: Vec::new(), models: Vec::new(),
mcp_servers: Vec::new(),
} }
} }
@@ -579,6 +611,7 @@ mod tests {
kind: DriverKind::ClaudeCli, kind: DriverKind::ClaudeCli,
command: Some("/usr/bin/claude".to_string()), command: Some("/usr/bin/claude".to_string()),
models: Vec::new(), models: Vec::new(),
mcp_servers: Vec::new(),
}, },
]), ]),
MachineConfig { MachineConfig {
@@ -597,6 +630,7 @@ mod tests {
kind: DriverKind::ClaudeCli, kind: DriverKind::ClaudeCli,
command: None, command: None,
models: vec!["haiku".to_string()], models: vec!["haiku".to_string()],
mcp_servers: Vec::new(),
}], }],
}, },
], ],
@@ -734,6 +768,7 @@ sessions: [(
kind: DriverKind::ClaudeCli, kind: DriverKind::ClaudeCli,
command: Some("/usr/bin/claude".to_string()), command: Some("/usr/bin/claude".to_string()),
models: Vec::new(), models: Vec::new(),
mcp_servers: Vec::new(),
}, },
]); ]);
assert_eq!( assert_eq!(
+359
View File
@@ -0,0 +1,359 @@
//! Just enough of the GGUF container to read a model's own name out of it.
//!
//! A `.gguf` file opens with a key/value table, and `general.name` in it is
//! what the people who published the model called it -- "Qwen3-0.6B",
//! "Qwen3.8-27B GSQ-RCO". Everything else this server knows a model by is
//! filesystem trivia: `owner/repo/file.gguf` is where it was downloaded from,
//! which is an address rather than a name, and on a phone it is a line of path
//! where a word would do.
//!
//! **Read as far as the answer and no further.** The same table holds the
//! tokenizer, which for a modern model is a 150,000-entry string array and
//! most of several megabytes; `general.*` is written first by every converter
//! in practice, so stopping at the name costs a few kilobytes instead. That is
//! what makes this affordable to run over every model in a directory, and what
//! lets the remote case work from a bounded prefix of the file rather than the
//! whole of it.
//!
//! Anything unreadable is [`None`] rather than an error, at every level. A
//! model with no name, a truncated prefix, a container version this does not
//! know and a file that is not GGUF at all are one answer here -- "this file
//! does not tell us" -- and the caller has a file name to fall back on. There
//! is nothing a reader could do with the distinction.
use std::io::Read;
/// How many bytes of a model file are worth fetching to look for its name.
///
/// Only the remote path needs a number: a local read stops when it finds the
/// key, but a file on another machine has to be asked for a fixed amount
/// before anything can be parsed. Measured 2026-09-19 against the two models
/// on this machine, `general.name` ends at byte **130** and **94** -- every
/// converter writes `general.*` before the tokenizer arrays that make up the
/// rest of the table. 8 KiB is two orders of magnitude of slack for that and
/// still makes listing a directory of models one round trip's worth of bytes
/// rather than a download, which is what decides the number: this is paid per
/// model every time a spawn screen opens.
pub const PREFIX_BYTES: u64 = 8 * 1024;
/// The longest string this will allocate for, so a corrupt length field
/// cannot ask for a gigabyte. Longer than any key or `general.*` value.
const MAX_STRING: u64 = 64 * 1024;
/// What `general.name` says, or `None` for every way of not finding out.
///
/// `read` is consumed only as far as the key: pass a file to read a local
/// model, or a cursor over a prefix to read one whose bytes came from
/// somewhere else.
pub fn name(read: &mut impl Read) -> Option<String> {
match find(read, |key| key == "general.name") {
Some((STRING, read)) => string(read),
_ => None,
}
}
/// Whether this model carries a multi-token-prediction head.
///
/// Worth asking because `llama-server` **exits** when told to use one that is
/// not there -- `--spec-type draft-mtp` on a plain model is "context type MTP
/// requested but model doesn't contain MTP layers" and then a server that
/// never comes up. So the flag can only be passed once this has said yes, and
/// a `false` here is the same answer as an unreadable file: don't ask for it.
///
/// Matched on the key's tail rather than its whole name, because the key is
/// prefixed with the architecture (`qwen35.nextn_predict_layers`) and the
/// architecture is whatever the next model is. The value is not read: a model
/// that declares the key at all is one whose tensors carry the head, and the
/// two disagreeing is a broken file rather than a state to handle.
pub fn has_mtp_head(read: &mut impl Read) -> bool {
find(read, |key| key.ends_with(".nextn_predict_layers")).is_some()
}
/// Steps through the metadata table to the first key `wanted` accepts,
/// returning its value's type tag and the reader positioned at the value.
fn find<R: Read>(read: &mut R, wanted: impl Fn(&str) -> bool) -> Option<(u32, &mut R)> {
let mut magic = [0u8; 4];
read.read_exact(&mut magic).ok()?;
if &magic != b"GGUF" {
return None;
}
let _version = u32s(read)?;
let _tensors = u64s(read)?;
let count = u64s(read)?;
for _ in 0..count {
let found = string(read)?;
let kind = u32s(read)?;
if wanted(&found) {
return Some((kind, read));
}
skip_value(kind, read)?;
}
None
}
// The value type tags, in the container's own numbering. Only the two this
// has to act on are named; the rest are widths, and `scalar_width` is where
// the numbering is written down once.
const STRING: u32 = 8;
const ARRAY: u32 = 9;
/// How many bytes a scalar of this type occupies, or `None` for a type that
/// is not a scalar -- which includes a tag this build does not know, since a
/// value of unknown length cannot be stepped over.
fn scalar_width(kind: u32) -> Option<u64> {
match kind {
// u8, i8, bool
0 | 1 | 7 => Some(1),
// u16, i16
2 | 3 => Some(2),
// u32, i32, f32
4..=6 => Some(4),
// u64, i64, f64
10..=12 => Some(8),
_ => None,
}
}
/// Steps over one value of `kind` without keeping it.
///
/// Recursive only in the sense that an array's elements are values; GGUF
/// arrays do not nest, so the recursion is one level deep by construction.
fn skip_value(kind: u32, read: &mut impl Read) -> Option<()> {
match kind {
STRING => {
let len = u64s(read)?;
skip(len, read)
}
ARRAY => {
let element = u32s(read)?;
let count = u64s(read)?;
match scalar_width(element) {
// The whole array at once: this is the tokenizer's scores and
// token types, and stepping over them one at a time is a
// syscall per token.
Some(width) => skip(count.checked_mul(width)?, read),
None if element == STRING => {
for _ in 0..count {
let len = u64s(read)?;
skip(len, read)?;
}
Some(())
}
// An array of arrays, or of something this build has no width
// for: the rest of the table can no longer be located.
None => None,
}
}
_ => skip(scalar_width(kind)?, read),
}
}
/// Discards `count` bytes, failing if the input ends first.
///
/// Chunked against a bounded buffer rather than read into a `Vec` of the
/// stated size: the sizes here come out of the file, and the file may be a
/// truncated prefix or not a GGUF at all.
fn skip(count: u64, read: &mut impl Read) -> Option<()> {
let mut scratch = [0u8; 8192];
let mut left = count;
while left > 0 {
let want = left.min(scratch.len() as u64) as usize;
read.read_exact(&mut scratch[..want]).ok()?;
left -= want as u64;
}
Some(())
}
fn string(read: &mut impl Read) -> Option<String> {
let len = u64s(read)?;
if len > MAX_STRING {
return None;
}
let mut bytes = vec![0u8; len as usize];
read.read_exact(&mut bytes).ok()?;
String::from_utf8(bytes).ok()
}
fn u32s(read: &mut impl Read) -> Option<u32> {
let mut bytes = [0u8; 4];
read.read_exact(&mut bytes).ok()?;
Some(u32::from_le_bytes(bytes))
}
fn u64s(read: &mut impl Read) -> Option<u64> {
let mut bytes = [0u8; 8];
read.read_exact(&mut bytes).ok()?;
Some(u64::from_le_bytes(bytes))
}
#[cfg(test)]
mod tests {
use super::*;
/// Builds a GGUF header holding exactly these keys, so the parser is
/// tested against the layout rather than against a fixture nobody here
/// can regenerate.
fn header(entries: &[(&str, Value)]) -> Vec<u8> {
let mut out = Vec::from(*b"GGUF");
out.extend(3u32.to_le_bytes());
out.extend(0u64.to_le_bytes());
out.extend((entries.len() as u64).to_le_bytes());
for (key, value) in entries {
put_string(&mut out, key);
value.write(&mut out);
}
out
}
enum Value {
Str(&'static str),
U32(u32),
Strings(Vec<&'static str>),
Floats(Vec<f32>),
}
impl Value {
fn write(&self, out: &mut Vec<u8>) {
match self {
Self::Str(text) => {
out.extend(STRING.to_le_bytes());
put_string(out, text);
}
Self::U32(number) => {
out.extend(4u32.to_le_bytes());
out.extend(number.to_le_bytes());
}
Self::Strings(items) => {
out.extend(ARRAY.to_le_bytes());
out.extend(STRING.to_le_bytes());
out.extend((items.len() as u64).to_le_bytes());
for item in items {
put_string(out, item);
}
}
Self::Floats(items) => {
out.extend(ARRAY.to_le_bytes());
out.extend(6u32.to_le_bytes());
out.extend((items.len() as u64).to_le_bytes());
for item in items {
out.extend(item.to_le_bytes());
}
}
}
}
}
fn put_string(out: &mut Vec<u8>, text: &str) {
out.extend((text.len() as u64).to_le_bytes());
out.extend(text.as_bytes());
}
#[test]
fn the_name_is_read_past_every_other_kind_of_value() {
let bytes = header(&[
("general.architecture", Value::Str("qwen3")),
("general.file_type", Value::U32(7)),
("qwen3.attention.head_count", Value::U32(16)),
("tokenizer.ggml.scores", Value::Floats(vec![0.5; 64])),
("tokenizer.ggml.tokens", Value::Strings(vec!["a", "b", "c"])),
("general.name", Value::Str("Qwen3-0.6B")),
]);
assert_eq!(
name(&mut bytes.as_slice()),
Some("Qwen3-0.6B".to_string()),
"every value before the name has to be steppable over",
);
}
#[test]
/// The remote case: a prefix is all there is, and running off the end of
/// it is "we don't know" rather than a failure worth reporting. The
/// caller has the file name.
fn a_truncated_file_has_no_name_rather_than_failing() {
let bytes = header(&[
("tokenizer.ggml.tokens", Value::Strings(vec!["a", "b", "c"])),
("general.name", Value::Str("Qwen3-0.6B")),
]);
for cut in [4, 12, 24, bytes.len() - 4] {
assert_eq!(name(&mut &bytes[..cut]), None, "cut at {cut}");
}
}
#[test]
fn a_file_that_is_not_gguf_has_no_name() {
assert_eq!(name(&mut b"not a model at all".as_slice()), None);
assert_eq!(name(&mut b"".as_slice()), None);
}
#[test]
/// A name that is not a string is not a name. The alternative is
/// rendering a number as one, which reads as a model called "7".
fn a_name_of_the_wrong_type_is_not_read() {
let bytes = header(&[("general.name", Value::U32(7))]);
assert_eq!(name(&mut bytes.as_slice()), None);
}
#[test]
/// The head is found by the tail of the key, because the whole key is
/// prefixed with whatever architecture the model is.
fn an_mtp_head_is_found_whatever_the_architecture_is_called() {
let with = header(&[
("general.architecture", Value::Str("qwen35")),
("qwen35.block_count", Value::U32(64)),
("qwen35.nextn_predict_layers", Value::U32(1)),
]);
assert!(has_mtp_head(&mut with.as_slice()));
let without = header(&[
("general.architecture", Value::Str("qwen3")),
("qwen3.block_count", Value::U32(28)),
]);
assert!(!has_mtp_head(&mut without.as_slice()));
}
#[test]
/// A prefix that stops short says no, and that is the direction it has to
/// fail in: `--spec-type draft-mtp` on a model with no head is a server
/// that exits, so "we could not tell" and "it has none" both mean don't
/// ask for it.
fn a_truncated_file_reports_no_mtp_head() {
let bytes = header(&[("qwen35.nextn_predict_layers", Value::U32(1))]);
assert!(!has_mtp_head(&mut &bytes[..12]));
}
#[test]
/// The real thing, when this machine happens to have one. Skipped rather
/// than failed where it does not: the models directory is not part of the
/// checkout, and a test that needs gigabytes to run is one nobody runs.
fn a_real_model_on_this_machine_reads_back_its_name() {
let Some(home) = std::env::var_os("HOME") else {
return;
};
let dir = std::path::Path::new(&home).join(".local/share/ai-app/models");
let mut found = Vec::new();
collect_gguf(&dir, &mut found);
for path in found {
let mut file = std::fs::File::open(&path).expect("open");
let read = name(&mut file);
assert!(
read.is_some_and(|name| !name.trim().is_empty()),
"{} has a name in it and this did not read one",
path.display(),
);
}
}
fn collect_gguf(dir: &std::path::Path, found: &mut Vec<std::path::PathBuf>) {
let Ok(entries) = std::fs::read_dir(dir) else {
return;
};
for entry in entries.flatten() {
let path = entry.path();
if path.is_dir() {
collect_gguf(&path, found);
} else if path.extension().is_some_and(|e| e == "gguf") {
found.push(path);
}
}
}
}
+76 -11
View File
@@ -15,6 +15,7 @@
//! phone is not being given. //! phone is not being given.
use anyhow::{Context, Result}; use anyhow::{Context, Result};
use serde::Serialize;
use serde_json::{Value, json}; use serde_json::{Value, json};
use crate::config::{DriverKind, ProviderConfig}; use crate::config::{DriverKind, ProviderConfig};
@@ -57,12 +58,7 @@ pub async fn discover(transport: &Transport) -> Result<Vec<ProviderConfig>> {
// and nowhere else. Offering it on a remote machine would be a choice that // and nowhere else. Offering it on a remote machine would be a choice that
// changes nothing. // changes nothing.
if matches!(transport, Transport::Here) { if matches!(transport, Transport::Here) {
providers.push(ProviderConfig { providers.push(crate::config::Config::echo_provider());
name: crate::config::ECHO_PROVIDER.to_string(),
kind: DriverKind::Echo,
command: None,
models: Vec::new(),
});
} }
for (name, binary, kind) in PROBES { for (name, binary, kind) in PROBES {
let path = found let path = found
@@ -83,22 +79,88 @@ pub async fn discover(transport: &Transport) -> Result<Vec<ProviderConfig>> {
DriverKind::ClaudeCli => CLAUDE_MODELS.iter().map(|m| (*m).to_string()).collect(), DriverKind::ClaudeCli => CLAUDE_MODELS.iter().map(|m| (*m).to_string()).collect(),
_ => Vec::new(), _ => Vec::new(),
}, },
mcp_servers: mcp_defaults(*kind),
}); });
} }
Ok(providers) Ok(providers)
} }
/// One model a picker can offer, and what to call it there.
///
/// Two fields rather than one string because for one provider they differ:
/// a llama.cpp model is chosen by the path it lives at and read as the name
/// its own metadata gives it. Every other provider's id is already the name,
/// and says so by repeating it -- which is what keeps the picker free of a
/// branch on the session kind.
#[derive(Debug, Clone, Serialize)]
#[serde(rename_all = "camelCase")]
pub struct OfferedModel {
/// What a spawn or a model change is given. Opaque to the phone.
pub id: String,
/// What a person reads on the chip.
pub label: String,
}
impl OfferedModel {
/// A model whose id is its own name, which is every provider but llama.
fn plain(id: impl Into<String>) -> Self {
let id = id.into();
Self {
label: id.clone(),
id,
}
}
}
/// The MCP servers a newly discovered provider of this kind starts with.
///
/// A default rather than something to be typed in: a llama session with no web
/// search is the state somebody would then have to find out how to leave, and
/// Exa is what llama.cpp's own web UI offers under the same name. It is an
/// ordinary config entry once written, so removing it is deleting a line.
///
/// Only llama.cpp, because only a driver that runs its own agent loop can use
/// one -- the coding CLIs configure MCP themselves and a second answer here
/// would quietly disagree with theirs.
fn mcp_defaults(kind: DriverKind) -> Vec<crate::config::McpServerConfig> {
match kind {
DriverKind::LlamaCpp => vec![crate::config::McpServerConfig {
name: "exa".to_string(),
url: crate::session::llama::EXA_MCP_URL.to_string(),
}],
_ => Vec::new(),
}
}
/// Models the selected provider currently offers on this machine. /// Models the selected provider currently offers on this machine.
/// ///
/// Codex's catalog is account- and CLI-version-specific, so it is asked at the /// Codex's catalog is account- and CLI-version-specific, so it is asked at the
/// moment the picker opens rather than copied into `config.ron`. Other /// moment the picker opens rather than copied into `config.ron`. A llama.cpp
/// providers retain the shortcut list discovery stored for them. /// provider offers the GGUFs on the machine it runs on, through this same call
/// -- there was a second route answering that alone, and it went when this one
/// learned to, because a picker offering a model the spawn screen does not, or
/// naming it differently, is two answers to one question. Other providers
/// retain the shortcut list discovery stored for them.
pub async fn provider_models( pub async fn provider_models(
transport: &Transport, transport: &Transport,
provider: &ProviderConfig, provider: &ProviderConfig,
) -> Result<Vec<String>> { models_dir: &std::path::Path,
) -> Result<Vec<OfferedModel>> {
if provider.kind == DriverKind::LlamaCpp {
let dir = crate::models::dir_on(transport, models_dir);
let found = crate::models::on_machine(transport, &dir).await?;
let labels = crate::models::labels(&found);
return Ok(found
.into_iter()
.zip(labels)
.map(|(model, label)| OfferedModel {
id: model.key,
label,
})
.collect());
}
if provider.kind != DriverKind::CodexCli { if provider.kind != DriverKind::CodexCli {
return Ok(provider.models.clone()); return Ok(provider.models.iter().map(OfferedModel::plain).collect());
} }
let transport = transport.clone(); let transport = transport.clone();
let program = provider.program().to_string(); let program = provider.program().to_string();
@@ -114,7 +176,10 @@ pub async fn provider_models(
json!({"id": 2, "method": "model/list", "params": {"includeHidden": false, "limit": 100}}), json!({"id": 2, "method": "model/list", "params": {"includeHidden": false, "limit": 100}}),
]; ];
let answer = transport.request_json_blocking(&launch, &initial, &requests, 2)?; let answer = transport.request_json_blocking(&launch, &initial, &requests, 2)?;
parse_codex_models(&answer) Ok(parse_codex_models(&answer)?
.into_iter()
.map(OfferedModel::plain)
.collect())
}) })
.await? .await?
} }
+1
View File
@@ -14,6 +14,7 @@
mod auth; mod auth;
mod config; mod config;
mod files; mod files;
mod gguf;
mod machines; mod machines;
mod media; mod media;
mod models; mod models;
+103 -12
View File
@@ -50,6 +50,71 @@ pub struct LocalModel {
pub repo: String, pub repo: String,
pub file: String, pub file: String,
pub bytes: u64, pub bytes: u64,
/// What the file says it is called (`general.name` in its own metadata),
/// absent when it does not say or could not be read. Not a label: see
/// [`labels`] for what a reader is actually shown, which needs the rest of
/// the list to decide.
#[serde(skip_serializing_if = "Option::is_none")]
pub name: Option<String>,
}
/// What each of these models should be called on screen, in the same order.
///
/// A model's own name is the best answer and is not always an answer at all:
/// two quantisations of one model carry the same `general.name`, and a chip
/// row with two identical chips is one you cannot choose from. So this is a
/// cascade -- the model's own name, else its file name, else its full key --
/// and each model takes the first rung that nothing else on this machine
/// shares. The last rung always terminates it, because the key is what makes
/// these unique in the first place.
///
/// Decided over the whole list rather than per model because ambiguity is a
/// property of the set: the same file is unambiguous on a machine holding one
/// quantisation and not on a machine holding three, and only the list knows
/// which machine this is.
pub fn labels(models: &[LocalModel]) -> Vec<String> {
let rungs = |model: &LocalModel| {
[
model.name.clone(),
Some(model.file.trim_end_matches(".gguf").to_string()),
Some(model.key.clone()),
]
};
let mut taken: Vec<HashMap<String, usize>> = vec![HashMap::new(); 3];
for model in models {
for (rung, candidate) in rungs(model).into_iter().enumerate() {
if let Some(candidate) = candidate {
*taken[rung].entry(candidate).or_insert(0) += 1;
}
}
}
models
.iter()
.map(|model| {
rungs(model)
.into_iter()
.enumerate()
.find_map(|(rung, candidate)| {
let candidate = candidate?;
(taken[rung].get(&candidate) == Some(&1)).then_some(candidate)
})
// Unreachable: the key rung is unique by construction. Said as
// the key rather than as a panic, because a duplicate key would
// mean the same file listed twice and a name is still the
// honest thing to draw for it.
.unwrap_or_else(|| model.key.clone())
})
.collect()
}
/// The model's own name, read out of the file itself.
///
/// Absent for every way of not finding out -- see [`crate::gguf`]. The file is
/// opened and read only as far as the name, which is the first few hundred
/// bytes, so this is affordable once per model per listing.
fn name_of(path: &Path) -> Option<String> {
let mut file = std::fs::File::open(path).ok()?;
crate::gguf::name(&mut file)
} }
/// What a run is doing, or did. Flat rather than a tagged enum carrying its /// What a run is doing, or did. Flat rather than a tagged enum carrying its
@@ -519,6 +584,7 @@ fn collect(root: &Path, dir: &Path, found: &mut Vec<LocalModel>) {
repo: repo.to_string(), repo: repo.to_string(),
file: file.to_string(), file: file.to_string(),
bytes: entry.metadata().map(|m| m.len()).unwrap_or(0), bytes: entry.metadata().map(|m| m.len()).unwrap_or(0),
name: name_of(&path),
}); });
} }
} }
@@ -566,33 +632,46 @@ pub fn dir_on(transport: &Transport, local: &Path) -> String {
/// directory that is not there is an empty list rather than a failure: a /// directory that is not there is an empty list rather than a failure: a
/// machine that has never had a model put on it is an ordinary state, and /// machine that has never had a model put on it is an ordinary state, and
/// the same one as a machine whose directory exists and is empty. /// the same one as a machine whose directory exists and is empty.
///
/// Each record carries the head of the file as well as its size, because a
/// model's own name is inside it (see [`crate::gguf`]) and the file is on the
/// far machine. The alternative is a second round trip per model, or naming
/// remote models by path while local ones get their proper names -- one
/// machine's models reading differently from another's is exactly the
/// confusion the name was added to remove. The prefix is bounded at
/// [`crate::gguf::PREFIX_BYTES`], which is what keeps this one round trip's
/// worth of bytes.
pub async fn on_machine(transport: &Transport, dir: &str) -> Result<Vec<LocalModel>> { pub async fn on_machine(transport: &Transport, dir: &str) -> Result<Vec<LocalModel>> {
let script = "p=$1; case $p in \"~\") p=$HOME;; \"~/\"*) p=$HOME/${p#\"~/\"};; esac; \ let script = format!(
[ -d \"$p\" ] || exit 0; \ "p=$1; case $p in \"~\") p=$HOME;; \"~/\"*) p=$HOME/${{p#\"~/\"}};; esac; \
find \"$p\" -type f -name '*.gguf' -printf '%s\\t%P\\0'"; [ -d \"$p\" ] || exit 0; cd \"$p\" || exit 0; \
find . -type f -name '*.gguf' -exec sh -c '\
for f do printf \"%s\\t%s\\t%s\\0\" \"$(wc -c < \"$f\")\" \
\"$(head -c {prefix} \"$f\" | base64 | tr -d \"\\n\")\" \"${{f#./}}\"; done\
' sh {{}} +",
prefix = crate::gguf::PREFIX_BYTES,
);
let launch = Launch::new( let launch = Launch::new(
"sh", "sh",
vec![ vec!["-c".to_string(), script, "sh".to_string(), dir.to_string()],
"-c".to_string(),
script.to_string(),
"sh".to_string(),
dir.to_string(),
],
None, None,
); );
let out = transport.capture(&launch).await?; let out = transport.capture(&launch).await?;
let mut found: Vec<LocalModel> = out let mut found: Vec<LocalModel> = out
.split('\0') .split('\0')
.filter(|record| !record.is_empty()) .filter(|record| !record.is_empty())
// Two fields, and the name last, so a `\t` in a filename survives. // Three fields, and the name last, so a `\t` in a filename survives.
.filter_map(|record| record.split_once('\t')) // The middle one is base64, which has no tab in its alphabet.
.filter_map(|(bytes, key)| { .filter_map(|record| {
let (bytes, rest) = record.split_once('\t')?;
let (head, key) = rest.split_once('\t')?;
let (repo, file) = key.rsplit_once('/')?; let (repo, file) = key.rsplit_once('/')?;
Some(LocalModel { Some(LocalModel {
key: key.to_string(), key: key.to_string(),
repo: repo.to_string(), repo: repo.to_string(),
file: file.to_string(), file: file.to_string(),
bytes: bytes.trim().parse().unwrap_or(0), bytes: bytes.trim().parse().unwrap_or(0),
name: name_in_prefix(head),
}) })
}) })
.collect(); .collect();
@@ -600,6 +679,18 @@ pub async fn on_machine(transport: &Transport, dir: &str) -> Result<Vec<LocalMod
Ok(found) Ok(found)
} }
/// The model's name out of a base64 prefix of its file.
///
/// The remote half of [`name_of`], and `None` for everything that half
/// answers `None` for, plus a prefix that did not survive the trip.
fn name_in_prefix(head: &str) -> Option<String> {
use base64::Engine as _;
let bytes = base64::engine::general_purpose::STANDARD
.decode(head.trim())
.ok()?;
crate::gguf::name(&mut bytes.as_slice())
}
/// A model repository on HuggingFace, as the browse screen shows it. /// A model repository on HuggingFace, as the browse screen shows it.
#[derive(Debug, Clone, Serialize)] #[derive(Debug, Clone, Serialize)]
#[serde(rename_all = "camelCase")] #[serde(rename_all = "camelCase")]
+1
View File
@@ -436,6 +436,7 @@ mod tests {
kind: DriverKind::ClaudeCli, kind: DriverKind::ClaudeCli,
command: Some(cli.display().to_string()), command: Some(cli.display().to_string()),
models: Vec::new(), models: Vec::new(),
mcp_servers: Vec::new(),
}; };
let started = logins.start(machine, provider); let started = logins.start(machine, provider);
+6 -29
View File
@@ -7,8 +7,7 @@
//! POST /machines add {name, ssh?} -- providers are discovered //! POST /machines add {name, ssh?} -- providers are discovered
//! POST /machines/probe dry run {ssh?}: what would be found there //! POST /machines/probe dry run {ssh?}: what would be found there
//! GET /machines/{id} one machine, for refetching after a change //! GET /machines/{id} one machine, for refetching after a change
//! GET /machines/{id}/models GGUFs on that machine, for a llama session //! GET /machines/{id}/providers/{provider}/models models that provider offers
//! GET /machines/{id}/providers/{provider}/models models a CLI currently offers
//! POST /machines/{id}/providers/{provider}/auth begin provider sign-in //! POST /machines/{id}/providers/{provider}/auth begin provider sign-in
//! GET /machines/{id}/providers/{provider}/auth/{attempt} sign-in state //! GET /machines/{id}/providers/{provider}/auth/{attempt} sign-in state
//! POST /machines/{id}/providers/{provider}/auth/{attempt}/code submit browser code //! POST /machines/{id}/providers/{provider}/auth/{attempt}/code submit browser code
@@ -131,7 +130,6 @@ pub fn router(manager: Arc<SessionManager>) -> Router {
get(read_machine).put(update_machine).delete(delete_machine), get(read_machine).put(update_machine).delete(delete_machine),
) )
// The models on a configured machine, for a llama session there. // The models on a configured machine, for a llama session there.
.route("/machines/{id}/models", get(machine_models))
.route( .route(
"/machines/{id}/providers/{provider}/models", "/machines/{id}/providers/{provider}/models",
get(provider_models), get(provider_models),
@@ -565,40 +563,19 @@ struct PathQuery {
path: String, path: String,
} }
/// The models **that machine** has, which is the list a llama.cpp session /// The models a provider currently offers on its configured machine.
/// on it can choose from. /// Codex answers from its live account catalog, llama.cpp from the GGUFs on
/// /// that machine, and providers with a configured shortcut list return it.
/// Not `GET /models`, which is this backend's own downloads: those are on
/// the machine a session runs on only when they are the same machine. A
/// spawn screen offering this backend's list for a remote machine would be
/// naming files that are not there, and the session would fail at the
/// point of loading rather than at the point of choosing.
async fn machine_models(
State(manager): State<Arc<SessionManager>>,
UrlPath(id): UrlPath<String>,
) -> Result<axum::Json<Vec<crate::models::LocalModel>>, ApiError> {
let machine = machine_by_id(&manager, &id)?;
let transport = crate::session::transport::Transport::for_machine(&machine);
let dir = crate::models::dir_on(&transport, manager.models_dir());
crate::models::on_machine(&transport, &dir)
.await
.map(axum::Json)
.map_err(from_machine)
}
/// The models a CLI provider currently offers on its configured machine.
/// Codex answers from its live account catalog; providers with a configured
/// shortcut list return that list.
async fn provider_models( async fn provider_models(
State(manager): State<Arc<SessionManager>>, State(manager): State<Arc<SessionManager>>,
UrlPath((id, provider_name)): UrlPath<(String, String)>, UrlPath((id, provider_name)): UrlPath<(String, String)>,
) -> Result<axum::Json<Vec<String>>, ApiError> { ) -> Result<axum::Json<Vec<crate::machines::OfferedModel>>, ApiError> {
let machine = machine_by_id(&manager, &id)?; let machine = machine_by_id(&manager, &id)?;
let provider = machine.provider(&provider_name).ok_or_else(|| { let provider = machine.provider(&provider_name).ok_or_else(|| {
ApiError::NotFound(format!("no provider {provider_name} on {}", machine.name)) ApiError::NotFound(format!("no provider {provider_name} on {}", machine.name))
})?; })?;
let transport = crate::session::transport::Transport::for_machine(&machine); let transport = crate::session::transport::Transport::for_machine(&machine);
crate::machines::provider_models(&transport, provider) crate::machines::provider_models(&transport, provider, manager.models_dir())
.await .await
.map(axum::Json) .map(axum::Json)
.map_err(from_machine) .map_err(from_machine)
+16
View File
@@ -564,6 +564,22 @@ pub enum SessionStatus {
Running, Running,
AwaitingInput, AwaitingInput,
Compacting, Compacting,
/// The session's process is up but cannot be spoken to yet.
///
/// Its own state because the two it would otherwise borrow are both
/// wrong in ways somebody notices. `Running` means the session is
/// answering, so a model taking a minute to load looks like a model
/// thinking for a minute -- and there is no way to tell from the screen
/// that the first message will be refused. `Idle` invites that message
/// and then loses it.
///
/// It exists for `llama-server`, which reads a multi-gigabyte file off
/// disk before it answers anything, and it is general because the
/// condition is: a process that is started and not yet ready is a state
/// any driver may have to report. Nothing is queued *because* of this
/// state -- a driver that reports it is responsible for holding what it
/// is sent until it can deliver it -- but this is what says so on screen.
Loading,
/// The session's own turn is over, but work it started is still going: /// The session's own turn is over, but work it started is still going:
/// a backgrounded subagent, or a command left running. /// a backgrounded subagent, or a command left running.
/// ///
-907
View File
@@ -1,907 +0,0 @@
//! The llama.cpp driver: a `llama-server` process per session, spoken to over
//! its OpenAI-compatible HTTP API and translated into the common event model.
//!
//! Two things make this shaped differently from the Claude driver.
//!
//! **It is spawned but not spoken to over stdio.** The process is started
//! through the same [`Transport`] as any other and then reached over HTTP on a
//! loopback port. That is the second half of what a transport is -- "run this"
//! plus "reach this port" -- and it is what lets a session run on another
//! machine: [`Transport::reserve_port`] hands back a port the server binds
//! *there* and one that reaches it *here*, and the ssh connection carrying the
//! command carries the tunnel between them. The far `llama-server` binds
//! loopback only, so a model is never served to that machine's network.
//!
//! **The model file is the far machine's, not this one's.** A remote machine
//! names its own models directory (`SshConfig::models_dir`, defaulting to where
//! this backend keeps its downloads), and the file is looked for *there* -- so
//! a session naming a model that machine does not have says so, instead of
//! starting a server that will never load one. Downloading to another machine
//! is not built; the model gets there however anything else does.
//!
//! **The server is stateless between requests**, so the whole conversation goes
//! with every one. It is rebuilt from the session's transcript rather than kept
//! in this struct, which is not tidiness: a copy in driver memory is invisible
//! to a second device and gone when this process restarts.
//!
//! That leaves the Claude driver as the odd one out rather than this one -- the
//! CLI's own memory of a conversation is a cache in front of the same
//! transcript. Resolve any inconsistency in this direction.
use std::path::{Path, PathBuf};
use std::sync::Arc;
use std::sync::atomic::{AtomicBool, Ordering};
use anyhow::{Context, Result, bail};
use serde::{Deserialize, Serialize};
use serde_json::json;
use super::driver::{AttachmentRef, Driver, Event, EventSink, SessionStatus};
use super::process;
use super::transport::{Launch, Streams, Transport};
use crate::config::{ProviderConfig, SessionConfig};
/// How long to wait for a model to load before giving up. Loading is mostly
/// disk, and a large quantised model on a cold cache is genuinely slow, so this
/// is generous -- the failure it exists for is a server that will never answer.
const READY_TIMEOUT: std::time::Duration = std::time::Duration::from_secs(300);
/// One turn in the conversation this driver keeps on the server's behalf.
#[derive(Debug, Clone, Serialize, Deserialize)]
struct Message {
role: String,
content: String,
}
pub struct LlamaDriver {
sink: EventSink,
/// Where this session's own llama-server answers.
endpoint: String,
/// Where the conversation is read back from, one line per event.
transcript: PathBuf,
/// Sampling settings chosen at spawn, sent with every request.
sampling: serde_json::Map<String, serde_json::Value>,
/// Set by [`Driver::interrupt`]; the streaming loop checks it between
/// chunks and stops, leaving what was generated in the transcript.
cancel: Arc<AtomicBool>,
/// Where this session's process record lives, so [`Driver::stop`] can find
/// the server it has to end.
session_dir: PathBuf,
}
impl LlamaDriver {
/// Takes charge of this session's `llama-server`: the one already loaded if
/// there is one, otherwise a new one.
///
/// One entry point, for the reason `ClaudeDriver::launch` gives, expensive
/// in a different currency: two servers holding the same model is twice the
/// memory, and the second would bind a different port while the phone kept
/// talking to the first.
#[allow(clippy::too_many_arguments)]
pub fn launch(
meta: &SessionConfig,
provider: &ProviderConfig,
transport: &Transport,
models_dir: &Path,
transcript: &Path,
session_dir: &Path,
sink: EventSink,
// llama.cpp has no notion of a Task call, so this is accepted only
// to keep one shape across every driver's launch -- see
// `SUBAGENTS.md`'s "Server layout".
_subagents: Arc<super::subagent::Subagents>,
) -> Result<Self> {
let model = meta.model.as_deref().context(
"a llama.cpp session needs a model -- one of the downloaded ones, by its key",
)?;
let path = model_on(transport, models_dir, model)?;
// Already loaded and still running: keep talking to it. The health poll
// below confirms it is really answering, so adopting a pid whose server
// has wedged still reports as a failure rather than as a session that
// silently never replies.
if let Some(process::Record {
detail: process::Detail::Http { port },
pid,
..
}) = process::live(session_dir)
{
tracing::info!(
"session {} reattaching to the llama-server it left loaded (pid {pid}, port {port})",
meta.id
);
return Ok(Self::attached(
format!("http://127.0.0.1:{port}"),
meta,
model,
transcript,
session_dir,
sink,
));
}
// Where it listens on its own machine, and where that is reached
// from here -- the same number when that machine is this one.
let forward = transport
.reserve_port()
.context("finding a port for llama-server")?;
let mut args: Vec<String> = vec![
"-m".into(),
path.clone(),
// Loopback there, whichever machine there is: what reaches it
// from outside that machine is the ssh tunnel and nothing
// else.
"--host".into(),
"127.0.0.1".into(),
"--port".into(),
forward.there.to_string(),
];
// Settings that belong to the server because they decide how the model
// is loaded; the sampling ones ride on each request instead, so changing
// them later needn't reload anything.
for (key, flag) in [
("contextSize", "-c"),
("gpuLayers", "-ngl"),
("threads", "-t"),
] {
if let Some(value) = meta.params.get(key) {
args.push(flag.to_string());
args.push(value.clone());
}
}
let program = provider.program();
let launch = Launch::new(program, args, meta.cwd.as_deref()).reaching(forward);
// Its output goes to files, not pipes. Not only so the process can
// outlive this server: nothing ever read those pipes, so a chatty
// llama-server filled the 64 KB buffer and blocked mid-load with no sign
// of why.
let child = transport.spawn(
&launch,
Streams::Detached {
stdin: std::process::Stdio::null(),
stdout: log_file(&session_dir.join(SERVER_LOG))?.into(),
stderr: log_file(&session_dir.join(SERVER_LOG))?.into(),
},
)?;
let pid = child
.id()
.context("llama-server exited before it could be recorded")?;
tracing::info!(
"session {} running {program} for {model} {} on 127.0.0.1:{} there, \
reached at 127.0.0.1:{} here, as pid {pid}",
meta.id,
transport.describe(),
forward.there,
forward.here,
);
// Reaped so it does not become a zombie while this server is still its
// parent; the health poll and the record are what say whether the
// session is alive, because after a restart there is no `Child` to ask.
tokio::spawn(async move {
let mut child = child;
let _ = child.wait().await;
});
// The *near* port, because that is the one anything reaching this
// server has to dial -- including a later run of this backend,
// which adopts the record without knowing which machine the server
// is on. For a remote session the recorded pid is the ssh
// client's, which is the process this machine owns and which holds
// the tunnel open for exactly as long as the far server lives.
let record = process::Record::of(pid, process::Detail::Http { port: forward.here })
.context("llama-server was gone before its start time could be read")?;
process::write(session_dir, &record);
Ok(Self::attached(
format!("http://127.0.0.1:{}", forward.here),
meta,
model,
transcript,
session_dir,
sink,
))
}
/// The driver for a `llama-server` at `endpoint`, however it got there.
///
/// Shared by starting one and adopting one, because everything after "there
/// is a server at this address" is identical -- including waiting for it to
/// answer, which an adopted one still owes: a recorded pid says a process
/// exists, not that its model is loaded.
fn attached(
endpoint: String,
meta: &SessionConfig,
model: &str,
transcript: &Path,
session_dir: &Path,
sink: EventSink,
) -> Self {
// Loading is slow enough to be worth saying so: the session shows as
// running until the model is in memory, rather than looking ready and
// refusing the first message.
let _ = sink.send(Event::Status {
state: SessionStatus::Running,
});
{
let sink = sink.clone();
let endpoint = endpoint.clone();
let model = model.to_string();
let session_dir = session_dir.to_path_buf();
std::thread::spawn(move || match wait_until_ready(&endpoint, &session_dir) {
Ok(()) => {
tracing::info!("{model} loaded and answering at {endpoint}");
let _ = sink.send(Event::Status {
state: SessionStatus::Idle,
});
watch(session_dir, sink);
}
Err(err) => {
let _ = sink.send(Event::Error {
message: format!("{model} never became ready: {err:#}"),
});
let _ = sink.send(Event::Status {
state: SessionStatus::Exited,
});
process::clear(&session_dir);
}
});
}
let mut sampling = serde_json::Map::new();
for (key, field) in [
("temperature", "temperature"),
("topP", "top_p"),
("topK", "top_k"),
("maxTokens", "max_tokens"),
] {
if let Some(raw) = meta.params.get(key)
&& let Ok(number) = raw.parse::<f64>()
{
sampling.insert(field.to_string(), json!(number));
}
}
Self {
sink,
endpoint,
transcript: transcript.to_path_buf(),
sampling,
cancel: Arc::new(AtomicBool::new(false)),
session_dir: session_dir.to_path_buf(),
}
}
}
/// Where llama-server's own output goes. One file for both streams: it is
/// diagnostics nobody parses, and interleaving them is how it reads in a
/// terminal anyway.
const SERVER_LOG: &str = "llama-server.log";
/// How often a loaded server is checked for still being there. Slower than the
/// Claude driver's stdout poll because nothing is waiting on it: this only has
/// to notice a server that has gone.
const WATCH_INTERVAL: std::time::Duration = std::time::Duration::from_secs(2);
/// An owner-only log opened for appending, so the two streams pointed at
/// it do not overwrite each other and a reattach keeps what came before.
fn log_file(path: &Path) -> Result<std::fs::File> {
use std::os::unix::fs::OpenOptionsExt;
std::fs::OpenOptions::new()
.create(true)
.append(true)
.mode(0o600)
.open(path)
.with_context(|| format!("opening {}", path.display()))
}
/// Reports the server going away, for as long as the session is there to report
/// it to.
///
/// Polled rather than waited on, for the reason the Claude driver gives: after a
/// restart this server is not the process's parent, so liveness has to be a
/// question asked of the record -- and asking it two different ways is how the
/// two answers come to disagree.
fn watch(session_dir: PathBuf, sink: EventSink) {
std::thread::spawn(move || {
loop {
std::thread::sleep(WATCH_INTERVAL);
match process::recorded(&session_dir) {
Some((_, process::Liveness::Alive)) => {}
// Nothing recorded means the session was stopped or deleted
// deliberately, and whoever did that has already said so.
None => return,
Some((_, process::Liveness::Dead)) => {
if !process::stopping(&session_dir) {
let _ = sink.send(Event::Error {
message: "llama-server exited".to_string(),
});
}
let _ = sink.send(Event::Status {
state: SessionStatus::Exited,
});
process::clear(&session_dir);
return;
}
Some((_, process::Liveness::Unknown)) => {
let _ = sink.send(Event::Status {
state: SessionStatus::Unknown,
});
}
}
if sink.is_closed() {
return;
}
}
});
}
impl Driver for LlamaDriver {
fn send_user_message(&self, text: String, attachments: Vec<AttachmentRef>) {
if !attachments.is_empty() {
let _ = self.sink.send(Event::Error {
message: "this model can't be sent attachments or files".to_string(),
});
}
let sink = self.sink.clone();
let endpoint = self.endpoint.clone();
let transcript = self.transcript.clone();
let sampling = self.sampling.clone();
let cancel = Arc::clone(&self.cancel);
cancel.store(false, Ordering::Relaxed);
// Its own thread: the request blocks for as long as the model takes to
// generate, which is the whole point of streaming it.
std::thread::spawn(move || {
// Nothing is ever held back here -- there is no queue to wait in --
// so the message is taken the moment it arrives. Said anyway,
// because this is what records it: see `MessageTaken`.
let _ = sink.send(Event::MessageTaken {
id: None,
text: text.clone(),
// Never any: this driver refuses attachments above.
attachments: Vec::new(),
});
let _ = sink.send(Event::Status {
state: SessionStatus::Running,
});
// Everything before this message, plus this message. Read rather
// than remembered, and `text` is appended here rather than waited
// for, because the message's own transcript entry is still on its
// way when this runs.
let mut messages = conversation(&transcript);
messages.push(Message {
role: "user".into(),
content: text,
});
// The reply is not stored: the deltas below are the durable record,
// so the next turn reads back exactly what the phone was shown --
// including a partial one that was interrupted.
if let Err(err) = generate(&endpoint, &messages, &sampling, &cancel, &sink) {
let _ = sink.send(Event::Error {
message: format!("{err:#}"),
});
}
let _ = sink.send(Event::Status {
state: SessionStatus::Idle,
});
});
}
fn answer_question(&self, _id: &str, _answers: &[String]) {
// Nothing here asks questions: this driver has no tools.
}
fn interrupt(&self) {
self.cancel.store(true, Ordering::Relaxed);
}
// Nothing to forward: this process has no notion of what the conversation
// is called, and the rename has already happened where the name lives.
fn set_title(&self, _title: &str) {}
fn set_permission_mode(&self, _mode: &str) {
let _ = self.sink.send(Event::Error {
message: "a llama.cpp session runs no tools, so there is nothing for a permission \
mode to govern."
.to_string(),
});
}
fn set_model(&self, _model: &str) {
let _ = self.sink.send(Event::Error {
message: "a llama.cpp session's model is fixed when it starts, because the server \
loads one model into memory. Spawn another session to use a different one."
.to_string(),
});
}
fn run_command(&self, text: &str) {
let _ = self.sink.send(Event::Error {
message: format!(
"a llama.cpp session has no commands of its own, so {text} means nothing to it."
),
});
}
fn compact(&self) {
let _ = self.sink.send(Event::Error {
message: "llama.cpp has no compaction. Clear the session instead, which costs nothing."
.to_string(),
});
}
fn clear(&self) {
// All of it. `conversation` folds from the last of these, so recording
// the marker *is* the reset -- there is no driver state to keep in step
// with it, which is the same property that makes a second device see the
// same conversation this one does.
let _ = self.sink.send(Event::Cleared);
}
/// Stops generating and leaves the server loaded.
///
/// Worth being deliberate about, because the cost points the other way from
/// the Claude driver's: a `llama-server` holds its whole model in memory, so
/// a leaked one is gigabytes nobody is using. It is left anyway, because the
/// alternative is unloading and reloading that model on every backend
/// restart -- minutes of disk, for a session somebody is in the middle of.
/// The record is what keeps it from being *nobody's*.
fn detach(&self) {
self.cancel.store(true, Ordering::Relaxed);
}
fn stop(&self) {
self.cancel.store(true, Ordering::Relaxed);
if let Some(record) = process::live(&self.session_dir) {
process::stop(&record, process::STOP_GRACE);
}
process::clear(&self.session_dir);
}
}
/// The conversation so far, folded out of the transcript.
///
/// Consecutive `AssistantText` deltas are one assistant turn, closed by the next
/// user message -- which is also what makes an interrupted reply come back as
/// the partial text the phone actually saw.
///
/// This must stay a pure function of the transcript and must never re-render
/// earlier turns. llama.cpp caches the prompt prefix, so a growing conversation
/// reprocesses almost nothing -- but only while every turn is byte-identical to
/// last time. Changing how an old turn is rendered silently reprocesses the
/// whole history on every message.
fn conversation(path: &Path) -> Vec<Message> {
let Ok(events) = crate::session::transcript::read_after(path, 0) else {
return Vec::new();
};
let mut messages: Vec<Message> = Vec::new();
let mut pending = String::new();
// Everything before the last clear is still in the transcript and is
// deliberately not in the conversation. Folding from zero would put it back,
// which is the whole of what clearing had to undo.
let events = match events.iter().rposition(|e| e.event == Event::Cleared) {
Some(at) => &events[at + 1..],
None => &events[..],
};
for event in events.iter().cloned() {
match event.event {
Event::UserMessage { text, .. } => {
if !pending.is_empty() {
messages.push(Message {
role: "assistant".into(),
content: std::mem::take(&mut pending),
});
}
messages.push(Message {
role: "user".into(),
content: text,
});
}
Event::AssistantText { delta } => pending.push_str(&delta),
Event::AssistantTextFinal { text } => {
pending = text;
}
_ => {}
}
}
if !pending.is_empty() {
messages.push(Message {
role: "assistant".into(),
content: pending,
});
}
messages
}
/// Where a model key resolves to on disk, refusing anything that climbs
/// out of the models directory -- the key arrives from a phone.
fn model_path(models_dir: &Path, key: &str) -> Result<PathBuf> {
let mut path = models_dir.to_path_buf();
for part in key.split('/') {
if part.is_empty() || part == "." || part == ".." {
bail!("\"{key}\" is not a model key this can resolve");
}
path.push(part);
}
if !path.is_file() {
bail!("no downloaded model called \"{key}\" -- download it first");
}
Ok(path)
}
/// The model file's path **on the machine that will serve it**, confirmed to be
/// there.
///
/// One function rather than a local check and hope for the other case: the same
/// question has to be asked of two filesystems. The remote answer is measured
/// for the reason the local one is -- a missing file otherwise becomes a
/// `llama-server` that starts, fails to load, and reports as a session that
/// never became ready, which reads as the machine being slow.
///
/// One blocking round trip on a remote spawn, which is what the spawn is
/// already paying to start ssh. The alternative is a path built here from a `~`
/// this machine cannot expand.
fn model_on(transport: &Transport, models_dir: &Path, key: &str) -> Result<String> {
let Transport::Ssh { name, .. } = transport else {
return Ok(model_path(models_dir, key)?.to_string_lossy().into_owned());
};
// The same directory the spawn screen listed for this machine, and one
// function for the same reason: a list from one place and a load from
// another is a model that appears and then fails.
let dir = crate::models::dir_on(transport, models_dir);
// Checked here rather than in the script: `..` in a key would walk out of
// the models directory on a machine this server can start processes on,
// and the phone is where the key comes from.
for part in key.split('/') {
if part.is_empty() || part == "." || part == ".." {
bail!("\"{key}\" is not a model key this can resolve");
}
}
let path = format!("{}/{key}", dir.trim_end_matches('/'));
// `$HOME` on the far side, which is the only machine that knows what it is,
// and the resolved path printed back so the launch hands `llama-server`
// something absolute. "Not there" is answered rather than failed, because a
// machine that could not be asked at all has to say so in its own words --
// it would otherwise arrive as this same sentence about a missing model.
let script = "p=$1; case $p in \"~\") p=$HOME;; \"~/\"*) p=$HOME/${p#\"~/\"};; esac; \
[ -f \"$p\" ] && printf 'at\\t%s\\n' \"$p\" || printf 'missing\\n'"
.to_string();
let launch = Launch::new(
"sh",
vec!["-c".to_string(), script, "sh".to_string(), path.clone()],
None,
);
let answer = transport
.capture_blocking(&launch)
.with_context(|| format!("couldn't ask {name} where its models are"))?;
match answer.trim().split_once('\t') {
Some(("at", resolved)) => Ok(resolved.to_string()),
_ => bail!(
"{name} has no model at {path}. A llama.cpp session serves the file from the \
machine it runs on, so the model has to be on {name} -- what this backend has \
downloaded is somewhere else."
),
}
}
/// Polls until the server says it is ready, or gives up.
///
/// Watches the process as well as the port, because the two failures need
/// different words and one of them is common: a model that will not load,
/// a port already taken on the far machine, a `llama-server` too old for
/// a flag. All of those exit within a second and none of them will ever
/// answer `/health`, so waiting out the timeout turns a server that said
/// exactly what was wrong into "gave up after 300s".
fn wait_until_ready(endpoint: &str, session_dir: &Path) -> Result<()> {
let deadline = std::time::Instant::now() + READY_TIMEOUT;
let url = format!("{endpoint}/health");
loop {
if let Ok(response) = ureq::get(&url).call()
&& response.status() == 200
{
return Ok(());
}
// `None` is the session having been stopped or deleted while this
// waited, which is nobody's fault and still not worth waiting on.
match process::recorded(session_dir) {
Some((_, process::Liveness::Alive | process::Liveness::Unknown)) => {}
Some((_, process::Liveness::Dead)) | None => {
bail!("it exited before it answered.{}", log_tail(session_dir));
}
}
if std::time::Instant::now() > deadline {
bail!(
"gave up after {}s.{}",
READY_TIMEOUT.as_secs(),
log_tail(session_dir)
);
}
std::thread::sleep(std::time::Duration::from_millis(250));
}
}
/// The end of `llama-server`'s own log, for a failure message.
///
/// Its account of what went wrong is the useful half -- "failed to load
/// model", "bind: Address already in use" -- and on a remote session it
/// is the only half, since nobody reading the phone can open a file on
/// that machine. Bounded, because this ends up in an event a phone draws.
fn log_tail(session_dir: &Path) -> String {
let Ok(text) = std::fs::read_to_string(session_dir.join(SERVER_LOG)) else {
return String::new();
};
let tail: Vec<&str> = text.lines().rev().take(LOG_TAIL_LINES).collect();
if tail.is_empty() {
return String::new();
}
format!(
" It last said: {}",
tail.into_iter().rev().collect::<Vec<_>>().join(" / ")
)
}
/// How much of that log to carry into a message somebody reads on a phone.
const LOG_TAIL_LINES: usize = 6;
/// One streamed completion: posts the conversation, emits each delta as it
/// arrives. Emits rather than returns, because the transcript those events land
/// in is what the next turn reads back.
fn generate(
endpoint: &str,
messages: &[Message],
sampling: &serde_json::Map<String, serde_json::Value>,
cancel: &AtomicBool,
sink: &EventSink,
) -> Result<()> {
let mut body = json!({
"messages": messages,
"stream": true,
"stream_options": {"include_usage": true},
});
let map = body.as_object_mut().expect("built as an object");
for (key, value) in sampling {
map.insert(key.clone(), value.clone());
}
let mut response = ureq::post(format!("{endpoint}/v1/chat/completions"))
.header("Content-Type", "application/json")
.send_json(&body)
.context("asking llama-server to generate")?;
let reader = std::io::BufReader::new(response.body_mut().as_reader());
let mut tokens = 0u64;
// The prompt side only, which is what the model is holding -- the same
// definition the other dialects report, so one word on the phone means one
// thing whichever kind of session it is.
let mut context = None;
for line in std::io::BufRead::lines(reader) {
if cancel.load(Ordering::Relaxed) {
break;
}
let line = line.context("reading the generation stream")?;
// Server-sent events: the payload lines are the ones that matter.
let Some(payload) = line.strip_prefix("data: ") else {
continue;
};
if payload.trim() == "[DONE]" {
break;
}
let Ok(chunk) = serde_json::from_str::<serde_json::Value>(payload) else {
continue;
};
if let Some(usage) = chunk.get("usage") {
if let Some(total) = usage
.get("total_tokens")
.and_then(serde_json::Value::as_u64)
{
tokens = total;
}
if let Some(prompt) = usage
.get("prompt_tokens")
.and_then(serde_json::Value::as_u64)
{
context = Some(prompt);
}
}
let delta = chunk
.get("choices")
.and_then(|c| c.get(0))
.and_then(|c| c.get("delta"))
.and_then(|d| d.get("content"))
.and_then(serde_json::Value::as_str)
.unwrap_or_default();
if !delta.is_empty() {
let _ = sink.send(Event::AssistantText {
delta: delta.to_string(),
});
}
}
if tokens > 0 {
let _ = sink.send(Event::UsageDelta { tokens, context });
}
Ok(())
}
#[cfg(test)]
mod tests {
use super::*;
use crate::session::transcript::Transcript;
/// Writes a transcript the way the pump does, so the fold is tested against
/// the real file format rather than a hand-built vector.
fn transcript_with(events: &[Event]) -> (tempfile::TempDir, PathBuf) {
let dir = tempfile::tempdir().expect("tempdir");
let path = dir.path().join("transcript.jsonl");
let mut transcript = Transcript::open(&path).expect("open");
for event in events {
transcript.append(event.clone(), 0.0).expect("append");
}
(dir, path)
}
#[test]
fn deltas_between_user_messages_are_one_assistant_turn() {
let (_dir, path) = transcript_with(&[
Event::UserMessage {
id: None,
text: "hello".into(),
attachments: Vec::new(),
},
Event::AssistantText {
delta: "hi ".into(),
},
Event::AssistantText {
delta: "there".into(),
},
Event::Status {
state: SessionStatus::Idle,
},
Event::UserMessage {
id: None,
text: "again".into(),
attachments: Vec::new(),
},
Event::AssistantText {
delta: "yes".into(),
},
]);
let messages = conversation(&path);
assert_eq!(
messages
.iter()
.map(|m| (m.role.as_str(), m.content.as_str()))
.collect::<Vec<_>>(),
[
("user", "hello"),
("assistant", "hi there"),
("user", "again"),
("assistant", "yes")
],
);
}
#[test]
/// The interrupted case, which decides what a resumed conversation is built
/// from: whatever the phone was shown. The deltas that arrived before the
/// stop are in the transcript, so they are in the prompt -- the model is
/// never told it said something the user did not see.
fn an_interrupted_reply_stays_in_the_conversation() {
let (_dir, path) = transcript_with(&[
Event::UserMessage {
id: None,
text: "count".into(),
attachments: Vec::new(),
},
Event::AssistantText {
delta: "one two".into(),
},
Event::Status {
state: SessionStatus::Idle,
},
]);
let messages = conversation(&path);
assert_eq!(messages.len(), 2);
assert_eq!(messages[1].content, "one two");
}
#[test]
/// Events this driver does not produce must not disturb the fold: a
/// transcript can carry errors and status changes from a session that
/// was, say, relaunched.
fn other_events_are_not_part_of_the_conversation() {
let (_dir, path) = transcript_with(&[
Event::Status {
state: SessionStatus::Running,
},
Event::UserMessage {
id: None,
text: "hello".into(),
attachments: Vec::new(),
},
Event::Error {
message: "something went wrong".into(),
},
Event::AssistantText {
delta: "still here".into(),
},
Event::UsageDelta {
tokens: 12,
context: Some(12),
},
]);
let messages = conversation(&path);
assert_eq!(messages.len(), 2);
assert_eq!(messages[0].content, "hello");
assert_eq!(messages[1].content, "still here");
}
#[test]
/// Clearing decides what the *model* is given, not just what the phone
/// draws. Everything above the marker stays in the transcript and none of it
/// is sent.
fn the_conversation_starts_after_the_last_clear() {
let (_dir, path) = transcript_with(&[
Event::UserMessage {
id: None,
text: "the long expensive conversation".into(),
attachments: Vec::new(),
},
Event::AssistantText {
delta: "at length".into(),
},
Event::Cleared,
Event::UserMessage {
id: None,
text: "a fresh start".into(),
attachments: Vec::new(),
},
Event::AssistantText {
delta: "cheaply".into(),
},
]);
let messages = conversation(&path);
assert_eq!(messages.len(), 2);
assert_eq!(messages[0].content, "a fresh start");
assert_eq!(messages[1].content, "cheaply");
}
#[test]
/// The *last* one, so clearing twice does not resurrect what the
/// first clear dropped.
fn only_the_newest_clear_counts() {
let (_dir, path) = transcript_with(&[
Event::UserMessage {
id: None,
text: "one".into(),
attachments: Vec::new(),
},
Event::Cleared,
Event::UserMessage {
id: None,
text: "two".into(),
attachments: Vec::new(),
},
Event::Cleared,
Event::UserMessage {
id: None,
text: "three".into(),
attachments: Vec::new(),
},
]);
let messages = conversation(&path);
assert_eq!(messages.len(), 1);
assert_eq!(messages[0].content, "three");
}
#[test]
fn a_model_key_cannot_climb_out_of_the_models_directory() {
let dir = tempfile::tempdir().expect("tempdir");
for attempt in ["../../etc/passwd", "unsloth/../../escape.gguf", ""] {
assert!(
model_path(dir.path(), attempt).is_err(),
"{attempt:?} should have been refused",
);
}
}
}
+375
View File
@@ -0,0 +1,375 @@
//! An MCP client, for the tools a llama session has that `llama-server` does
//! not provide itself.
//!
//! **Why this is here and not a flag on `llama-server`.** That server can host
//! MCP servers (`--mcp-servers-json`), but only ones it can *spawn*: its
//! configuration is Cursor's, and an entry without a `command` is skipped with
//! "MCP server 'exa' has no command". Exa's is a remote HTTP endpoint with
//! nothing to spawn, so reaching it that way means a local process bridging
//! stdio to HTTP -- a Node install on the machine serving the model, and a
//! package to keep current, for what is three JSON-RPC calls.
//!
//! llama.cpp's own web UI does not do that either. It ships Exa in a
//! "recommended servers" list and connects to `https://mcp.exa.ai/mcp`
//! *itself*, from the browser. This is the same arrangement with this server
//! in the browser's place, and it is the right one for a second reason: it
//! puts the search on the machine running the backend rather than on whichever
//! machine happens to be serving the model, which may have no route out at
//! all.
//!
//! **Only the three calls a tool needs.** `initialize`, `tools/list`,
//! `tools/call`. Nothing here implements resources, prompts, sampling or the
//! server-to-client stream, because nothing here uses them; a session's tools
//! are a list fetched once and a call made on demand. That is why this is a
//! file rather than a dependency on a protocol crate -- there is no spec
//! surface to get subtly wrong, only a request and its reply.
//!
//! Transport is "streamable HTTP": every message is a POST, and the reply is
//! either JSON or a one-event SSE stream carrying the same JSON. Both are
//! accepted because which one arrives is the server's choice, not ours.
use anyhow::{Context, Result, bail};
use serde_json::{Value, json};
/// Identifies this client to an MCP server.
///
/// Not politeness: Exa's endpoint is behind Cloudflare, which answers **403**
/// to a request with no `User-Agent` at all (measured 2026-09-19 -- the same
/// request with one succeeds). A client that omitted it would look exactly
/// like a server that was refusing us.
const USER_AGENT: &str = concat!("ai-server/", env!("CARGO_PKG_VERSION"));
/// The protocol version this speaks. Sent at `initialize`; a server that
/// prefers another says so in its answer and this goes along with whatever it
/// then sends, since none of the three calls here has changed between
/// versions.
const PROTOCOL_VERSION: &str = "2025-06-18";
/// How long any one call may take.
///
/// Generous because a web search is a search: Exa fetches and cleans pages
/// before answering. Bounded at all because this blocks a turn, and a tool
/// that never returns is a session that never speaks again.
const CALL_TIMEOUT: std::time::Duration = std::time::Duration::from_secs(120);
/// A connected MCP server, and the tools it offered.
pub struct McpServer {
/// The name this server is configured under. It prefixes every tool, so
/// two servers offering `search` are two different tools.
name: String,
url: String,
/// What the server called this conversation, when it named one. Sent back
/// on every later request; a server that keeps no session sends no header
/// and this stays `None`.
session: Option<String>,
/// The tool names this server answers to, without the prefix, keyed by the
/// prefixed name the model is given.
tools: Vec<McpTool>,
}
/// One tool an MCP server offers, in both the names it has.
pub struct McpTool {
/// `{server}_{tool}` -- what the model calls it, and what comes back in a
/// tool call. Prefixed the way `llama-server` prefixes the MCP tools it
/// hosts itself, so a reader sees one naming convention whichever side a
/// tool came from.
pub qualified: String,
/// What the server calls it.
bare: String,
/// The OpenAI-shaped function definition sent to the model.
pub definition: Value,
}
impl McpServer {
/// Connects, handshakes, and asks what it can do.
///
/// All three steps or none: a server that answered `initialize` and then
/// failed to list its tools is not a server with no tools, and returning
/// an empty list for it would put a session on screen that silently
/// cannot search.
pub fn connect(name: &str, url: &str) -> Result<Self> {
let mut server = Self {
name: name.to_string(),
url: url.to_string(),
session: None,
tools: Vec::new(),
};
server
.request(
1,
"initialize",
json!({
"protocolVersion": PROTOCOL_VERSION,
"capabilities": {},
"clientInfo": {"name": "ai-server", "title": "AI Sessions", "version": env!("CARGO_PKG_VERSION")},
}),
)
.with_context(|| format!("handshaking with the {name} MCP server at {url}"))?;
// A notification: no id, and the server answers with no body. Sent
// because the specification requires it before any other call, and
// Exa's server does enforce it.
server.notify("notifications/initialized")?;
let listed = server
.request(2, "tools/list", json!({}))
.with_context(|| format!("asking the {name} MCP server what it offers"))?;
server.tools = listed
.get("tools")
.and_then(Value::as_array)
.map(|tools| {
tools
.iter()
.filter_map(|tool| server.describe(tool))
.collect()
})
.unwrap_or_default();
Ok(server)
}
/// Turns one entry of `tools/list` into the function definition a model is
/// given, or `None` for one this cannot name or call.
fn describe(&self, tool: &Value) -> Option<McpTool> {
let bare = tool.get("name").and_then(Value::as_str)?.to_string();
let qualified = format!("{}_{bare}", self.name);
let mut function = serde_json::Map::new();
function.insert("name".into(), json!(qualified));
if let Some(description) = tool.get("description").and_then(Value::as_str) {
function.insert("description".into(), json!(description));
}
// `inputSchema` in MCP, `parameters` in the OpenAI shape: the same
// JSON Schema under two names. A tool that declares none takes no
// arguments, which is an empty object rather than an absent key --
// some templates render the key unconditionally.
function.insert(
"parameters".into(),
tool.get("inputSchema")
.cloned()
.unwrap_or_else(|| json!({"type": "object", "properties": {}})),
);
Some(McpTool {
qualified,
bare,
definition: json!({"type": "function", "function": function}),
})
}
pub fn tools(&self) -> &[McpTool] {
&self.tools
}
/// Runs one of this server's tools, named as the model named it.
///
/// The result is the text a model is shown. A tool the server reports as
/// failing is **not** an error here: `isError` means the tool ran and went
/// wrong -- a search that found nothing, a page that would not fetch --
/// and the model is the one that has to know, so it comes back as its own
/// message. An error is reserved for not having reached the server at all.
pub fn call(&mut self, qualified: &str, arguments: &Value) -> Result<String> {
// The bare name is taken before the call, because the call needs the
// whole of `self` and the tool list is part of it.
let bare = self
.tools
.iter()
.find(|tool| tool.qualified == qualified)
.map(|tool| tool.bare.clone())
.with_context(|| format!("{} does not offer {qualified}", self.name))?;
let result = self.request(
3,
"tools/call",
json!({"name": bare, "arguments": arguments}),
)?;
Ok(rendered(&result))
}
/// One request, and its result.
///
/// `&self` rather than `&mut self` everywhere but the handshake would be
/// tidier and is wrong: the session header is assigned by the server on
/// the first reply and has to be kept.
fn request(&mut self, id: u64, method: &str, params: Value) -> Result<Value> {
let body = json!({"jsonrpc": "2.0", "id": id, "method": method, "params": params});
let answer = self.post(&body)?.with_context(|| {
format!(
"the {} MCP server answered {method} with nothing",
self.name
)
})?;
if let Some(message) = answer.pointer("/error/message").and_then(Value::as_str) {
bail!("{} refused {method}: {message}", self.name);
}
answer
.get("result")
.cloned()
.with_context(|| format!("the {} MCP server's {method} carried no result", self.name))
}
/// A message with no id, which is answered with no body.
fn notify(&mut self, method: &str) -> Result<()> {
self.post(&json!({"jsonrpc": "2.0", "method": method}))?;
Ok(())
}
/// Posts one JSON-RPC message and returns whatever came back, which for a
/// notification is nothing.
fn post(&mut self, body: &Value) -> Result<Option<Value>> {
let mut request = ureq::post(&self.url)
.config()
.timeout_global(Some(CALL_TIMEOUT))
.build()
.header("content-type", "application/json")
// Both, because which one a server replies with is its choice.
.header("accept", "application/json, text/event-stream")
.header("user-agent", USER_AGENT);
if let Some(session) = &self.session {
request = request.header("mcp-session-id", session);
}
let mut response = request
.send_json(body)
.with_context(|| format!("reaching the {} MCP server at {}", self.name, self.url))?;
if let Some(session) = response
.headers()
.get("mcp-session-id")
.and_then(|value| value.to_str().ok())
{
self.session = Some(session.to_string());
}
let streamed = response
.headers()
.get("content-type")
.and_then(|value| value.to_str().ok())
.is_some_and(|value| value.contains("text/event-stream"));
let text = response
.body_mut()
.read_to_string()
.with_context(|| format!("reading the {} MCP server's answer", self.name))?;
Ok(first_message(&text, streamed))
}
}
/// The first JSON-RPC message in a reply body.
///
/// One, not all: every call here carries a single id and the server answers it
/// once. Server-sent events are unwrapped to their payload lines; a plain JSON
/// body is itself.
fn first_message(text: &str, streamed: bool) -> Option<Value> {
if streamed {
return text
.lines()
.filter_map(|line| line.strip_prefix("data: "))
.find_map(|payload| serde_json::from_str(payload).ok());
}
serde_json::from_str(text.trim()).ok()
}
/// A `tools/call` result as the text a model is given.
///
/// MCP answers with a list of content blocks; the text ones are joined and the
/// rest are named rather than dropped, because a model told nothing came back
/// will try again. `structuredContent` is used when there is no text at all,
/// which is how some servers answer entirely.
fn rendered(result: &Value) -> String {
let blocks = result.get("content").and_then(Value::as_array);
let mut parts: Vec<String> = Vec::new();
for block in blocks.into_iter().flatten() {
match block.get("type").and_then(Value::as_str) {
Some("text") => parts.push(
block
.get("text")
.and_then(Value::as_str)
.unwrap_or_default()
.to_string(),
),
Some(kind) => parts.push(format!("[{kind} content, which this session cannot show]")),
None => {}
}
}
if parts.iter().all(|part| part.trim().is_empty())
&& let Some(structured) = result.get("structuredContent")
{
return structured.to_string();
}
parts.join("\n")
}
#[cfg(test)]
mod tests {
use super::*;
#[test]
fn an_event_stream_body_is_unwrapped_to_its_payload() {
let body =
"event: message\ndata: {\"jsonrpc\":\"2.0\",\"id\":1,\"result\":{\"ok\":true}}\n\n";
assert_eq!(
first_message(body, true),
Some(json!({"jsonrpc": "2.0", "id": 1, "result": {"ok": true}})),
);
}
#[test]
fn a_plain_json_body_is_the_message() {
let body = " {\"jsonrpc\":\"2.0\",\"id\":1,\"result\":{}}\n";
assert_eq!(
first_message(body, false),
Some(json!({"jsonrpc": "2.0", "id": 1, "result": {}})),
);
}
#[test]
/// A notification's reply, which is nothing at all.
fn an_empty_body_is_no_message() {
assert_eq!(first_message("", false), None);
assert_eq!(first_message("event: ping\n", true), None);
}
#[test]
fn text_blocks_are_joined_and_other_kinds_are_named() {
let result = json!({"content": [
{"type": "text", "text": "first"},
{"type": "image", "data": ""},
{"type": "text", "text": "second"},
]});
assert_eq!(
rendered(&result),
"first\n[image content, which this session cannot show]\nsecond",
);
}
#[test]
/// A server that answers only in structured form. Rendering "" for it
/// would tell the model the search came back empty, which is a different
/// fact from the one that is true.
fn a_result_with_no_text_falls_back_to_its_structured_form() {
let result = json!({"content": [], "structuredContent": {"hits": 2}});
assert_eq!(rendered(&result), "{\"hits\":2}");
}
#[test]
/// The real endpoint, which is the only thing that can confirm the
/// handshake, the session header and the SSE unwrapping all agree with a
/// server nobody here wrote. Skipped without network rather than failed:
/// `./run-tests.sh` has to pass on a machine with no route out.
fn exa_answers_a_search_over_the_real_protocol() {
let Ok(mut server) = McpServer::connect("exa", super::super::EXA_MCP_URL) else {
eprintln!("skipping: could not reach Exa");
return;
};
assert!(
server
.tools()
.iter()
.any(|tool| tool.qualified == "exa_web_search_exa"),
"Exa offered {:?}",
server
.tools()
.iter()
.map(|tool| &tool.qualified)
.collect::<Vec<_>>(),
);
let answer = server
.call(
"exa_web_search_exa",
&json!({"query": "llama.cpp server", "numResults": 1}),
)
.expect("search");
assert!(!answer.trim().is_empty(), "a search returned nothing");
}
}
File diff suppressed because it is too large. Load diff
+263
View File
@@ -0,0 +1,263 @@
//! What a llama session can do besides talk, and who runs it.
//!
//! Two sources, one list. `llama-server` started with `--tools` runs a set of
//! its own -- reading, searching, editing, a shell -- and publishes them at
//! `GET /tools` in the shape a model is given, with `POST /tools` to run one.
//! Anything else comes from an MCP server this backend is connected to (see
//! [`super::mcp`]). Both arrive here as a definition to offer and a way to
//! call, and nothing downstream of [`Tools::execute`] knows which a tool was.
//!
//! **The built-in tools run where the model does, and that is the point.** A
//! session on another machine edits that machine's files, because that is the
//! machine `llama-server` is on -- the same rule the model file already
//! follows. MCP tools run here instead, which is right for the opposite
//! reason: a web search wants the machine with a route out, not the one with
//! the GPU.
//!
//! **A tool's failure is a result, not an error.** A missing file, a command
//! that exited non-zero, a search that found nothing: all of those are things
//! the model has to read and act on, so they come back as the tool's output.
//! [`Tools::execute`] returns `Err` only when the tool could not be reached at
//! all, which is a fact about this server rather than about the work.
use std::collections::HashMap;
use std::sync::{Arc, Mutex};
use anyhow::{Context, Result};
use serde_json::{Value, json};
use super::mcp::McpServer;
/// How long one tool call may take.
///
/// This is the shell tool's budget as much as anything: a build, a test run,
/// a `find` over a large tree. Bounded because it blocks the turn, and a
/// session stuck behind a command that will never finish cannot even be told
/// to stop.
const EXECUTE_TIMEOUT: std::time::Duration = std::time::Duration::from_secs(300);
/// The tools one session has, and how to run each of them.
pub struct Tools {
/// The `llama-server` these belong to. Replaced when the session's model
/// changes, because that is a different server on a different port.
endpoint: String,
/// What the model is given, in the order it is offered: the server's own
/// tools first, then each MCP server's.
definitions: Vec<Value>,
/// The server's tools, and whether each is run relative to a working
/// directory. Only the ones that say so are sent one -- a tool that
/// ignores it would still have its cache keyed on it.
server: HashMap<String, bool>,
/// Connected MCP servers, each of which knows its own tools by the
/// prefixed names they were offered under.
mcp: Vec<Arc<Mutex<McpServer>>>,
}
impl Tools {
/// Asks a ready `llama-server` what it offers and adds what the MCP
/// servers offered.
///
/// A server started without `--tools` answers with an empty list, and a
/// session with only MCP tools is a perfectly good session -- so nothing
/// here treats "no tools" as a failure. What *is* a failure is not being
/// able to ask, because that is the same server the conversation is about
/// to go to.
pub fn discover(endpoint: &str, mcp: Vec<Arc<Mutex<McpServer>>>) -> Result<Self> {
let catalog: Vec<Value> = ureq::get(format!("{endpoint}/tools"))
.call()
.context("asking llama-server which tools it has")?
.body_mut()
.read_json()
.context("reading llama-server's tool list")?;
let mut definitions = Vec::new();
let mut server = HashMap::new();
for entry in &catalog {
let Some(name) = entry
.pointer("/definition/function/name")
.and_then(Value::as_str)
else {
continue;
};
let Some(definition) = entry.get("definition") else {
continue;
};
server.insert(
name.to_string(),
entry
.get("uses_cwd")
.and_then(Value::as_bool)
.unwrap_or(false),
);
definitions.push(definition.clone());
}
for connected in &mcp {
for tool in connected.lock().unwrap().tools() {
definitions.push(tool.definition.clone());
}
}
Ok(Self {
endpoint: endpoint.to_string(),
definitions,
server,
mcp,
})
}
/// What goes in the request's `tools`, or `None` when there is nothing to
/// offer.
///
/// Absent rather than empty for a reason that shows on screen: a chat
/// template branches on whether tools were given, and an empty list
/// renders the whole "you may call one or more functions" preamble with no
/// functions under it.
pub fn offered(&self) -> Option<&[Value]> {
(!self.definitions.is_empty()).then_some(&self.definitions)
}
/// Whether this is a tool at all, which decides what to do about a call
/// naming something else.
pub fn knows(&self, name: &str) -> bool {
self.server.contains_key(name) || self.mcp_for(name).is_some()
}
/// The MCP server that offered `name`, if one did.
fn mcp_for(&self, name: &str) -> Option<&Arc<Mutex<McpServer>>> {
self.mcp.iter().find(|server| {
server
.lock()
.unwrap()
.tools()
.iter()
.any(|tool| tool.qualified == name)
})
}
/// Runs one call and returns what the model should read.
///
/// `cwd` is the session's working directory, sent only to the tools that
/// say they use one. A session with no working directory sends none, and
/// `llama-server` falls back to its own -- which is the honest outcome:
/// this server has no better answer for where "here" is.
pub fn execute(&self, name: &str, arguments: &Value, cwd: Option<&str>) -> Result<String> {
if let Some(server) = self.mcp_for(name) {
return server.lock().unwrap().call(name, arguments);
}
let uses_cwd = *self
.server
.get(name)
.with_context(|| format!("no tool called {name}"))?;
let mut request = ureq::post(format!("{}/tools", self.endpoint))
.config()
.timeout_global(Some(EXECUTE_TIMEOUT))
// The refusal is a sentence the model can act on, so it is read as
// one rather than discarded in favour of its status code -- the
// same reason `super::refusal` exists for generation.
.http_status_as_error(false)
.build()
.header("content-type", "application/json");
if let (true, Some(cwd)) = (uses_cwd, cwd) {
request = request.header("x-tool-cwd", cwd);
}
let mut response = request
.send_json(json!({"tool": name, "params": arguments}))
.with_context(|| format!("asking llama-server to run {name}"))?;
let body = response
.body_mut()
.read_to_string()
.with_context(|| format!("reading what {name} produced"))?;
Ok(match serde_json::from_str::<Value>(&body) {
Ok(answer) => result_text(&answer),
// Not JSON at all: hand over what was said rather than a parse
// error about it, since the model is what has to carry on.
Err(_) => body,
})
}
}
/// `POST /tools`'s answer as the text a model is given.
///
/// The server answers `plain_text_response` for a tool that ran and `error`
/// for one that did not, and both are the model's business -- see this
/// module's note on failures being results. Anything else is handed over as
/// itself rather than discarded, since a tool this build has not seen before
/// is exactly the case where guessing is worst.
fn result_text(answer: &Value) -> String {
if let Some(text) = answer.get("plain_text_response").and_then(Value::as_str) {
return text.to_string();
}
if let Some(message) = answer.get("error").and_then(Value::as_str) {
return message.to_string();
}
answer.to_string()
}
/// How much a session asks before it acts.
///
/// Two, because two is what the mechanism underneath actually has. The web UI
/// that ships with `llama-server` asks before every call and remembers the
/// tools you said "always" to, and that pair -- a prompt and a growing set of
/// exceptions -- is the whole of its permission model. A third mode sitting
/// between them would have to invent a rule about which tools are "edits",
/// and the rule would be this app's opinion rather than anything the tools
/// declare.
pub const MODES: &[&str] = &["manual", "bypassPermissions"];
/// What a new llama session asks by default.
///
/// The cautious one, matching the web UI: a model with a shell on somebody's
/// own machine is the case to be wrong about in this direction, and one tap
/// on "always allow" is what makes it bearable afterwards.
pub const DEFAULT_MODE: &str = "manual";
/// The answer that makes an allowance permanent for the session. The tool's
/// name follows it, which is what makes the transcript alone enough to
/// rebuild the set -- see `super::allowed`.
pub const ALWAYS_PREFIX: &str = "Always allow ";
pub const ALLOW_ONCE: &str = "Allow once";
pub const REFUSE: &str = "Don't allow";
/// What the model is told when a call was refused.
///
/// Addressed to the model, not to the reader: it has to understand that the
/// work did not happen and that trying the same call again is not the way
/// round it, or it retries in a loop.
pub const REFUSED: &str = "The person using this session did not allow this call, so it was not run. Do not try it \
again -- say what you were going to do and why it needed that, and let them decide.";
/// What stands in for a call that never finished, when the transcript is read
/// back into a conversation.
///
/// Every tool call in the history owes a result, because that is the shape a
/// chat template renders; a turn stopped between the call and its result
/// leaves one that has none. Saying so is better than inventing an outcome,
/// and better than dropping the call -- which would tell the model it never
/// asked.
pub const UNFINISHED: &str = "This call was interrupted before it produced anything.";
#[cfg(test)]
mod tests {
use super::*;
#[test]
fn a_tool_that_ran_reads_as_its_output() {
assert_eq!(
result_text(&json!({"plain_text_response": "hello\n"})),
"hello\n",
);
}
#[test]
/// The model is told what went wrong, because the model is what has to do
/// something about it -- read a different path, fix the command.
fn a_tool_that_failed_reads_as_its_message() {
assert_eq!(
result_text(&json!({"error": "cannot stat file: /tmp/nope"})),
"cannot stat file: /tmp/nope",
);
}
#[test]
fn anything_else_is_handed_over_as_itself() {
assert_eq!(result_text(&json!({"rows": 2})), "{\"rows\":2}");
}
}
+1
View File
@@ -3980,6 +3980,7 @@ mod tests {
kind: DriverKind::ClaudeCli, kind: DriverKind::ClaudeCli,
command: Some(command.to_string_lossy().into_owned()), command: Some(command.to_string_lossy().into_owned()),
models: Vec::new(), models: Vec::new(),
mcp_servers: Vec::new(),
}, },
])], ])],
..Config::default() ..Config::default()
+2
View File
@@ -1108,6 +1108,7 @@ mod tests {
kind: DriverKind::ClaudeCli, kind: DriverKind::ClaudeCli,
command: None, command: None,
models: vec![], models: vec![],
mcp_servers: Vec::new(),
}], }],
} }
} }
@@ -1171,6 +1172,7 @@ mod tests {
kind: DriverKind::Echo, kind: DriverKind::Echo,
command: None, command: None,
models: vec![], models: vec![],
mcp_servers: Vec::new(),
}]; }];
// A machine with no Claude on it has no Claude limits, and a row // A machine with no Claude on it has no Claude limits, and a row
// reporting on it would be a fact about nothing. Echo included: // reporting on it would be a fact about nothing. Echo included: