Give llama.cpp sessions tools, web search and a model picker
A llama session was a chat box: no tools, a fixed model, no permission
mode, and a model name drawn as the path the file sits at. It now runs the
agent loop itself, which is what the pieces below all hang off.
Tools are `llama-server`'s own (`--tools all`), which that server both
publishes and runs -- `GET /tools` for the definitions, `POST /tools` to
call one. Web search is Exa's MCP server, reached from this backend rather
than from the machine serving the model: that is what llama.cpp's own web
UI does, and it puts the search on the machine with a route out instead of
the one with the GPU. `llama-server`'s `--mcp-servers-json` can only spawn
local commands, so using it would have meant a Node bridge on every
machine that serves a model.
Driving the loop is what makes the permission gate ours. Two modes,
`manual` and `bypassPermissions`, which is what the mechanism has: the web
UI asks before every call and remembers the tools you say "always" to. The
allowances fold back out of the transcript's own answers, so they survive
a restart and a model change without being stored anywhere else.
Also here, because tools made each of them matter:
- **Loading is a state.** A 12 GB model takes twenty seconds to reach
memory and refuses everything until it has; the session used to report
`running` for that whole time, and a message sent meanwhile came back as
an error. It is `loading` now, and the message waits.
- **The model can be changed.** A `llama-server` holds one model, so this
stops it and starts another. The conversation survives because it was
never in the server.
- **Models are named, not pathed.** `general.name` read out of the file
itself -- over ssh too, in the round trip the spawn was already making.
Where two models share a name the file name breaks the tie.
- **`-np 1`, and the MTP draft head where the file has one.** Measured on
the 27B here: 41.5 tok/s plain, 61.4 with `--spec-type draft-mtp` at one
slot, and 28 with it at four -- speculating against a split KV cache is
worse than not speculating. The flag is conditional because asking for a
head that is not there makes `llama-server` exit.
- **A refusal says what to do.** Tool results are thousands of tokens, so
an overrun context is now ordinary; it was "http status: 400" and is now
the server's own "exceeds the available context size, try increasing it".
`GET /machines/{id}/models` is gone: the provider models route answers the
same question, and two answers to one question is how a picker comes to
offer a model the spawn screen does not.
Verified end to end against real models: a tool call asked and allowed, an
Exa search, a shell command, a 27B loaded while a message waited on it, a
model switch mid-session, a second message queued behind a running turn,
and the whole of it again on a session running over ssh.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
1 parent
392cc5413d
commit
ac476ab0c9
23 files changed
+3627
-1096
No files matched your search
@@ -143,18 +143,42 @@ moment you use it — `ANDROID_SERIAL=$(emu serial) ./gradlew …`.
|
||||
|
||||
### Testing llama.cpp and ssh here
|
||||
|
||||
**Both are set up here as of 2026-09-04** and need nothing typed. The
|
||||
prebuilt CPU llama.cpp lives outside the repo at `~/.local/opt/llama.cpp`
|
||||
(the 15 MB `ubuntu-x64` release asset) and is symlinked as
|
||||
`/usr/local/bin/llama-server`, which is what makes **discovery find it over
|
||||
ssh**: `~/.local/bin` is not on the PATH a non-interactive ssh session gets.
|
||||
It resolves its own libraries through `$ORIGIN`, so no `LD_LIBRARY_PATH` is
|
||||
needed. One model is downloaded — `unsloth/Qwen3-0.6B-GGUF/Qwen3-0.6B-Q8_0.gguf`,
|
||||
639 MB under `~/.local/share/ai-app/models` — and answers at usable speed on
|
||||
this VM's 8 cores. **Do not test with a 2-bit quant**: the
|
||||
IQ2_XXS of that model produces fluent nonsense, which reads exactly like a
|
||||
broken driver — `llama-cli` produces the same from the file directly, which
|
||||
is how to tell the two apart in a hurry.
|
||||
**Both are set up here** and need nothing typed. The prebuilt llama.cpp lives
|
||||
outside the repo at `~/.local/opt/llama.cpp-vk` — a **Vulkan** build as of
|
||||
2026-09-19, replacing the CPU one that was there before — and is symlinked as
|
||||
both `~/.local/bin/llama-server` and `/usr/local/bin/llama-server`. The second
|
||||
is what makes **discovery find it over ssh**: `~/.local/bin` is not on the
|
||||
PATH a non-interactive ssh session gets. It resolves its own libraries through
|
||||
`$ORIGIN`, so no `LD_LIBRARY_PATH` is needed.
|
||||
|
||||
Two models are downloaded under `~/.local/share/ai-app/models`:
|
||||
|
||||
- `unsloth/Qwen3-0.6B-GGUF/Qwen3-0.6B-Q8_0.gguf`, 639 MB, loads in ~4s. It
|
||||
calls tools correctly and is the right rig for the driver's shape. Do not
|
||||
judge *answers* by it — asked for the second line of a file it read from
|
||||
line 2 and then named the third.
|
||||
- `ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF/Qwen3.8-27B-GSQ-RCO-IQ3_S-mtp.gguf`,
|
||||
12 GB, ~20s to load, and the only one here with a multi-token-prediction
|
||||
head. It is the rig for anything about `loading` being a state of its own,
|
||||
since 20s is long enough to send into.
|
||||
|
||||
**Do not test with a 2-bit quant**: the IQ2_XXS of the 0.6B produces fluent
|
||||
nonsense, which reads exactly like a broken driver — `llama-cli` produces the
|
||||
same from the file directly, which is how to tell the two apart in a hurry.
|
||||
|
||||
**The GPU is shared and llama-server dies loudly when it runs out.** A second
|
||||
server loading a model while the 27B holds VRAM fails with `radv/amdgpu:
|
||||
Failed to allocate a buffer` / `MESA: error: buffer allocation failed` and
|
||||
exits mid-request. `-ngl 0` runs it on the 8 cores instead, which is the way
|
||||
to test the driver while something else holds the card.
|
||||
|
||||
**Testing tools and MCP without the app**: `llama-server --tools all` publishes
|
||||
its built-in tools at `GET /tools` and runs one at `POST /tools` with
|
||||
`{"tool": …, "params": …}` and an `x-tool-cwd` header — so a whole agent loop
|
||||
is drivable with `curl` and no model at all. The Exa MCP server at
|
||||
`https://mcp.exa.ai/mcp` answers **without an API key** and needs a
|
||||
`User-Agent` header (Cloudflare answers 403 without one, which reads as a
|
||||
refusal rather than a missing header).
|
||||
|
||||
There is no second machine, so **ssh this VM to itself**. That is set up
|
||||
too: the key is `~/.config/ai-app/ssh-self` (its public half is in
|
||||
@@ -212,6 +236,30 @@ where it was instead of half-deleted.
|
||||
|
||||
## Measurements worth not re-taking
|
||||
|
||||
- **`-np 1` is what makes the MTP draft head pay.** Taken 2026-09-19 on the
|
||||
27B above, decode speed for a 300-token reply, from `llama-server`'s own
|
||||
timings rather than the clock:
|
||||
|
||||
| flags | tok/s |
|
||||
| --- | --- |
|
||||
| plain, any `-np` | 41.5 |
|
||||
| `--spec-type draft-mtp -np 1` | 61.4 |
|
||||
| `--spec-type draft-mtp -np 2` (n-max 2) | 65.9 |
|
||||
| `--spec-type draft-mtp`, default `-np` (4 slots) | 28 |
|
||||
|
||||
Draft acceptance is 0.53–0.73 in every case, so the head is working in all
|
||||
of them: what changes is that speculating against a KV cache split four ways
|
||||
is slower than not speculating. The driver passes `-np 1` always, so this is
|
||||
recorded for whoever next sees MTP look broken. `--spec-draft-n-max 2` was
|
||||
worth another 7% in a single sample and is deliberately *not* passed — one
|
||||
sample on a virtualised GPU is not a number to hardcode.
|
||||
|
||||
- **Asking for the head when the file has none is fatal**, not ignored:
|
||||
`context type MTP requested but model doesn't contain MTP layers` and the
|
||||
server exits. Without the flag the same file logs `unused tensor
|
||||
blk.N.nextn.* — ignoring` and runs normally, which is the state to look for
|
||||
when MTP is silently not happening.
|
||||
|
||||
- **What the transcript screen costs to scroll.** Taken 2026-08-30 on the GPU
|
||||
emulator against a real imported transcript with the server at
|
||||
`--delay 120`. Settled and flinging fast, both into fresh history and back
|
||||
|
||||
@@ -47,6 +47,21 @@ Module-by-module intent is in PLAN.md's "Backend layout".
|
||||
readiness poll watches the process as well as the port, since a model that
|
||||
will not load exits in a second and was being reported as "gave up after
|
||||
300s". See PLAN.md's "Transport" and "llama-server management".
|
||||
**A llama session has tools and runs the loop itself** (2026-09-19):
|
||||
`--tools all` gives it `llama-server`'s built-in set, which that server also
|
||||
*runs* (`GET /tools` for the definitions, `POST /tools` to call one), while
|
||||
web search comes from an MCP server this backend connects to directly
|
||||
(`session/llama/mcp.rs`, Exa preset in a discovered provider's
|
||||
`mcpServers`). Driving the loop is what makes the permission gate ours:
|
||||
`manual` asks before every call and remembers a tool you answer
|
||||
"Always allow …" to, `bypassPermissions` never asks, and the allowances are
|
||||
folded back out of the transcript. Three more things fall out of it and are
|
||||
easy to get wrong again — a model change **reloads the server** rather than
|
||||
being refused, since the conversation lives in the transcript rather than in
|
||||
`llama-server`; `-np 1` is always passed, and it is what decides whether the
|
||||
MTP draft head is a 50% speed-up or a 33% loss; and `--spec-type draft-mtp`
|
||||
is conditional on the file actually having a head, because asking for one
|
||||
that is not there makes `llama-server` **exit**.
|
||||
Codex is one persistent `codex app-server --stdio` process per session; its
|
||||
driver uses native turn steering and interruption, persists the protocol
|
||||
state and thread id, and reads subscription limits through the same CLI
|
||||
@@ -310,6 +325,15 @@ written, and the fold uses that same predicate to decide a reply is settled.
|
||||
|
||||
## Things that have bitten
|
||||
|
||||
- **A llama session reports `loading`, and a message sent into it waits.**
|
||||
Before 2026-09-19 the session showed `running` from the moment the process
|
||||
started, so a minute of reading a model off disk was indistinguishable from
|
||||
a minute of thinking -- and anything sent in that window came back as an
|
||||
error, because `llama-server` refuses everything until the model is in
|
||||
memory. `SessionStatus::Loading` is the state and `Shared::await_ready` is
|
||||
the waiting. A driver that reports `Loading` owes the holding as well as the
|
||||
word.
|
||||
|
||||
- **A transcript outlives the enum.** Removing `Event::TaskNote` hours after
|
||||
adding it made every transcript that had recorded one unreadable, so
|
||||
`launch` failed for those sessions and `SessionManager::new` skipped them —
|
||||
|
||||
@@ -345,6 +345,55 @@ deliberate and easy to undo by accident:
|
||||
out the 300s timeout turned the server's own account of the problem into
|
||||
"gave up". The failure carries the tail of `llama-server.log`, which on a
|
||||
remote session is the only copy anybody reading the phone can see.
|
||||
- **Loading is a state of its own** (2026-09-19, `SessionStatus::Loading`).
|
||||
A multi-gigabyte model takes tens of seconds to reach memory and refuses
|
||||
everything until it has, and the session used to report `running` for that
|
||||
whole time — indistinguishable from a model thinking, with the added
|
||||
detail that any message sent meanwhile came back as an error. It is now
|
||||
`loading` on both screens, and a message sent into a load **waits** for it
|
||||
rather than failing. The waiting is the driver's (`Shared::await_ready`, a
|
||||
condvar on a three-state `Serving`), because "there is a process and it is
|
||||
not ready" is a fact only a driver can have. The third state matters as
|
||||
much as the first two: a model that will never load has to answer a waiting
|
||||
message with what went wrong rather than holding it for ever.
|
||||
- **The driver runs the agent loop, and therefore owns the permission gate**
|
||||
(2026-09-19). `llama-server --tools all` *hosts* the built-in tools —
|
||||
`GET /tools` is their definitions, `POST /tools` runs one — but it does not
|
||||
drive a conversation: a completion comes back with tool calls in it and
|
||||
stops. So the loop is here, which is what puts "may I run this?" somewhere
|
||||
a phone can answer it. Two modes, `manual` and `bypassPermissions`, which
|
||||
is what the mechanism actually has: llama.cpp's own web UI asks before
|
||||
every call and remembers the tools you said "always" to, and a third mode
|
||||
between them would have to invent a rule about which tools count as edits.
|
||||
The allowances are folded out of the transcript's `Answered` events, like
|
||||
everything else this driver remembers, which is why the answer carries the
|
||||
tool's name in it.
|
||||
- **Tools run where the model does; MCP runs here** (2026-09-19). The
|
||||
built-in tools are the far machine's, for the same reason the model file
|
||||
is — they act on that machine's disk. An MCP server is reached from *this*
|
||||
backend instead (`session/llama/mcp.rs`), which is both what llama.cpp's
|
||||
own web UI does (it connects to `https://mcp.exa.ai/mcp` from the browser)
|
||||
and the right side to be on: a web search wants the machine with a route
|
||||
out, not the machine with the GPU. `llama-server`'s own `--mcp-servers-json`
|
||||
is deliberately not used — it can only spawn local commands, so a remote
|
||||
server would mean a Node bridge on whichever machine serves the model.
|
||||
- **A model change reloads the server rather than being refused** (2026-09-19).
|
||||
A `llama-server` holds one model, so switching stops it and starts another;
|
||||
the conversation survives because the conversation was never in the server.
|
||||
What is lost is the prompt cache, which is exactly what the phone already
|
||||
warns about before a switch.
|
||||
- **One slot, and the draft head where the file has one** (measured
|
||||
2026-09-19). `-np 1` always: a session is one conversation making one
|
||||
request at a time, so the other three slots `llama-server` picks on its own
|
||||
are context this session could have had. It is also what decides whether
|
||||
multi-token prediction pays — on the 27B here, **41.5 tok/s** plain at any
|
||||
slot count, **61.4** with `--spec-type draft-mtp` at one slot, and **28**
|
||||
with the head at four. Speculating against a split KV cache is worse than
|
||||
not speculating, and it reads exactly like the head being broken.
|
||||
The flag is conditional because it must be: asked for on a model without a
|
||||
head, `llama-server` exits. `crate::gguf::has_mtp_head` reads the answer out
|
||||
of the file — on the machine that will serve it, in the round trip the spawn
|
||||
was already making — and `params["speculative"] = "off"` is the way out.
|
||||
|
||||
### Models (2026-08-28)
|
||||
|
||||
@@ -1386,8 +1435,10 @@ verified by running it, matching dev-updater's posture.
|
||||
it truncates old KV cache entries, which is silent forgetting with no
|
||||
summary, and it corrupts the harness's view of what the model knows. Fine
|
||||
as a server-side safety net; not memory management.
|
||||
- **Remote llama-server** needs its port forwarded (`ssh -L`) and is not
|
||||
built; such a session is refused rather than misdirected.
|
||||
- **MCP servers are configured in `config.ron`, not from the phone**
|
||||
(2026-09-19). `mcpServers` on a llama provider, with Exa preset on a newly
|
||||
discovered one. A phone screen for them is the obvious next step and was
|
||||
deliberately left out of the change that added them.
|
||||
- **Claude sessions over ssh need the remote machine logged in to Claude.**
|
||||
Usage reporting reads each machine's own credentials, and the Machines tab
|
||||
can run that machine's CLI login without requiring an interactive SSH shell.
|
||||
|
||||
@@ -1334,7 +1334,18 @@ fun deleteSession(settings: ServerSettings, sessionId: String, deleteForeign: Bo
|
||||
// Browsing is proxied by the server rather than done here, because this app trusts exactly one
|
||||
// certificate and has no general internet trust to spend on huggingface.co.
|
||||
|
||||
data class LocalModel(val key: String, val repo: String, val file: String, val bytes: Long)
|
||||
data class LocalModel(
|
||||
val key: String,
|
||||
val repo: String,
|
||||
val file: String,
|
||||
val bytes: Long,
|
||||
/**
|
||||
* What the file itself says it is called, or null when it does not say. Not what to draw: see
|
||||
* the server's `models::labels`, which needs the whole list to decide -- two quantisations of
|
||||
* one model share a name.
|
||||
*/
|
||||
val name: String?,
|
||||
)
|
||||
|
||||
/**
|
||||
* A download in flight or finished. [total] is null when the server never said how big the file is
|
||||
@@ -1371,36 +1382,28 @@ private fun parseDownload(o: JSONObject) =
|
||||
)
|
||||
|
||||
/**
|
||||
* The models on one machine, which is the list a llama.cpp session there can choose from.
|
||||
* One model a picker can offer.
|
||||
*
|
||||
* Not [fetchModels], which is what the *backend* has downloaded. A session serves its model from
|
||||
* the machine it runs on, so for a machine reached over ssh those are two different lists -- and
|
||||
* offering the backend's would name files that are not there, turning a choice that cannot work
|
||||
* into a session that fails when it tries to load one.
|
||||
* Two fields because for one provider they differ: a llama.cpp session names its model by the path
|
||||
* it lives at and reads it as the name its own metadata gives it. Every other provider's [id] is
|
||||
* already what a person calls it, and the server says so by repeating it -- which is what keeps
|
||||
* every picker here free of a branch on the session kind.
|
||||
*/
|
||||
fun fetchMachineModels(settings: ServerSettings, machineId: String): List<LocalModel> =
|
||||
requestFromServer(settings, "/machines/${machineId.urlEncoded()}/models") { connection ->
|
||||
JSONArray(connection.inputStream.bufferedReader().readText()).mapObjects { m ->
|
||||
LocalModel(
|
||||
key = m.getString("key"),
|
||||
repo = m.getString("repo"),
|
||||
file = m.getString("file"),
|
||||
bytes = m.getLong("bytes"),
|
||||
)
|
||||
}
|
||||
}
|
||||
data class OfferedModel(val id: String, val label: String)
|
||||
|
||||
/** The current model catalog for one CLI provider on the machine where it runs. */
|
||||
/** The current model catalog for one provider on the machine where it runs. */
|
||||
fun fetchProviderModels(
|
||||
settings: ServerSettings,
|
||||
machineId: String,
|
||||
provider: String,
|
||||
): List<String> =
|
||||
): List<OfferedModel> =
|
||||
requestFromServer(
|
||||
settings,
|
||||
"/machines/${machineId.urlEncoded()}/providers/${provider.urlEncoded()}/models",
|
||||
) { connection ->
|
||||
JSONArray(connection.inputStream.bufferedReader().readText()).strings()
|
||||
JSONArray(connection.inputStream.bufferedReader().readText()).mapObjects { m ->
|
||||
OfferedModel(id = m.getString("id"), label = m.getString("label"))
|
||||
}
|
||||
}
|
||||
|
||||
fun fetchModels(settings: ServerSettings): Models =
|
||||
@@ -1414,6 +1417,7 @@ fun fetchModels(settings: ServerSettings): Models =
|
||||
repo = m.getString("repo"),
|
||||
file = m.getString("file"),
|
||||
bytes = m.getLong("bytes"),
|
||||
name = m.optString("name").ifEmpty { null },
|
||||
)
|
||||
},
|
||||
downloads = body.getJSONArray("downloads").mapObjects(::parseDownload),
|
||||
|
||||
@@ -325,7 +325,8 @@ fun parseSeqEvent(json: String): SeqEvent {
|
||||
* first time the server grows a state, and the drift would be a reply that never splits or one
|
||||
* split mid-stream.
|
||||
*/
|
||||
fun sessionWorking(state: String): Boolean = state == "running" || state == "compacting"
|
||||
fun sessionWorking(state: String): Boolean =
|
||||
state == "running" || state == "compacting" || state == "loading"
|
||||
|
||||
/** Whether the latest events still say this session needs an explicit provider login. */
|
||||
internal fun authenticationPromptAfter(open: Boolean, event: SessionEvent): Boolean =
|
||||
|
||||
@@ -21,12 +21,25 @@ const val DEFAULT_MODEL = "default"
|
||||
* one model rather than one model from another. Anything that does not look like that is returned
|
||||
* untouched.
|
||||
*
|
||||
* A llama.cpp session's model is not an identifier at all -- it is `owner/repo/file.gguf`, where
|
||||
* the file was downloaded from -- so what is kept is the file, which is the part that tells two
|
||||
* models apart, and the extension goes with the directories. The model's *own* name is better still
|
||||
* and is not derivable here: it is inside the file, and only the server has ever opened it. Where a
|
||||
* screen has the server's answer it should prefer it; this is the floor under every screen that
|
||||
* does not.
|
||||
*
|
||||
* A display decision, not a correction: the full name is what the session reports.
|
||||
*/
|
||||
fun modelLabel(model: String?): String {
|
||||
val name = model?.takeIf { it.isNotBlank() } ?: return DEFAULT_MODEL
|
||||
if (name.endsWith(GGUF)) {
|
||||
return name.substringAfterLast('/').removeSuffix(GGUF)
|
||||
}
|
||||
return name.removePrefix("claude-").replace(DATED_SUFFIX, "")
|
||||
}
|
||||
|
||||
/** A trailing `-YYYYMMDD`, which is how these identifiers carry their release date. */
|
||||
private val DATED_SUFFIX = Regex("""-\d{8}$""")
|
||||
|
||||
/** What every model a llama.cpp session can run is stored as. */
|
||||
private const val GGUF = ".gguf"
|
||||
@@ -302,7 +302,7 @@ fun SessionScreen(
|
||||
var permissionMode by remember { mutableStateOf(summary.permissionMode ?: "auto") }
|
||||
// The models this provider actually offers, asked of the server rather than listed here: a
|
||||
// hardcoded list is a claim about a machine.
|
||||
var offeredModels by remember { mutableStateOf<List<String>>(emptyList()) }
|
||||
var offeredModels by remember { mutableStateOf<List<OfferedModel>>(emptyList()) }
|
||||
var offeredPermissionModes by remember { mutableStateOf<List<String>>(emptyList()) }
|
||||
val lifecycleOwner = LocalLifecycleOwner.current
|
||||
// The resume cursor, written from the stream's IO thread.
|
||||
@@ -1194,6 +1194,17 @@ fun SessionScreen(
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* What to call a model on screen.
|
||||
*
|
||||
* The provider's own answer where it has one, because only the server can have it: a llama
|
||||
* model is identified by the path it lives at and named by what is written inside the file, and
|
||||
* the phone has never opened that file. [modelLabel] is the fallback and the right one for the
|
||||
* rest -- a coding CLI's identifier already is its name.
|
||||
*/
|
||||
fun label(id: String?): String =
|
||||
offeredModels.firstOrNull { it.id == id }?.label ?: modelLabel(id)
|
||||
|
||||
// Only for the model picker, which a subagent does not have.
|
||||
if (!isSubagent) {
|
||||
LaunchedEffect(summary.machine, summary.provider) {
|
||||
@@ -1889,8 +1900,8 @@ fun SessionScreen(
|
||||
) {
|
||||
pendingModel?.let { chosen ->
|
||||
ModelSwitchWarning(
|
||||
from = modelLabel(model),
|
||||
to = modelLabel(chosen),
|
||||
from = label(model),
|
||||
to = label(chosen),
|
||||
onDismiss = { pendingModel = null },
|
||||
onConfirm = {
|
||||
pendingModel = null
|
||||
@@ -2013,22 +2024,28 @@ fun SessionScreen(
|
||||
) {
|
||||
if (offeredModels.isNotEmpty()) {
|
||||
PickerButton(
|
||||
current = modelLabel(model),
|
||||
current = label(model),
|
||||
// What the machine offers, plus the state a session is in when
|
||||
// it has chosen none of them. The button has always been able
|
||||
// to
|
||||
// say "default"; until this the list could not, so leaving it
|
||||
// was a one-way trip.
|
||||
options = listOf(DEFAULT_MODEL) + offeredModels,
|
||||
options =
|
||||
listOf(DEFAULT_MODEL) + offeredModels.map { it.label },
|
||||
// Not set here. The button follows what the session reports it
|
||||
// is set to, which arrives a moment later and is sometimes a
|
||||
// different answer -- a name the CLI resolved, or no change at
|
||||
// all on a provider whose model is fixed. Asked about first,
|
||||
// unless there is nothing to lose by it -- see
|
||||
// [ModelSwitchWarning].
|
||||
onPick = { chosen ->
|
||||
onPick = { picked ->
|
||||
// Back to the id, because that is what the server resolves
|
||||
// and it is not always the word on the chip.
|
||||
val chosen =
|
||||
offeredModels.firstOrNull { it.label == picked }?.id
|
||||
?: picked
|
||||
if (
|
||||
modelLabel(chosen) == modelLabel(model) ||
|
||||
label(chosen) == label(model) ||
|
||||
!worthWarningAbout(status, contextTokens, items)
|
||||
) {
|
||||
act { setSessionModel(settings, summary.id, chosen) }
|
||||
|
||||
@@ -22,6 +22,10 @@ fun sessionStatusWord(status: String, subagent: Boolean = false): String =
|
||||
"idle" -> "idle"
|
||||
"running" -> "running"
|
||||
"compacting" -> "compacting"
|
||||
// Not "running": a model coming off disk is not a model answering, and the difference is
|
||||
// minutes. Said in its own word so a first message that waits is explained rather than
|
||||
// looking like a session that has stopped responding. See `SessionStatus::Loading`.
|
||||
"loading" -> "loading"
|
||||
// Its own word, because the state it is easily mistaken for means the opposite: "idle"
|
||||
// invites the reader to type something, and a waiting session is going to carry on without
|
||||
// them. See `SessionStatus::Waiting`.
|
||||
@@ -53,6 +57,9 @@ fun sessionStatusColour(status: String): Color =
|
||||
"awaitingInput" -> awaitingColor
|
||||
"running" -> runningColor
|
||||
"compacting" -> commandColor
|
||||
// The same accent as the other states that are busy on their own account, because that is
|
||||
// what this is: something is happening and nothing is wanted from the reader.
|
||||
"loading" -> commandColor
|
||||
"waiting" -> waitingColor
|
||||
else -> MaterialTheme.colorScheme.onSurfaceVariant
|
||||
}
|
||||
@@ -58,7 +58,7 @@ fun SpawnScreen(
|
||||
var providerName by remember { mutableStateOf<String?>(null) }
|
||||
var title by remember { mutableStateOf("") }
|
||||
var model by remember { mutableStateOf("") }
|
||||
var providerModels by remember { mutableStateOf<List<String>>(emptyList()) }
|
||||
var providerModels by remember { mutableStateOf<List<OfferedModel>>(emptyList()) }
|
||||
var providerModelsLoading by remember { mutableStateOf(false) }
|
||||
var providerModelsError by remember { mutableStateOf<String?>(null) }
|
||||
var cwd by remember { mutableStateOf("") }
|
||||
@@ -73,12 +73,13 @@ fun SpawnScreen(
|
||||
// Only the spawn's own failure. The fetch's lives in `options`: this one leaves a filled-in
|
||||
// form worth keeping, and that one leaves nothing to fill in.
|
||||
var spawnError by remember { mutableStateOf<String?>(null) }
|
||||
// GGUFs on the chosen machine, for a llama provider to choose between. Kept separate from a
|
||||
// coding CLI's provider catalog and refetched when the machine changes.
|
||||
var models by remember { mutableStateOf<List<LocalModel>>(emptyList()) }
|
||||
var modelKey by remember { mutableStateOf<String?>(null) }
|
||||
var contextSize by remember { mutableStateOf("") }
|
||||
var temperature by remember { mutableStateOf("") }
|
||||
// Whether a model that carries a multi-token-prediction head drafts with it. Left to the
|
||||
// server by default, which turns it on exactly where the file has one -- see `SPECULATIVE` in
|
||||
// the llama driver. Here so a machine where drafting does not pay has a way out that is not
|
||||
// an edit to config.ron.
|
||||
var speculative by remember { mutableStateOf(SPECULATIVE_AUTO) }
|
||||
|
||||
LaunchedEffect(Unit) {
|
||||
// Separate from the machines fetch below and deliberately not fatal: failing to learn the
|
||||
@@ -126,17 +127,6 @@ fun SpawnScreen(
|
||||
is LoadState.Loaded -> state.value
|
||||
}
|
||||
val machine = machines.firstOrNull { it.name == machineName }
|
||||
// Whichever machine is chosen now, asked again when that changes. The old machine's list
|
||||
// is dropped first rather than left on screen: a file name from another machine looks
|
||||
// exactly like one from this one.
|
||||
LaunchedEffect(machine?.id) {
|
||||
models = emptyList()
|
||||
modelKey = null
|
||||
val id = machine?.id ?: return@LaunchedEffect
|
||||
models =
|
||||
runCatching { withContext(Dispatchers.IO) { fetchMachineModels(settings, id) } }
|
||||
.getOrDefault(emptyList())
|
||||
}
|
||||
val current = machine?.providers?.firstOrNull { it.name == providerName }
|
||||
// Coding CLIs take a working directory, model, permission mode and thinking level. Keying
|
||||
// the extra fields on the kind rather than the provider name keeps a second installation
|
||||
@@ -145,25 +135,35 @@ fun SpawnScreen(
|
||||
val isCodex = current?.kind == "codex_cli"
|
||||
val isCodingCli = isClaude || isCodex
|
||||
val isLlama = current?.kind == "llama_cpp"
|
||||
// Echo is the only kind with nothing to choose between.
|
||||
val offersModels = isCodingCli || isLlama
|
||||
// Where a session's tools act, which is the only thing a working directory decides.
|
||||
val takesCwd = isCodingCli || isLlama
|
||||
|
||||
// Whichever machine and provider are chosen now, asked again when either changes. The
|
||||
// previous answer is dropped first rather than left on screen: a model name from another
|
||||
// machine looks exactly like one from this one.
|
||||
LaunchedEffect(machine?.id, current?.name) {
|
||||
model = ""
|
||||
providerModels = emptyList()
|
||||
providerModelsError = null
|
||||
permissionMode = current?.defaultPermissionMode.orEmpty()
|
||||
if (isCodingCli) {
|
||||
providerModelsLoading = true
|
||||
try {
|
||||
providerModels =
|
||||
withContext(Dispatchers.IO) {
|
||||
fetchProviderModels(settings, machine.id, current.name)
|
||||
}
|
||||
} catch (e: ApiException) {
|
||||
providerModelsError = e.message
|
||||
} finally {
|
||||
providerModelsLoading = false
|
||||
}
|
||||
} else {
|
||||
// Every kind that offers models at all, not only the coding CLIs: a llama provider
|
||||
// answers with the GGUFs on the machine it runs on, through the same call. One
|
||||
// question with one answer is what keeps the picker free of a branch on the kind.
|
||||
if (machine == null || current == null || !offersModels) {
|
||||
providerModelsLoading = false
|
||||
return@LaunchedEffect
|
||||
}
|
||||
providerModelsLoading = true
|
||||
try {
|
||||
providerModels =
|
||||
withContext(Dispatchers.IO) {
|
||||
fetchProviderModels(settings, machine.id, current.name)
|
||||
}
|
||||
} catch (e: ApiException) {
|
||||
providerModelsError = e.message
|
||||
} finally {
|
||||
providerModelsLoading = false
|
||||
}
|
||||
}
|
||||
@@ -220,29 +220,70 @@ fun SpawnScreen(
|
||||
modifier = Modifier.fillMaxWidth(),
|
||||
)
|
||||
|
||||
if (isLlama) {
|
||||
// A llama session names one of the models on the machine it will run on, so the
|
||||
// choice is that list rather than free text -- a name that is not on that machine's
|
||||
// disk is a session that cannot start.
|
||||
if (models.isEmpty()) {
|
||||
Text(
|
||||
"No models on ${machine.name}. The Models screen downloads " +
|
||||
"to the backend; another machine needs the file put there itself.",
|
||||
style = MaterialTheme.typography.bodyMedium,
|
||||
color = MaterialTheme.colorScheme.onSurfaceVariant,
|
||||
)
|
||||
} else {
|
||||
ChipGroup(
|
||||
label = "Model",
|
||||
// The file, not the whole key: the repository is the same for every
|
||||
// quantisation of a model, so the file name is what tells two of them apart.
|
||||
options = models.map { it.file },
|
||||
selected = models.firstOrNull { it.key == modelKey }?.file,
|
||||
onSelect = { file -> modelKey = models.first { it.file == file }.key },
|
||||
)
|
||||
if (offersModels) {
|
||||
when {
|
||||
providerModelsLoading ->
|
||||
Text(
|
||||
"Loading model choices…",
|
||||
style = MaterialTheme.typography.bodySmall,
|
||||
color = MaterialTheme.colorScheme.onSurfaceVariant,
|
||||
)
|
||||
providerModelsError != null ->
|
||||
Text(
|
||||
"Model choices unavailable: $providerModelsError",
|
||||
style = MaterialTheme.typography.bodySmall,
|
||||
color = MaterialTheme.colorScheme.error,
|
||||
)
|
||||
// A llama session cannot start without one, so this says what to do about it
|
||||
// rather than only that there is nothing -- the models it needs are on the
|
||||
// machine that will serve them, which is not always this backend.
|
||||
providerModels.isEmpty() && isLlama ->
|
||||
Text(
|
||||
"No models on ${machine?.name}. The Models screen downloads " +
|
||||
"to the backend; another machine needs the file put there itself.",
|
||||
style = MaterialTheme.typography.bodyMedium,
|
||||
color = MaterialTheme.colorScheme.onSurfaceVariant,
|
||||
)
|
||||
providerModels.isEmpty() ->
|
||||
Text(
|
||||
"This machine reported no selectable models.",
|
||||
style = MaterialTheme.typography.bodySmall,
|
||||
color = MaterialTheme.colorScheme.onSurfaceVariant,
|
||||
)
|
||||
else -> {
|
||||
Spacer(Modifier.height(16.dp))
|
||||
ChipGroup(
|
||||
label = "Model",
|
||||
// The label, and the id is what is sent: for a llama model those differ,
|
||||
// since it is chosen by path and named by what is inside the file.
|
||||
options = providerModels.map { it.label },
|
||||
selected = providerModels.firstOrNull { it.id == model }?.label,
|
||||
onSelect = { chosen ->
|
||||
val id = providerModels.first { it.label == chosen }.id
|
||||
// A llama session has to have one, so choosing the same chip twice
|
||||
// must not clear it -- there is nothing to fall back to.
|
||||
model = if (model == id && !isLlama) "" else id
|
||||
},
|
||||
)
|
||||
}
|
||||
}
|
||||
Spacer(Modifier.height(16.dp))
|
||||
}
|
||||
|
||||
if (isCodingCli) {
|
||||
// Free text as well as the chips above: the catalog is a shortcut, and a CLI will
|
||||
// take a name it did not list.
|
||||
OutlinedTextField(
|
||||
value = model,
|
||||
onValueChange = { model = it },
|
||||
label = { Text("Model (blank = the CLI's default)") },
|
||||
singleLine = true,
|
||||
modifier = Modifier.fillMaxWidth(),
|
||||
)
|
||||
Spacer(Modifier.height(16.dp))
|
||||
}
|
||||
|
||||
if (isLlama) {
|
||||
OutlinedTextField(
|
||||
value = contextSize,
|
||||
onValueChange = { contextSize = it },
|
||||
@@ -260,48 +301,21 @@ fun SpawnScreen(
|
||||
modifier = Modifier.fillMaxWidth(),
|
||||
)
|
||||
Spacer(Modifier.height(16.dp))
|
||||
}
|
||||
|
||||
if (isCodingCli) {
|
||||
when {
|
||||
providerModelsLoading ->
|
||||
Text(
|
||||
"Loading model choices…",
|
||||
style = MaterialTheme.typography.bodySmall,
|
||||
color = MaterialTheme.colorScheme.onSurfaceVariant,
|
||||
)
|
||||
providerModelsError != null ->
|
||||
Text(
|
||||
"Model choices unavailable: $providerModelsError",
|
||||
style = MaterialTheme.typography.bodySmall,
|
||||
color = MaterialTheme.colorScheme.error,
|
||||
)
|
||||
providerModels.isEmpty() ->
|
||||
Text(
|
||||
"This machine reported no selectable models.",
|
||||
style = MaterialTheme.typography.bodySmall,
|
||||
color = MaterialTheme.colorScheme.onSurfaceVariant,
|
||||
)
|
||||
else -> {
|
||||
Spacer(Modifier.height(16.dp))
|
||||
ChipGroup(
|
||||
label = "Model",
|
||||
options = providerModels,
|
||||
selected = model.ifEmpty { null },
|
||||
onSelect = { chosen -> model = if (model == chosen) "" else chosen },
|
||||
)
|
||||
}
|
||||
}
|
||||
Spacer(Modifier.height(8.dp))
|
||||
OutlinedTextField(
|
||||
value = model,
|
||||
onValueChange = { model = it },
|
||||
label = { Text("Model (blank = the CLI's default)") },
|
||||
singleLine = true,
|
||||
modifier = Modifier.fillMaxWidth(),
|
||||
// Said as what it is rather than as "MTP": the reader is choosing whether the session
|
||||
// goes faster, and most models have nothing to turn on here at all.
|
||||
ChipGroup(
|
||||
label = "Speculative decoding (models that carry a draft head)",
|
||||
options = listOf(SPECULATIVE_AUTO, SPECULATIVE_OFF),
|
||||
selected = speculative,
|
||||
onSelect = { speculative = it },
|
||||
)
|
||||
Spacer(Modifier.height(16.dp))
|
||||
}
|
||||
|
||||
// Every session whose tools act on files needs one, which is both kinds that have
|
||||
// tools -- a llama session's built-in tools run in it exactly as a CLI's do.
|
||||
if (takesCwd) {
|
||||
OutlinedTextField(
|
||||
value = cwd,
|
||||
onValueChange = { cwd = it },
|
||||
@@ -311,7 +325,11 @@ fun SpawnScreen(
|
||||
modifier = Modifier.fillMaxWidth(),
|
||||
)
|
||||
Spacer(Modifier.height(16.dp))
|
||||
}
|
||||
|
||||
// Offered wherever the provider has modes, rather than where this screen believes it
|
||||
// does: the server is what knows, and llama.cpp grew them without this line changing.
|
||||
if (current != null && current.permissionModes.isNotEmpty()) {
|
||||
ChipGroup(
|
||||
label = "Permissions",
|
||||
options = current.permissionModes,
|
||||
@@ -319,7 +337,9 @@ fun SpawnScreen(
|
||||
onSelect = { permissionMode = it },
|
||||
)
|
||||
Spacer(Modifier.height(16.dp))
|
||||
}
|
||||
|
||||
if (isCodingCli) {
|
||||
// Says what it does to *later* spawns as well, because it does: the level chosen here
|
||||
// is stored as the default, which is the whole way that default is set. A picker that
|
||||
// quietly changed a global would be the same control with the fact left out.
|
||||
@@ -363,11 +383,9 @@ fun SpawnScreen(
|
||||
machine = machine.id,
|
||||
provider = chosen.name,
|
||||
title = title.trim(),
|
||||
model =
|
||||
if (isLlama) modelKey
|
||||
else model.trim().takeIf { isCodingCli },
|
||||
cwd = cwd.trim().takeIf { isCodingCli },
|
||||
permissionMode = permissionMode.takeIf { isCodingCli },
|
||||
model = model.trim().takeIf { offersModels },
|
||||
cwd = cwd.trim().takeIf { takesCwd },
|
||||
permissionMode = permissionMode.takeIf { it.isNotEmpty() },
|
||||
effort = effort.takeIf { isCodingCli },
|
||||
// Sent only when set, so blank means "whatever llama.cpp does
|
||||
// by default" rather than a zero.
|
||||
@@ -382,6 +400,13 @@ fun SpawnScreen(
|
||||
.trim()
|
||||
.takeIf { it.isNotEmpty() }
|
||||
?.let { put("temperature", it) }
|
||||
// Only the choice that changes anything: "auto"
|
||||
// is the absence of the setting, not a value of
|
||||
// it, so a session spawned without an opinion
|
||||
// carries none.
|
||||
if (speculative == SPECULATIVE_OFF) {
|
||||
put("speculative", "off")
|
||||
}
|
||||
}
|
||||
},
|
||||
)
|
||||
@@ -393,13 +418,23 @@ fun SpawnScreen(
|
||||
}
|
||||
}
|
||||
},
|
||||
enabled = !busy && current != null && !(isLlama && modelKey == null),
|
||||
// A llama session names the file to load, so there is nothing to spawn without one.
|
||||
enabled = !busy && current != null && !(isLlama && model.isEmpty()),
|
||||
) {
|
||||
Text(if (busy) "Spawning..." else "Spawn")
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Leave the draft head to the server, which uses one wherever the model file has one. Spelled the
|
||||
* same as the absence of the `speculative` parameter, because that is what it means.
|
||||
*/
|
||||
private const val SPECULATIVE_AUTO = "auto"
|
||||
|
||||
/** The `speculative` parameter's only other value; see the llama driver's `SPECULATIVE`. */
|
||||
private const val SPECULATIVE_OFF = "off"
|
||||
|
||||
/**
|
||||
* A labeled row of choices that wraps onto as many lines as it needs.
|
||||
*
|
||||
|
||||
+37
-2
@@ -92,6 +92,32 @@ pub struct ProviderConfig {
|
||||
/// too; this is a shortcut list, not a restriction.
|
||||
#[serde(default, skip_serializing_if = "Vec::is_empty")]
|
||||
pub models: Vec<String>,
|
||||
/// MCP servers whose tools this provider's sessions can use, on top of
|
||||
/// whatever the provider runs itself.
|
||||
///
|
||||
/// On the provider rather than the machine, because it is a statement
|
||||
/// about what a session can do rather than about where it runs -- and
|
||||
/// because only a driver that runs its own agent loop can use one. Today
|
||||
/// that is llama.cpp; the coding CLIs have their own MCP configuration
|
||||
/// and this would be a second, quieter answer to the same question.
|
||||
#[serde(default, skip_serializing_if = "Vec::is_empty")]
|
||||
pub mcp_servers: Vec<McpServerConfig>,
|
||||
}
|
||||
|
||||
/// An MCP server reached over HTTP.
|
||||
///
|
||||
/// A URL and nothing else: this backend connects to remote servers rather than
|
||||
/// spawning local ones, so there is no command, no arguments and no
|
||||
/// environment to configure. See `session::llama::mcp` for why that is the
|
||||
/// shape -- in short, it is what llama.cpp's own web UI does, and it keeps the
|
||||
/// tools on the machine with a route out rather than the machine with the GPU.
|
||||
#[derive(Debug, Clone, Serialize, Deserialize)]
|
||||
#[serde(rename_all = "camelCase")]
|
||||
pub struct McpServerConfig {
|
||||
/// Prefixes every tool this server offers, so two servers with a `search`
|
||||
/// are two tools. Also what a failure to connect is named by.
|
||||
pub name: String,
|
||||
pub url: String,
|
||||
}
|
||||
|
||||
impl ProviderConfig {
|
||||
@@ -278,7 +304,11 @@ impl DriverKind {
|
||||
match self {
|
||||
Self::ClaudeCli => &["manual", "acceptEdits", "auto", "bypassPermissions", "plan"],
|
||||
Self::CodexCli => &["workspace-write", "read-only", "danger-full-access"],
|
||||
Self::Echo | Self::LlamaCpp => &[],
|
||||
// Named by the driver that enforces them rather than repeated
|
||||
// here: this list and the one the gate matches on being two
|
||||
// literals is how a mode comes to be offered and then refused.
|
||||
Self::LlamaCpp => crate::session::llama::PERMISSION_MODES,
|
||||
Self::Echo => &[],
|
||||
}
|
||||
}
|
||||
|
||||
@@ -287,7 +317,8 @@ impl DriverKind {
|
||||
match self {
|
||||
Self::ClaudeCli => Some("auto"),
|
||||
Self::CodexCli => Some("workspace-write"),
|
||||
Self::Echo | Self::LlamaCpp => None,
|
||||
Self::LlamaCpp => Some(crate::session::llama::DEFAULT_PERMISSION_MODE),
|
||||
Self::Echo => None,
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -499,6 +530,7 @@ impl Config {
|
||||
kind: DriverKind::Echo,
|
||||
command: None,
|
||||
models: Vec::new(),
|
||||
mcp_servers: Vec::new(),
|
||||
}
|
||||
}
|
||||
|
||||
@@ -579,6 +611,7 @@ mod tests {
|
||||
kind: DriverKind::ClaudeCli,
|
||||
command: Some("/usr/bin/claude".to_string()),
|
||||
models: Vec::new(),
|
||||
mcp_servers: Vec::new(),
|
||||
},
|
||||
]),
|
||||
MachineConfig {
|
||||
@@ -597,6 +630,7 @@ mod tests {
|
||||
kind: DriverKind::ClaudeCli,
|
||||
command: None,
|
||||
models: vec!["haiku".to_string()],
|
||||
mcp_servers: Vec::new(),
|
||||
}],
|
||||
},
|
||||
],
|
||||
@@ -734,6 +768,7 @@ sessions: [(
|
||||
kind: DriverKind::ClaudeCli,
|
||||
command: Some("/usr/bin/claude".to_string()),
|
||||
models: Vec::new(),
|
||||
mcp_servers: Vec::new(),
|
||||
},
|
||||
]);
|
||||
assert_eq!(
|
||||
|
||||
@@ -0,0 +1,359 @@
|
||||
//! Just enough of the GGUF container to read a model's own name out of it.
|
||||
//!
|
||||
//! A `.gguf` file opens with a key/value table, and `general.name` in it is
|
||||
//! what the people who published the model called it -- "Qwen3-0.6B",
|
||||
//! "Qwen3.8-27B GSQ-RCO". Everything else this server knows a model by is
|
||||
//! filesystem trivia: `owner/repo/file.gguf` is where it was downloaded from,
|
||||
//! which is an address rather than a name, and on a phone it is a line of path
|
||||
//! where a word would do.
|
||||
//!
|
||||
//! **Read as far as the answer and no further.** The same table holds the
|
||||
//! tokenizer, which for a modern model is a 150,000-entry string array and
|
||||
//! most of several megabytes; `general.*` is written first by every converter
|
||||
//! in practice, so stopping at the name costs a few kilobytes instead. That is
|
||||
//! what makes this affordable to run over every model in a directory, and what
|
||||
//! lets the remote case work from a bounded prefix of the file rather than the
|
||||
//! whole of it.
|
||||
//!
|
||||
//! Anything unreadable is [`None`] rather than an error, at every level. A
|
||||
//! model with no name, a truncated prefix, a container version this does not
|
||||
//! know and a file that is not GGUF at all are one answer here -- "this file
|
||||
//! does not tell us" -- and the caller has a file name to fall back on. There
|
||||
//! is nothing a reader could do with the distinction.
|
||||
|
||||
use std::io::Read;
|
||||
|
||||
/// How many bytes of a model file are worth fetching to look for its name.
|
||||
///
|
||||
/// Only the remote path needs a number: a local read stops when it finds the
|
||||
/// key, but a file on another machine has to be asked for a fixed amount
|
||||
/// before anything can be parsed. Measured 2026-09-19 against the two models
|
||||
/// on this machine, `general.name` ends at byte **130** and **94** -- every
|
||||
/// converter writes `general.*` before the tokenizer arrays that make up the
|
||||
/// rest of the table. 8 KiB is two orders of magnitude of slack for that and
|
||||
/// still makes listing a directory of models one round trip's worth of bytes
|
||||
/// rather than a download, which is what decides the number: this is paid per
|
||||
/// model every time a spawn screen opens.
|
||||
pub const PREFIX_BYTES: u64 = 8 * 1024;
|
||||
|
||||
/// The longest string this will allocate for, so a corrupt length field
|
||||
/// cannot ask for a gigabyte. Longer than any key or `general.*` value.
|
||||
const MAX_STRING: u64 = 64 * 1024;
|
||||
|
||||
/// What `general.name` says, or `None` for every way of not finding out.
|
||||
///
|
||||
/// `read` is consumed only as far as the key: pass a file to read a local
|
||||
/// model, or a cursor over a prefix to read one whose bytes came from
|
||||
/// somewhere else.
|
||||
pub fn name(read: &mut impl Read) -> Option<String> {
|
||||
match find(read, |key| key == "general.name") {
|
||||
Some((STRING, read)) => string(read),
|
||||
_ => None,
|
||||
}
|
||||
}
|
||||
|
||||
/// Whether this model carries a multi-token-prediction head.
|
||||
///
|
||||
/// Worth asking because `llama-server` **exits** when told to use one that is
|
||||
/// not there -- `--spec-type draft-mtp` on a plain model is "context type MTP
|
||||
/// requested but model doesn't contain MTP layers" and then a server that
|
||||
/// never comes up. So the flag can only be passed once this has said yes, and
|
||||
/// a `false` here is the same answer as an unreadable file: don't ask for it.
|
||||
///
|
||||
/// Matched on the key's tail rather than its whole name, because the key is
|
||||
/// prefixed with the architecture (`qwen35.nextn_predict_layers`) and the
|
||||
/// architecture is whatever the next model is. The value is not read: a model
|
||||
/// that declares the key at all is one whose tensors carry the head, and the
|
||||
/// two disagreeing is a broken file rather than a state to handle.
|
||||
pub fn has_mtp_head(read: &mut impl Read) -> bool {
|
||||
find(read, |key| key.ends_with(".nextn_predict_layers")).is_some()
|
||||
}
|
||||
|
||||
/// Steps through the metadata table to the first key `wanted` accepts,
|
||||
/// returning its value's type tag and the reader positioned at the value.
|
||||
fn find<R: Read>(read: &mut R, wanted: impl Fn(&str) -> bool) -> Option<(u32, &mut R)> {
|
||||
let mut magic = [0u8; 4];
|
||||
read.read_exact(&mut magic).ok()?;
|
||||
if &magic != b"GGUF" {
|
||||
return None;
|
||||
}
|
||||
let _version = u32s(read)?;
|
||||
let _tensors = u64s(read)?;
|
||||
let count = u64s(read)?;
|
||||
for _ in 0..count {
|
||||
let found = string(read)?;
|
||||
let kind = u32s(read)?;
|
||||
if wanted(&found) {
|
||||
return Some((kind, read));
|
||||
}
|
||||
skip_value(kind, read)?;
|
||||
}
|
||||
None
|
||||
}
|
||||
|
||||
// The value type tags, in the container's own numbering. Only the two this
|
||||
// has to act on are named; the rest are widths, and `scalar_width` is where
|
||||
// the numbering is written down once.
|
||||
const STRING: u32 = 8;
|
||||
const ARRAY: u32 = 9;
|
||||
|
||||
/// How many bytes a scalar of this type occupies, or `None` for a type that
|
||||
/// is not a scalar -- which includes a tag this build does not know, since a
|
||||
/// value of unknown length cannot be stepped over.
|
||||
fn scalar_width(kind: u32) -> Option<u64> {
|
||||
match kind {
|
||||
// u8, i8, bool
|
||||
0 | 1 | 7 => Some(1),
|
||||
// u16, i16
|
||||
2 | 3 => Some(2),
|
||||
// u32, i32, f32
|
||||
4..=6 => Some(4),
|
||||
// u64, i64, f64
|
||||
10..=12 => Some(8),
|
||||
_ => None,
|
||||
}
|
||||
}
|
||||
|
||||
/// Steps over one value of `kind` without keeping it.
|
||||
///
|
||||
/// Recursive only in the sense that an array's elements are values; GGUF
|
||||
/// arrays do not nest, so the recursion is one level deep by construction.
|
||||
fn skip_value(kind: u32, read: &mut impl Read) -> Option<()> {
|
||||
match kind {
|
||||
STRING => {
|
||||
let len = u64s(read)?;
|
||||
skip(len, read)
|
||||
}
|
||||
ARRAY => {
|
||||
let element = u32s(read)?;
|
||||
let count = u64s(read)?;
|
||||
match scalar_width(element) {
|
||||
// The whole array at once: this is the tokenizer's scores and
|
||||
// token types, and stepping over them one at a time is a
|
||||
// syscall per token.
|
||||
Some(width) => skip(count.checked_mul(width)?, read),
|
||||
None if element == STRING => {
|
||||
for _ in 0..count {
|
||||
let len = u64s(read)?;
|
||||
skip(len, read)?;
|
||||
}
|
||||
Some(())
|
||||
}
|
||||
// An array of arrays, or of something this build has no width
|
||||
// for: the rest of the table can no longer be located.
|
||||
None => None,
|
||||
}
|
||||
}
|
||||
_ => skip(scalar_width(kind)?, read),
|
||||
}
|
||||
}
|
||||
|
||||
/// Discards `count` bytes, failing if the input ends first.
|
||||
///
|
||||
/// Chunked against a bounded buffer rather than read into a `Vec` of the
|
||||
/// stated size: the sizes here come out of the file, and the file may be a
|
||||
/// truncated prefix or not a GGUF at all.
|
||||
fn skip(count: u64, read: &mut impl Read) -> Option<()> {
|
||||
let mut scratch = [0u8; 8192];
|
||||
let mut left = count;
|
||||
while left > 0 {
|
||||
let want = left.min(scratch.len() as u64) as usize;
|
||||
read.read_exact(&mut scratch[..want]).ok()?;
|
||||
left -= want as u64;
|
||||
}
|
||||
Some(())
|
||||
}
|
||||
|
||||
fn string(read: &mut impl Read) -> Option<String> {
|
||||
let len = u64s(read)?;
|
||||
if len > MAX_STRING {
|
||||
return None;
|
||||
}
|
||||
let mut bytes = vec![0u8; len as usize];
|
||||
read.read_exact(&mut bytes).ok()?;
|
||||
String::from_utf8(bytes).ok()
|
||||
}
|
||||
|
||||
fn u32s(read: &mut impl Read) -> Option<u32> {
|
||||
let mut bytes = [0u8; 4];
|
||||
read.read_exact(&mut bytes).ok()?;
|
||||
Some(u32::from_le_bytes(bytes))
|
||||
}
|
||||
|
||||
fn u64s(read: &mut impl Read) -> Option<u64> {
|
||||
let mut bytes = [0u8; 8];
|
||||
read.read_exact(&mut bytes).ok()?;
|
||||
Some(u64::from_le_bytes(bytes))
|
||||
}
|
||||
|
||||
#[cfg(test)]
|
||||
mod tests {
|
||||
use super::*;
|
||||
|
||||
/// Builds a GGUF header holding exactly these keys, so the parser is
|
||||
/// tested against the layout rather than against a fixture nobody here
|
||||
/// can regenerate.
|
||||
fn header(entries: &[(&str, Value)]) -> Vec<u8> {
|
||||
let mut out = Vec::from(*b"GGUF");
|
||||
out.extend(3u32.to_le_bytes());
|
||||
out.extend(0u64.to_le_bytes());
|
||||
out.extend((entries.len() as u64).to_le_bytes());
|
||||
for (key, value) in entries {
|
||||
put_string(&mut out, key);
|
||||
value.write(&mut out);
|
||||
}
|
||||
out
|
||||
}
|
||||
|
||||
enum Value {
|
||||
Str(&'static str),
|
||||
U32(u32),
|
||||
Strings(Vec<&'static str>),
|
||||
Floats(Vec<f32>),
|
||||
}
|
||||
|
||||
impl Value {
|
||||
fn write(&self, out: &mut Vec<u8>) {
|
||||
match self {
|
||||
Self::Str(text) => {
|
||||
out.extend(STRING.to_le_bytes());
|
||||
put_string(out, text);
|
||||
}
|
||||
Self::U32(number) => {
|
||||
out.extend(4u32.to_le_bytes());
|
||||
out.extend(number.to_le_bytes());
|
||||
}
|
||||
Self::Strings(items) => {
|
||||
out.extend(ARRAY.to_le_bytes());
|
||||
out.extend(STRING.to_le_bytes());
|
||||
out.extend((items.len() as u64).to_le_bytes());
|
||||
for item in items {
|
||||
put_string(out, item);
|
||||
}
|
||||
}
|
||||
Self::Floats(items) => {
|
||||
out.extend(ARRAY.to_le_bytes());
|
||||
out.extend(6u32.to_le_bytes());
|
||||
out.extend((items.len() as u64).to_le_bytes());
|
||||
for item in items {
|
||||
out.extend(item.to_le_bytes());
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
fn put_string(out: &mut Vec<u8>, text: &str) {
|
||||
out.extend((text.len() as u64).to_le_bytes());
|
||||
out.extend(text.as_bytes());
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn the_name_is_read_past_every_other_kind_of_value() {
|
||||
let bytes = header(&[
|
||||
("general.architecture", Value::Str("qwen3")),
|
||||
("general.file_type", Value::U32(7)),
|
||||
("qwen3.attention.head_count", Value::U32(16)),
|
||||
("tokenizer.ggml.scores", Value::Floats(vec![0.5; 64])),
|
||||
("tokenizer.ggml.tokens", Value::Strings(vec!["a", "b", "c"])),
|
||||
("general.name", Value::Str("Qwen3-0.6B")),
|
||||
]);
|
||||
assert_eq!(
|
||||
name(&mut bytes.as_slice()),
|
||||
Some("Qwen3-0.6B".to_string()),
|
||||
"every value before the name has to be steppable over",
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
/// The remote case: a prefix is all there is, and running off the end of
|
||||
/// it is "we don't know" rather than a failure worth reporting. The
|
||||
/// caller has the file name.
|
||||
fn a_truncated_file_has_no_name_rather_than_failing() {
|
||||
let bytes = header(&[
|
||||
("tokenizer.ggml.tokens", Value::Strings(vec!["a", "b", "c"])),
|
||||
("general.name", Value::Str("Qwen3-0.6B")),
|
||||
]);
|
||||
for cut in [4, 12, 24, bytes.len() - 4] {
|
||||
assert_eq!(name(&mut &bytes[..cut]), None, "cut at {cut}");
|
||||
}
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn a_file_that_is_not_gguf_has_no_name() {
|
||||
assert_eq!(name(&mut b"not a model at all".as_slice()), None);
|
||||
assert_eq!(name(&mut b"".as_slice()), None);
|
||||
}
|
||||
|
||||
#[test]
|
||||
/// A name that is not a string is not a name. The alternative is
|
||||
/// rendering a number as one, which reads as a model called "7".
|
||||
fn a_name_of_the_wrong_type_is_not_read() {
|
||||
let bytes = header(&[("general.name", Value::U32(7))]);
|
||||
assert_eq!(name(&mut bytes.as_slice()), None);
|
||||
}
|
||||
|
||||
#[test]
|
||||
/// The head is found by the tail of the key, because the whole key is
|
||||
/// prefixed with whatever architecture the model is.
|
||||
fn an_mtp_head_is_found_whatever_the_architecture_is_called() {
|
||||
let with = header(&[
|
||||
("general.architecture", Value::Str("qwen35")),
|
||||
("qwen35.block_count", Value::U32(64)),
|
||||
("qwen35.nextn_predict_layers", Value::U32(1)),
|
||||
]);
|
||||
assert!(has_mtp_head(&mut with.as_slice()));
|
||||
|
||||
let without = header(&[
|
||||
("general.architecture", Value::Str("qwen3")),
|
||||
("qwen3.block_count", Value::U32(28)),
|
||||
]);
|
||||
assert!(!has_mtp_head(&mut without.as_slice()));
|
||||
}
|
||||
|
||||
#[test]
|
||||
/// A prefix that stops short says no, and that is the direction it has to
|
||||
/// fail in: `--spec-type draft-mtp` on a model with no head is a server
|
||||
/// that exits, so "we could not tell" and "it has none" both mean don't
|
||||
/// ask for it.
|
||||
fn a_truncated_file_reports_no_mtp_head() {
|
||||
let bytes = header(&[("qwen35.nextn_predict_layers", Value::U32(1))]);
|
||||
assert!(!has_mtp_head(&mut &bytes[..12]));
|
||||
}
|
||||
|
||||
#[test]
|
||||
/// The real thing, when this machine happens to have one. Skipped rather
|
||||
/// than failed where it does not: the models directory is not part of the
|
||||
/// checkout, and a test that needs gigabytes to run is one nobody runs.
|
||||
fn a_real_model_on_this_machine_reads_back_its_name() {
|
||||
let Some(home) = std::env::var_os("HOME") else {
|
||||
return;
|
||||
};
|
||||
let dir = std::path::Path::new(&home).join(".local/share/ai-app/models");
|
||||
let mut found = Vec::new();
|
||||
collect_gguf(&dir, &mut found);
|
||||
for path in found {
|
||||
let mut file = std::fs::File::open(&path).expect("open");
|
||||
let read = name(&mut file);
|
||||
assert!(
|
||||
read.is_some_and(|name| !name.trim().is_empty()),
|
||||
"{} has a name in it and this did not read one",
|
||||
path.display(),
|
||||
);
|
||||
}
|
||||
}
|
||||
|
||||
fn collect_gguf(dir: &std::path::Path, found: &mut Vec<std::path::PathBuf>) {
|
||||
let Ok(entries) = std::fs::read_dir(dir) else {
|
||||
return;
|
||||
};
|
||||
for entry in entries.flatten() {
|
||||
let path = entry.path();
|
||||
if path.is_dir() {
|
||||
collect_gguf(&path, found);
|
||||
} else if path.extension().is_some_and(|e| e == "gguf") {
|
||||
found.push(path);
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
+76
-11
@@ -15,6 +15,7 @@
|
||||
//! phone is not being given.
|
||||
|
||||
use anyhow::{Context, Result};
|
||||
use serde::Serialize;
|
||||
use serde_json::{Value, json};
|
||||
|
||||
use crate::config::{DriverKind, ProviderConfig};
|
||||
@@ -57,12 +58,7 @@ pub async fn discover(transport: &Transport) -> Result<Vec<ProviderConfig>> {
|
||||
// and nowhere else. Offering it on a remote machine would be a choice that
|
||||
// changes nothing.
|
||||
if matches!(transport, Transport::Here) {
|
||||
providers.push(ProviderConfig {
|
||||
name: crate::config::ECHO_PROVIDER.to_string(),
|
||||
kind: DriverKind::Echo,
|
||||
command: None,
|
||||
models: Vec::new(),
|
||||
});
|
||||
providers.push(crate::config::Config::echo_provider());
|
||||
}
|
||||
for (name, binary, kind) in PROBES {
|
||||
let path = found
|
||||
@@ -83,22 +79,88 @@ pub async fn discover(transport: &Transport) -> Result<Vec<ProviderConfig>> {
|
||||
DriverKind::ClaudeCli => CLAUDE_MODELS.iter().map(|m| (*m).to_string()).collect(),
|
||||
_ => Vec::new(),
|
||||
},
|
||||
mcp_servers: mcp_defaults(*kind),
|
||||
});
|
||||
}
|
||||
Ok(providers)
|
||||
}
|
||||
|
||||
/// One model a picker can offer, and what to call it there.
|
||||
///
|
||||
/// Two fields rather than one string because for one provider they differ:
|
||||
/// a llama.cpp model is chosen by the path it lives at and read as the name
|
||||
/// its own metadata gives it. Every other provider's id is already the name,
|
||||
/// and says so by repeating it -- which is what keeps the picker free of a
|
||||
/// branch on the session kind.
|
||||
#[derive(Debug, Clone, Serialize)]
|
||||
#[serde(rename_all = "camelCase")]
|
||||
pub struct OfferedModel {
|
||||
/// What a spawn or a model change is given. Opaque to the phone.
|
||||
pub id: String,
|
||||
/// What a person reads on the chip.
|
||||
pub label: String,
|
||||
}
|
||||
|
||||
impl OfferedModel {
|
||||
/// A model whose id is its own name, which is every provider but llama.
|
||||
fn plain(id: impl Into<String>) -> Self {
|
||||
let id = id.into();
|
||||
Self {
|
||||
label: id.clone(),
|
||||
id,
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
/// The MCP servers a newly discovered provider of this kind starts with.
|
||||
///
|
||||
/// A default rather than something to be typed in: a llama session with no web
|
||||
/// search is the state somebody would then have to find out how to leave, and
|
||||
/// Exa is what llama.cpp's own web UI offers under the same name. It is an
|
||||
/// ordinary config entry once written, so removing it is deleting a line.
|
||||
///
|
||||
/// Only llama.cpp, because only a driver that runs its own agent loop can use
|
||||
/// one -- the coding CLIs configure MCP themselves and a second answer here
|
||||
/// would quietly disagree with theirs.
|
||||
fn mcp_defaults(kind: DriverKind) -> Vec<crate::config::McpServerConfig> {
|
||||
match kind {
|
||||
DriverKind::LlamaCpp => vec![crate::config::McpServerConfig {
|
||||
name: "exa".to_string(),
|
||||
url: crate::session::llama::EXA_MCP_URL.to_string(),
|
||||
}],
|
||||
_ => Vec::new(),
|
||||
}
|
||||
}
|
||||
|
||||
/// Models the selected provider currently offers on this machine.
|
||||
///
|
||||
/// Codex's catalog is account- and CLI-version-specific, so it is asked at the
|
||||
/// moment the picker opens rather than copied into `config.ron`. Other
|
||||
/// providers retain the shortcut list discovery stored for them.
|
||||
/// moment the picker opens rather than copied into `config.ron`. A llama.cpp
|
||||
/// provider offers the GGUFs on the machine it runs on, through this same call
|
||||
/// -- there was a second route answering that alone, and it went when this one
|
||||
/// learned to, because a picker offering a model the spawn screen does not, or
|
||||
/// naming it differently, is two answers to one question. Other providers
|
||||
/// retain the shortcut list discovery stored for them.
|
||||
pub async fn provider_models(
|
||||
transport: &Transport,
|
||||
provider: &ProviderConfig,
|
||||
) -> Result<Vec<String>> {
|
||||
models_dir: &std::path::Path,
|
||||
) -> Result<Vec<OfferedModel>> {
|
||||
if provider.kind == DriverKind::LlamaCpp {
|
||||
let dir = crate::models::dir_on(transport, models_dir);
|
||||
let found = crate::models::on_machine(transport, &dir).await?;
|
||||
let labels = crate::models::labels(&found);
|
||||
return Ok(found
|
||||
.into_iter()
|
||||
.zip(labels)
|
||||
.map(|(model, label)| OfferedModel {
|
||||
id: model.key,
|
||||
label,
|
||||
})
|
||||
.collect());
|
||||
}
|
||||
if provider.kind != DriverKind::CodexCli {
|
||||
return Ok(provider.models.clone());
|
||||
return Ok(provider.models.iter().map(OfferedModel::plain).collect());
|
||||
}
|
||||
let transport = transport.clone();
|
||||
let program = provider.program().to_string();
|
||||
@@ -114,7 +176,10 @@ pub async fn provider_models(
|
||||
json!({"id": 2, "method": "model/list", "params": {"includeHidden": false, "limit": 100}}),
|
||||
];
|
||||
let answer = transport.request_json_blocking(&launch, &initial, &requests, 2)?;
|
||||
parse_codex_models(&answer)
|
||||
Ok(parse_codex_models(&answer)?
|
||||
.into_iter()
|
||||
.map(OfferedModel::plain)
|
||||
.collect())
|
||||
})
|
||||
.await?
|
||||
}
|
||||
|
||||
@@ -14,6 +14,7 @@
|
||||
mod auth;
|
||||
mod config;
|
||||
mod files;
|
||||
mod gguf;
|
||||
mod machines;
|
||||
mod media;
|
||||
mod models;
|
||||
|
||||
+103
-12
@@ -50,6 +50,71 @@ pub struct LocalModel {
|
||||
pub repo: String,
|
||||
pub file: String,
|
||||
pub bytes: u64,
|
||||
/// What the file says it is called (`general.name` in its own metadata),
|
||||
/// absent when it does not say or could not be read. Not a label: see
|
||||
/// [`labels`] for what a reader is actually shown, which needs the rest of
|
||||
/// the list to decide.
|
||||
#[serde(skip_serializing_if = "Option::is_none")]
|
||||
pub name: Option<String>,
|
||||
}
|
||||
|
||||
/// What each of these models should be called on screen, in the same order.
|
||||
///
|
||||
/// A model's own name is the best answer and is not always an answer at all:
|
||||
/// two quantisations of one model carry the same `general.name`, and a chip
|
||||
/// row with two identical chips is one you cannot choose from. So this is a
|
||||
/// cascade -- the model's own name, else its file name, else its full key --
|
||||
/// and each model takes the first rung that nothing else on this machine
|
||||
/// shares. The last rung always terminates it, because the key is what makes
|
||||
/// these unique in the first place.
|
||||
///
|
||||
/// Decided over the whole list rather than per model because ambiguity is a
|
||||
/// property of the set: the same file is unambiguous on a machine holding one
|
||||
/// quantisation and not on a machine holding three, and only the list knows
|
||||
/// which machine this is.
|
||||
pub fn labels(models: &[LocalModel]) -> Vec<String> {
|
||||
let rungs = |model: &LocalModel| {
|
||||
[
|
||||
model.name.clone(),
|
||||
Some(model.file.trim_end_matches(".gguf").to_string()),
|
||||
Some(model.key.clone()),
|
||||
]
|
||||
};
|
||||
let mut taken: Vec<HashMap<String, usize>> = vec![HashMap::new(); 3];
|
||||
for model in models {
|
||||
for (rung, candidate) in rungs(model).into_iter().enumerate() {
|
||||
if let Some(candidate) = candidate {
|
||||
*taken[rung].entry(candidate).or_insert(0) += 1;
|
||||
}
|
||||
}
|
||||
}
|
||||
models
|
||||
.iter()
|
||||
.map(|model| {
|
||||
rungs(model)
|
||||
.into_iter()
|
||||
.enumerate()
|
||||
.find_map(|(rung, candidate)| {
|
||||
let candidate = candidate?;
|
||||
(taken[rung].get(&candidate) == Some(&1)).then_some(candidate)
|
||||
})
|
||||
// Unreachable: the key rung is unique by construction. Said as
|
||||
// the key rather than as a panic, because a duplicate key would
|
||||
// mean the same file listed twice and a name is still the
|
||||
// honest thing to draw for it.
|
||||
.unwrap_or_else(|| model.key.clone())
|
||||
})
|
||||
.collect()
|
||||
}
|
||||
|
||||
/// The model's own name, read out of the file itself.
|
||||
///
|
||||
/// Absent for every way of not finding out -- see [`crate::gguf`]. The file is
|
||||
/// opened and read only as far as the name, which is the first few hundred
|
||||
/// bytes, so this is affordable once per model per listing.
|
||||
fn name_of(path: &Path) -> Option<String> {
|
||||
let mut file = std::fs::File::open(path).ok()?;
|
||||
crate::gguf::name(&mut file)
|
||||
}
|
||||
|
||||
/// What a run is doing, or did. Flat rather than a tagged enum carrying its
|
||||
@@ -519,6 +584,7 @@ fn collect(root: &Path, dir: &Path, found: &mut Vec<LocalModel>) {
|
||||
repo: repo.to_string(),
|
||||
file: file.to_string(),
|
||||
bytes: entry.metadata().map(|m| m.len()).unwrap_or(0),
|
||||
name: name_of(&path),
|
||||
});
|
||||
}
|
||||
}
|
||||
@@ -566,33 +632,46 @@ pub fn dir_on(transport: &Transport, local: &Path) -> String {
|
||||
/// directory that is not there is an empty list rather than a failure: a
|
||||
/// machine that has never had a model put on it is an ordinary state, and
|
||||
/// the same one as a machine whose directory exists and is empty.
|
||||
///
|
||||
/// Each record carries the head of the file as well as its size, because a
|
||||
/// model's own name is inside it (see [`crate::gguf`]) and the file is on the
|
||||
/// far machine. The alternative is a second round trip per model, or naming
|
||||
/// remote models by path while local ones get their proper names -- one
|
||||
/// machine's models reading differently from another's is exactly the
|
||||
/// confusion the name was added to remove. The prefix is bounded at
|
||||
/// [`crate::gguf::PREFIX_BYTES`], which is what keeps this one round trip's
|
||||
/// worth of bytes.
|
||||
pub async fn on_machine(transport: &Transport, dir: &str) -> Result<Vec<LocalModel>> {
|
||||
let script = "p=$1; case $p in \"~\") p=$HOME;; \"~/\"*) p=$HOME/${p#\"~/\"};; esac; \
|
||||
[ -d \"$p\" ] || exit 0; \
|
||||
find \"$p\" -type f -name '*.gguf' -printf '%s\\t%P\\0'";
|
||||
let script = format!(
|
||||
"p=$1; case $p in \"~\") p=$HOME;; \"~/\"*) p=$HOME/${{p#\"~/\"}};; esac; \
|
||||
[ -d \"$p\" ] || exit 0; cd \"$p\" || exit 0; \
|
||||
find . -type f -name '*.gguf' -exec sh -c '\
|
||||
for f do printf \"%s\\t%s\\t%s\\0\" \"$(wc -c < \"$f\")\" \
|
||||
\"$(head -c {prefix} \"$f\" | base64 | tr -d \"\\n\")\" \"${{f#./}}\"; done\
|
||||
' sh {{}} +",
|
||||
prefix = crate::gguf::PREFIX_BYTES,
|
||||
);
|
||||
let launch = Launch::new(
|
||||
"sh",
|
||||
vec![
|
||||
"-c".to_string(),
|
||||
script.to_string(),
|
||||
"sh".to_string(),
|
||||
dir.to_string(),
|
||||
],
|
||||
vec!["-c".to_string(), script, "sh".to_string(), dir.to_string()],
|
||||
None,
|
||||
);
|
||||
let out = transport.capture(&launch).await?;
|
||||
let mut found: Vec<LocalModel> = out
|
||||
.split('\0')
|
||||
.filter(|record| !record.is_empty())
|
||||
// Two fields, and the name last, so a `\t` in a filename survives.
|
||||
.filter_map(|record| record.split_once('\t'))
|
||||
.filter_map(|(bytes, key)| {
|
||||
// Three fields, and the name last, so a `\t` in a filename survives.
|
||||
// The middle one is base64, which has no tab in its alphabet.
|
||||
.filter_map(|record| {
|
||||
let (bytes, rest) = record.split_once('\t')?;
|
||||
let (head, key) = rest.split_once('\t')?;
|
||||
let (repo, file) = key.rsplit_once('/')?;
|
||||
Some(LocalModel {
|
||||
key: key.to_string(),
|
||||
repo: repo.to_string(),
|
||||
file: file.to_string(),
|
||||
bytes: bytes.trim().parse().unwrap_or(0),
|
||||
name: name_in_prefix(head),
|
||||
})
|
||||
})
|
||||
.collect();
|
||||
@@ -600,6 +679,18 @@ pub async fn on_machine(transport: &Transport, dir: &str) -> Result<Vec<LocalMod
|
||||
Ok(found)
|
||||
}
|
||||
|
||||
/// The model's name out of a base64 prefix of its file.
|
||||
///
|
||||
/// The remote half of [`name_of`], and `None` for everything that half
|
||||
/// answers `None` for, plus a prefix that did not survive the trip.
|
||||
fn name_in_prefix(head: &str) -> Option<String> {
|
||||
use base64::Engine as _;
|
||||
let bytes = base64::engine::general_purpose::STANDARD
|
||||
.decode(head.trim())
|
||||
.ok()?;
|
||||
crate::gguf::name(&mut bytes.as_slice())
|
||||
}
|
||||
|
||||
/// A model repository on HuggingFace, as the browse screen shows it.
|
||||
#[derive(Debug, Clone, Serialize)]
|
||||
#[serde(rename_all = "camelCase")]
|
||||
|
||||
@@ -436,6 +436,7 @@ mod tests {
|
||||
kind: DriverKind::ClaudeCli,
|
||||
command: Some(cli.display().to_string()),
|
||||
models: Vec::new(),
|
||||
mcp_servers: Vec::new(),
|
||||
};
|
||||
|
||||
let started = logins.start(machine, provider);
|
||||
|
||||
+6
-29
@@ -7,8 +7,7 @@
|
||||
//! POST /machines add {name, ssh?} -- providers are discovered
|
||||
//! POST /machines/probe dry run {ssh?}: what would be found there
|
||||
//! GET /machines/{id} one machine, for refetching after a change
|
||||
//! GET /machines/{id}/models GGUFs on that machine, for a llama session
|
||||
//! GET /machines/{id}/providers/{provider}/models models a CLI currently offers
|
||||
//! GET /machines/{id}/providers/{provider}/models models that provider offers
|
||||
//! POST /machines/{id}/providers/{provider}/auth begin provider sign-in
|
||||
//! GET /machines/{id}/providers/{provider}/auth/{attempt} sign-in state
|
||||
//! POST /machines/{id}/providers/{provider}/auth/{attempt}/code submit browser code
|
||||
@@ -131,7 +130,6 @@ pub fn router(manager: Arc<SessionManager>) -> Router {
|
||||
get(read_machine).put(update_machine).delete(delete_machine),
|
||||
)
|
||||
// The models on a configured machine, for a llama session there.
|
||||
.route("/machines/{id}/models", get(machine_models))
|
||||
.route(
|
||||
"/machines/{id}/providers/{provider}/models",
|
||||
get(provider_models),
|
||||
@@ -565,40 +563,19 @@ struct PathQuery {
|
||||
path: String,
|
||||
}
|
||||
|
||||
/// The models **that machine** has, which is the list a llama.cpp session
|
||||
/// on it can choose from.
|
||||
///
|
||||
/// Not `GET /models`, which is this backend's own downloads: those are on
|
||||
/// the machine a session runs on only when they are the same machine. A
|
||||
/// spawn screen offering this backend's list for a remote machine would be
|
||||
/// naming files that are not there, and the session would fail at the
|
||||
/// point of loading rather than at the point of choosing.
|
||||
async fn machine_models(
|
||||
State(manager): State<Arc<SessionManager>>,
|
||||
UrlPath(id): UrlPath<String>,
|
||||
) -> Result<axum::Json<Vec<crate::models::LocalModel>>, ApiError> {
|
||||
let machine = machine_by_id(&manager, &id)?;
|
||||
let transport = crate::session::transport::Transport::for_machine(&machine);
|
||||
let dir = crate::models::dir_on(&transport, manager.models_dir());
|
||||
crate::models::on_machine(&transport, &dir)
|
||||
.await
|
||||
.map(axum::Json)
|
||||
.map_err(from_machine)
|
||||
}
|
||||
|
||||
/// The models a CLI provider currently offers on its configured machine.
|
||||
/// Codex answers from its live account catalog; providers with a configured
|
||||
/// shortcut list return that list.
|
||||
/// The models a provider currently offers on its configured machine.
|
||||
/// Codex answers from its live account catalog, llama.cpp from the GGUFs on
|
||||
/// that machine, and providers with a configured shortcut list return it.
|
||||
async fn provider_models(
|
||||
State(manager): State<Arc<SessionManager>>,
|
||||
UrlPath((id, provider_name)): UrlPath<(String, String)>,
|
||||
) -> Result<axum::Json<Vec<String>>, ApiError> {
|
||||
) -> Result<axum::Json<Vec<crate::machines::OfferedModel>>, ApiError> {
|
||||
let machine = machine_by_id(&manager, &id)?;
|
||||
let provider = machine.provider(&provider_name).ok_or_else(|| {
|
||||
ApiError::NotFound(format!("no provider {provider_name} on {}", machine.name))
|
||||
})?;
|
||||
let transport = crate::session::transport::Transport::for_machine(&machine);
|
||||
crate::machines::provider_models(&transport, provider)
|
||||
crate::machines::provider_models(&transport, provider, manager.models_dir())
|
||||
.await
|
||||
.map(axum::Json)
|
||||
.map_err(from_machine)
|
||||
|
||||
@@ -564,6 +564,22 @@ pub enum SessionStatus {
|
||||
Running,
|
||||
AwaitingInput,
|
||||
Compacting,
|
||||
/// The session's process is up but cannot be spoken to yet.
|
||||
///
|
||||
/// Its own state because the two it would otherwise borrow are both
|
||||
/// wrong in ways somebody notices. `Running` means the session is
|
||||
/// answering, so a model taking a minute to load looks like a model
|
||||
/// thinking for a minute -- and there is no way to tell from the screen
|
||||
/// that the first message will be refused. `Idle` invites that message
|
||||
/// and then loses it.
|
||||
///
|
||||
/// It exists for `llama-server`, which reads a multi-gigabyte file off
|
||||
/// disk before it answers anything, and it is general because the
|
||||
/// condition is: a process that is started and not yet ready is a state
|
||||
/// any driver may have to report. Nothing is queued *because* of this
|
||||
/// state -- a driver that reports it is responsible for holding what it
|
||||
/// is sent until it can deliver it -- but this is what says so on screen.
|
||||
Loading,
|
||||
/// The session's own turn is over, but work it started is still going:
|
||||
/// a backgrounded subagent, or a command left running.
|
||||
///
|
||||
|
||||
@@ -1,907 +0,0 @@
|
||||
//! The llama.cpp driver: a `llama-server` process per session, spoken to over
|
||||
//! its OpenAI-compatible HTTP API and translated into the common event model.
|
||||
//!
|
||||
//! Two things make this shaped differently from the Claude driver.
|
||||
//!
|
||||
//! **It is spawned but not spoken to over stdio.** The process is started
|
||||
//! through the same [`Transport`] as any other and then reached over HTTP on a
|
||||
//! loopback port. That is the second half of what a transport is -- "run this"
|
||||
//! plus "reach this port" -- and it is what lets a session run on another
|
||||
//! machine: [`Transport::reserve_port`] hands back a port the server binds
|
||||
//! *there* and one that reaches it *here*, and the ssh connection carrying the
|
||||
//! command carries the tunnel between them. The far `llama-server` binds
|
||||
//! loopback only, so a model is never served to that machine's network.
|
||||
//!
|
||||
//! **The model file is the far machine's, not this one's.** A remote machine
|
||||
//! names its own models directory (`SshConfig::models_dir`, defaulting to where
|
||||
//! this backend keeps its downloads), and the file is looked for *there* -- so
|
||||
//! a session naming a model that machine does not have says so, instead of
|
||||
//! starting a server that will never load one. Downloading to another machine
|
||||
//! is not built; the model gets there however anything else does.
|
||||
//!
|
||||
//! **The server is stateless between requests**, so the whole conversation goes
|
||||
//! with every one. It is rebuilt from the session's transcript rather than kept
|
||||
//! in this struct, which is not tidiness: a copy in driver memory is invisible
|
||||
//! to a second device and gone when this process restarts.
|
||||
//!
|
||||
//! That leaves the Claude driver as the odd one out rather than this one -- the
|
||||
//! CLI's own memory of a conversation is a cache in front of the same
|
||||
//! transcript. Resolve any inconsistency in this direction.
|
||||
|
||||
use std::path::{Path, PathBuf};
|
||||
use std::sync::Arc;
|
||||
use std::sync::atomic::{AtomicBool, Ordering};
|
||||
|
||||
use anyhow::{Context, Result, bail};
|
||||
use serde::{Deserialize, Serialize};
|
||||
use serde_json::json;
|
||||
|
||||
use super::driver::{AttachmentRef, Driver, Event, EventSink, SessionStatus};
|
||||
use super::process;
|
||||
use super::transport::{Launch, Streams, Transport};
|
||||
use crate::config::{ProviderConfig, SessionConfig};
|
||||
|
||||
/// How long to wait for a model to load before giving up. Loading is mostly
|
||||
/// disk, and a large quantised model on a cold cache is genuinely slow, so this
|
||||
/// is generous -- the failure it exists for is a server that will never answer.
|
||||
const READY_TIMEOUT: std::time::Duration = std::time::Duration::from_secs(300);
|
||||
|
||||
/// One turn in the conversation this driver keeps on the server's behalf.
|
||||
#[derive(Debug, Clone, Serialize, Deserialize)]
|
||||
struct Message {
|
||||
role: String,
|
||||
content: String,
|
||||
}
|
||||
|
||||
pub struct LlamaDriver {
|
||||
sink: EventSink,
|
||||
/// Where this session's own llama-server answers.
|
||||
endpoint: String,
|
||||
/// Where the conversation is read back from, one line per event.
|
||||
transcript: PathBuf,
|
||||
/// Sampling settings chosen at spawn, sent with every request.
|
||||
sampling: serde_json::Map<String, serde_json::Value>,
|
||||
/// Set by [`Driver::interrupt`]; the streaming loop checks it between
|
||||
/// chunks and stops, leaving what was generated in the transcript.
|
||||
cancel: Arc<AtomicBool>,
|
||||
/// Where this session's process record lives, so [`Driver::stop`] can find
|
||||
/// the server it has to end.
|
||||
session_dir: PathBuf,
|
||||
}
|
||||
|
||||
impl LlamaDriver {
|
||||
/// Takes charge of this session's `llama-server`: the one already loaded if
|
||||
/// there is one, otherwise a new one.
|
||||
///
|
||||
/// One entry point, for the reason `ClaudeDriver::launch` gives, expensive
|
||||
/// in a different currency: two servers holding the same model is twice the
|
||||
/// memory, and the second would bind a different port while the phone kept
|
||||
/// talking to the first.
|
||||
#[allow(clippy::too_many_arguments)]
|
||||
pub fn launch(
|
||||
meta: &SessionConfig,
|
||||
provider: &ProviderConfig,
|
||||
transport: &Transport,
|
||||
models_dir: &Path,
|
||||
transcript: &Path,
|
||||
session_dir: &Path,
|
||||
sink: EventSink,
|
||||
// llama.cpp has no notion of a Task call, so this is accepted only
|
||||
// to keep one shape across every driver's launch -- see
|
||||
// `SUBAGENTS.md`'s "Server layout".
|
||||
_subagents: Arc<super::subagent::Subagents>,
|
||||
) -> Result<Self> {
|
||||
let model = meta.model.as_deref().context(
|
||||
"a llama.cpp session needs a model -- one of the downloaded ones, by its key",
|
||||
)?;
|
||||
let path = model_on(transport, models_dir, model)?;
|
||||
|
||||
// Already loaded and still running: keep talking to it. The health poll
|
||||
// below confirms it is really answering, so adopting a pid whose server
|
||||
// has wedged still reports as a failure rather than as a session that
|
||||
// silently never replies.
|
||||
if let Some(process::Record {
|
||||
detail: process::Detail::Http { port },
|
||||
pid,
|
||||
..
|
||||
}) = process::live(session_dir)
|
||||
{
|
||||
tracing::info!(
|
||||
"session {} reattaching to the llama-server it left loaded (pid {pid}, port {port})",
|
||||
meta.id
|
||||
);
|
||||
return Ok(Self::attached(
|
||||
format!("http://127.0.0.1:{port}"),
|
||||
meta,
|
||||
model,
|
||||
transcript,
|
||||
session_dir,
|
||||
sink,
|
||||
));
|
||||
}
|
||||
|
||||
// Where it listens on its own machine, and where that is reached
|
||||
// from here -- the same number when that machine is this one.
|
||||
let forward = transport
|
||||
.reserve_port()
|
||||
.context("finding a port for llama-server")?;
|
||||
let mut args: Vec<String> = vec![
|
||||
"-m".into(),
|
||||
path.clone(),
|
||||
// Loopback there, whichever machine there is: what reaches it
|
||||
// from outside that machine is the ssh tunnel and nothing
|
||||
// else.
|
||||
"--host".into(),
|
||||
"127.0.0.1".into(),
|
||||
"--port".into(),
|
||||
forward.there.to_string(),
|
||||
];
|
||||
// Settings that belong to the server because they decide how the model
|
||||
// is loaded; the sampling ones ride on each request instead, so changing
|
||||
// them later needn't reload anything.
|
||||
for (key, flag) in [
|
||||
("contextSize", "-c"),
|
||||
("gpuLayers", "-ngl"),
|
||||
("threads", "-t"),
|
||||
] {
|
||||
if let Some(value) = meta.params.get(key) {
|
||||
args.push(flag.to_string());
|
||||
args.push(value.clone());
|
||||
}
|
||||
}
|
||||
|
||||
let program = provider.program();
|
||||
let launch = Launch::new(program, args, meta.cwd.as_deref()).reaching(forward);
|
||||
// Its output goes to files, not pipes. Not only so the process can
|
||||
// outlive this server: nothing ever read those pipes, so a chatty
|
||||
// llama-server filled the 64 KB buffer and blocked mid-load with no sign
|
||||
// of why.
|
||||
let child = transport.spawn(
|
||||
&launch,
|
||||
Streams::Detached {
|
||||
stdin: std::process::Stdio::null(),
|
||||
stdout: log_file(&session_dir.join(SERVER_LOG))?.into(),
|
||||
stderr: log_file(&session_dir.join(SERVER_LOG))?.into(),
|
||||
},
|
||||
)?;
|
||||
let pid = child
|
||||
.id()
|
||||
.context("llama-server exited before it could be recorded")?;
|
||||
tracing::info!(
|
||||
"session {} running {program} for {model} {} on 127.0.0.1:{} there, \
|
||||
reached at 127.0.0.1:{} here, as pid {pid}",
|
||||
meta.id,
|
||||
transport.describe(),
|
||||
forward.there,
|
||||
forward.here,
|
||||
);
|
||||
// Reaped so it does not become a zombie while this server is still its
|
||||
// parent; the health poll and the record are what say whether the
|
||||
// session is alive, because after a restart there is no `Child` to ask.
|
||||
tokio::spawn(async move {
|
||||
let mut child = child;
|
||||
let _ = child.wait().await;
|
||||
});
|
||||
|
||||
// The *near* port, because that is the one anything reaching this
|
||||
// server has to dial -- including a later run of this backend,
|
||||
// which adopts the record without knowing which machine the server
|
||||
// is on. For a remote session the recorded pid is the ssh
|
||||
// client's, which is the process this machine owns and which holds
|
||||
// the tunnel open for exactly as long as the far server lives.
|
||||
let record = process::Record::of(pid, process::Detail::Http { port: forward.here })
|
||||
.context("llama-server was gone before its start time could be read")?;
|
||||
process::write(session_dir, &record);
|
||||
|
||||
Ok(Self::attached(
|
||||
format!("http://127.0.0.1:{}", forward.here),
|
||||
meta,
|
||||
model,
|
||||
transcript,
|
||||
session_dir,
|
||||
sink,
|
||||
))
|
||||
}
|
||||
|
||||
/// The driver for a `llama-server` at `endpoint`, however it got there.
|
||||
///
|
||||
/// Shared by starting one and adopting one, because everything after "there
|
||||
/// is a server at this address" is identical -- including waiting for it to
|
||||
/// answer, which an adopted one still owes: a recorded pid says a process
|
||||
/// exists, not that its model is loaded.
|
||||
fn attached(
|
||||
endpoint: String,
|
||||
meta: &SessionConfig,
|
||||
model: &str,
|
||||
transcript: &Path,
|
||||
session_dir: &Path,
|
||||
sink: EventSink,
|
||||
) -> Self {
|
||||
// Loading is slow enough to be worth saying so: the session shows as
|
||||
// running until the model is in memory, rather than looking ready and
|
||||
// refusing the first message.
|
||||
let _ = sink.send(Event::Status {
|
||||
state: SessionStatus::Running,
|
||||
});
|
||||
{
|
||||
let sink = sink.clone();
|
||||
let endpoint = endpoint.clone();
|
||||
let model = model.to_string();
|
||||
let session_dir = session_dir.to_path_buf();
|
||||
std::thread::spawn(move || match wait_until_ready(&endpoint, &session_dir) {
|
||||
Ok(()) => {
|
||||
tracing::info!("{model} loaded and answering at {endpoint}");
|
||||
let _ = sink.send(Event::Status {
|
||||
state: SessionStatus::Idle,
|
||||
});
|
||||
watch(session_dir, sink);
|
||||
}
|
||||
Err(err) => {
|
||||
let _ = sink.send(Event::Error {
|
||||
message: format!("{model} never became ready: {err:#}"),
|
||||
});
|
||||
let _ = sink.send(Event::Status {
|
||||
state: SessionStatus::Exited,
|
||||
});
|
||||
process::clear(&session_dir);
|
||||
}
|
||||
});
|
||||
}
|
||||
|
||||
let mut sampling = serde_json::Map::new();
|
||||
for (key, field) in [
|
||||
("temperature", "temperature"),
|
||||
("topP", "top_p"),
|
||||
("topK", "top_k"),
|
||||
("maxTokens", "max_tokens"),
|
||||
] {
|
||||
if let Some(raw) = meta.params.get(key)
|
||||
&& let Ok(number) = raw.parse::<f64>()
|
||||
{
|
||||
sampling.insert(field.to_string(), json!(number));
|
||||
}
|
||||
}
|
||||
|
||||
Self {
|
||||
sink,
|
||||
endpoint,
|
||||
transcript: transcript.to_path_buf(),
|
||||
sampling,
|
||||
cancel: Arc::new(AtomicBool::new(false)),
|
||||
session_dir: session_dir.to_path_buf(),
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
/// Where llama-server's own output goes. One file for both streams: it is
|
||||
/// diagnostics nobody parses, and interleaving them is how it reads in a
|
||||
/// terminal anyway.
|
||||
const SERVER_LOG: &str = "llama-server.log";
|
||||
|
||||
/// How often a loaded server is checked for still being there. Slower than the
|
||||
/// Claude driver's stdout poll because nothing is waiting on it: this only has
|
||||
/// to notice a server that has gone.
|
||||
const WATCH_INTERVAL: std::time::Duration = std::time::Duration::from_secs(2);
|
||||
|
||||
/// An owner-only log opened for appending, so the two streams pointed at
|
||||
/// it do not overwrite each other and a reattach keeps what came before.
|
||||
fn log_file(path: &Path) -> Result<std::fs::File> {
|
||||
use std::os::unix::fs::OpenOptionsExt;
|
||||
std::fs::OpenOptions::new()
|
||||
.create(true)
|
||||
.append(true)
|
||||
.mode(0o600)
|
||||
.open(path)
|
||||
.with_context(|| format!("opening {}", path.display()))
|
||||
}
|
||||
|
||||
/// Reports the server going away, for as long as the session is there to report
|
||||
/// it to.
|
||||
///
|
||||
/// Polled rather than waited on, for the reason the Claude driver gives: after a
|
||||
/// restart this server is not the process's parent, so liveness has to be a
|
||||
/// question asked of the record -- and asking it two different ways is how the
|
||||
/// two answers come to disagree.
|
||||
fn watch(session_dir: PathBuf, sink: EventSink) {
|
||||
std::thread::spawn(move || {
|
||||
loop {
|
||||
std::thread::sleep(WATCH_INTERVAL);
|
||||
match process::recorded(&session_dir) {
|
||||
Some((_, process::Liveness::Alive)) => {}
|
||||
// Nothing recorded means the session was stopped or deleted
|
||||
// deliberately, and whoever did that has already said so.
|
||||
None => return,
|
||||
Some((_, process::Liveness::Dead)) => {
|
||||
if !process::stopping(&session_dir) {
|
||||
let _ = sink.send(Event::Error {
|
||||
message: "llama-server exited".to_string(),
|
||||
});
|
||||
}
|
||||
let _ = sink.send(Event::Status {
|
||||
state: SessionStatus::Exited,
|
||||
});
|
||||
process::clear(&session_dir);
|
||||
return;
|
||||
}
|
||||
Some((_, process::Liveness::Unknown)) => {
|
||||
let _ = sink.send(Event::Status {
|
||||
state: SessionStatus::Unknown,
|
||||
});
|
||||
}
|
||||
}
|
||||
if sink.is_closed() {
|
||||
return;
|
||||
}
|
||||
}
|
||||
});
|
||||
}
|
||||
|
||||
impl Driver for LlamaDriver {
|
||||
fn send_user_message(&self, text: String, attachments: Vec<AttachmentRef>) {
|
||||
if !attachments.is_empty() {
|
||||
let _ = self.sink.send(Event::Error {
|
||||
message: "this model can't be sent attachments or files".to_string(),
|
||||
});
|
||||
}
|
||||
let sink = self.sink.clone();
|
||||
let endpoint = self.endpoint.clone();
|
||||
let transcript = self.transcript.clone();
|
||||
let sampling = self.sampling.clone();
|
||||
let cancel = Arc::clone(&self.cancel);
|
||||
cancel.store(false, Ordering::Relaxed);
|
||||
|
||||
// Its own thread: the request blocks for as long as the model takes to
|
||||
// generate, which is the whole point of streaming it.
|
||||
std::thread::spawn(move || {
|
||||
// Nothing is ever held back here -- there is no queue to wait in --
|
||||
// so the message is taken the moment it arrives. Said anyway,
|
||||
// because this is what records it: see `MessageTaken`.
|
||||
let _ = sink.send(Event::MessageTaken {
|
||||
id: None,
|
||||
text: text.clone(),
|
||||
// Never any: this driver refuses attachments above.
|
||||
attachments: Vec::new(),
|
||||
});
|
||||
let _ = sink.send(Event::Status {
|
||||
state: SessionStatus::Running,
|
||||
});
|
||||
// Everything before this message, plus this message. Read rather
|
||||
// than remembered, and `text` is appended here rather than waited
|
||||
// for, because the message's own transcript entry is still on its
|
||||
// way when this runs.
|
||||
let mut messages = conversation(&transcript);
|
||||
messages.push(Message {
|
||||
role: "user".into(),
|
||||
content: text,
|
||||
});
|
||||
// The reply is not stored: the deltas below are the durable record,
|
||||
// so the next turn reads back exactly what the phone was shown --
|
||||
// including a partial one that was interrupted.
|
||||
if let Err(err) = generate(&endpoint, &messages, &sampling, &cancel, &sink) {
|
||||
let _ = sink.send(Event::Error {
|
||||
message: format!("{err:#}"),
|
||||
});
|
||||
}
|
||||
let _ = sink.send(Event::Status {
|
||||
state: SessionStatus::Idle,
|
||||
});
|
||||
});
|
||||
}
|
||||
|
||||
fn answer_question(&self, _id: &str, _answers: &[String]) {
|
||||
// Nothing here asks questions: this driver has no tools.
|
||||
}
|
||||
|
||||
fn interrupt(&self) {
|
||||
self.cancel.store(true, Ordering::Relaxed);
|
||||
}
|
||||
|
||||
// Nothing to forward: this process has no notion of what the conversation
|
||||
// is called, and the rename has already happened where the name lives.
|
||||
fn set_title(&self, _title: &str) {}
|
||||
|
||||
fn set_permission_mode(&self, _mode: &str) {
|
||||
let _ = self.sink.send(Event::Error {
|
||||
message: "a llama.cpp session runs no tools, so there is nothing for a permission \
|
||||
mode to govern."
|
||||
.to_string(),
|
||||
});
|
||||
}
|
||||
|
||||
fn set_model(&self, _model: &str) {
|
||||
let _ = self.sink.send(Event::Error {
|
||||
message: "a llama.cpp session's model is fixed when it starts, because the server \
|
||||
loads one model into memory. Spawn another session to use a different one."
|
||||
.to_string(),
|
||||
});
|
||||
}
|
||||
|
||||
fn run_command(&self, text: &str) {
|
||||
let _ = self.sink.send(Event::Error {
|
||||
message: format!(
|
||||
"a llama.cpp session has no commands of its own, so {text} means nothing to it."
|
||||
),
|
||||
});
|
||||
}
|
||||
|
||||
fn compact(&self) {
|
||||
let _ = self.sink.send(Event::Error {
|
||||
message: "llama.cpp has no compaction. Clear the session instead, which costs nothing."
|
||||
.to_string(),
|
||||
});
|
||||
}
|
||||
|
||||
fn clear(&self) {
|
||||
// All of it. `conversation` folds from the last of these, so recording
|
||||
// the marker *is* the reset -- there is no driver state to keep in step
|
||||
// with it, which is the same property that makes a second device see the
|
||||
// same conversation this one does.
|
||||
let _ = self.sink.send(Event::Cleared);
|
||||
}
|
||||
|
||||
/// Stops generating and leaves the server loaded.
|
||||
///
|
||||
/// Worth being deliberate about, because the cost points the other way from
|
||||
/// the Claude driver's: a `llama-server` holds its whole model in memory, so
|
||||
/// a leaked one is gigabytes nobody is using. It is left anyway, because the
|
||||
/// alternative is unloading and reloading that model on every backend
|
||||
/// restart -- minutes of disk, for a session somebody is in the middle of.
|
||||
/// The record is what keeps it from being *nobody's*.
|
||||
fn detach(&self) {
|
||||
self.cancel.store(true, Ordering::Relaxed);
|
||||
}
|
||||
|
||||
fn stop(&self) {
|
||||
self.cancel.store(true, Ordering::Relaxed);
|
||||
if let Some(record) = process::live(&self.session_dir) {
|
||||
process::stop(&record, process::STOP_GRACE);
|
||||
}
|
||||
process::clear(&self.session_dir);
|
||||
}
|
||||
}
|
||||
|
||||
/// The conversation so far, folded out of the transcript.
|
||||
///
|
||||
/// Consecutive `AssistantText` deltas are one assistant turn, closed by the next
|
||||
/// user message -- which is also what makes an interrupted reply come back as
|
||||
/// the partial text the phone actually saw.
|
||||
///
|
||||
/// This must stay a pure function of the transcript and must never re-render
|
||||
/// earlier turns. llama.cpp caches the prompt prefix, so a growing conversation
|
||||
/// reprocesses almost nothing -- but only while every turn is byte-identical to
|
||||
/// last time. Changing how an old turn is rendered silently reprocesses the
|
||||
/// whole history on every message.
|
||||
fn conversation(path: &Path) -> Vec<Message> {
|
||||
let Ok(events) = crate::session::transcript::read_after(path, 0) else {
|
||||
return Vec::new();
|
||||
};
|
||||
let mut messages: Vec<Message> = Vec::new();
|
||||
let mut pending = String::new();
|
||||
// Everything before the last clear is still in the transcript and is
|
||||
// deliberately not in the conversation. Folding from zero would put it back,
|
||||
// which is the whole of what clearing had to undo.
|
||||
let events = match events.iter().rposition(|e| e.event == Event::Cleared) {
|
||||
Some(at) => &events[at + 1..],
|
||||
None => &events[..],
|
||||
};
|
||||
for event in events.iter().cloned() {
|
||||
match event.event {
|
||||
Event::UserMessage { text, .. } => {
|
||||
if !pending.is_empty() {
|
||||
messages.push(Message {
|
||||
role: "assistant".into(),
|
||||
content: std::mem::take(&mut pending),
|
||||
});
|
||||
}
|
||||
messages.push(Message {
|
||||
role: "user".into(),
|
||||
content: text,
|
||||
});
|
||||
}
|
||||
Event::AssistantText { delta } => pending.push_str(&delta),
|
||||
Event::AssistantTextFinal { text } => {
|
||||
pending = text;
|
||||
}
|
||||
_ => {}
|
||||
}
|
||||
}
|
||||
if !pending.is_empty() {
|
||||
messages.push(Message {
|
||||
role: "assistant".into(),
|
||||
content: pending,
|
||||
});
|
||||
}
|
||||
messages
|
||||
}
|
||||
|
||||
/// Where a model key resolves to on disk, refusing anything that climbs
|
||||
/// out of the models directory -- the key arrives from a phone.
|
||||
fn model_path(models_dir: &Path, key: &str) -> Result<PathBuf> {
|
||||
let mut path = models_dir.to_path_buf();
|
||||
for part in key.split('/') {
|
||||
if part.is_empty() || part == "." || part == ".." {
|
||||
bail!("\"{key}\" is not a model key this can resolve");
|
||||
}
|
||||
path.push(part);
|
||||
}
|
||||
if !path.is_file() {
|
||||
bail!("no downloaded model called \"{key}\" -- download it first");
|
||||
}
|
||||
Ok(path)
|
||||
}
|
||||
|
||||
/// The model file's path **on the machine that will serve it**, confirmed to be
|
||||
/// there.
|
||||
///
|
||||
/// One function rather than a local check and hope for the other case: the same
|
||||
/// question has to be asked of two filesystems. The remote answer is measured
|
||||
/// for the reason the local one is -- a missing file otherwise becomes a
|
||||
/// `llama-server` that starts, fails to load, and reports as a session that
|
||||
/// never became ready, which reads as the machine being slow.
|
||||
///
|
||||
/// One blocking round trip on a remote spawn, which is what the spawn is
|
||||
/// already paying to start ssh. The alternative is a path built here from a `~`
|
||||
/// this machine cannot expand.
|
||||
fn model_on(transport: &Transport, models_dir: &Path, key: &str) -> Result<String> {
|
||||
let Transport::Ssh { name, .. } = transport else {
|
||||
return Ok(model_path(models_dir, key)?.to_string_lossy().into_owned());
|
||||
};
|
||||
// The same directory the spawn screen listed for this machine, and one
|
||||
// function for the same reason: a list from one place and a load from
|
||||
// another is a model that appears and then fails.
|
||||
let dir = crate::models::dir_on(transport, models_dir);
|
||||
// Checked here rather than in the script: `..` in a key would walk out of
|
||||
// the models directory on a machine this server can start processes on,
|
||||
// and the phone is where the key comes from.
|
||||
for part in key.split('/') {
|
||||
if part.is_empty() || part == "." || part == ".." {
|
||||
bail!("\"{key}\" is not a model key this can resolve");
|
||||
}
|
||||
}
|
||||
let path = format!("{}/{key}", dir.trim_end_matches('/'));
|
||||
// `$HOME` on the far side, which is the only machine that knows what it is,
|
||||
// and the resolved path printed back so the launch hands `llama-server`
|
||||
// something absolute. "Not there" is answered rather than failed, because a
|
||||
// machine that could not be asked at all has to say so in its own words --
|
||||
// it would otherwise arrive as this same sentence about a missing model.
|
||||
let script = "p=$1; case $p in \"~\") p=$HOME;; \"~/\"*) p=$HOME/${p#\"~/\"};; esac; \
|
||||
[ -f \"$p\" ] && printf 'at\\t%s\\n' \"$p\" || printf 'missing\\n'"
|
||||
.to_string();
|
||||
let launch = Launch::new(
|
||||
"sh",
|
||||
vec!["-c".to_string(), script, "sh".to_string(), path.clone()],
|
||||
None,
|
||||
);
|
||||
let answer = transport
|
||||
.capture_blocking(&launch)
|
||||
.with_context(|| format!("couldn't ask {name} where its models are"))?;
|
||||
match answer.trim().split_once('\t') {
|
||||
Some(("at", resolved)) => Ok(resolved.to_string()),
|
||||
_ => bail!(
|
||||
"{name} has no model at {path}. A llama.cpp session serves the file from the \
|
||||
machine it runs on, so the model has to be on {name} -- what this backend has \
|
||||
downloaded is somewhere else."
|
||||
),
|
||||
}
|
||||
}
|
||||
|
||||
/// Polls until the server says it is ready, or gives up.
|
||||
///
|
||||
/// Watches the process as well as the port, because the two failures need
|
||||
/// different words and one of them is common: a model that will not load,
|
||||
/// a port already taken on the far machine, a `llama-server` too old for
|
||||
/// a flag. All of those exit within a second and none of them will ever
|
||||
/// answer `/health`, so waiting out the timeout turns a server that said
|
||||
/// exactly what was wrong into "gave up after 300s".
|
||||
fn wait_until_ready(endpoint: &str, session_dir: &Path) -> Result<()> {
|
||||
let deadline = std::time::Instant::now() + READY_TIMEOUT;
|
||||
let url = format!("{endpoint}/health");
|
||||
loop {
|
||||
if let Ok(response) = ureq::get(&url).call()
|
||||
&& response.status() == 200
|
||||
{
|
||||
return Ok(());
|
||||
}
|
||||
// `None` is the session having been stopped or deleted while this
|
||||
// waited, which is nobody's fault and still not worth waiting on.
|
||||
match process::recorded(session_dir) {
|
||||
Some((_, process::Liveness::Alive | process::Liveness::Unknown)) => {}
|
||||
Some((_, process::Liveness::Dead)) | None => {
|
||||
bail!("it exited before it answered.{}", log_tail(session_dir));
|
||||
}
|
||||
}
|
||||
if std::time::Instant::now() > deadline {
|
||||
bail!(
|
||||
"gave up after {}s.{}",
|
||||
READY_TIMEOUT.as_secs(),
|
||||
log_tail(session_dir)
|
||||
);
|
||||
}
|
||||
std::thread::sleep(std::time::Duration::from_millis(250));
|
||||
}
|
||||
}
|
||||
|
||||
/// The end of `llama-server`'s own log, for a failure message.
|
||||
///
|
||||
/// Its account of what went wrong is the useful half -- "failed to load
|
||||
/// model", "bind: Address already in use" -- and on a remote session it
|
||||
/// is the only half, since nobody reading the phone can open a file on
|
||||
/// that machine. Bounded, because this ends up in an event a phone draws.
|
||||
fn log_tail(session_dir: &Path) -> String {
|
||||
let Ok(text) = std::fs::read_to_string(session_dir.join(SERVER_LOG)) else {
|
||||
return String::new();
|
||||
};
|
||||
let tail: Vec<&str> = text.lines().rev().take(LOG_TAIL_LINES).collect();
|
||||
if tail.is_empty() {
|
||||
return String::new();
|
||||
}
|
||||
format!(
|
||||
" It last said: {}",
|
||||
tail.into_iter().rev().collect::<Vec<_>>().join(" / ")
|
||||
)
|
||||
}
|
||||
|
||||
/// How much of that log to carry into a message somebody reads on a phone.
|
||||
const LOG_TAIL_LINES: usize = 6;
|
||||
|
||||
/// One streamed completion: posts the conversation, emits each delta as it
|
||||
/// arrives. Emits rather than returns, because the transcript those events land
|
||||
/// in is what the next turn reads back.
|
||||
fn generate(
|
||||
endpoint: &str,
|
||||
messages: &[Message],
|
||||
sampling: &serde_json::Map<String, serde_json::Value>,
|
||||
cancel: &AtomicBool,
|
||||
sink: &EventSink,
|
||||
) -> Result<()> {
|
||||
let mut body = json!({
|
||||
"messages": messages,
|
||||
"stream": true,
|
||||
"stream_options": {"include_usage": true},
|
||||
});
|
||||
let map = body.as_object_mut().expect("built as an object");
|
||||
for (key, value) in sampling {
|
||||
map.insert(key.clone(), value.clone());
|
||||
}
|
||||
|
||||
let mut response = ureq::post(format!("{endpoint}/v1/chat/completions"))
|
||||
.header("Content-Type", "application/json")
|
||||
.send_json(&body)
|
||||
.context("asking llama-server to generate")?;
|
||||
|
||||
let reader = std::io::BufReader::new(response.body_mut().as_reader());
|
||||
let mut tokens = 0u64;
|
||||
// The prompt side only, which is what the model is holding -- the same
|
||||
// definition the other dialects report, so one word on the phone means one
|
||||
// thing whichever kind of session it is.
|
||||
let mut context = None;
|
||||
for line in std::io::BufRead::lines(reader) {
|
||||
if cancel.load(Ordering::Relaxed) {
|
||||
break;
|
||||
}
|
||||
let line = line.context("reading the generation stream")?;
|
||||
// Server-sent events: the payload lines are the ones that matter.
|
||||
let Some(payload) = line.strip_prefix("data: ") else {
|
||||
continue;
|
||||
};
|
||||
if payload.trim() == "[DONE]" {
|
||||
break;
|
||||
}
|
||||
let Ok(chunk) = serde_json::from_str::<serde_json::Value>(payload) else {
|
||||
continue;
|
||||
};
|
||||
if let Some(usage) = chunk.get("usage") {
|
||||
if let Some(total) = usage
|
||||
.get("total_tokens")
|
||||
.and_then(serde_json::Value::as_u64)
|
||||
{
|
||||
tokens = total;
|
||||
}
|
||||
if let Some(prompt) = usage
|
||||
.get("prompt_tokens")
|
||||
.and_then(serde_json::Value::as_u64)
|
||||
{
|
||||
context = Some(prompt);
|
||||
}
|
||||
}
|
||||
let delta = chunk
|
||||
.get("choices")
|
||||
.and_then(|c| c.get(0))
|
||||
.and_then(|c| c.get("delta"))
|
||||
.and_then(|d| d.get("content"))
|
||||
.and_then(serde_json::Value::as_str)
|
||||
.unwrap_or_default();
|
||||
if !delta.is_empty() {
|
||||
let _ = sink.send(Event::AssistantText {
|
||||
delta: delta.to_string(),
|
||||
});
|
||||
}
|
||||
}
|
||||
if tokens > 0 {
|
||||
let _ = sink.send(Event::UsageDelta { tokens, context });
|
||||
}
|
||||
Ok(())
|
||||
}
|
||||
|
||||
#[cfg(test)]
|
||||
mod tests {
|
||||
use super::*;
|
||||
use crate::session::transcript::Transcript;
|
||||
|
||||
/// Writes a transcript the way the pump does, so the fold is tested against
|
||||
/// the real file format rather than a hand-built vector.
|
||||
fn transcript_with(events: &[Event]) -> (tempfile::TempDir, PathBuf) {
|
||||
let dir = tempfile::tempdir().expect("tempdir");
|
||||
let path = dir.path().join("transcript.jsonl");
|
||||
let mut transcript = Transcript::open(&path).expect("open");
|
||||
for event in events {
|
||||
transcript.append(event.clone(), 0.0).expect("append");
|
||||
}
|
||||
(dir, path)
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn deltas_between_user_messages_are_one_assistant_turn() {
|
||||
let (_dir, path) = transcript_with(&[
|
||||
Event::UserMessage {
|
||||
id: None,
|
||||
text: "hello".into(),
|
||||
attachments: Vec::new(),
|
||||
},
|
||||
Event::AssistantText {
|
||||
delta: "hi ".into(),
|
||||
},
|
||||
Event::AssistantText {
|
||||
delta: "there".into(),
|
||||
},
|
||||
Event::Status {
|
||||
state: SessionStatus::Idle,
|
||||
},
|
||||
Event::UserMessage {
|
||||
id: None,
|
||||
text: "again".into(),
|
||||
attachments: Vec::new(),
|
||||
},
|
||||
Event::AssistantText {
|
||||
delta: "yes".into(),
|
||||
},
|
||||
]);
|
||||
let messages = conversation(&path);
|
||||
assert_eq!(
|
||||
messages
|
||||
.iter()
|
||||
.map(|m| (m.role.as_str(), m.content.as_str()))
|
||||
.collect::<Vec<_>>(),
|
||||
[
|
||||
("user", "hello"),
|
||||
("assistant", "hi there"),
|
||||
("user", "again"),
|
||||
("assistant", "yes")
|
||||
],
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
/// The interrupted case, which decides what a resumed conversation is built
|
||||
/// from: whatever the phone was shown. The deltas that arrived before the
|
||||
/// stop are in the transcript, so they are in the prompt -- the model is
|
||||
/// never told it said something the user did not see.
|
||||
fn an_interrupted_reply_stays_in_the_conversation() {
|
||||
let (_dir, path) = transcript_with(&[
|
||||
Event::UserMessage {
|
||||
id: None,
|
||||
text: "count".into(),
|
||||
attachments: Vec::new(),
|
||||
},
|
||||
Event::AssistantText {
|
||||
delta: "one two".into(),
|
||||
},
|
||||
Event::Status {
|
||||
state: SessionStatus::Idle,
|
||||
},
|
||||
]);
|
||||
let messages = conversation(&path);
|
||||
assert_eq!(messages.len(), 2);
|
||||
assert_eq!(messages[1].content, "one two");
|
||||
}
|
||||
|
||||
#[test]
|
||||
/// Events this driver does not produce must not disturb the fold: a
|
||||
/// transcript can carry errors and status changes from a session that
|
||||
/// was, say, relaunched.
|
||||
fn other_events_are_not_part_of_the_conversation() {
|
||||
let (_dir, path) = transcript_with(&[
|
||||
Event::Status {
|
||||
state: SessionStatus::Running,
|
||||
},
|
||||
Event::UserMessage {
|
||||
id: None,
|
||||
text: "hello".into(),
|
||||
attachments: Vec::new(),
|
||||
},
|
||||
Event::Error {
|
||||
message: "something went wrong".into(),
|
||||
},
|
||||
Event::AssistantText {
|
||||
delta: "still here".into(),
|
||||
},
|
||||
Event::UsageDelta {
|
||||
tokens: 12,
|
||||
context: Some(12),
|
||||
},
|
||||
]);
|
||||
let messages = conversation(&path);
|
||||
assert_eq!(messages.len(), 2);
|
||||
assert_eq!(messages[0].content, "hello");
|
||||
assert_eq!(messages[1].content, "still here");
|
||||
}
|
||||
|
||||
#[test]
|
||||
/// Clearing decides what the *model* is given, not just what the phone
|
||||
/// draws. Everything above the marker stays in the transcript and none of it
|
||||
/// is sent.
|
||||
fn the_conversation_starts_after_the_last_clear() {
|
||||
let (_dir, path) = transcript_with(&[
|
||||
Event::UserMessage {
|
||||
id: None,
|
||||
text: "the long expensive conversation".into(),
|
||||
attachments: Vec::new(),
|
||||
},
|
||||
Event::AssistantText {
|
||||
delta: "at length".into(),
|
||||
},
|
||||
Event::Cleared,
|
||||
Event::UserMessage {
|
||||
id: None,
|
||||
text: "a fresh start".into(),
|
||||
attachments: Vec::new(),
|
||||
},
|
||||
Event::AssistantText {
|
||||
delta: "cheaply".into(),
|
||||
},
|
||||
]);
|
||||
let messages = conversation(&path);
|
||||
assert_eq!(messages.len(), 2);
|
||||
assert_eq!(messages[0].content, "a fresh start");
|
||||
assert_eq!(messages[1].content, "cheaply");
|
||||
}
|
||||
|
||||
#[test]
|
||||
/// The *last* one, so clearing twice does not resurrect what the
|
||||
/// first clear dropped.
|
||||
fn only_the_newest_clear_counts() {
|
||||
let (_dir, path) = transcript_with(&[
|
||||
Event::UserMessage {
|
||||
id: None,
|
||||
text: "one".into(),
|
||||
attachments: Vec::new(),
|
||||
},
|
||||
Event::Cleared,
|
||||
Event::UserMessage {
|
||||
id: None,
|
||||
text: "two".into(),
|
||||
attachments: Vec::new(),
|
||||
},
|
||||
Event::Cleared,
|
||||
Event::UserMessage {
|
||||
id: None,
|
||||
text: "three".into(),
|
||||
attachments: Vec::new(),
|
||||
},
|
||||
]);
|
||||
let messages = conversation(&path);
|
||||
assert_eq!(messages.len(), 1);
|
||||
assert_eq!(messages[0].content, "three");
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn a_model_key_cannot_climb_out_of_the_models_directory() {
|
||||
let dir = tempfile::tempdir().expect("tempdir");
|
||||
for attempt in ["../../etc/passwd", "unsloth/../../escape.gguf", ""] {
|
||||
assert!(
|
||||
model_path(dir.path(), attempt).is_err(),
|
||||
"{attempt:?} should have been refused",
|
||||
);
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,375 @@
|
||||
//! An MCP client, for the tools a llama session has that `llama-server` does
|
||||
//! not provide itself.
|
||||
//!
|
||||
//! **Why this is here and not a flag on `llama-server`.** That server can host
|
||||
//! MCP servers (`--mcp-servers-json`), but only ones it can *spawn*: its
|
||||
//! configuration is Cursor's, and an entry without a `command` is skipped with
|
||||
//! "MCP server 'exa' has no command". Exa's is a remote HTTP endpoint with
|
||||
//! nothing to spawn, so reaching it that way means a local process bridging
|
||||
//! stdio to HTTP -- a Node install on the machine serving the model, and a
|
||||
//! package to keep current, for what is three JSON-RPC calls.
|
||||
//!
|
||||
//! llama.cpp's own web UI does not do that either. It ships Exa in a
|
||||
//! "recommended servers" list and connects to `https://mcp.exa.ai/mcp`
|
||||
//! *itself*, from the browser. This is the same arrangement with this server
|
||||
//! in the browser's place, and it is the right one for a second reason: it
|
||||
//! puts the search on the machine running the backend rather than on whichever
|
||||
//! machine happens to be serving the model, which may have no route out at
|
||||
//! all.
|
||||
//!
|
||||
//! **Only the three calls a tool needs.** `initialize`, `tools/list`,
|
||||
//! `tools/call`. Nothing here implements resources, prompts, sampling or the
|
||||
//! server-to-client stream, because nothing here uses them; a session's tools
|
||||
//! are a list fetched once and a call made on demand. That is why this is a
|
||||
//! file rather than a dependency on a protocol crate -- there is no spec
|
||||
//! surface to get subtly wrong, only a request and its reply.
|
||||
//!
|
||||
//! Transport is "streamable HTTP": every message is a POST, and the reply is
|
||||
//! either JSON or a one-event SSE stream carrying the same JSON. Both are
|
||||
//! accepted because which one arrives is the server's choice, not ours.
|
||||
|
||||
use anyhow::{Context, Result, bail};
|
||||
use serde_json::{Value, json};
|
||||
|
||||
/// Identifies this client to an MCP server.
|
||||
///
|
||||
/// Not politeness: Exa's endpoint is behind Cloudflare, which answers **403**
|
||||
/// to a request with no `User-Agent` at all (measured 2026-09-19 -- the same
|
||||
/// request with one succeeds). A client that omitted it would look exactly
|
||||
/// like a server that was refusing us.
|
||||
const USER_AGENT: &str = concat!("ai-server/", env!("CARGO_PKG_VERSION"));
|
||||
|
||||
/// The protocol version this speaks. Sent at `initialize`; a server that
|
||||
/// prefers another says so in its answer and this goes along with whatever it
|
||||
/// then sends, since none of the three calls here has changed between
|
||||
/// versions.
|
||||
const PROTOCOL_VERSION: &str = "2025-06-18";
|
||||
|
||||
/// How long any one call may take.
|
||||
///
|
||||
/// Generous because a web search is a search: Exa fetches and cleans pages
|
||||
/// before answering. Bounded at all because this blocks a turn, and a tool
|
||||
/// that never returns is a session that never speaks again.
|
||||
const CALL_TIMEOUT: std::time::Duration = std::time::Duration::from_secs(120);
|
||||
|
||||
/// A connected MCP server, and the tools it offered.
|
||||
pub struct McpServer {
|
||||
/// The name this server is configured under. It prefixes every tool, so
|
||||
/// two servers offering `search` are two different tools.
|
||||
name: String,
|
||||
url: String,
|
||||
/// What the server called this conversation, when it named one. Sent back
|
||||
/// on every later request; a server that keeps no session sends no header
|
||||
/// and this stays `None`.
|
||||
session: Option<String>,
|
||||
/// The tool names this server answers to, without the prefix, keyed by the
|
||||
/// prefixed name the model is given.
|
||||
tools: Vec<McpTool>,
|
||||
}
|
||||
|
||||
/// One tool an MCP server offers, in both the names it has.
|
||||
pub struct McpTool {
|
||||
/// `{server}_{tool}` -- what the model calls it, and what comes back in a
|
||||
/// tool call. Prefixed the way `llama-server` prefixes the MCP tools it
|
||||
/// hosts itself, so a reader sees one naming convention whichever side a
|
||||
/// tool came from.
|
||||
pub qualified: String,
|
||||
/// What the server calls it.
|
||||
bare: String,
|
||||
/// The OpenAI-shaped function definition sent to the model.
|
||||
pub definition: Value,
|
||||
}
|
||||
|
||||
impl McpServer {
|
||||
/// Connects, handshakes, and asks what it can do.
|
||||
///
|
||||
/// All three steps or none: a server that answered `initialize` and then
|
||||
/// failed to list its tools is not a server with no tools, and returning
|
||||
/// an empty list for it would put a session on screen that silently
|
||||
/// cannot search.
|
||||
pub fn connect(name: &str, url: &str) -> Result<Self> {
|
||||
let mut server = Self {
|
||||
name: name.to_string(),
|
||||
url: url.to_string(),
|
||||
session: None,
|
||||
tools: Vec::new(),
|
||||
};
|
||||
server
|
||||
.request(
|
||||
1,
|
||||
"initialize",
|
||||
json!({
|
||||
"protocolVersion": PROTOCOL_VERSION,
|
||||
"capabilities": {},
|
||||
"clientInfo": {"name": "ai-server", "title": "AI Sessions", "version": env!("CARGO_PKG_VERSION")},
|
||||
}),
|
||||
)
|
||||
.with_context(|| format!("handshaking with the {name} MCP server at {url}"))?;
|
||||
// A notification: no id, and the server answers with no body. Sent
|
||||
// because the specification requires it before any other call, and
|
||||
// Exa's server does enforce it.
|
||||
server.notify("notifications/initialized")?;
|
||||
let listed = server
|
||||
.request(2, "tools/list", json!({}))
|
||||
.with_context(|| format!("asking the {name} MCP server what it offers"))?;
|
||||
server.tools = listed
|
||||
.get("tools")
|
||||
.and_then(Value::as_array)
|
||||
.map(|tools| {
|
||||
tools
|
||||
.iter()
|
||||
.filter_map(|tool| server.describe(tool))
|
||||
.collect()
|
||||
})
|
||||
.unwrap_or_default();
|
||||
Ok(server)
|
||||
}
|
||||
|
||||
/// Turns one entry of `tools/list` into the function definition a model is
|
||||
/// given, or `None` for one this cannot name or call.
|
||||
fn describe(&self, tool: &Value) -> Option<McpTool> {
|
||||
let bare = tool.get("name").and_then(Value::as_str)?.to_string();
|
||||
let qualified = format!("{}_{bare}", self.name);
|
||||
let mut function = serde_json::Map::new();
|
||||
function.insert("name".into(), json!(qualified));
|
||||
if let Some(description) = tool.get("description").and_then(Value::as_str) {
|
||||
function.insert("description".into(), json!(description));
|
||||
}
|
||||
// `inputSchema` in MCP, `parameters` in the OpenAI shape: the same
|
||||
// JSON Schema under two names. A tool that declares none takes no
|
||||
// arguments, which is an empty object rather than an absent key --
|
||||
// some templates render the key unconditionally.
|
||||
function.insert(
|
||||
"parameters".into(),
|
||||
tool.get("inputSchema")
|
||||
.cloned()
|
||||
.unwrap_or_else(|| json!({"type": "object", "properties": {}})),
|
||||
);
|
||||
Some(McpTool {
|
||||
qualified,
|
||||
bare,
|
||||
definition: json!({"type": "function", "function": function}),
|
||||
})
|
||||
}
|
||||
|
||||
pub fn tools(&self) -> &[McpTool] {
|
||||
&self.tools
|
||||
}
|
||||
|
||||
/// Runs one of this server's tools, named as the model named it.
|
||||
///
|
||||
/// The result is the text a model is shown. A tool the server reports as
|
||||
/// failing is **not** an error here: `isError` means the tool ran and went
|
||||
/// wrong -- a search that found nothing, a page that would not fetch --
|
||||
/// and the model is the one that has to know, so it comes back as its own
|
||||
/// message. An error is reserved for not having reached the server at all.
|
||||
pub fn call(&mut self, qualified: &str, arguments: &Value) -> Result<String> {
|
||||
// The bare name is taken before the call, because the call needs the
|
||||
// whole of `self` and the tool list is part of it.
|
||||
let bare = self
|
||||
.tools
|
||||
.iter()
|
||||
.find(|tool| tool.qualified == qualified)
|
||||
.map(|tool| tool.bare.clone())
|
||||
.with_context(|| format!("{} does not offer {qualified}", self.name))?;
|
||||
let result = self.request(
|
||||
3,
|
||||
"tools/call",
|
||||
json!({"name": bare, "arguments": arguments}),
|
||||
)?;
|
||||
Ok(rendered(&result))
|
||||
}
|
||||
|
||||
/// One request, and its result.
|
||||
///
|
||||
/// `&self` rather than `&mut self` everywhere but the handshake would be
|
||||
/// tidier and is wrong: the session header is assigned by the server on
|
||||
/// the first reply and has to be kept.
|
||||
fn request(&mut self, id: u64, method: &str, params: Value) -> Result<Value> {
|
||||
let body = json!({"jsonrpc": "2.0", "id": id, "method": method, "params": params});
|
||||
let answer = self.post(&body)?.with_context(|| {
|
||||
format!(
|
||||
"the {} MCP server answered {method} with nothing",
|
||||
self.name
|
||||
)
|
||||
})?;
|
||||
if let Some(message) = answer.pointer("/error/message").and_then(Value::as_str) {
|
||||
bail!("{} refused {method}: {message}", self.name);
|
||||
}
|
||||
answer
|
||||
.get("result")
|
||||
.cloned()
|
||||
.with_context(|| format!("the {} MCP server's {method} carried no result", self.name))
|
||||
}
|
||||
|
||||
/// A message with no id, which is answered with no body.
|
||||
fn notify(&mut self, method: &str) -> Result<()> {
|
||||
self.post(&json!({"jsonrpc": "2.0", "method": method}))?;
|
||||
Ok(())
|
||||
}
|
||||
|
||||
/// Posts one JSON-RPC message and returns whatever came back, which for a
|
||||
/// notification is nothing.
|
||||
fn post(&mut self, body: &Value) -> Result<Option<Value>> {
|
||||
let mut request = ureq::post(&self.url)
|
||||
.config()
|
||||
.timeout_global(Some(CALL_TIMEOUT))
|
||||
.build()
|
||||
.header("content-type", "application/json")
|
||||
// Both, because which one a server replies with is its choice.
|
||||
.header("accept", "application/json, text/event-stream")
|
||||
.header("user-agent", USER_AGENT);
|
||||
if let Some(session) = &self.session {
|
||||
request = request.header("mcp-session-id", session);
|
||||
}
|
||||
let mut response = request
|
||||
.send_json(body)
|
||||
.with_context(|| format!("reaching the {} MCP server at {}", self.name, self.url))?;
|
||||
if let Some(session) = response
|
||||
.headers()
|
||||
.get("mcp-session-id")
|
||||
.and_then(|value| value.to_str().ok())
|
||||
{
|
||||
self.session = Some(session.to_string());
|
||||
}
|
||||
let streamed = response
|
||||
.headers()
|
||||
.get("content-type")
|
||||
.and_then(|value| value.to_str().ok())
|
||||
.is_some_and(|value| value.contains("text/event-stream"));
|
||||
let text = response
|
||||
.body_mut()
|
||||
.read_to_string()
|
||||
.with_context(|| format!("reading the {} MCP server's answer", self.name))?;
|
||||
Ok(first_message(&text, streamed))
|
||||
}
|
||||
}
|
||||
|
||||
/// The first JSON-RPC message in a reply body.
|
||||
///
|
||||
/// One, not all: every call here carries a single id and the server answers it
|
||||
/// once. Server-sent events are unwrapped to their payload lines; a plain JSON
|
||||
/// body is itself.
|
||||
fn first_message(text: &str, streamed: bool) -> Option<Value> {
|
||||
if streamed {
|
||||
return text
|
||||
.lines()
|
||||
.filter_map(|line| line.strip_prefix("data: "))
|
||||
.find_map(|payload| serde_json::from_str(payload).ok());
|
||||
}
|
||||
serde_json::from_str(text.trim()).ok()
|
||||
}
|
||||
|
||||
/// A `tools/call` result as the text a model is given.
|
||||
///
|
||||
/// MCP answers with a list of content blocks; the text ones are joined and the
|
||||
/// rest are named rather than dropped, because a model told nothing came back
|
||||
/// will try again. `structuredContent` is used when there is no text at all,
|
||||
/// which is how some servers answer entirely.
|
||||
fn rendered(result: &Value) -> String {
|
||||
let blocks = result.get("content").and_then(Value::as_array);
|
||||
let mut parts: Vec<String> = Vec::new();
|
||||
for block in blocks.into_iter().flatten() {
|
||||
match block.get("type").and_then(Value::as_str) {
|
||||
Some("text") => parts.push(
|
||||
block
|
||||
.get("text")
|
||||
.and_then(Value::as_str)
|
||||
.unwrap_or_default()
|
||||
.to_string(),
|
||||
),
|
||||
Some(kind) => parts.push(format!("[{kind} content, which this session cannot show]")),
|
||||
None => {}
|
||||
}
|
||||
}
|
||||
if parts.iter().all(|part| part.trim().is_empty())
|
||||
&& let Some(structured) = result.get("structuredContent")
|
||||
{
|
||||
return structured.to_string();
|
||||
}
|
||||
parts.join("\n")
|
||||
}
|
||||
|
||||
#[cfg(test)]
|
||||
mod tests {
|
||||
use super::*;
|
||||
|
||||
#[test]
|
||||
fn an_event_stream_body_is_unwrapped_to_its_payload() {
|
||||
let body =
|
||||
"event: message\ndata: {\"jsonrpc\":\"2.0\",\"id\":1,\"result\":{\"ok\":true}}\n\n";
|
||||
assert_eq!(
|
||||
first_message(body, true),
|
||||
Some(json!({"jsonrpc": "2.0", "id": 1, "result": {"ok": true}})),
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn a_plain_json_body_is_the_message() {
|
||||
let body = " {\"jsonrpc\":\"2.0\",\"id\":1,\"result\":{}}\n";
|
||||
assert_eq!(
|
||||
first_message(body, false),
|
||||
Some(json!({"jsonrpc": "2.0", "id": 1, "result": {}})),
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
/// A notification's reply, which is nothing at all.
|
||||
fn an_empty_body_is_no_message() {
|
||||
assert_eq!(first_message("", false), None);
|
||||
assert_eq!(first_message("event: ping\n", true), None);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn text_blocks_are_joined_and_other_kinds_are_named() {
|
||||
let result = json!({"content": [
|
||||
{"type": "text", "text": "first"},
|
||||
{"type": "image", "data": "…"},
|
||||
{"type": "text", "text": "second"},
|
||||
]});
|
||||
assert_eq!(
|
||||
rendered(&result),
|
||||
"first\n[image content, which this session cannot show]\nsecond",
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
/// A server that answers only in structured form. Rendering "" for it
|
||||
/// would tell the model the search came back empty, which is a different
|
||||
/// fact from the one that is true.
|
||||
fn a_result_with_no_text_falls_back_to_its_structured_form() {
|
||||
let result = json!({"content": [], "structuredContent": {"hits": 2}});
|
||||
assert_eq!(rendered(&result), "{\"hits\":2}");
|
||||
}
|
||||
|
||||
#[test]
|
||||
/// The real endpoint, which is the only thing that can confirm the
|
||||
/// handshake, the session header and the SSE unwrapping all agree with a
|
||||
/// server nobody here wrote. Skipped without network rather than failed:
|
||||
/// `./run-tests.sh` has to pass on a machine with no route out.
|
||||
fn exa_answers_a_search_over_the_real_protocol() {
|
||||
let Ok(mut server) = McpServer::connect("exa", super::super::EXA_MCP_URL) else {
|
||||
eprintln!("skipping: could not reach Exa");
|
||||
return;
|
||||
};
|
||||
assert!(
|
||||
server
|
||||
.tools()
|
||||
.iter()
|
||||
.any(|tool| tool.qualified == "exa_web_search_exa"),
|
||||
"Exa offered {:?}",
|
||||
server
|
||||
.tools()
|
||||
.iter()
|
||||
.map(|tool| &tool.qualified)
|
||||
.collect::<Vec<_>>(),
|
||||
);
|
||||
let answer = server
|
||||
.call(
|
||||
"exa_web_search_exa",
|
||||
&json!({"query": "llama.cpp server", "numResults": 1}),
|
||||
)
|
||||
.expect("search");
|
||||
assert!(!answer.trim().is_empty(), "a search returned nothing");
|
||||
}
|
||||
}
|
||||
File diff suppressed because it is too large.
Load diff
@@ -0,0 +1,263 @@
|
||||
//! What a llama session can do besides talk, and who runs it.
|
||||
//!
|
||||
//! Two sources, one list. `llama-server` started with `--tools` runs a set of
|
||||
//! its own -- reading, searching, editing, a shell -- and publishes them at
|
||||
//! `GET /tools` in the shape a model is given, with `POST /tools` to run one.
|
||||
//! Anything else comes from an MCP server this backend is connected to (see
|
||||
//! [`super::mcp`]). Both arrive here as a definition to offer and a way to
|
||||
//! call, and nothing downstream of [`Tools::execute`] knows which a tool was.
|
||||
//!
|
||||
//! **The built-in tools run where the model does, and that is the point.** A
|
||||
//! session on another machine edits that machine's files, because that is the
|
||||
//! machine `llama-server` is on -- the same rule the model file already
|
||||
//! follows. MCP tools run here instead, which is right for the opposite
|
||||
//! reason: a web search wants the machine with a route out, not the one with
|
||||
//! the GPU.
|
||||
//!
|
||||
//! **A tool's failure is a result, not an error.** A missing file, a command
|
||||
//! that exited non-zero, a search that found nothing: all of those are things
|
||||
//! the model has to read and act on, so they come back as the tool's output.
|
||||
//! [`Tools::execute`] returns `Err` only when the tool could not be reached at
|
||||
//! all, which is a fact about this server rather than about the work.
|
||||
|
||||
use std::collections::HashMap;
|
||||
use std::sync::{Arc, Mutex};
|
||||
|
||||
use anyhow::{Context, Result};
|
||||
use serde_json::{Value, json};
|
||||
|
||||
use super::mcp::McpServer;
|
||||
|
||||
/// How long one tool call may take.
|
||||
///
|
||||
/// This is the shell tool's budget as much as anything: a build, a test run,
|
||||
/// a `find` over a large tree. Bounded because it blocks the turn, and a
|
||||
/// session stuck behind a command that will never finish cannot even be told
|
||||
/// to stop.
|
||||
const EXECUTE_TIMEOUT: std::time::Duration = std::time::Duration::from_secs(300);
|
||||
|
||||
/// The tools one session has, and how to run each of them.
|
||||
pub struct Tools {
|
||||
/// The `llama-server` these belong to. Replaced when the session's model
|
||||
/// changes, because that is a different server on a different port.
|
||||
endpoint: String,
|
||||
/// What the model is given, in the order it is offered: the server's own
|
||||
/// tools first, then each MCP server's.
|
||||
definitions: Vec<Value>,
|
||||
/// The server's tools, and whether each is run relative to a working
|
||||
/// directory. Only the ones that say so are sent one -- a tool that
|
||||
/// ignores it would still have its cache keyed on it.
|
||||
server: HashMap<String, bool>,
|
||||
/// Connected MCP servers, each of which knows its own tools by the
|
||||
/// prefixed names they were offered under.
|
||||
mcp: Vec<Arc<Mutex<McpServer>>>,
|
||||
}
|
||||
|
||||
impl Tools {
|
||||
/// Asks a ready `llama-server` what it offers and adds what the MCP
|
||||
/// servers offered.
|
||||
///
|
||||
/// A server started without `--tools` answers with an empty list, and a
|
||||
/// session with only MCP tools is a perfectly good session -- so nothing
|
||||
/// here treats "no tools" as a failure. What *is* a failure is not being
|
||||
/// able to ask, because that is the same server the conversation is about
|
||||
/// to go to.
|
||||
pub fn discover(endpoint: &str, mcp: Vec<Arc<Mutex<McpServer>>>) -> Result<Self> {
|
||||
let catalog: Vec<Value> = ureq::get(format!("{endpoint}/tools"))
|
||||
.call()
|
||||
.context("asking llama-server which tools it has")?
|
||||
.body_mut()
|
||||
.read_json()
|
||||
.context("reading llama-server's tool list")?;
|
||||
let mut definitions = Vec::new();
|
||||
let mut server = HashMap::new();
|
||||
for entry in &catalog {
|
||||
let Some(name) = entry
|
||||
.pointer("/definition/function/name")
|
||||
.and_then(Value::as_str)
|
||||
else {
|
||||
continue;
|
||||
};
|
||||
let Some(definition) = entry.get("definition") else {
|
||||
continue;
|
||||
};
|
||||
server.insert(
|
||||
name.to_string(),
|
||||
entry
|
||||
.get("uses_cwd")
|
||||
.and_then(Value::as_bool)
|
||||
.unwrap_or(false),
|
||||
);
|
||||
definitions.push(definition.clone());
|
||||
}
|
||||
for connected in &mcp {
|
||||
for tool in connected.lock().unwrap().tools() {
|
||||
definitions.push(tool.definition.clone());
|
||||
}
|
||||
}
|
||||
Ok(Self {
|
||||
endpoint: endpoint.to_string(),
|
||||
definitions,
|
||||
server,
|
||||
mcp,
|
||||
})
|
||||
}
|
||||
|
||||
/// What goes in the request's `tools`, or `None` when there is nothing to
|
||||
/// offer.
|
||||
///
|
||||
/// Absent rather than empty for a reason that shows on screen: a chat
|
||||
/// template branches on whether tools were given, and an empty list
|
||||
/// renders the whole "you may call one or more functions" preamble with no
|
||||
/// functions under it.
|
||||
pub fn offered(&self) -> Option<&[Value]> {
|
||||
(!self.definitions.is_empty()).then_some(&self.definitions)
|
||||
}
|
||||
|
||||
/// Whether this is a tool at all, which decides what to do about a call
|
||||
/// naming something else.
|
||||
pub fn knows(&self, name: &str) -> bool {
|
||||
self.server.contains_key(name) || self.mcp_for(name).is_some()
|
||||
}
|
||||
|
||||
/// The MCP server that offered `name`, if one did.
|
||||
fn mcp_for(&self, name: &str) -> Option<&Arc<Mutex<McpServer>>> {
|
||||
self.mcp.iter().find(|server| {
|
||||
server
|
||||
.lock()
|
||||
.unwrap()
|
||||
.tools()
|
||||
.iter()
|
||||
.any(|tool| tool.qualified == name)
|
||||
})
|
||||
}
|
||||
|
||||
/// Runs one call and returns what the model should read.
|
||||
///
|
||||
/// `cwd` is the session's working directory, sent only to the tools that
|
||||
/// say they use one. A session with no working directory sends none, and
|
||||
/// `llama-server` falls back to its own -- which is the honest outcome:
|
||||
/// this server has no better answer for where "here" is.
|
||||
pub fn execute(&self, name: &str, arguments: &Value, cwd: Option<&str>) -> Result<String> {
|
||||
if let Some(server) = self.mcp_for(name) {
|
||||
return server.lock().unwrap().call(name, arguments);
|
||||
}
|
||||
let uses_cwd = *self
|
||||
.server
|
||||
.get(name)
|
||||
.with_context(|| format!("no tool called {name}"))?;
|
||||
let mut request = ureq::post(format!("{}/tools", self.endpoint))
|
||||
.config()
|
||||
.timeout_global(Some(EXECUTE_TIMEOUT))
|
||||
// The refusal is a sentence the model can act on, so it is read as
|
||||
// one rather than discarded in favour of its status code -- the
|
||||
// same reason `super::refusal` exists for generation.
|
||||
.http_status_as_error(false)
|
||||
.build()
|
||||
.header("content-type", "application/json");
|
||||
if let (true, Some(cwd)) = (uses_cwd, cwd) {
|
||||
request = request.header("x-tool-cwd", cwd);
|
||||
}
|
||||
let mut response = request
|
||||
.send_json(json!({"tool": name, "params": arguments}))
|
||||
.with_context(|| format!("asking llama-server to run {name}"))?;
|
||||
let body = response
|
||||
.body_mut()
|
||||
.read_to_string()
|
||||
.with_context(|| format!("reading what {name} produced"))?;
|
||||
Ok(match serde_json::from_str::<Value>(&body) {
|
||||
Ok(answer) => result_text(&answer),
|
||||
// Not JSON at all: hand over what was said rather than a parse
|
||||
// error about it, since the model is what has to carry on.
|
||||
Err(_) => body,
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
/// `POST /tools`'s answer as the text a model is given.
|
||||
///
|
||||
/// The server answers `plain_text_response` for a tool that ran and `error`
|
||||
/// for one that did not, and both are the model's business -- see this
|
||||
/// module's note on failures being results. Anything else is handed over as
|
||||
/// itself rather than discarded, since a tool this build has not seen before
|
||||
/// is exactly the case where guessing is worst.
|
||||
fn result_text(answer: &Value) -> String {
|
||||
if let Some(text) = answer.get("plain_text_response").and_then(Value::as_str) {
|
||||
return text.to_string();
|
||||
}
|
||||
if let Some(message) = answer.get("error").and_then(Value::as_str) {
|
||||
return message.to_string();
|
||||
}
|
||||
answer.to_string()
|
||||
}
|
||||
|
||||
/// How much a session asks before it acts.
|
||||
///
|
||||
/// Two, because two is what the mechanism underneath actually has. The web UI
|
||||
/// that ships with `llama-server` asks before every call and remembers the
|
||||
/// tools you said "always" to, and that pair -- a prompt and a growing set of
|
||||
/// exceptions -- is the whole of its permission model. A third mode sitting
|
||||
/// between them would have to invent a rule about which tools are "edits",
|
||||
/// and the rule would be this app's opinion rather than anything the tools
|
||||
/// declare.
|
||||
pub const MODES: &[&str] = &["manual", "bypassPermissions"];
|
||||
|
||||
/// What a new llama session asks by default.
|
||||
///
|
||||
/// The cautious one, matching the web UI: a model with a shell on somebody's
|
||||
/// own machine is the case to be wrong about in this direction, and one tap
|
||||
/// on "always allow" is what makes it bearable afterwards.
|
||||
pub const DEFAULT_MODE: &str = "manual";
|
||||
|
||||
/// The answer that makes an allowance permanent for the session. The tool's
|
||||
/// name follows it, which is what makes the transcript alone enough to
|
||||
/// rebuild the set -- see `super::allowed`.
|
||||
pub const ALWAYS_PREFIX: &str = "Always allow ";
|
||||
pub const ALLOW_ONCE: &str = "Allow once";
|
||||
pub const REFUSE: &str = "Don't allow";
|
||||
|
||||
/// What the model is told when a call was refused.
|
||||
///
|
||||
/// Addressed to the model, not to the reader: it has to understand that the
|
||||
/// work did not happen and that trying the same call again is not the way
|
||||
/// round it, or it retries in a loop.
|
||||
pub const REFUSED: &str = "The person using this session did not allow this call, so it was not run. Do not try it \
|
||||
again -- say what you were going to do and why it needed that, and let them decide.";
|
||||
|
||||
/// What stands in for a call that never finished, when the transcript is read
|
||||
/// back into a conversation.
|
||||
///
|
||||
/// Every tool call in the history owes a result, because that is the shape a
|
||||
/// chat template renders; a turn stopped between the call and its result
|
||||
/// leaves one that has none. Saying so is better than inventing an outcome,
|
||||
/// and better than dropping the call -- which would tell the model it never
|
||||
/// asked.
|
||||
pub const UNFINISHED: &str = "This call was interrupted before it produced anything.";
|
||||
|
||||
#[cfg(test)]
|
||||
mod tests {
|
||||
use super::*;
|
||||
|
||||
#[test]
|
||||
fn a_tool_that_ran_reads_as_its_output() {
|
||||
assert_eq!(
|
||||
result_text(&json!({"plain_text_response": "hello\n"})),
|
||||
"hello\n",
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
/// The model is told what went wrong, because the model is what has to do
|
||||
/// something about it -- read a different path, fix the command.
|
||||
fn a_tool_that_failed_reads_as_its_message() {
|
||||
assert_eq!(
|
||||
result_text(&json!({"error": "cannot stat file: /tmp/nope"})),
|
||||
"cannot stat file: /tmp/nope",
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn anything_else_is_handed_over_as_itself() {
|
||||
assert_eq!(result_text(&json!({"rows": 2})), "{\"rows\":2}");
|
||||
}
|
||||
}
|
||||
@@ -3980,6 +3980,7 @@ mod tests {
|
||||
kind: DriverKind::ClaudeCli,
|
||||
command: Some(command.to_string_lossy().into_owned()),
|
||||
models: Vec::new(),
|
||||
mcp_servers: Vec::new(),
|
||||
},
|
||||
])],
|
||||
..Config::default()
|
||||
|
||||
@@ -1108,6 +1108,7 @@ mod tests {
|
||||
kind: DriverKind::ClaudeCli,
|
||||
command: None,
|
||||
models: vec![],
|
||||
mcp_servers: Vec::new(),
|
||||
}],
|
||||
}
|
||||
}
|
||||
@@ -1171,6 +1172,7 @@ mod tests {
|
||||
kind: DriverKind::Echo,
|
||||
command: None,
|
||||
models: vec![],
|
||||
mcp_servers: Vec::new(),
|
||||
}];
|
||||
// A machine with no Claude on it has no Claude limits, and a row
|
||||
// reporting on it would be a fact about nothing. Echo included:
|
||||
|
||||
Reference in new issue
Block a user