Give llama.cpp sessions tools, web search and a model picker
A llama session was a chat box: no tools, a fixed model, no permission
mode, and a model name drawn as the path the file sits at. It now runs the
agent loop itself, which is what the pieces below all hang off.
Tools are `llama-server`'s own (`--tools all`), which that server both
publishes and runs -- `GET /tools` for the definitions, `POST /tools` to
call one. Web search is Exa's MCP server, reached from this backend rather
than from the machine serving the model: that is what llama.cpp's own web
UI does, and it puts the search on the machine with a route out instead of
the one with the GPU. `llama-server`'s `--mcp-servers-json` can only spawn
local commands, so using it would have meant a Node bridge on every
machine that serves a model.
Driving the loop is what makes the permission gate ours. Two modes,
`manual` and `bypassPermissions`, which is what the mechanism has: the web
UI asks before every call and remembers the tools you say "always" to. The
allowances fold back out of the transcript's own answers, so they survive
a restart and a model change without being stored anywhere else.
Also here, because tools made each of them matter:
- **Loading is a state.** A 12 GB model takes twenty seconds to reach
memory and refuses everything until it has; the session used to report
`running` for that whole time, and a message sent meanwhile came back as
an error. It is `loading` now, and the message waits.
- **The model can be changed.** A `llama-server` holds one model, so this
stops it and starts another. The conversation survives because it was
never in the server.
- **Models are named, not pathed.** `general.name` read out of the file
itself -- over ssh too, in the round trip the spawn was already making.
Where two models share a name the file name breaks the tie.
- **`-np 1`, and the MTP draft head where the file has one.** Measured on
the 27B here: 41.5 tok/s plain, 61.4 with `--spec-type draft-mtp` at one
slot, and 28 with it at four -- speculating against a split KV cache is
worse than not speculating. The flag is conditional because asking for a
head that is not there makes `llama-server` exit.
- **A refusal says what to do.** Tool results are thousands of tokens, so
an overrun context is now ordinary; it was "http status: 400" and is now
the server's own "exceeds the available context size, try increasing it".
`GET /machines/{id}/models` is gone: the provider models route answers the
same question, and two answers to one question is how a picker comes to
offer a model the spawn screen does not.
Verified end to end against real models: a tool call asked and allowed, an
Exa search, a shell command, a 27B loaded while a message waited on it, a
model switch mid-session, a second message queued behind a running turn,
and the whole of it again on a session running over ssh.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
1 parent
392cc5413d
commit
ac476ab0c9
23 files changed
+3627
-1096
No files matched your search
@@ -143,18 +143,42 @@ moment you use it — `ANDROID_SERIAL=$(emu serial) ./gradlew …`.
|
||||
|
||||
### Testing llama.cpp and ssh here
|
||||
|
||||
**Both are set up here as of 2026-09-04** and need nothing typed. The
|
||||
prebuilt CPU llama.cpp lives outside the repo at `~/.local/opt/llama.cpp`
|
||||
(the 15 MB `ubuntu-x64` release asset) and is symlinked as
|
||||
`/usr/local/bin/llama-server`, which is what makes **discovery find it over
|
||||
ssh**: `~/.local/bin` is not on the PATH a non-interactive ssh session gets.
|
||||
It resolves its own libraries through `$ORIGIN`, so no `LD_LIBRARY_PATH` is
|
||||
needed. One model is downloaded — `unsloth/Qwen3-0.6B-GGUF/Qwen3-0.6B-Q8_0.gguf`,
|
||||
639 MB under `~/.local/share/ai-app/models` — and answers at usable speed on
|
||||
this VM's 8 cores. **Do not test with a 2-bit quant**: the
|
||||
IQ2_XXS of that model produces fluent nonsense, which reads exactly like a
|
||||
broken driver — `llama-cli` produces the same from the file directly, which
|
||||
is how to tell the two apart in a hurry.
|
||||
**Both are set up here** and need nothing typed. The prebuilt llama.cpp lives
|
||||
outside the repo at `~/.local/opt/llama.cpp-vk` — a **Vulkan** build as of
|
||||
2026-09-19, replacing the CPU one that was there before — and is symlinked as
|
||||
both `~/.local/bin/llama-server` and `/usr/local/bin/llama-server`. The second
|
||||
is what makes **discovery find it over ssh**: `~/.local/bin` is not on the
|
||||
PATH a non-interactive ssh session gets. It resolves its own libraries through
|
||||
`$ORIGIN`, so no `LD_LIBRARY_PATH` is needed.
|
||||
|
||||
Two models are downloaded under `~/.local/share/ai-app/models`:
|
||||
|
||||
- `unsloth/Qwen3-0.6B-GGUF/Qwen3-0.6B-Q8_0.gguf`, 639 MB, loads in ~4s. It
|
||||
calls tools correctly and is the right rig for the driver's shape. Do not
|
||||
judge *answers* by it — asked for the second line of a file it read from
|
||||
line 2 and then named the third.
|
||||
- `ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF/Qwen3.8-27B-GSQ-RCO-IQ3_S-mtp.gguf`,
|
||||
12 GB, ~20s to load, and the only one here with a multi-token-prediction
|
||||
head. It is the rig for anything about `loading` being a state of its own,
|
||||
since 20s is long enough to send into.
|
||||
|
||||
**Do not test with a 2-bit quant**: the IQ2_XXS of the 0.6B produces fluent
|
||||
nonsense, which reads exactly like a broken driver — `llama-cli` produces the
|
||||
same from the file directly, which is how to tell the two apart in a hurry.
|
||||
|
||||
**The GPU is shared and llama-server dies loudly when it runs out.** A second
|
||||
server loading a model while the 27B holds VRAM fails with `radv/amdgpu:
|
||||
Failed to allocate a buffer` / `MESA: error: buffer allocation failed` and
|
||||
exits mid-request. `-ngl 0` runs it on the 8 cores instead, which is the way
|
||||
to test the driver while something else holds the card.
|
||||
|
||||
**Testing tools and MCP without the app**: `llama-server --tools all` publishes
|
||||
its built-in tools at `GET /tools` and runs one at `POST /tools` with
|
||||
`{"tool": …, "params": …}` and an `x-tool-cwd` header — so a whole agent loop
|
||||
is drivable with `curl` and no model at all. The Exa MCP server at
|
||||
`https://mcp.exa.ai/mcp` answers **without an API key** and needs a
|
||||
`User-Agent` header (Cloudflare answers 403 without one, which reads as a
|
||||
refusal rather than a missing header).
|
||||
|
||||
There is no second machine, so **ssh this VM to itself**. That is set up
|
||||
too: the key is `~/.config/ai-app/ssh-self` (its public half is in
|
||||
@@ -212,6 +236,30 @@ where it was instead of half-deleted.
|
||||
|
||||
## Measurements worth not re-taking
|
||||
|
||||
- **`-np 1` is what makes the MTP draft head pay.** Taken 2026-09-19 on the
|
||||
27B above, decode speed for a 300-token reply, from `llama-server`'s own
|
||||
timings rather than the clock:
|
||||
|
||||
| flags | tok/s |
|
||||
| --- | --- |
|
||||
| plain, any `-np` | 41.5 |
|
||||
| `--spec-type draft-mtp -np 1` | 61.4 |
|
||||
| `--spec-type draft-mtp -np 2` (n-max 2) | 65.9 |
|
||||
| `--spec-type draft-mtp`, default `-np` (4 slots) | 28 |
|
||||
|
||||
Draft acceptance is 0.53–0.73 in every case, so the head is working in all
|
||||
of them: what changes is that speculating against a KV cache split four ways
|
||||
is slower than not speculating. The driver passes `-np 1` always, so this is
|
||||
recorded for whoever next sees MTP look broken. `--spec-draft-n-max 2` was
|
||||
worth another 7% in a single sample and is deliberately *not* passed — one
|
||||
sample on a virtualised GPU is not a number to hardcode.
|
||||
|
||||
- **Asking for the head when the file has none is fatal**, not ignored:
|
||||
`context type MTP requested but model doesn't contain MTP layers` and the
|
||||
server exits. Without the flag the same file logs `unused tensor
|
||||
blk.N.nextn.* — ignoring` and runs normally, which is the state to look for
|
||||
when MTP is silently not happening.
|
||||
|
||||
- **What the transcript screen costs to scroll.** Taken 2026-08-30 on the GPU
|
||||
emulator against a real imported transcript with the server at
|
||||
`--delay 120`. Settled and flinging fast, both into fresh history and back
|
||||
|
||||
Reference in new issue
Block a user