Give llama.cpp sessions tools, web search and a model picker

A llama session was a chat box: no tools, a fixed model, no permission
mode, and a model name drawn as the path the file sits at. It now runs the
agent loop itself, which is what the pieces below all hang off.

Tools are `llama-server`'s own (`--tools all`), which that server both
publishes and runs -- `GET /tools` for the definitions, `POST /tools` to
call one. Web search is Exa's MCP server, reached from this backend rather
than from the machine serving the model: that is what llama.cpp's own web
UI does, and it puts the search on the machine with a route out instead of
the one with the GPU. `llama-server`'s `--mcp-servers-json` can only spawn
local commands, so using it would have meant a Node bridge on every
machine that serves a model.

Driving the loop is what makes the permission gate ours. Two modes,
`manual` and `bypassPermissions`, which is what the mechanism has: the web
UI asks before every call and remembers the tools you say "always" to. The
allowances fold back out of the transcript's own answers, so they survive
a restart and a model change without being stored anywhere else.

Also here, because tools made each of them matter:

- **Loading is a state.** A 12 GB model takes twenty seconds to reach
  memory and refuses everything until it has; the session used to report
  `running` for that whole time, and a message sent meanwhile came back as
  an error. It is `loading` now, and the message waits.
- **The model can be changed.** A `llama-server` holds one model, so this
  stops it and starts another. The conversation survives because it was
  never in the server.
- **Models are named, not pathed.** `general.name` read out of the file
  itself -- over ssh too, in the round trip the spawn was already making.
  Where two models share a name the file name breaks the tie.
- **`-np 1`, and the MTP draft head where the file has one.** Measured on
  the 27B here: 41.5 tok/s plain, 61.4 with `--spec-type draft-mtp` at one
  slot, and 28 with it at four -- speculating against a split KV cache is
  worse than not speculating. The flag is conditional because asking for a
  head that is not there makes `llama-server` exit.
- **A refusal says what to do.** Tool results are thousands of tokens, so
  an overrun context is now ordinary; it was "http status: 400" and is now
  the server's own "exceeds the available context size, try increasing it".

`GET /machines/{id}/models` is gone: the provider models route answers the
same question, and two answers to one question is how a picker comes to
offer a model the spawn screen does not.

Verified end to end against real models: a tool call asked and allowed, an
Exa search, a shell command, a 27B loaded while a message waited on it, a
model switch mid-session, a second message queued behind a running turn,
and the whole of it again on a session running over ssh.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
iris-aiandClaude Opus 5 committed 2026-09-19 08:11:49 -04:00
1 parent 392cc5413d
commit ac476ab0c9
23 files changed
+3627 -1096

No files matched your search

+60 -12
View File
@@ -143,18 +143,42 @@ moment you use it — `ANDROID_SERIAL=$(emu serial) ./gradlew …`.
### Testing llama.cpp and ssh here
**Both are set up here as of 2026-09-04** and need nothing typed. The
prebuilt CPU llama.cpp lives outside the repo at `~/.local/opt/llama.cpp`
(the 15 MB `ubuntu-x64` release asset) and is symlinked as
`/usr/local/bin/llama-server`, which is what makes **discovery find it over
ssh**: `~/.local/bin` is not on the PATH a non-interactive ssh session gets.
It resolves its own libraries through `$ORIGIN`, so no `LD_LIBRARY_PATH` is
needed. One model is downloaded — `unsloth/Qwen3-0.6B-GGUF/Qwen3-0.6B-Q8_0.gguf`,
639 MB under `~/.local/share/ai-app/models` — and answers at usable speed on
this VM's 8 cores. **Do not test with a 2-bit quant**: the
IQ2_XXS of that model produces fluent nonsense, which reads exactly like a
broken driver — `llama-cli` produces the same from the file directly, which
is how to tell the two apart in a hurry.
**Both are set up here** and need nothing typed. The prebuilt llama.cpp lives
outside the repo at `~/.local/opt/llama.cpp-vk` — a **Vulkan** build as of
2026-09-19, replacing the CPU one that was there before — and is symlinked as
both `~/.local/bin/llama-server` and `/usr/local/bin/llama-server`. The second
is what makes **discovery find it over ssh**: `~/.local/bin` is not on the
PATH a non-interactive ssh session gets. It resolves its own libraries through
`$ORIGIN`, so no `LD_LIBRARY_PATH` is needed.
Two models are downloaded under `~/.local/share/ai-app/models`:
- `unsloth/Qwen3-0.6B-GGUF/Qwen3-0.6B-Q8_0.gguf`, 639 MB, loads in ~4s. It
calls tools correctly and is the right rig for the driver's shape. Do not
judge *answers* by it — asked for the second line of a file it read from
line 2 and then named the third.
- `ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF/Qwen3.8-27B-GSQ-RCO-IQ3_S-mtp.gguf`,
12 GB, ~20s to load, and the only one here with a multi-token-prediction
head. It is the rig for anything about `loading` being a state of its own,
since 20s is long enough to send into.
**Do not test with a 2-bit quant**: the IQ2_XXS of the 0.6B produces fluent
nonsense, which reads exactly like a broken driver — `llama-cli` produces the
same from the file directly, which is how to tell the two apart in a hurry.
**The GPU is shared and llama-server dies loudly when it runs out.** A second
server loading a model while the 27B holds VRAM fails with `radv/amdgpu:
Failed to allocate a buffer` / `MESA: error: buffer allocation failed` and
exits mid-request. `-ngl 0` runs it on the 8 cores instead, which is the way
to test the driver while something else holds the card.
**Testing tools and MCP without the app**: `llama-server --tools all` publishes
its built-in tools at `GET /tools` and runs one at `POST /tools` with
`{"tool": …, "params": …}` and an `x-tool-cwd` header — so a whole agent loop
is drivable with `curl` and no model at all. The Exa MCP server at
`https://mcp.exa.ai/mcp` answers **without an API key** and needs a
`User-Agent` header (Cloudflare answers 403 without one, which reads as a
refusal rather than a missing header).
There is no second machine, so **ssh this VM to itself**. That is set up
too: the key is `~/.config/ai-app/ssh-self` (its public half is in
@@ -212,6 +236,30 @@ where it was instead of half-deleted.
## Measurements worth not re-taking
- **`-np 1` is what makes the MTP draft head pay.** Taken 2026-09-19 on the
27B above, decode speed for a 300-token reply, from `llama-server`'s own
timings rather than the clock:
| flags | tok/s |
| --- | --- |
| plain, any `-np` | 41.5 |
| `--spec-type draft-mtp -np 1` | 61.4 |
| `--spec-type draft-mtp -np 2` (n-max 2) | 65.9 |
| `--spec-type draft-mtp`, default `-np` (4 slots) | 28 |
Draft acceptance is 0.530.73 in every case, so the head is working in all
of them: what changes is that speculating against a KV cache split four ways
is slower than not speculating. The driver passes `-np 1` always, so this is
recorded for whoever next sees MTP look broken. `--spec-draft-n-max 2` was
worth another 7% in a single sample and is deliberately *not* passed — one
sample on a virtualised GPU is not a number to hardcode.
- **Asking for the head when the file has none is fatal**, not ignored:
`context type MTP requested but model doesn't contain MTP layers` and the
server exits. Without the flag the same file logs `unused tensor
blk.N.nextn.* — ignoring` and runs normally, which is the state to look for
when MTP is silently not happening.
- **What the transcript screen costs to scroll.** Taken 2026-08-30 on the GPU
emulator against a real imported transcript with the server at
`--delay 120`. Settled and flinging fast, both into fresh history and back