Commit Graph
6 Commits
Author SHA1 Message Date
iris-aiandClaude Opus 5 7371af8e36 Ask a provider for its models, and let the app update its CLI
The Claude model list was four words in machines.rs, copied into every
machine's config.ron at discovery -- so a model the CLI had gained was
unreachable from the phone, which is how Opus 5.5 came to be invisible.
Every catalog is now read from the provider itself when a picker opens:
Claude Code over its control channel (control_request{subtype:list_models},
the same channel the driver sends set_model down, since the CLI has no
listing command), Codex over app-server, llama.cpp from the GGUFs on its
machine. Echo is the one case config.ron still answers, having nothing to
ask. A row the CLI marks disabled -- a model the installed version is too
old to run -- is dropped rather than offered as a chip that fails.

The reading of each catalog lives in that driver's own module, and
provider_models is one line per driver: core code says which driver to ask
and never what an answer looks like. parse_codex_models moved out for the
same reason.

Beside it, a provider now reports the version of its program and can be
told to update it. The version comes from --version on its own machine,
read from whichever stream carried it but only from a run that exited 0 --
llama-server prints its version to stderr, and so does "command not found".
It is never compared against a latest release, which nothing here can know.
Update runs the driver's own updater, or an updateCommand the config file
names for an install those will not touch; its whole output comes back,
because every install on these machines is package-managed and the
updater's refusal is the sentence worth reading. Nothing on its stdin, so a
password prompt fails rather than hangs.

A catalog that cannot be read no longer fails the whole provider view: that
is exactly the provider somebody came there to update, and refusing the
screen took the version and the Update button away with it.

Verified against a scratch server and on the emulator: the live catalog
(1.06s, disabled row absent), all four providers' versions including the
unknown state, the confirmation and output dialogs, a configured command
returning in 87ms with its stdin closed, and Update correctly disabled for
llama.cpp.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-23 18:32:51 -04:00
iris-ai 3dbf04f5ec One control per kind of setting, and no paragraphs under any of them
Three explanatory paragraphs were left on the session settings screen: what
auto-resume does, what Move costs, what changing the thinking level costs.
The last two are consequences of an action, so they are `RestartDialog` now --
the same question the machines tab already asks before it reloads a model, and
now in one file rather than private to that screen. The first is gone; the
switch beside it says what it is.

Thinking was two controls: a row of chips for a llama session, where it is a
provider-declared choice, and a picker button for a Claude session, where it
is the session's own effort. One kind of information drawn two ways, decided
by which code path the value came down. A choice is `PickerRow` in a list of
settings and `ChipGroup` on a form being filled in, and which of the two is
the screen's to say rather than the provider's.

`warnAboutRestart` went with them. It marked a control "(on restart)" and
wrote a sentence under the form, and it had nothing to mark: every session
param is `restart: false` and every model param is `restart: true`. The field
now decides whether saving a model's settings stops to ask, which is the
question it was always about.

The wait's bar keeps the status row's own margin instead of running to the
edges of the glass.

Checked on the emulator against the sandbox: an echo, a claude-cli and a
llama session, each with no paragraph left under a setting and thinking drawn
the same way in all three; "Think high?", "Move to /tmp?" and the model
settings dialog all stop to ask. ktfmt, compile, lint and the unit tests are
clean.
2026-09-21 12:25:49 -04:00
iris-ai 1aac22bfc9 Say a transcript's size, name a model by something, and ask before a reload
Five things asked for on the phone, and one trap behind the first of them.

A llama prompt carries `<__media__>` where a picture was, and llama.cpp
pairs each marker with a decoded image when it tokenizes -- so a marker in
words nobody attached a picture to fails the turn, and then fails every
later one, since the conversation is folded out of a transcript that holds
it for good. A model saying the marker back is enough to do it. Every
message with words in it now goes through `without_marker`.

A model's `general.name` is filled in by whatever converted the file, and
`convert_hf_to_gguf.py` fills it from the directory it converted: Prism ML's
Bonsai publishes `general.name = "Hf"`, which is unique and so passed the
label cascade and told a reader nothing. A name is now used only where it
shares a word with the repo or the file it came from; one that does not
drops to the file name.

The model settings dialog and the server card said what a save would cost in
a paragraph under the control, read after the decision if at all. Both ask
instead, in the shape the rest of that screen already uses for Stop and
Delete -- and the dialog's question is asked over the edits, so Cancel comes
back to them.

A field's label was `labelMedium` in the variant colour while every setting
beside it was body text, which on one form read as two ranks of setting.

`GET /sessions/{id}`'s `transcriptFile` now carries the file's size, and the
settings screen draws it beside what this phone has cached.

Checked on the emulator against the sandbox: the labels line up, the row
says "14 kB · 14 kB cached", and both confirmations appear over a loaded
Qwen3-0.6B. `cargo test` 275 passed, clippy and fmt clean, lint clean.
2026-09-21 11:59:26 -04:00
iris-ai 7278a58387 Settings as a screen with two tabs, and fields that cost one line
The settings dialog had outgrown a dialog: it scrolled inside itself,
covered the session it is about, and had nowhere to put a second tab. It
is a screen now, drawn over the session like the file explorer so the
session under it stays composed, with the back gesture to leave it. The
second tab is ProviderScreen itself -- the same composable the machines
tab opens -- so a provider's settings have two ways in and one
implementation.

Every text field in the app goes through LabelledField: the label is a
line above the box rather than a thing floating inside it, the hint says
what leaving it blank means, and the padding is one line's worth.
Material's outlined field spends the height of three lines to hold one,
which on a form of a dozen settings is a screen and a half of scrolling.
The value's own text is unchanged -- the framing was what cost.

A session also gets a system prompt, which for llama.cpp is one entry in
the params table and no app change: it rides in front of the conversation
on every request rather than being recorded as the first thing in it, so
changing it takes effect on the next message. ParamKind::Prose is new
because a paragraph in a one-line box shows six words of itself.
2026-09-21 03:35:54 -04:00
iris-aiandClaude Opus 5 81c30dcda1 Download a model onto the machine that will serve it
The Models tab was about this backend's own disk, which is the wrong disk
for every session that runs anywhere else: llama.cpp reads the file where
it runs. So the models of a machine live under that machine's llama.cpp
provider now, beside the settings deciding how each is loaded, and the
download that produces one happens there.

A download is a detached `curl` on that machine, started by a script this
server writes and never spoken to again. Its state is a file beside the
partial, so nothing about it is held here: it survives the app closing,
this backend restarting and a second device watching, and the progress is
`wc -c` of the partial against the size HuggingFace published rather than
anything remembered. A run whose process is gone is reported failed, since
`kill -0` is asked at each listing, and there is no "finished" state -- a
download that finished is a model, in the list beside the ones still
going. Resuming is guarded by the published sha256, which is also checked
before the file takes its real name.

Two other things the same screens wanted:

A provider is drawn as a card rather than as a line of text, bordered
against the machine card it sits in -- the tint it had was one step along
the surface ladder and rendered as one flat block -- with room to tap and
no chevron.

Nothing in a raw block wraps any more; the block scrolls sideways
instead, one offset for all its lines, so a diff or a column-aligned test
run still reads as one.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-19 18:55:51 -04:00
iris-ai 8c323fc7a9 Serve a machine's models from one shared llama-server
A llama.cpp session had its own `llama-server`: two sessions on one model
held two copies of it in memory, a model change bought a load only that
session benefited from, and the process was a session's to end. A machine's
models are now served by one `llama-server` in **router mode** -- no `-m`,
a preset file naming models and their flags, a child server per model asked
for, and each request routed by its `model` field. So one server per model
with that model's own settings is what a machine runs, while this backend
has one process, one port and one record per machine to keep track of.

The record is the mechanism every other driver already uses, so a restart
adopts it; a session records the same pid in its own directory as
`Detail::Shared`, and `process::signal` refuses to signal one of those --
which is what keeps stopping, deleting or cleaning up after one session
from unloading a model every other session is using. Nothing stops a router
on its own. That is deliberate (a loaded model is minutes of disk) and it is
why the machines tab now has a card per provider that opens its own screen:
how each model is loaded, how many stay in memory, Unload, and Stop.

How a model is *loaded* therefore belongs to the model on its machine rather
than to a session -- context size, GPU layers, threads, slots, speculative
decoding -- written into the preset as llama-server's own argument names.
Saving them re-reads that file, which unloads the model; that is the change
taking effect, and the dialog says so before you save. What stays a
session's is everything that rides on a request, including which tools it
offers: the router hosts one set for the machine and the choice is a filter
applied here, so it costs no reload (2,181 tokens of prompt with all seven,
698 with none).

Verified end to end against the scratch backend and the emulator: two
sessions sharing one loaded model with one child process, a second session
joining it with a 26ms prefill, a backend restart adopting the router and
answering with the prompt cache intact, the same over ssh to this VM, a
model's settings reaching the running server, Unload, and Stop leaving every
session `exited` with no error line.
2026-09-19 17:37:31 -04:00