Let a llama session be shown a picture where the model reads one

A multimodal model is loaded with the `mmproj` found beside its weights --
which is how a repository publishes the pair -- and an attached image rides
in the request as an `image_url` data URI, so it reaches a model on another
machine without the file going there. Nothing is done for a model without a
projector: no captioning, no OCR, no second model.

Whether a session takes pictures is measured rather than assumed:
`/props`'s `modalities.vision` from the server that loaded the model, in
three states, because a model still coming off disk has genuinely not said.
Unknown is offered rather than refused -- a control withheld because nobody
could ask goes missing from sessions that would have taken it. The answer
reaches the phone twice per model as `Event::Images`, so the photo button is
withdrawn the moment a model with vision is left rather than at whatever
later point the session row is fetched again.

A message carrying an image a model cannot read is stopped rather than
stripped: `llama-server` refuses the whole request over one image part, and
a message sent without its picture would be answered as though the picture
had never been mentioned. The phone will not attach one, and the driver
refuses it again at the three moments the answer can first exist -- at the
door, when a message queued behind a loading model is read, and at the tool
boundary a steer enters by. An earlier turn's image folds into a line of
words for a model without vision, so switching a conversation onto one does
not end it.

A projector is filtered out of the models a provider *offers*, since a
session started on one is a server that cannot load it; it stays in the
machine's own model list, where a file on a disk is managed.

Verified against ggml-org/SmolVLM-256M-Instruct-GGUF, local and over ssh:
"In this picture there is a red circle." Switching that session to
Qwen3-0.6B reports `refused`, refuses the next picture with the reason, and
still answers an ordinary message.
This commit is contained in:
iris-ai committed 2026-09-20 16:19:53 -04:00
1 parent b7fd18b195
commit bd9596d782
17 files changed
+835 -97

No files matched your search

+17
View File
@@ -84,6 +84,23 @@ Module-by-module intent is in PLAN.md's "Backend layout".
decides whether the MTP draft head is a 50% speed-up or a 33% loss; and
`spec-type = draft-mtp` is conditional on the file actually having a head,
because asking for one that is not there makes `llama-server` **exit**.
**A llama session takes a picture only where the model natively reads one**
(2026-09-20): a multimodal model is loaded with the `mmproj` found beside
its weights (overridable per model, `off` included), an attachment rides in
the request as an `image_url` data URI, and nothing at all is done for a
model without a projector. Four things fall out of it and are easy to get
wrong again -- whether a session takes pictures is `/props`'s
`modalities.vision` from the loaded server and never a guess from this side,
with three states because a loading model has not answered yet
(`Images::Unknown` is *offered*, since a control withheld because nobody
could ask is missing from sessions that would have taken it); a message
carrying an image a model cannot read is **stopped rather than stripped**,
refused at the door, at the queue and at the steering boundary, because
`llama-server` refuses the whole request over one part and a message sent
without its picture is a different message; an earlier turn's image folds
into a line of words for a model without vision, so switching models does
not end the conversation; and a projector is filtered out of the models a
provider *offers*, while staying in the machine's own model list.
**A llama session's thinking is drawn** (2026-09-19): `reasoning_content`
becomes `Event::Thinking` deltas closed by an `Event::ThinkingDone` carrying
the span the *driver* measured, and the phone draws a card that spins while