diff --git a/AGENTS.md b/AGENTS.md index 4ef5492..b7efec5 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -80,6 +80,13 @@ Module-by-module intent is in PLAN.md's "Backend layout". 14k), and as `running` it looked exactly like thinking. The phone draws both with the working spinner and its own words, "loading model" and "reading prompt". + **Thinking effort is a param, and which levels exist is the model's answer** + (2026-09-19): the `thinking` param rides on the request as a chat-template + argument (`reasoning_effort`, or `enable_thinking: false` for `off`), so it + needs no restart -- and the driver asks the loaded server which levels its + template actually takes rather than trusting the offered list, because the + 27B raises on `high` and answers to `xhigh`. A level it cannot take is + dropped and said in the transcript, naming the ones it can. **Every one of those is a default rather than a constant** (2026-09-19): `DriverKind::params` declares what a provider takes — key, label, shape, and whether a change waits for a restart — and the phone renders whatever diff --git a/PLAN.md b/PLAN.md index 325b4de..194e792 100644 --- a/PLAN.md +++ b/PLAN.md @@ -357,6 +357,22 @@ deliberate and easy to undo by accident: ephemeral one, because no portable way to ask a machine for a free port avoids racing the bind anyway; a collision is not silent, since the server fails to bind and the readiness poll reports what its log said. +- **A llama session's thinking level is the model's, asked of the model** + (2026-09-19). Thinking effort is a chat-template argument rather than a + server flag, so it rides on the next request and changes nothing about the + loaded model -- which is why it is a `params` entry (`thinking`) and not the + `effort` a coding CLI reads at launch. The levels templates use disagree: + the 27B here takes `low`, `medium` and `xhigh` and **raises** on `high` and + `max`, so a fixed list would be a turn that fails on send. The driver asks + the loaded server instead (`thinking_options`): `chat_template_caps. + supports_reasoning_effort` says whether levels mean anything at all -- the + gate that stops the control silently doing nothing on a template that + ignores the argument -- and `/apply-template` says which of them render, one + cheap round trip each at load time. `off` is a separate question and a + separate argument (`enable_thinking: false`), taken as supported when + turning it off renders a different prompt. A level the loaded model cannot + take is dropped from the request and said out loud, naming what it does + take. - **The model file lives on the machine that serves it** (2026-09-04). Each machine has its own models directory (`SshConfig::models_dir`, default `~/.local/share/ai-app/models` expanded *there*), and a spawn resolves the diff --git a/server/src/config.rs b/server/src/config.rs index b66d368..039c3d7 100644 --- a/server/src/config.rs +++ b/server/src/config.rs @@ -443,6 +443,23 @@ const LLAMA_PARAMS: &[ParamSpec] = &[ kind: ParamKind::Integer, restart: true, }, + ParamSpec { + key: "thinking", + label: "Thinking", + // Rides on the request, like the sampling settings below it: the level + // is a chat-template argument rather than a server flag, so a session + // changes how hard it thinks without reloading its model. + unset: "however hard the model thinks by default", + // Every level any of these templates uses, because which of them a + // *particular* model takes is the model's business and only the loaded + // one can answer it -- the driver asks it (`thinking_options`) and says + // what it takes when a level it cannot is chosen. "off" is the one that + // is not a level: it asks the template for no thinking at all. + kind: ParamKind::Choice { + options: &["auto", "off", "low", "medium", "high", "xhigh", "max"], + }, + restart: false, + }, ParamSpec { key: "temperature", label: "Temperature", diff --git a/server/src/session/llama/mod.rs b/server/src/session/llama/mod.rs index 9963fa7..d639419 100644 --- a/server/src/session/llama/mod.rs +++ b/server/src/session/llama/mod.rs @@ -91,6 +91,19 @@ fn sampling_from( sampling } +/// The thinking setting in [`params`], or `None` for the model's own default. +/// +/// `"auto"` is that default rather than a value, and the phone clears the +/// setting when it is chosen -- so it is here only for a config written by +/// hand. +fn chosen_thinking(params: &std::collections::BTreeMap) -> Option { + params + .get(THINKING) + .map(String::as_str) + .filter(|value| !value.is_empty() && *value != "auto") + .map(str::to_string) +} + /// Exa's own hosted MCP server, which is what a llama session searches the web /// with. The address llama.cpp's web UI offers under "Exa" in its recommended /// servers, so a session here reaches the same thing that UI does. @@ -102,6 +115,20 @@ pub const EXA_MCP_URL: &str = "https://mcp.exa.ai/mcp"; /// is the escape for a machine where drafting turns out not to pay. const SPECULATIVE: &str = "speculative"; +/// The parameter naming how hard the model should think, which is a chat +/// template argument rather than a server flag -- so unlike the flags it takes +/// effect on the next request. `"off"` asks the template for no thinking at +/// all; anything else is one of the levels a template names. +const THINKING: &str = "thinking"; + +/// The levels [`THINKING`] offers, which is every one any of these templates +/// uses. Which of them a given model takes is [`thinking_options`]'s answer. +const THINKING_LEVELS: &[&str] = &["low", "medium", "high", "xhigh", "max"]; + +/// [`THINKING`]'s value for no thinking at all, which is not a level: it is +/// `enable_thinking: false`, a different argument to the template. +const THINKING_OFF: &str = "off"; + /// The spawn parameter naming which built-in tools a session gets, as /// `llama-server`'s own comma-separated list. Absent is all of them, and /// `"none"` is the way to ask for a session that only talks. @@ -289,6 +316,20 @@ struct Shared { allowed: Mutex>, /// Questions a turn is blocked on, by question id. asked: Mutex>>>, + /// How hard this session asks the model to think: a level, `"off"`, or + /// `None` for the model's own default. Live like the sampling settings, + /// and for the same reason -- it rides on the next request. + thinking: Mutex>, + /// What the loaded model's chat template actually takes, asked of the + /// server that loaded it (see [`thinking_options`]). + /// + /// `None` until one has been asked, and `None` again while a model is + /// being replaced: a setting is only meaningful against a template, and + /// the previous model's answer says nothing about this one's. Not an empty + /// list, because that is the model that genuinely takes none -- and saying + /// so about a model nobody has asked yet is the one mistake here that + /// looks exactly like an answer. + thinking_options: Mutex>>, } pub struct LlamaDriver { @@ -350,6 +391,8 @@ impl LlamaDriver { ), allowed: Mutex::new(allowances(transcript)), asked: Mutex::new(HashMap::new()), + thinking: Mutex::new(chosen_thinking(&meta.params)), + thinking_options: Mutex::new(None), }), respawn: Respawn { meta: meta.clone(), @@ -383,6 +426,9 @@ impl LlamaDriver { // as loading until the model is in memory, rather than looking ready // and refusing the first message. *shared.serving.lock().unwrap() = Serving::Loading; + // The answer belonged to whatever was loaded before this; the model + // starting now gets asked for itself. + *shared.thinking_options.lock().unwrap() = None; let _ = shared.sink.send(Event::Status { state: SessionStatus::Loading, }); @@ -433,6 +479,13 @@ impl LlamaDriver { if let Some(window) = context_window(&endpoint) { shared.emit(Event::ContextWindow { tokens: window }); } + // Asked of the model that is now loaded, for the same + // reason the window is: only it knows. A level the session + // is carrying that this model cannot take is said here + // rather than at the next turn, which is where it would + // otherwise surface as the model simply not doing it. + *shared.thinking_options.lock().unwrap() = Some(thinking_options(&endpoint)); + shared.note_unusable_thinking(); Serving::Ready { endpoint: endpoint.clone(), tools: Arc::new(tools), @@ -464,6 +517,40 @@ impl LlamaDriver { } impl Shared { + /// Says so when this session is set to think in a way the loaded model + /// cannot, naming what it *does* take. + /// + /// Said rather than silently dropped, because the two are the same on + /// screen otherwise: a picker sitting on "xhigh" and a model answering as + /// though nothing had been chosen. A model that takes none of it says that + /// instead, which is the honest end of the same sentence. + fn note_unusable_thinking(&self) { + let Some(chosen) = self.thinking.lock().unwrap().clone() else { + return; + }; + // Nothing is said about a model nobody has asked yet -- a session + // whose server is not up is told what it takes when one is. + let Some(options) = self.thinking_options.lock().unwrap().clone() else { + return; + }; + if options.contains(&chosen) { + return; + } + self.emit(Event::Error { + message: if options.is_empty() { + format!( + "this model's chat template takes no thinking setting, so \ + \"{chosen}\" is doing nothing here." + ) + } else { + format!( + "this model does not take \"{chosen}\" -- it takes {}.", + options.join(", "), + ) + }, + }); + } + /// Blocks until the server can be spoken to, and says what to talk to. /// /// The whole of what [`SessionStatus::Loading`] means in practice: a @@ -1113,6 +1200,8 @@ impl Driver for LlamaDriver { /// will read all of them. fn set_params(&self, params: &std::collections::BTreeMap) { *self.shared.sampling.lock().unwrap() = sampling_from(params); + *self.shared.thinking.lock().unwrap() = chosen_thinking(params); + self.shared.note_unusable_thinking(); // Only the settings that actually differ from what this session's // server was started with. Listing every restart-only one on every // save would be a wall of text about nothing having changed. @@ -1553,6 +1642,100 @@ fn context_window(endpoint: &str) -> Option { .and_then(Value::as_u64) } +/// What the loaded model's chat template takes for thinking, asked of it. +/// +/// Two different questions, because they are two different template arguments. +/// **Levels** are meaningful only where the server says the template reads a +/// reasoning effort at all (`chat_template_caps.supports_reasoning_effort`, +/// which llama.cpp works out by rendering) -- without that gate a template that +/// ignores the argument accepts every level and the control silently does +/// nothing. Which levels it then takes is asked one at a time, because the +/// vocabularies disagree: the 27B here raises on `high` and `max` and answers +/// to `xhigh`, so a fixed list would be a turn that fails on send. **Off** is +/// its own question -- a template honours `enable_thinking` if turning it off +/// renders a different prompt. +/// +/// An empty answer is a model that takes neither, which is a thing to say +/// rather than a failure; a server that will not answer gives the same, since +/// a setting nobody can check is one nobody should be told worked. +fn thinking_options(endpoint: &str) -> Vec { + let mut options = Vec::new(); + let Some(props) = ureq::get(format!("{endpoint}/props")) + .call() + .ok() + .and_then(|mut response| response.body_mut().read_json::().ok()) + else { + return options; + }; + // Both renders have to have worked: a template that *refuses* + // `enable_thinking` also differs from the plain render, and reading that as + // support would offer an "off" that fails every turn. + let plain = render_template(endpoint, &json!({})); + let off = render_template(endpoint, &json!({"enable_thinking": false})); + if plain.is_some() && off.is_some() && off != plain { + options.push(THINKING_OFF.to_string()); + } + if props + .pointer("/chat_template_caps/supports_reasoning_effort") + .and_then(Value::as_bool) + == Some(true) + { + options.extend( + THINKING_LEVELS + .iter() + .filter(|level| { + render_template(endpoint, &json!({"reasoning_effort": level})).is_some() + }) + .map(|level| (*level).to_string()), + ); + } + options +} + +/// One prompt as this model's template renders it under `kwargs`, or `None` +/// where the template refused them. +/// +/// `/apply-template` is the cheap half of a request: it renders and returns, +/// with no model involved, so asking it seven questions at load time costs +/// nothing anybody waits for. +fn render_template(endpoint: &str, kwargs: &Value) -> Option { + ureq::post(format!("{endpoint}/apply-template")) + .config() + .http_status_as_error(false) + .build() + .send_json(json!({ + "messages": [{"role": "user", "content": "hi"}], + "chat_template_kwargs": kwargs, + })) + .ok() + .filter(|response| response.status().is_success())? + .body_mut() + .read_json::() + .ok()? + .get("prompt")? + .as_str() + .map(str::to_string) +} + +/// The template arguments for how hard this session has asked the model to +/// think, or `None` where it has asked for nothing the model can do. +/// +/// Filtered against what the *loaded* model takes rather than trusted from the +/// config: a session keeps its setting across a model change, and sending a +/// level the new template raises on would fail every turn. +fn thinking_kwargs(shared: &Shared) -> Option { + let chosen = shared.thinking.lock().unwrap().clone()?; + let options = shared.thinking_options.lock().unwrap().clone(); + if !options.is_some_and(|options| options.contains(&chosen)) { + return None; + } + Some(if chosen == THINKING_OFF { + json!({"enable_thinking": false}) + } else { + json!({"reasoning_effort": chosen}) + }) +} + /// One streamed completion: posts the conversation, emits each text delta as /// it arrives, and assembles whatever tool calls came with it. /// @@ -1587,6 +1770,9 @@ fn generate( for (key, value) in sampling { map.insert(key.clone(), value.clone()); } + if let Some(kwargs) = thinking_kwargs(shared) { + map.insert("chat_template_kwargs".to_string(), kwargs); + } // Prompt processing starts the moment this is sent and nothing comes back // until it is done, so this is where the wait somebody is watching begins. @@ -1903,6 +2089,24 @@ mod tests { ); } + #[test] + /// `auto` is the model's own default rather than a value to send, which is + /// what the first option of a `Choice` means everywhere -- and the one + /// thing here that is not obvious from the type, since it arrives as the + /// string "auto" like any other level. + fn auto_is_no_thinking_setting_at_all() { + let chosen = |value: &str| { + chosen_thinking(&std::collections::BTreeMap::from([( + THINKING.to_string(), + value.to_string(), + )])) + }; + assert_eq!(chosen("auto"), None); + assert_eq!(chosen(""), None); + assert_eq!(chosen("xhigh"), Some("xhigh".to_string())); + assert_eq!(chosen_thinking(&std::collections::BTreeMap::new()), None); + } + #[test] /// A model's working is drawn and is deliberately not sent back to it: the /// prompt is what was said, and feeding reasoning back costs the whole of