Files
ai-app/server/src/models.rs
T
iris-ai 1aac22bfc9 Say a transcript's size, name a model by something, and ask before a reload
Five things asked for on the phone, and one trap behind the first of them.

A llama prompt carries `<__media__>` where a picture was, and llama.cpp
pairs each marker with a decoded image when it tokenizes -- so a marker in
words nobody attached a picture to fails the turn, and then fails every
later one, since the conversation is folded out of a transcript that holds
it for good. A model saying the marker back is enough to do it. Every
message with words in it now goes through `without_marker`.

A model's `general.name` is filled in by whatever converted the file, and
`convert_hf_to_gguf.py` fills it from the directory it converted: Prism ML's
Bonsai publishes `general.name = "Hf"`, which is unique and so passed the
label cascade and told a reader nothing. A name is now used only where it
shares a word with the repo or the file it came from; one that does not
drops to the file name.

The model settings dialog and the server card said what a save would cost in
a paragraph under the control, read after the decision if at all. Both ask
instead, in the shape the rest of that screen already uses for Stop and
Delete -- and the dialog's question is asked over the edits, so Cancel comes
back to them.

A field's label was `labelMedium` in the variant colour while every setting
beside it was body text, which on one form read as two ranks of setting.

`GET /sessions/{id}`'s `transcriptFile` now carries the file's size, and the
settings screen draws it beside what this phone has cached.

Checked on the emulator against the sandbox: the labels line up, the row
says "14 kB · 14 kB cached", and both confirmations appear over a loaded
Qwen3-0.6B. `cargo test` 275 passed, clippy and fmt clean, lint clean.
2026-09-21 11:59:26 -04:00

798 lines
32 KiB
Rust

//! GGUF models on the machines this backend can run them on, and the
//! downloads that produce them.
//!
//! **A model file belongs to the machine that will serve it.** Everything
//! here is addressed by a transport and a directory rather than by this
//! process's own disk, because `llama-server` reads the file where it runs
//! and a list or a download naming anywhere else is a claim about the wrong
//! filesystem.
//!
//! **So a download runs there too** (2026-09-19, replacing the fetch this
//! server used to perform itself onto its own disk). It is a detached `curl`
//! on that machine, started by a script written here and never spoken to
//! again. Three things follow, and all three are why it is shaped this way:
//!
//! - **Its state is a file beside the model**, `x.gguf.download`, written by
//! that script at each step. Nothing about it is held in this process, so
//! one download survives the phone locking, this backend restarting, and
//! being watched from a second device -- and what is on that disk is the
//! answer rather than something this server remembers about it.
//! - **Progress is measured, never estimated.** The bytes are `wc -c` of the
//! partial file and the total is the size HuggingFace published for it. A
//! download whose process is gone is reported failed rather than left
//! saying "running" for ever, which is what `kill -0` at each listing is
//! for.
//! - **Resume is guarded by identity.** The partial records the sha256 the
//! file is meant to have; one written against a different sha is discarded
//! rather than resumed onto, and the finished file is checked against it
//! before it takes the real name.
use std::collections::HashMap;
use std::path::Path;
use anyhow::{Context, Result};
use serde::Serialize;
use crate::session::transport::{Launch, Transport};
/// Identifies this client to HuggingFace. They ask for one, and a request
/// without it is more likely to be rate-limited.
const USER_AGENT: &str = concat!("ai-server/", env!("CARGO_PKG_VERSION"));
/// A model file sitting on a machine, ready to run there.
#[derive(Debug, Clone, Serialize)]
#[serde(rename_all = "camelCase")]
pub struct LocalModel {
/// `owner/repo/file.gguf` -- the HuggingFace coordinates, which are already
/// unique, so nothing has to invent an id.
pub key: String,
pub repo: String,
pub file: String,
pub bytes: u64,
/// What the file says it is called (`general.name` in its own metadata),
/// absent when it does not say or could not be read. Not a label: see
/// [`labels`] for what a reader is actually shown, which needs the rest of
/// the list to decide.
#[serde(skip_serializing_if = "Option::is_none")]
pub name: Option<String>,
}
/// What each of these models should be called on screen, in the same order.
///
/// A model's own name is the best answer and is not always an answer at all:
/// two quantisations of one model carry the same `general.name`, and a chip
/// row with two identical chips is one you cannot choose from. So this is a
/// cascade -- the model's own name, else its file name, else its full key --
/// and each model takes the first rung that nothing else on this machine
/// shares. The last rung always terminates it, because the key is what makes
/// these unique in the first place.
///
/// Decided over the whole list rather than per model because ambiguity is a
/// property of the set: the same file is unambiguous on a machine holding one
/// quantisation and not on a machine holding three, and only the list knows
/// which machine this is.
pub fn labels(models: &[LocalModel]) -> Vec<String> {
let rungs = |model: &LocalModel| {
[
model
.name
.clone()
.filter(|name| names_the_file(name, model)),
Some(model.file.trim_end_matches(".gguf").to_string()),
Some(model.key.clone()),
]
};
let mut taken: Vec<HashMap<String, usize>> = vec![HashMap::new(); 3];
for model in models {
for (rung, candidate) in rungs(model).into_iter().enumerate() {
if let Some(candidate) = candidate {
*taken[rung].entry(candidate).or_insert(0) += 1;
}
}
}
models
.iter()
.map(|model| {
rungs(model)
.into_iter()
.enumerate()
.find_map(|(rung, candidate)| {
let candidate = candidate?;
(taken[rung].get(&candidate) == Some(&1)).then_some(candidate)
})
// Unreachable: the key rung is unique by construction. Said as
// the key rather than as a panic, because a duplicate key would
// mean the same file listed twice and a name is still the
// honest thing to draw for it.
.unwrap_or_else(|| model.key.clone())
})
.collect()
}
/// Whether a model's own `general.name` is a name for *this* model, or the
/// converter's working directory wearing the same field.
///
/// The field is filled in by whatever produced the file, and
/// `convert_hf_to_gguf.py` fills it from the directory it converted -- so a
/// model converted out of a folder called `hf` publishes `general.name = "hf"`,
/// which is unique, passes the cascade above, and tells a reader nothing at
/// all. Prism ML's Bonsai is the one here (reported 2026-09-21, drawn as "hf"
/// in the model picker).
///
/// The test is corroboration rather than a list of words to distrust: a real
/// name shares something with where the file came from, both being about the
/// same model, and a directory name picked up in passing does not. One word in
/// common is enough. A name that fails it drops to the next rung, which is the
/// file name the reader downloaded.
fn names_the_file(name: &str, model: &LocalModel) -> bool {
let words = |text: &str| -> Vec<String> {
text.split(|c: char| !c.is_ascii_alphanumeric())
.filter(|word| word.len() > 1)
.map(str::to_ascii_lowercase)
.collect()
};
let from = words(&model.repo);
let from_file = words(&model.file);
words(name)
.iter()
.any(|word| from.contains(word) || from_file.contains(word))
}
/// What a download is doing, or did.
///
/// Flat rather than a tagged enum carrying its message, because the phone
/// switches on this and a word it can compare is easier to render than a
/// variant it has to destructure. There is no `finished`: a download that
/// finished is a model, and it is in the list beside this one -- an outcome
/// kept here as well would be the same fact stated twice, in two places that
/// could disagree.
#[derive(Debug, Clone, Copy, PartialEq, Eq, Serialize)]
#[serde(rename_all = "snake_case")]
pub enum DownloadState {
Running,
/// Reading the finished file back to check it against the hash
/// HuggingFace publishes. Its own state because it takes real time on a
/// multi-gigabyte file, and "still working" is the honest thing to show
/// where a bar sitting at 100% for half a minute is not.
Verifying,
Failed,
Cancelled,
}
impl DownloadState {
fn from_word(word: &str) -> Option<Self> {
match word {
"running" => Some(Self::Running),
"verifying" => Some(Self::Verifying),
"failed" => Some(Self::Failed),
"cancelled" => Some(Self::Cancelled),
_ => None,
}
}
}
/// One download on one machine, as the phone sees it.
#[derive(Debug, Clone, Serialize)]
#[serde(rename_all = "camelCase")]
pub struct Download {
/// `owner/repo/file.gguf`, the same key the finished model will have.
/// One download per model at a time, so this identifies it.
pub key: String,
pub repo: String,
pub file: String,
pub state: DownloadState,
pub done: u64,
/// What HuggingFace published as the file's size, or absent when it
/// published none. Absent means "not known", never "zero".
#[serde(skip_serializing_if = "Option::is_none")]
pub total: Option<u64>,
#[serde(skip_serializing_if = "Option::is_none")]
pub error: Option<String>,
}
/// What HuggingFace says about one file before anything is fetched: how big
/// it is, and what it should hash to.
///
/// Both read here rather than on the machine doing the download, because this
/// is the side with internet trust to spend (see [`search`]) and because the
/// sha is what makes a resume safe -- the machine is handed the answer rather
/// than asked to find it.
#[derive(Debug, Clone)]
pub struct Published {
pub bytes: Option<u64>,
/// The LFS object id, which for these repositories is the sha256 of the
/// content -- a free integrity check rather than a second source of
/// truth. Absent for a repository that publishes none, and then the file
/// is not verified and a partial is never resumed onto.
pub sha256: Option<String>,
}
/// Refuses a key that is anything but `owner/repo/file.gguf`.
///
/// The key comes from a phone, so it is treated as hostile: it reaches a
/// shell as a positional argument, which is what stops it being read as a
/// command, but nothing there stops `../..` from naming a file outside the
/// models directory. Rejected rather than rewritten, since a silently
/// corrected path downloads the right bytes to the wrong place.
fn checked(key: &str) -> Result<()> {
let bad = |why: &str| anyhow::anyhow!("\"{key}\" is not a model this can store: {why}");
for part in key.split('/') {
if part.is_empty() || part == "." || part == ".." {
return Err(bad(
"a path component that would leave the models directory",
));
}
if part.contains('\\') || part.chars().any(char::is_control) {
return Err(bad("a path component with a character no file here has"));
}
}
let Some((_, file)) = key.rsplit_once('/') else {
return Err(bad("a key is owner/repo/file.gguf"));
};
if !file.ends_with(".gguf") {
return Err(bad("only .gguf files are models"));
}
Ok(())
}
pub fn key_for(repo: &str, file: &str) -> String {
format!("{repo}/{file}")
}
/// Expands a leading `~` the way that machine's own shell would, which is the
/// only place that knows what it stands for. Prefixed to every script here;
/// each takes the directory as `$1`.
const EXPAND: &str = "d=$1; case $d in \"~\") d=$HOME;; \"~/\"*) d=$HOME/${d#\"~/\"};; esac;";
/// The download itself, run detached on the machine that will hold the file.
///
/// Passed to [`START`] as an argument rather than written inside it, so that
/// neither script has to quote the other: `sh -c "$worker"` takes this whole
/// text as one string the outer shell never parses.
///
/// `-sS` rather than plain: curl's progress meter goes to stderr, and the
/// same file is what a failure is read out of -- a message buried in three
/// screens of redrawn bar is one the phone cannot show. Progress is read off
/// the partial file instead, which is the measurement rather than a report
/// of it.
///
/// The state file is written whole and moved into place, because a listing
/// can arrive in the middle of any of these writes and a truncated one is a
/// download that blinks out of the phone's list for a poll.
///
/// `curl` is waited on rather than run in the foreground, because a cancel
/// arrives as a signal and a shell blocked in `wait` is the one that can run
/// a trap for it. The trap writes nothing when the partial has already gone:
/// that is how [`REMOVE`] deletes a running download without the dying worker
/// recreating the state file behind it.
const WORKER: &str = r#"
t=$1; url=$2; sha=$3; total=$4; p=$t.part; s=$t.download; e=$t.part.err
st() { printf 'pid=%s\nsha=%s\ntotal=%s\nstate=%s\nerror=%s\n' "$$" "$sha" "$total" "$1" "$2" > "$s.tmp" && mv "$s.tmp" "$s"; }
trap '[ -n "$c" ] && kill "$c" 2>/dev/null; [ -f "$p" ] && st cancelled ""; exit 0' TERM INT
c=
st running ""
have=$(wc -c 2>/dev/null < "$p" | tr -d " ")
if [ -z "$total" ] || [ "$have" != "$total" ]; then
curl -fsSL --retry 3 --retry-delay 2 -C - -o "$p" "$url" 2> "$e" &
c=$!
if ! wait "$c"; then
st failed "$(tr -s '\n\t\r' ' ' < "$e" | tail -c 300)"
exit 1
fi
fi
rm -f "$e"
if [ -n "$sha" ] && command -v sha256sum > /dev/null 2>&1; then
st verifying ""
got=$(sha256sum "$p" | cut -d' ' -f1)
if [ "$got" != "$sha" ]; then
rm -f "$p"
st failed "it arrived corrupted -- HuggingFace publishes sha256 $sha, what arrived hashes to $got. It has been deleted; downloading again starts clean."
exit 1
fi
fi
mv "$p" "$t" && rm -f "$s"
"#;
/// Starts [`WORKER`], or leaves alone whatever is already happening.
///
/// Idempotent on purpose, in both directions: a phone that lost its
/// connection presses the button again, and that must join the run already
/// going rather than start a second one writing the same file. A model
/// already downloaded is likewise the answer rather than an error, which is
/// the ordinary case once two devices can both see the list.
const START: &str = r#"
k=$2; url=$3; sha=$4; total=$5; worker=$6
t=$d/$k; s=$t.download
[ -f "$t" ] && exit 0
pid=$(sed -n 's/^pid=//p' "$s" 2>/dev/null)
if [ -n "$pid" ] && kill -0 "$pid" 2>/dev/null; then exit 0; fi
command -v curl > /dev/null 2>&1 || { echo "there is no curl on this machine to download with" >&2; exit 1; }
mkdir -p "${t%/*}" || { echo "cannot make ${t%/*} to download into" >&2; exit 1; }
had=$(sed -n 's/^sha=//p' "$s" 2>/dev/null)
if [ -z "$sha" ] || [ "$had" != "$sha" ]; then rm -f "$t.part"; fi
nohup sh -c "$worker" sh "$t" "$url" "$sha" "$total" > /dev/null 2>&1 &
"#;
/// Every download the machine is holding state for, with the bytes read off
/// the partial file rather than from anything remembered here.
///
/// `kill -0` is why a stopped machine does not leave a download claiming to
/// be running for ever: a state file whose process is gone is the failure it
/// actually is.
///
/// A state file beside a finished model is skipped, because the file is the
/// truth and the state file is only ever a report about producing it. That is
/// also what keeps the one race here self-correcting: two devices asking for
/// the same model in the same second get two workers writing one partial,
/// whose sha will not match -- and whichever of them loses the rename would
/// otherwise leave a failure recorded against a model that is sitting there.
const LIST: &str = r#"
[ -d "$d" ] || exit 0
cd "$d" || exit 0
find . -type f -name '*.download' | while read -r f; do
k=${f#./}; k=${k%.download}
[ -f "$k" ] && continue
size=$(wc -c 2>/dev/null < "$k.part" | tr -d " ")
[ -n "$size" ] || size=0
pid=$(sed -n 's/^pid=//p' "$f")
alive=no
if [ -n "$pid" ] && kill -0 "$pid" 2>/dev/null; then alive=yes; fi
printf '%s\037%s\037%s\037%s\000' "$size" "$alive" "$(tr '\n' '\036' < "$f")" "$k"
done
"#;
/// Asks a running download to stop; the partial stays, so starting again
/// resumes rather than refetching.
///
/// A record whose process has already gone is written off here rather than
/// left to be reported as a failure, because somebody asking for it to stop
/// and it being stopped is not a fault.
const CANCEL: &str = r#"
k=$2; s=$d/$k.download
pid=$(sed -n 's/^pid=//p' "$s" 2>/dev/null)
[ -n "$pid" ] || { echo "nothing is downloading $k" >&2; exit 1; }
if ! kill -TERM "$pid" 2>/dev/null; then
sed 's/^state=.*/state=cancelled/' "$s" > "$s.tmp" && mv "$s.tmp" "$s"
fi
"#;
/// Takes a model off a machine: the file, a partial, and the state beside it.
///
/// The partial goes *before* the signal, so that a worker dying from it finds
/// nothing to write about and leaves no state file behind -- see [`WORKER`]'s
/// trap.
const REMOVE: &str = r#"
k=$2; t=$d/$k
if [ ! -f "$t" ] && [ ! -f "$t.part" ] && [ ! -f "$t.download" ]; then
echo "$k is not on this machine" >&2
exit 1
fi
rm -f "$t.part"
pid=$(sed -n 's/^pid=//p' "$t.download" 2>/dev/null)
[ -n "$pid" ] && kill -TERM "$pid" 2>/dev/null
rm -f "$t" "$t.part.err" "$t.download" "$t.download.tmp"
# The repository's own directory, and only that one: it was made by the
# download and is litter once empty. Never its parent, which for a key with
# one component would be the models directory itself.
rmdir "${t%/*}" 2>/dev/null
exit 0
"#;
/// Runs one of the scripts above on `transport`, against `dir`.
async fn run(transport: &Transport, script: &str, dir: &str, rest: &[&str]) -> Result<String> {
let mut args = vec![
"-c".to_string(),
format!("{EXPAND}{script}"),
"sh".to_string(),
dir.to_string(),
];
args.extend(rest.iter().map(|arg| (*arg).to_string()));
transport.capture(&Launch::new("sh", args, None)).await
}
/// Starts fetching `file` from `repo` onto the machine `transport` reaches.
pub async fn start(
transport: &Transport,
dir: &str,
repo: &str,
file: &str,
published: &Published,
) -> Result<()> {
let key = key_for(repo, file);
checked(&key)?;
let url = format!(
"https://huggingface.co/{repo}/resolve/main/{}",
file.replace(' ', "%20")
);
run(
transport,
START,
dir,
&[
&key,
&url,
published.sha256.as_deref().unwrap_or(""),
&published.bytes.map(|b| b.to_string()).unwrap_or_default(),
WORKER,
],
)
.await?;
Ok(())
}
/// Every download in flight or stopped on that machine, in key order --
/// which is the order the models they become are listed in.
pub async fn downloads(transport: &Transport, dir: &str) -> Result<Vec<Download>> {
let out = run(transport, LIST, dir, &[]).await?;
let mut found: Vec<Download> = out
.split('\0')
.filter(|record| !record.is_empty())
// The key last, so a name with a separator in it cannot eat a field,
// and the state file's own lines separated by RS inside field three.
.filter_map(|record| {
let mut fields = record.splitn(4, '\u{1f}');
let done: u64 = fields.next()?.trim().parse().unwrap_or(0);
let alive = fields.next()? == "yes";
let status: HashMap<&str, &str> = fields
.next()?
.split('\u{1e}')
.filter_map(|line| line.split_once('='))
.collect();
let key = fields.next()?.trim().to_string();
let (repo, file) = key.rsplit_once('/')?;
let recorded = status
.get("state")
.and_then(|word| DownloadState::from_word(word.trim()));
let error = status
.get("error")
.map(|text| text.trim())
.filter(|text| !text.is_empty())
.map(str::to_string);
// A word this server does not know means a state file written by
// a different version of the script. Whether there is a process
// behind it is measured either way, so that is what it is
// reported as -- dropping the record instead would make a
// download that is using disk invisible.
let recorded = recorded.unwrap_or(if alive {
DownloadState::Running
} else {
DownloadState::Failed
});
// A record that says it is working and has no process behind it
// is the one case this has to correct: the machine was restarted,
// or the worker was killed by something other than a cancel.
let stopped = matches!(recorded, DownloadState::Running | DownloadState::Verifying);
Some(Download {
state: if stopped && !alive {
DownloadState::Failed
} else {
recorded
},
error: error.or_else(|| {
(stopped && !alive).then(|| {
"it stopped before it finished -- the machine may have been restarted. \
Downloading again carries on from here."
.to_string()
})
}),
done,
total: status.get("total").and_then(|t| t.trim().parse().ok()),
repo: repo.to_string(),
file: file.to_string(),
key,
})
})
.collect();
found.sort_by(|a, b| a.key.cmp(&b.key));
Ok(found)
}
/// Asks the machine to stop one download. The partial stays behind.
pub async fn cancel(transport: &Transport, dir: &str, key: &str) -> Result<()> {
checked(key)?;
run(transport, CANCEL, dir, &[key]).await?;
Ok(())
}
/// Removes a model from a machine, downloaded or half-downloaded.
pub async fn remove(transport: &Transport, dir: &str, key: &str) -> Result<()> {
checked(key)?;
run(transport, REMOVE, dir, &[key]).await?;
Ok(())
}
/// Where a machine reached over ssh keeps its models, when its machine does
/// not say.
///
/// The same place this backend puts its own downloads, written out rather
/// than derived: `$XDG_DATA_HOME` here describes *this* machine's
/// environment, and the far machine's is the far machine's business. A
/// machine whose models are elsewhere says so (`SshConfig::models_dir`).
const FAR_MODELS_DIR: &str = "~/.local/share/ai-app/models";
/// Which directory holds the models on the machine `transport` reaches.
///
/// One answer, because two things ask: the list a spawn screen offers,
/// and the path a session hands `llama-server`. A machine that listed one
/// directory and served from another would offer models that then failed
/// to load, which reads as the model being broken.
pub fn dir_on(transport: &Transport, local: &Path) -> String {
match transport {
Transport::Here => local.to_string_lossy().into_owned(),
Transport::Ssh { ssh, .. } => ssh
.models_dir
.as_ref()
.map_or(FAR_MODELS_DIR.to_string(), |dir| {
dir.to_string_lossy().into_owned()
}),
}
}
/// Every GGUF on a configured machine, which is the machine that
/// would have to serve it.
///
/// The local half of this is [`ModelStore::list`], reading the same shape
/// off this machine's disk; a caller picks by transport, since a machine
/// with no ssh *is* this machine and asking a shell about it would be a
/// slower way to the same answer. What must not happen is offering this
/// backend's downloads for a session on another machine: the file has to
/// be where `llama-server` runs, and a list that says otherwise is a
/// claim about the wrong filesystem.
///
/// `dir` is that machine's models directory, `~` included -- expanded on
/// the far side, which is the only place that knows what it is. A
/// directory that is not there is an empty list rather than a failure: a
/// machine that has never had a model put on it is an ordinary state, and
/// the same one as a machine whose directory exists and is empty.
///
/// Each record carries the head of the file as well as its size, because a
/// model's own name is inside it (see [`crate::gguf`]) and the file is on the
/// far machine. The alternative is a second round trip per model, or naming
/// remote models by path while local ones get their proper names -- one
/// machine's models reading differently from another's is exactly the
/// confusion the name was added to remove. The prefix is bounded at
/// [`crate::gguf::PREFIX_BYTES`], which is what keeps this one round trip's
/// worth of bytes.
pub async fn on_machine(transport: &Transport, dir: &str) -> Result<Vec<LocalModel>> {
let script = format!(
"p=$1; case $p in \"~\") p=$HOME;; \"~/\"*) p=$HOME/${{p#\"~/\"}};; esac; \
[ -d \"$p\" ] || exit 0; cd \"$p\" || exit 0; \
find . -type f -name '*.gguf' -exec sh -c '\
for f do printf \"%s\\t%s\\t%s\\0\" \"$(wc -c < \"$f\")\" \
\"$(head -c {prefix} \"$f\" | base64 | tr -d \"\\n\")\" \"${{f#./}}\"; done\
' sh {{}} +",
prefix = crate::gguf::PREFIX_BYTES,
);
let launch = Launch::new(
"sh",
vec!["-c".to_string(), script, "sh".to_string(), dir.to_string()],
None,
);
let out = transport.capture(&launch).await?;
let mut found: Vec<LocalModel> = out
.split('\0')
.filter(|record| !record.is_empty())
// Three fields, and the name last, so a `\t` in a filename survives.
// The middle one is base64, which has no tab in its alphabet.
.filter_map(|record| {
let (bytes, rest) = record.split_once('\t')?;
let (head, key) = rest.split_once('\t')?;
let (repo, file) = key.rsplit_once('/')?;
Some(LocalModel {
key: key.to_string(),
repo: repo.to_string(),
file: file.to_string(),
bytes: bytes.trim().parse().unwrap_or(0),
name: name_in_prefix(head),
})
})
.collect();
found.sort_by(|a, b| a.key.cmp(&b.key));
Ok(found)
}
/// The model's name out of a base64 prefix of its file.
///
/// The remote half of [`name_of`], and `None` for everything that half
/// answers `None` for, plus a prefix that did not survive the trip.
fn name_in_prefix(head: &str) -> Option<String> {
use base64::Engine as _;
let bytes = base64::engine::general_purpose::STANDARD
.decode(head.trim())
.ok()?;
crate::gguf::name(&mut bytes.as_slice())
}
/// A model repository on HuggingFace, as the browse screen shows it.
#[derive(Debug, Clone, Serialize)]
#[serde(rename_all = "camelCase")]
pub struct RemoteRepo {
/// `owner/name`, which is what everything else here is keyed by.
pub id: String,
pub downloads: u64,
pub likes: u64,
}
/// One downloadable file inside a repository.
#[derive(Debug, Clone, Serialize)]
#[serde(rename_all = "camelCase")]
pub struct RemoteFile {
pub path: String,
pub bytes: u64,
/// Already on the machine being looked at, so the phone can say so rather
/// than offering to fetch it again.
pub have: bool,
}
/// Searches HuggingFace for GGUF repositories matching `query`.
///
/// Proxied through this server rather than called from the phone, for two
/// reasons that both matter: the app trusts exactly one certificate -- this
/// server's -- and has no general internet trust to spend on huggingface.co,
/// and the machine that has to do the downloading is this one.
pub fn search(query: &str) -> Result<Vec<RemoteRepo>> {
let url = format!(
"https://huggingface.co/api/models?search={}&filter=gguf&limit=25&sort=downloads&direction=-1",
urlencode(query)
);
let body = get_json(&url)?;
let list = body
.as_array()
.context("HuggingFace returned something that is not a list")?;
Ok(list
.iter()
.filter_map(|m| {
Some(RemoteRepo {
id: m.get("id")?.as_str()?.to_string(),
downloads: m
.get("downloads")
.and_then(serde_json::Value::as_u64)
.unwrap_or(0),
likes: m
.get("likes")
.and_then(serde_json::Value::as_u64)
.unwrap_or(0),
})
})
.collect())
}
/// Asks HuggingFace for both halves of [`Published`] in the one call they
/// come from.
pub fn published(repo: &str, file: &str) -> Result<Published> {
let url = format!("https://huggingface.co/api/models/{repo}/tree/main?expand=true");
let body = get_json(&url)?;
let entry = body
.as_array()
.and_then(|files| {
files
.iter()
.find(|f| f.get("path").and_then(serde_json::Value::as_str) == Some(file))
})
.with_context(|| format!("{repo} does not have a file called {file}"))?;
Ok(Published {
bytes: entry.get("size").and_then(serde_json::Value::as_u64),
sha256: entry
.get("lfs")
.and_then(|lfs| lfs.get("oid"))
.and_then(serde_json::Value::as_str)
.map(str::to_string),
})
}
/// The GGUF files in one repository, largest last, with the ones already
/// downloaded marked.
pub fn files(repo: &str, have: &std::collections::HashSet<String>) -> Result<Vec<RemoteFile>> {
let url = format!("https://huggingface.co/api/models/{repo}/tree/main");
let body = get_json(&url)?;
let list = body
.as_array()
.context("HuggingFace returned something that is not a list")?;
let mut files: Vec<RemoteFile> = list
.iter()
.filter_map(|f| {
let path = f.get("path")?.as_str()?.to_string();
if !path.ends_with(".gguf") {
return None;
}
Some(RemoteFile {
bytes: f
.get("size")
.and_then(serde_json::Value::as_u64)
.unwrap_or(0),
have: have.contains(&key_for(repo, &path)),
path,
})
})
.collect();
files.sort_by_key(|f| f.bytes);
Ok(files)
}
fn get_json(url: &str) -> Result<serde_json::Value> {
let text = ureq::get(url)
.header("User-Agent", USER_AGENT)
.call()
.and_then(|mut r| r.body_mut().read_to_string())
.with_context(|| format!("GET {url}"))?;
serde_json::from_str(&text).with_context(|| format!("{url} did not return JSON"))
}
/// Percent-encodes a query string. Deliberately minimal -- this escapes what a
/// model search and a model key actually contain rather than implementing the
/// whole rule set, and anything unexpected becomes `%XX` rather than being
/// passed through.
pub fn urlencode(value: &str) -> String {
value
.bytes()
.map(|b| match b {
b'A'..=b'Z' | b'a'..=b'z' | b'0'..=b'9' | b'-' | b'_' | b'.' | b'~' => {
(b as char).to_string()
}
b' ' => "+".to_string(),
other => format!("%{other:02X}"),
})
.collect()
}
#[cfg(test)]
mod tests {
use super::*;
fn model(repo: &str, file: &str, name: Option<&str>) -> LocalModel {
LocalModel {
key: format!("{repo}/{file}"),
repo: repo.to_string(),
file: file.to_string(),
bytes: 1,
name: name.map(str::to_string),
}
}
/// The cascade, and the one rung that is not simply "is it unique": a name
/// the file it came from says nothing about is the converter's working
/// directory rather than the model's name.
#[test]
fn a_model_is_labelled_by_the_first_thing_that_identifies_it() {
let models = [
// Its own name, which is what a reader recognises.
model(
"Qwen/Qwen3-0.6B-GGUF",
"Qwen3-0.6B-Q8_0.gguf",
Some("Qwen3-0.6B"),
),
// Two quantisations carry one name, so both fall to the file.
model("org/Big-GGUF", "Big-Q4_K_M.gguf", Some("Big")),
model("org/Big-GGUF", "Big-Q8_0.gguf", Some("Big")),
// `convert_hf_to_gguf.py` naming the directory it converted.
model("PrismML/Bonsai-GGUF", "Bonsai-1B-TQ1_0.gguf", Some("hf")),
// Nothing to go on but where it came from.
model("org/Quiet-GGUF", "weights.gguf", None),
];
assert_eq!(
labels(&models),
[
"Qwen3-0.6B",
"Big-Q4_K_M",
"Big-Q8_0",
"Bonsai-1B-TQ1_0",
"weights",
],
);
}
/// Two repos holding a file of the same name: the file rung collides as
/// well, and the key is what always terminates the cascade.
#[test]
fn a_file_name_two_repos_share_falls_through_to_the_key() {
let models = [
model("one/GGUF", "model.gguf", None),
model("two/GGUF", "model.gguf", None),
];
assert_eq!(
labels(&models),
["one/GGUF/model.gguf", "two/GGUF/model.gguf"]
);
}
}