ff39ef5cf97a77da7850b4d38274c3abf8e5b475
145
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
d7c692a4ec |
Show a replayed session's images instead of dropping them
An imported session showed no screenshots. `text_of` kept only `text` blocks, so every image in the replayed tail was silently discarded -- while the *live* translator has always saved them into the session's `files/` and referenced them. Two readings of the same records, and the one used for history was the lesser. `save_image` moves out of `Translator` to a free function both paths call, since the naming scheme for that directory should exist once. `events_from` now takes the session directory to write into, which means the conversion has to happen where that directory exists -- so `Seed` carries the raw JSONL and `launch` turns it into events, rather than `routes` doing it before the session is created. Costs nothing in tokens, which is the point worth recording: this writes into ai-app's own session directory and the phone fetches a reference only when it draws one. Nothing here is ever written to the CLI's stdin -- it reads its own session file, and the only things this app sends it are typed messages, control requests and `/compact`. Verified against the 133 MB session behind the 2026-08-29 incident: 45 images in the replayed tail, written as real PNGs and served over the files route, with the transcript itself staying at 756 KB of references. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01VETa8afmpWaYezLCqJhDB8 |
||
|
|
d96bc041a7 |
Say how big a session is before it is imported
The import list reported a line count, which is the wrong axis: these transcripts embed screenshots as base64, so one line can be a megabyte. On this machine a 69 MB session has 3,427 lines while a 44 MB one has 6,792 — the number on the row said nothing about what continuing the session would cost, and size is the only thing there that predicts it. The session behind the 2026-08-29 incident was 65 MB across 13,000 lines, a line count that looks unremarkable. Shown beside the line count rather than instead of it, since a short file of long lines is exactly the expensive case. Not warned about and not marked: importing a large session is a choice somebody is entitled to make, and flagging it would be the interface nagging about a decision already taken. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01VETa8afmpWaYezLCqJhDB8 |
||
|
|
362d436d4f |
Let sessions outlive the backend, and never resume one twice
Three `claude` processes ended up running against this checkout on 2026-08-29, and the account hit its session limit. One cause, several ways in. An agent imported the Claude Code session it was *itself* running in. That is an ordinary import, and importing runs `--resume` -- so a second CLI attached to a file the first was still writing. The whole 65 MB conversation, 154 embedded screenshots included, was re-appended to the transcript under a new prompt id; both copies then read each other's writes as work done elsewhere, and the adopted one was billed for re-reading all of it. Meanwhile `shutdown_all` asked each session to stop and the process exited immediately, so the SIGKILL timer died with the runtime, the stop was unreliable, and whatever survived was orphaned with nothing written down to find it by. The processes leaked either way. So leak them on purpose, and be able to pick them back up. A session's process now outlives the backend and is adopted again on the way up, which is worth having for its own sake: restarting the server no longer ends a turn somebody is waiting on. Its stdio lives in the session directory -- a fifo opened read-write so the process is its own last writer and never reads EOF, plus stdout/stderr logs read from a byte offset. `session::process` records the pid *and* the kernel's start time for it, because a pid alone is reused and adopting a stranger's would mean never resuming the real conversation. That makes the fix structural rather than a check: everything goes through `ClaudeDriver::launch`, which adopts if it can and starts if it cannot, and `--resume` is reachable only on the second path. `Driver` gains two ways out where it had one -- `detach` (coming back) and `stop` (the session is being deleted, so the process must not survive). Importing a session that is open is now refused outright. Claude Code keeps `~/.claude/sessions/<pid>.json` for every live session, so this is a measurement rather than a guess; it reports no/yes/unknown, because a machine that keeps no such record cannot answer and "could not check" is not "nobody is using it". `SessionStatus` gains `Unknown` for the same reason. Also here, found on the way: - A reconnecting phone was sent the entire backlog. Opening a session was bounded to a page but reconnecting was not, so a long disconnect delivered thousands of events one frame at a time. Past `CATCH_UP_LIMIT` the stream sends a `reset` frame and the newest window, and the client rebuilds from it as it does on open -- without the reset the window is spliced onto rows no longer adjacent to it. - A session's status was assumed idle at launch. Read from the transcript instead, so a restart stops claiming an exited session is waiting for you. - `llama-server`'s stdout was piped and never drained, so a chatty one blocked on a full pipe buffer mid-load. It goes to a log now. - A turn that exited or errored never emitted `Idle`, so the queue stayed "running" for good: every later message was held forever and, since a message is only recorded when taken, vanished with nothing on screen. - Two doc comments had drifted onto the wrong functions. Verified by killing the server mid-turn: the process survived, finished its turn unattended (12.8 KB of output nothing was reading), and the restarted server adopted it -- one process, all 700 lines in the transcript, no hole, and it still took a new message afterwards. Deleting a session stops its process; a 266-event backlog resets while a 16-event one streams. 46 tests, clippy and rustfmt clean, app compiles and lints. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01VETa8afmpWaYezLCqJhDB8 |
||
|
|
9791afcfd6 |
Hold a queued message below the indicator until the turn takes it
The backend records a message the moment it is sent, so one sent into a running turn landed in the transcript at the time *we* spoke -- above the working indicator, in among things the session had already read. It had not read it. Showing it there says otherwise. Messages sent while a turn is in flight are now held below the indicator, drawn quieter, and take their place in the conversation when the turn ends. That is an approximation and worth naming: the CLI injects a queued message at a tool boundary, and tells nobody when it does, so the end of the turn is the first moment anything here can honestly say the message was taken. It errs toward "not yet read", which is the direction that cannot mislead. Also in this change, from the same pass over the screen: the app draws above the gesture strip rather than under it, and the three status colours that were literals in two other files -- an amber, a green and a red off Material's defaults -- are now Catppuccin members in Theme.kt beside the rest of the scheme. The ordering is verified by construction rather than photographed: the list is bottom-anchored, so the first item emitted is the lowest on screen, and the queued block is emitted before the indicator. Staging a real long-running turn to photograph cost four model turns and never produced one, because the model kept declining to sleep -- which is its own finding, and the reason the next change is a test command in the echo driver. |
||
|
|
f402a1f6c8 |
Hold the bottom while typing, and put the status where the answer goes
**Typing moved the newest message out of sight.** The list re-pinned on a new item and nothing else, but the two things that shrink it while somebody writes a reply are the field growing from one line to four and the keyboard opening under it -- neither of which is a new item. So the message being replied to drifted upward, and the view only came back when the reply was finally sent, which is the one moment it did not matter. It now watches the viewport as well as the item count, and only acts while already pinned. **The status has moved out of the corner.** As a label up there it said something about the session; at the end of the transcript it says something about a place -- this is where the next thing appears -- and that is where the reader is already looking, because it is where the last message is. `exited` is still said, in the same place. Removing the corner label without it would have left a session whose process is gone looking exactly like one waiting for input, on a screen whose whole purpose is typing at it. Verified on the phone-sized case rather than reasoned about: three lines of text with the keyboard open keeps the newest message directly above the field, and sending shows "working" immediately under the sent message. |
||
|
|
e03b757dec |
Stop losing tool calls at the seam between transcript pages
A tool's start and end are two events folded into one row, and the fold only ever *updated* an existing row -- so an end whose start was not in the same fold changed nothing and vanished. Not a broken row: no row at all, which is indistinguishable from a tool that never ran, and is what Iris saw as gaps where she remembered work happening. Paging made it routine. Each page was folded on its own and prepended, so every seam split whatever spanned it: 30 tool ends in the first page of this conversation, one of them already orphaned before a single scroll. Two changes, because the two halves fail differently. Pages are re-folded from the events they came from rather than folded apart and stitched together. That needs the events kept beside the rows, since folding is one-way. One pass over everything loaded, paid only when somebody scrolls back, which is the moment they asked for it. And an end with no start now creates a row instead of disappearing. Its name is unknown from an end alone, so it says "tool" until the earlier page arrives and replaces it -- a row that admits what it does not know beats silence, because silence is a claim that nothing happened. Verified across a real seam: scrolling back through the 80-event boundary of an 864-event import is continuous, with no gaps where tool calls were. |
||
|
|
2fe34176c0 |
Give the transcript a way back, and stop losing an import's name
**Every imported session was called "claude-cli session".** The app has nothing to say about the title -- the server names it after the session it is continuing -- so it sent `""`. That is `Some`, which satisfied the `or_else` meant to catch "no title given", so the imported name was computed and then thrown away in favour of the `<provider> session` fallback. Normalised at the boundary instead: blank means absent, because that is what it means to the person who left it blank. Both the client's value and the imported one go through the same trim, so neither can be a string of spaces standing in for a name. **And a jump-to-latest button**, shown only while the newest message is off-screen. Reading back through a conversation is a place to be rather than a state to be rescued from, so it waits to be wanted and leaves once there is nowhere to jump to. It says where it goes instead of drawing an arrow. The list is laid out from the bottom, so "down" in the data is up on the screen, and an arrow would be asking the reader to hold that in their head to press a button. Verified on screen: an import now arrives titled "ai-app" rather than "claude-cli session", and the button appears on scrolling back, returns to the newest message, and disappears on arrival. |
||
|
|
dcb158ee44 |
Fetch the transcript instead of replaying it one event at a time
**The five seconds of loading top-down.** Opening a session subscribed to the event stream from sequence zero, so the backlog arrived as one SSE frame per event -- 864 of them for an imported conversation, rendered as they landed. That is not a slow list; it is a conversation being replayed at network speed, and it looks like loading from the top because it is. The newest page now comes as one request, and the stream starts from where that page ended, carrying live events only -- which is what a stream is good at. Scrolling back fetches the page before it, so history costs something only when somebody actually reads it. 80 events instead of 864, and the first frame is already the end of the conversation. I had called this fixed after anchoring the list at the bottom, on the strength of an emulator on the same machine as the server. That test could not have shown the problem: the whole backlog arrived in one frame's worth of time over loopback. Iris's phone, over a tunnel, took five seconds. **Send disappearing while running.** It was never conditional -- the row simply ran out of width. A Row hands out intrinsic widths in order and clips the overflow, so when Stop appeared the pickers I had added pushed Send off the screen: the app's central control, gone at exactly the moment the app is most in use. The settings now share what is left after the actions have taken what they need. While a turn is in flight the button says **Queue**, because that is what sending then does -- the message is injected at the next tool boundary rather than starting a turn of its own. The backend has always done this; the button was describing something else. **And the model picker no longer dismisses the keyboard**, which it did by taking focus. Changing the model mid-sentence is an aside, not a departure from what you were typing. Verified on the 864-event import: at the newest message within a second, history paging back continuously past the first page, and Stop beside Queue while running. |
||
|
|
ba25a5cacf |
Open the transcript at the bottom instead of travelling there
The list was built oldest-first and then scrolled to the end, so opening a session started at the top and raced downward through everything in it. On an imported conversation that is nine hundred items measured before a word is readable, and it was visible every single time. Laying the list out from the bottom removes the journey rather than hiding it. The newest message is index 0, so the first frame is already the right one, and older items are composed only when somebody scrolls back to them -- which is also what makes a long history cheap rather than something to load up front. Following the tail gets simpler as a result. There is no longer a moment where new content pushes the anchor away, so "am I pinned" is read straight from the scroll position instead of being remembered across scrolls, and a new message is one step back to index 0 rather than a jump across the transcript. Checked on the 863-event import of this very conversation: a screenshot one second after opening is already at the newest message, and scrolling back reaches older ones in the order they happened. |
||
|
|
3f8805a610 |
Say which kind of delete this is, and stop offering what is already open
**Deleting was one word for two different acts.** An imported session's real transcript belongs to Claude Code and outlives anything this app does, so removing it here undoes a view. A session started here has no copy anywhere, and removing it ends the conversation. The dialog warned "this can't be undone" of both, which makes the warning worthless on the one where it is true -- and frightening on the one where it is not, since what it actually deletes is a cache of a conversation still sitting on the machine. Sessions now report whether they were imported, and the dialog says which act this is. No new mechanism: the soft delete already existed, it was just indistinguishable from the hard one. **And a session already open here is no longer offered for import.** Importing one twice would leave two `--resume` processes appending to the same transcript, each seeing the other's writes as work done elsewhere and replaying them -- both sessions then showing a conversation neither is having. The route refuses it as well, so the rule holds for anything not going through the app. Left out of the list rather than shown and disabled. The usual argument says absence is ambiguous, and it is wrong here: an imported session has not disappeared, it has moved to the session list, which is where it now belongs. Absence means "already somewhere you can reach it". Deleting the app's copy puts it straight back -- verified: 68 offered, 67 after importing one, 68 again after the soft delete, which is also the clearest demonstration that a soft delete keeps the conversation. |
||
|
|
d2915c12fa |
Record how the import sync tells its own writes apart
The reasoning belongs where somebody would look before changing the poll: status is the obvious discriminator and is wrong, and the failure it produces reads as the model repeating itself rather than as a bug. |
||
|
|
a9ea84c96c |
Stop choosing a model, and keep an imported session up to date
**Why the model became fable.** `spawn_session` fell back to the
provider's first listed model when none was given. That list is a shortcut
for the spawn screen, written in whatever order somebody typed it, and its
first entry is `fable` -- so every session spawned without a model, which
is every import, silently became a fable session. It looked like a default
and was an artefact of list order. Absent now means absent: no `--model`
flag, and the CLI uses whatever the person configured for themselves.
**Model and permission mode are now visible and changeable** from the
session, as buttons that read as their current value rather than labels
beside one. The mode was spawn-only; the CLI turns out to accept
`control_request{subtype:set_permission_mode}` and echo the mode back,
probed against 2.1.237 the same way the rest of the protocol record was.
Both default to `auto` -- on a phone every ask is a round trip to a
question card, which is how "allow Bash?" became the most-answered
question in the app.
The mode is reported by the API so the picker shows what the session is
actually set to, and it is kept in the live session beside the model for
the reason the model already was: `meta` is the shape a session was
*launched* with, so reporting from it shows the value a change replaced.
**And an imported session keeps itself level with its source file**, so
work done at a terminal arrives without a button. `--resume` appends to
the same transcript rather than forking -- measured, not assumed -- so the
only hard question is which new lines came from here.
Answered by counting the events this session has recorded. Status is the
obvious signal and is wrong, which cost a round trip to find: a turn that
starts and finishes between two polls reads as idle at both, so its output
is replayed on top of itself. It showed up on screen as `donedone`, and
only because the reply was one word -- with a longer answer it would have
looked like the model repeating itself.
Verified against both halves: text appended to the source file the way a
terminal writes it appears within one interval, and a message sent through
the app appears exactly once, before and after a turn.
|
||
|
|
c3e7f07a5d |
Pin the transcript to its tail, and stop asking about every command
**The scroll.** The transcript scrolled on new items and nothing else, which missed the two cases that matter most. An imported session's history arrived and left the view wherever it landed; the keyboard opening shrank the viewport and slid the newest messages under the IME, so typing meant typing into a view showing the middle of something. The view is now pinned to the tail, and it is the reader's scroll that decides: settling anywhere above the bottom releases the pin, settling back at the bottom re-arms it. The pin is written only when a scroll *ends*, so it survives the moment when new content has just pushed the bottom away but the reader never moved -- deriving it continuously from "is the bottom visible" would release it on every append, which is the race that makes naive follow-the-tail implementations let go at random. New items and viewport resizes both re-scroll; the jump is instant rather than animated, because an imported session appends hundreds of items at once and animating through them is a light show. **The input field gets a row of its own**, above the buttons. Sharing one row put the full width behind three controls, so the thing being typed into was the narrowest thing on the row. **Permissions default to auto, and importing can choose.** The spawn screen defaulted to "manual" and imports passed no mode at all, so the CLI asked about everything -- and on a phone every ask is a round trip to a question card, which is how "allow Bash?" became the most-answered question in the app. Both paths now default to auto, with the other modes one tap away for a session that warrants caution. Looked at running, all three: an imported session opens at its bottom, the tail stays visible while typing with the keyboard open, and the mode picker shows auto selected. |
||
|
|
233689ced6 |
Name a session in the import list, and let one be deleted
Three things about finding a session in a list of seventy, and one about getting rid of it. **A name beats anything inferred.** `/rename` appends a `custom-title` record, so if somebody has said what a session is, that is the row. Eleven of the seventy here turned out to be named already and none of it showed. **Otherwise the last thing said, not the first.** The question this list answers is "which one was I just in", and a session's opening line is the least distinctive thing about it -- several of these begin with the same slash command. Finding that last message took three tries, and the two wrong ones are worth recording because they failed in opposite directions. Grepping the user record type caught tool results, which are *also* user records -- so a session that ended mid-tool showed a tail of empty records and a row saying nothing was said, when plenty had been. Narrowing to a string `content` fixed those two and broke twenty others, because a message carrying an attachment stores its text in a list. Excluding `tool_use_id` keeps both shapes of a real message and drops the one that is not: seven rows still have nothing to show, and those are sessions that really are empty. **Sorted by when it was last used**, and the time is on the row. Naming was tried as the first sort key and is a worse list -- it buries what somebody was just doing under everything they ever named. A name is for recognising a row, not for ordering it, so it stays as the title and as a word beside it. **And a session can be deleted**, which is asked before it is done. The transcript *is* the session, so this ends any chance of resuming that conversation, and the dialog says exactly that rather than "are you sure?". Deletion resolves the id against what the machine reported, like importing, so no path crosses the wire in either direction. Looked at on the emulator, including the dialog -- which is where the delete button turned out to be missing entirely after a patch that compiled fine, and where the row layout got its second look. |
||
|
|
6bbc829a3e |
Import a Claude Code session the machine already has
Claude Code keeps every session as JSONL under `~/.claude/projects/`, and the CLI continues one with `--resume <id>`. `claude.rs` already resumes whenever it finds a resume token in the session directory, for crash recovery -- so importing is that same path with the token written before the driver starts, and there is deliberately no second way to begin a session. The seed goes through `launch` with the ordinary spawn, so the driver never learns which kind it got. Two things the machine answers and the phone does not. **Which sessions exist.** One command per setup rather than one per file, for the reason discovery already gives: over ssh each would be its own connection. Titles come from the first few user records rather than the first, because a session opens with records the CLI injected -- slash commands, caveats around local command output -- which are stored as ordinary user records without the meta flag, so titling by "first user record" produced a list where most rows read `<command-name>/clear`. **Which file an id names.** The phone sends an id and never a path; the server looks it up again among the sessions it enumerated. An enrolled token must not be able to turn a spawn into "read me this file", which is the same rule that keeps a provider's command out of `POST /setups`. Only the tail is replayed. The imported conversation is for reading -- continuing it is the CLI's job, and it reads the whole file itself -- so this is a display budget, and it has to be one: the session this was written in is 39 MB, and all of it would otherwise cross a tunnel to a phone. A recorded working directory can outlive itself, which this found immediately: every session from before the checkouts moved to `~/repos` still records `~/host/repos/...`. Resuming into one fails at `cd` before the CLI starts -- a confusing way to meet a feature whose promise is "carry on where you left off" -- so the directory is checked, and a missing one is dropped with a log line naming it rather than being passed on to fail. Verified against this very session: 905 events replayed from the tail (351 tool calls, 350 results, 185 assistant messages, 19 mine), the resume token pointing at its id, and the stale directory reported and dropped. The list was read on the emulator, where the top row is that session under its opening sentence. |
||
|
|
2a1bc84c1e |
Expand a leading ~, and say what the failing command said
Three things, two of which are the same failure seen from opposite ends. **A working directory of `~/repos/ai-app` never worked.** Everything crossing to the remote side is single-quoted, which is right for paths, model names and prompts alike -- unquoted they would be shell syntax rather than data. It is wrong for exactly one character: `~` means "expand me", and quoting is what stops expansion. So the remote shell was handed the literal four-character directory `~` and correctly said it did not exist, which reads as the path being wrong rather than the quoting. Paths now go through `quote_path`, which emits `"$HOME"` for a leading `~/` and single-quotes the rest. The variable expands, the expansion is not re-split or globbed because it is double-quoted, and nothing after it gains a meaning -- there is a test that pushes a quote-and-semicolon injection through the tilde branch and gets back one absurd path rather than three commands. `$HOME` is set by every shell this can land in, so this does not depend on the remote side being POSIX; verified by running the generated script under both sh and fish, which is what the dev VM actually uses. **The phone could not have told you any of that.** The exit report kept the last line of stderr, and a shell's error message ends with a blank line -- so the last line was empty, the report was a bare exit status, and the seven lines of fish complaining sat in the server's log where nobody holding a phone is looking. It now keeps the last 50 lines in a ring and reports them with blank lines trimmed from both ends. The tests use the real fish `cd` failure as their fixture. **The status bar was unreadable.** `isAppearanceLightStatusBars` was hardcoded to `true` -- dark icons -- which was right against the default light surface and wrong the moment the app wore Mocha. It now asks the scheme's own background for its luminance, so changing the palette cannot reintroduce it. **And the address field takes `user@host:port`.** One field rather than two, because that is how an address is written everywhere else and a port that is nearly always 22 does not deserve its own box on a phone keyboard. Absent means absent rather than 22: the backend already decides that default, and writing it here would be a second answer in a second place. A colon only means "port" when it can -- brackets for IPv6 as ssh writes them, otherwise exactly one colon followed by digits. Looked at on the emulator: the status bar, and the form, whose label I then shortened because it wrapped onto a second line and made that field taller than the two beside it. |
||
|
|
31135e3f22 |
Wear Catppuccin Mocha, and move Usage to where the provider is
Two changes to the app, plus the one they turned up. **The theme.** Catppuccin Mocha, copied from dev-updater rather than shared: wg-app-link is the *link* -- the tunnel, the pinned CA, enrollment -- and a palette is not that. The two apps looking alike is a preference rather than a contract, and the moment one wants a different accent a shared version becomes a thing to fight. dev-updater's ActionTone and its ANSI table did not come across; nothing here draws a log or a destructive button yet, and copying a vocabulary with no speakers is how a file starts lying about what the app does. **Usage is no longer a global button.** It belongs to the provider, and the session view is the only place a provider is currently named, so that is where the control sits -- beside the line that names it, rather than collected with the app-wide controls where its scope had to be guessed. It carries the session back with it, so Back returns to that session rather than dumping the reader on the list. Its real home is that provider's own settings, which do not exist yet. **And the bit that only running it could find.** I first subtitled the usage screen with the session's provider, which on an echo session put "echo" directly above a card reading "claude" -- Claude's account-wide numbers labelled as echo's, a claim about echo that nothing measured. The subtitle is gone; each card names the service that answered, which is the true scope, and the reason is written where the subtitle was so it does not get re-added. Looked at on the emulator rather than read: the palette, the session header, the usage screen, and Back landing on the session it came from. ktfmt, compile and Lint clean. |
||
|
|
4370c467ca |
Discover this machine's providers instead of asserting them
A fresh install wrote a `claude-cli` provider into the local setup unconditionally. Nothing looked for `claude`; the list was hardcoded in `Config::seed`, so on any machine without it -- which is every machine but the dev VM -- the phone was offered a provider that cannot spawn, stated with exactly the confidence of one that had been checked. Discovery already existed and was already right: `setups::discover` probes with `command -v` over the transport, includes echo for the local one because it runs in-process, and records the resolved path rather than the bare name. Only the local setup skipped it, which is the one place the answer felt obvious enough not to ask. So `seed` now takes the providers it is given, and seeding asks this machine the same question it asks any other. It moved out of `SessionManager::new` into an awaited step in main, because asking is I/O and a constructor that quietly spawns a subprocess surprises every caller. A discovery that fails seeds `echo` alone and says so, since echo is true wherever this server runs -- falling back to the hardcoded list would be the same bug with an extra step. The test that covered this agreed with the bug, because both were written from the same assumption: it asserted the seed contains `claude-cli`. It now asserts the opposite -- that the seed invents nothing -- and the session tests seed echo explicitly rather than relying on a constructor that would make them pass or fail on whether `claude` happens to be installed on whoever runs them. Verified by running a server on a PATH holding only `sh`: it seeds `echo` alone. With claude and llama-server present it finds both. 34 tests. |
||
|
|
3cf6925d90 |
Let Dev Updater supervise the backend, now that the switch can be sequenced
Re-applies the change reverted in
|
||
|
|
a8b6c13b01 |
Report a crashed service as failed, which OpenRC was never telling us
The OpenRC branch of this script had never run anywhere. A guest built to reproduce the host says it was wrong in the way the `failed` state exists to prevent: a service that fell over reported `stopped`, which reads as a decision somebody made. Two causes, both measured rather than reasoned about. `rc-service status` prints `* status: crashed` to **stderr**. The check was `status 2>/dev/null | grep -qw crashed`, which discards precisely the word it is searching for, finds nothing, and falls through to `stopped`. The old comment argued for reading the word rather than the exit code, and that argument was sound except that the code turns out to be specific rather than merely non-zero. So it now reads the code, which says more than the text did: 0 started, 3 stopped, 32 crashed, and 1 for every way the question cannot be answered -- an unknown service, XDG_RUNTIME_DIR unset, or a user softlevel that was never initialised. That last group is a real state and not one of the other three, so it exits non-zero and says so instead of guessing. `set -e` is the second cause, found by running the first fix: every answer except "running" is a non-zero exit, so a bare invocation killed the script before the code could be looked at. It prints nothing and exits 3, which is indistinguishable from a crash of this script itself. Verified in the guest, all four states: not-installed, stopped, failed after a real crash, and a non-zero exit with nothing on stdout when the softlevel is removed. Before the fix the crashed case printed `stopped`. |
||
|
|
297e85c68d |
Delete the pre-setups migration, which has done its job
The rule is that migration code goes once the update carrying it has been received, because there is one backend and one phone: once they are past a shape, nothing anywhere is still on it, and a second parsing path that nothing exercises only constrains later changes to the schema. The module's own comment said to delete it "once the host has started on a build containing it", and that has happened -- it is in the pushed commit the host reports itself up to date with, and the server has been starting on it. Out: the `legacy` module, `migrate_from_pre_setups`, the branch in `Config::load` that reached it, and the test. `Config::load` is now one expression. PLAN.md keeps the history rather than reverting to what it said before, because the interesting part is not the migration but the decision it replaced: refusing to start on an old config was the wrong trade and proved it on Iris's host, as a crash loop that could not explain itself because the crashing process is how the phone reaches the machine at all. Verified by running it, since the point of this change is what happens at startup: a server with no existing state starts, generates its CA, prints its enrollment QR and writes a config that reads back. 34 tests, clippy silent, rustfmt clean. |
||
|
|
af41d86186 |
Pin the crate at the local-network note
Nothing in this repo changes; the submodule moves to the commit carrying the measured finding that ACCESS_LOCAL_NETWORK is still required through the tunnel, contrary to Android's own documentation. |
||
|
|
d5a0f67a3a |
Put the service script back until the switch can be sequenced
Reverts the switch to `service: Managed(...)`. The switch is still right and the reasoning in that commit still holds; what was wrong was doing it now, unilaterally, to a checkout something is reading live. A dev-updater is running against this working tree, so deleting `server/service` did not wait for a pull to take effect -- the backend card went to "couldn't check -- failed to run the service script: No such file or directory" immediately, and the pushed declaration still names the script, so the tree and the declaration disagreed in the one direction that breaks things. My own commit message had said this change was not safe to pull blind; it turned out not to need a pull at all. The switch needs three steps in order, and only the middle one is mine: Uninstall from the backend card while the script is still declared, then take the change, then Install. Re-apply when Iris is ready to do that, which is also when dev-updater's conversion path can be deleted. |
||
|
|
295602adfe |
Save the config through the shared crate as well
`Config::save` was the same nine lines as dev-updater's, so it is now `format::write(path, self)`. The reasoning that made those nine lines correct -- the leftover temp file that keeps its old mode and is then renamed over the token hashes -- lives with the code and its test rather than in two places that could stop agreeing. Verified: 35 tests, clippy silent, rustfmt clean. |
||
|
|
b6b33dc9c5 |
Take the app half from wg-app-link as well
The four Kotlin files that were the link rather than this product now come from the submodule: the pinned TrustManager, the enrollment store and its Keystore sealing, the QR capture activity, and the local-network permission check. `:link` is a subproject resolved by path, so the app half is version-locked to the same commit the Rust half already was. What stays here is the two facts that are actually about this app, and both are load-bearing in a way that would fail quietly if got wrong: the `aiapp` URI scheme, and the Keystore alias `aiapp-token-key` that every enrolled phone's token is already sealed under. A wrong alias would leave those phones reading as not enrolled with nothing on screen to explain it, so the value is carried over exactly and the reason is written beside it. Call sites are unchanged. `ServerSettings` stays available unqualified as a typealias and `applyPinnedTls()` stays an extension, so the diff is the three files that bind the product-specific values plus two imports -- rather than every screen that happens to use a setting. Also clears a warning the build had been printing: `setup?.id.orEmpty()` where the compiler already knows `setup` is non-null, because `chosen` came from that setup's own provider list. Verified by running the build, not only by reading it: ktfmt, Kotlin compile and Android Lint are all clean with no warnings, and the APK still builds -- which exercises the pinned-CA generator, since that is the step that reads the CA off this machine. Still unpushed, per the hold until the rebuild bug is proven fixed. Note the submodule: a checkout of this commit needs `git submodule update --init` before `app/` or `server/` will build. |
||
|
|
c2dfaab349 |
Let Dev Updater supervise the backend instead of shipping a script
dev-updater now carries a built-in service implementation, generated from a template and driven through the identical interface a project-supplied script uses, so a project whose service is unremarkable no longer writes one. ai-app's was unremarkable: `ExecStart=$BINARY` and `command="$BINARY"` with no arguments and no environment. 233 lines of it, and the half that matters most -- the OpenRC branch, which neither project can exercise from a systemd machine -- existed twice, so a fix found by testing would have had two places to land and no way to notice the second. The field keeps its name; `Managed` takes the command, resolved against the component's `cwd`. The one thing the script said that the built-in cannot is kept, in AGENTS.md rather than lost: Stop on this card takes down the server a phone reaches through the tunnel, while Dev Updater itself is unaffected because it uses its own port -- which is exactly what makes that button easy to press and easy to regret. NOT SAFE TO PULL BLIND. A managed service is named after the component, so this one becomes `app-backend` while the installed one still has the name the script gave it. Uninstall from the backend card *before* taking this change, then Install after; pulling first orphans a service that stays enabled and starts at boot with nothing pointing at it. |
||
|
|
2c925a679f |
Take XDG resolution from the shared crate too
Sixth and last of the modules that were the link rather than this product. main.rs loses config_home, data_home and xdg_dir, and its test module with them -- it held one test, which moved to the crate that now holds the code. The helpers gained a `product` parameter, matching certs::ensure and netif::wg_address, which is what keeps two products' state apart while resolving it identically. Verified by running it: with only XDG_CONFIG_HOME and XDG_DATA_HOME set and no --config or --data-dir, the server puts its certificates in $XDG_CONFIG_HOME/ai-app/certs and its sessions in $XDG_DATA_HOME/ai-app/sessions, and still prints an aiapp:// enrollment URI. 35 tests here and 19 in the crate, clippy silent, rustfmt clean. |
||
|
|
aa05ff9336 |
Take the link from wg-app-link instead of keeping a second copy
The five modules underneath this backend that were never about AI sessions -- the pinned CA and leaf, QR enrollment and the bearer token, wg0 binding and the certificate's SANs, owner-only files, and the RON house rules -- were written twice, once here and once in dev-updater, and had drifted. They now come from the submodule, as a path dependency so both projects stay locked to one commit. What stayed is what makes this project itself: the routes, the drivers, the config schema, and the auth middleware, which is generic over this server's state. Sharing a transport is worth doing; sharing an API would mean inventing a vocabulary neither project wants. Four dependencies go with the code -- rcgen, qrcode, subtle and if-addrs are no longer named here at all -- and the three that remain are now described by what still uses them rather than by what used to. Verified by running it, not only by building: a fresh server generates its CA, prints an `aiapp://enroll` QR with the scheme now passed as a parameter, covers 127.0.0.1, 10.0.2.2 and wg0's 10.66.0.1 in the leaf, answers an enrolled token and returns 401 without one, and writes config.ron in the house rules with every file owner-only. 36 tests pass, clippy is silent, rustfmt is clean. |
||
|
|
a83dbcff6a |
Say when the server fell over, and where to read why
Iris found the backend crash-looping by checking rc-service by hand, because the card could only say `stopped` -- which reads as a state somebody chose. dev-updater's contract now has a fourth word, `failed`, and this script implements it. The OpenRC detail is the one worth not rederiving: it prints `crashed` *and* exits non-zero, so the word is read rather than the exit code. Leaning on the code would report "couldn't check", which is a different and less useful claim. The `running` check stays on its exit code, which already worked and does not depend on wording. The script also arranges the logging rather than only reporting it, because neither unit wrote a file: systemd went to the journal and OpenRC's `command_background=true` discarded output entirely, which is why a crash left nothing to read. Output now goes to $XDG_DATA_HOME/ai-server/ai-server.log -- generated data, outliving any one build, and not in a repository shared with a machine that should not read it. `start` rotates one generation aside, so what is kept is exactly this run and the one before: the pair worth having after a crash and a restart. `logs` prints the paths, newest first, and nothing else. Verified on systemd by causing the failure rather than reasoning about it: installed, started, confirmed `running` on the wg0 bind, wrote an unparseable config, restarted, and watched status settle on **failed** rather than stopped -- with the reason, line and column, in the file `logs` points at, and the crash preserved in .1 after recovery. Then restored, confirmed `running` again, and uninstalled. **The OpenRC branch is written from the documentation and is untested**, here and in dev-updater, since neither machine that can run it is one either of us can test on. It is also the branch that actually matters, since the backend runs under OpenRC on the host. `output_log`/`error_log` in the openrc-run script are the parts to distrust first. One thing that bit while writing it: the systemd heredoc is unquoted so $LOG expands, which makes a backtick in a comment inside it run as command substitution. A comment saying "`start` rotates" executed `start`, and the unit was written without ever being valid. There is now a note in the heredoc saying why it contains no backticks. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017xn8nHw1tw1R6PtiY1eEtw |
||
|
|
0c55b809b1 |
Drop the reader for the kebab-case driver kind
Iris has already moved past that spelling, so nothing will ever present it again -- there is one backend and one phone, and both are past it. The alias and the enum that carried it are gone; the legacy provider is just a `ProviderConfig` now. The rest of the migration stays until it has actually run on the host, because deleting it before then would strand the install it was written for. Its doc now says that outright, along with what to delete and when: this module and the branch in `Config::load` that reaches it, once the host has started on a build containing it. That is the general rule Iris gave, not a judgement about this migration: a reader for a superseded format has a defined end, because his population is one machine he controls, and leaving it keeps a second parsing path alive that nothing exercises and that constrains every later change to the schema. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017xn8nHw1tw1R6PtiY1eEtw |
||
|
|
5ccadffeaa |
Migrate an old config instead of refusing to start
The AI Sessions backend was crash-looping on the host, and I caused it. A config written before setups existed makes `Config::load` bail, the process exits 1 immediately, and under OpenRC's `command_background=true` that presents as a service that will not stay up. The refusal was deliberate and it was the wrong trade. I chose it to avoid silently emptying a config and re-seeding over it -- a real hazard -- but weighed it against the wrong cost. This process is how a phone reaches that machine at all, so refusing to run strands the person who would have to fix it, at a terminal, on the machine they were trying to avoid needing. And what it was protecting is the cheap half: providers and hosts are rediscoverable now, while the half that genuinely cannot be recovered -- the enrolled token hashes -- survives a migration untouched. So it migrates. Each old host becomes a setup keeping its name, since that is what sessions referenced; the top-level providers belong to the machine this server runs on; and every session's host becomes its setup, so conversations keep working. The original is copied to `config.ron.pre-setups` first, because this is a one-way conversion of the only record of what was configured and one file makes it reversible by hand. **Migrated hosts arrive with no providers, deliberately.** The old file never recorded which machine had which program -- that was the flaw the setups model exists to fix -- so inventing an answer would recreate exactly the impossible pairings it was meant to end. Rediscover asks the machine. Both driver-kind spellings are read. The kebab rename and the RON move landed on the same day, so a file written that morning says `r#claude-cli` and one from the afternoon says `claude_cli`; reading only one would have turned this fix into a different crash. Verified against a host-shaped config: the server starts, the token and both sessions survive, the remote session points at the migrated setup and the local one at `local`, the original is kept, and a second start is an ordinary load that neither migrates again nor overwrites the backup. Found by Iris, who had to check `rc-service` by hand because the card reported it as merely stopped -- dev-updater's session is adding a `failed` state for that separately. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017xn8nHw1tw1R6PtiY1eEtw |
||
|
|
4d162331e3 |
Say that llama.cpp reaches the phone now
The status section still said "server side" and "no app screen yet", both of which stopped being true today. Also records the discovery trade that would otherwise be rediscovered: command -v follows a non-interactive ssh PATH, so llama.cpp unpacked into ~/.local/opt is invisible until it is symlinked onto PATH. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017xn8nHw1tw1R6PtiY1eEtw |
||
|
|
6187958de3 |
Let a llama session actually be started from the phone
The driver worked and the models could be downloaded, but the spawn screen had no idea llama.cpp existed: the model field and every extra setting were gated behind `isClaude`, so a llama provider offered nothing, `model` arrived null, and the driver refused with "a llama.cpp session needs a model". The feature was reachable only by curl, which is not what was asked for. A llama provider now gets the models this backend has downloaded, as a picker rather than free text -- there is nothing sensible to type, and a name that is not on disk is a session that cannot start. Context size and temperature are there too, blank meaning llama.cpp's own default rather than a zero. Spawn stays disabled until a model is chosen, because without one the button could only fail. **Two bugs that only appeared by pressing the button**, both mine, both from changing the server without re-driving the app: - The app sent the setup's *label* where the server had started resolving by *id*. The failure was almost self-diagnosing -- `no setup named "this machine" -- configured: this machine` -- and that message now says "no setup with id" and lists ids, since listing labels was what made it read as a contradiction. - The session header showed `on local`, the id, because the app read `setup` where the server had begun sending both `setup` (id) and `setupName` (label). The app now carries only the label: nothing in it addresses a setup, and holding both is what let it show the wrong one. Verified by doing it: rediscovered the local machine from the phone so `local-llama` appeared, spawned a session on Qwen3-0.6B-Q8_0 with a 4096 context, sent "Reply with exactly one word: ready", and it replied "ready" with 125 tokens counted. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017xn8nHw1tw1R6PtiY1eEtw |
||
|
|
d0b6b66a44 |
Narrow a file this server rewrites, not only one it creates
`OpenOptions::mode` applies to a file the call creates and to nothing else, so rewriting a file that already existed kept whatever permissions it had. Three functions above, `create_dir` has carried a comment about exactly this hazard since it was written -- the file path never got the same treatment. This is not hypothetical here. `certs.rs` reissues the TLS leaf and rewrites its **private key on every start**, so a key that ever existed world-readable would have stayed that way for the rest of its life, with every subsequent start looking like it was setting the mode. The config's temp file is the other one: normally fresh, but a leftover from a crashed save would be reused with its old mode and then renamed over the real config, which holds the enrolled token hashes. Set through the open handle rather than the path, deliberately: `set_permissions` on a path re-resolves it, so between the open and the chmod something could put a different file -- or a symlink to one -- where this was, and the mode would land there instead. A handle cannot be redirected. Three tests, and the first was checked against the bug rather than only against the fix: with the new line commented out it fails with "rewriting left it at 644". Found by dev-updater's session, which had taken this module for a shared crate and read it as a unit. I had spotted the same line being wrong in their new `append_file` and missed that `create_file` -- the one I wrote -- had it too. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017xn8nHw1tw1R6PtiY1eEtw |
||
|
|
9f54a80ca4 |
Tell the difference between a blocked app and a missing server
Two findings from dev-updater's session reading this codebase, both checked against the code here before acting on them, and both real. **A denied local-network permission was invisible.** The manifest requests ACCESS_LOCAL_NETWORK and MainActivity asks for it, but nothing ever checked whether it was granted -- and on Android 17 a denial is indistinguishable from an unreachable server at the socket, because the OS simply drops the traffic. So every screen would have shown "is ai-server running, and is this device able to reach that address (WireGuard up)?", blaming two things that were both fine. Stated once at the root as a standing condition rather than appended to each failure it might have caused: it is not a property of any one request, and repeating it per error is how a message ends up saying the same thing twice, which this app has already done once today. **The Keystore read path was creating keys.** `unseal` called the get-or-create key function, so a sealed token whose key had been lost -- a device reset, or the app's data restored onto a device the key cannot travel to -- generated a fresh key, then failed to decrypt with it, leaving a key nothing had ever sealed with. The behaviour was already right by accident (it fails soft to "not enrolled"), but the read side now asks for the key without making one, which is what it meant all along. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017xn8nHw1tw1R6PtiY1eEtw |
||
|
|
3c144f8070 |
A screen for the machines, and failures a phone can act on
The other half of making setups editable: add, rename, rediscover and remove, with a Test that tries a machine before anything is saved. The screen cannot name a program, which is the point rather than an omission -- providers are what the server found when it asked, so this app has no way to introduce something to run. The dialog says so, because "what it can run is discovered, not typed" is the answer to the question a person will otherwise ask when they look for a command field. Two things running it changed. The card showed "this machine / this machine", because the seeded setup is *called* that and my fallback line for a local setup said the same -- the line now says something the name cannot also be. And the header row absorbed a fifth action without complaint, which is the earlier title-and-actions split paying off exactly as its comment predicted. **Host key verification is the failure that would have made this look broken.** Every machine fails it the first time, because its key is not in known_hosts yet, and ssh's own words -- "Host key verification failed." -- are written for somebody at a terminal on the backend, which is exactly who is not reading a phone. It now says what to do: ssh to it once from the backend and try again. Permission denied gets the same treatment. Deliberately *not* fixed by relaxing StrictHostKeyChecking. Accepting a new key is a decision somebody should make with the key in front of them, not something this app does quietly on their behalf while adding a machine. Verified on the emulator against a running server: the seeded setup renders with what was discovered on it, the add dialog explains itself, and Test against an untrusted machine produces the full explanation rather than ssh's four words. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017xn8nHw1tw1R6PtiY1eEtw |
||
|
|
19e3531c5d |
Add, rename and remove machines from the phone -- without letting it name commands
The gap PLAN.md recorded: setups were readable but only hand-editable, so
adding a machine meant a shell on the backend.
**The design decision, made with Bryan, is that the phone never composes a
command.** A setup carries providers, and a provider carries something to
run -- so a route that accepted a command from the request body would make
the enrolled token arbitrary code execution on every machine a setup names,
and the transport already reaches those over ssh. Instead the phone sends
connection details, and the server asks the machine itself what it has:
one `command -v` round trip per setup, matched against a table of the
drivers this server knows. The phone's authority is "add this machine",
never "run this".
Worth recording that this was a narrower change than it first appeared: the
token could already run anything on the backend, because the spawn screen
offers `bypassPermissions` with a free-text working directory. Discovery
does not close that door. What it does is keep the *list of what can run*
out of the phone's reach, and make adding a machine a thing you cannot get
wrong by typing.
It is also simply better to use. Nobody wants to type an absolute path on a
phone keyboard, and a machine whose binaries have moved answers correctly
on the next probe. The cost is that a program somewhere unusual is
invisible -- `command -v` follows PATH under a non-interactive ssh session,
which is not the PATH a person sees when they log in. That is the trade,
and the escape hatch is editing config.ron on the backend, which is exactly
the authority the phone is not being given.
Setups now have an **id separate from their label**, so renaming a machine
does not orphan the sessions that name it; a session stores the id, and
every row resolves the current label when it is built. `POST /setups/probe`
tries a machine without saving anything, so a wrong address or an
unauthorised key is caught while the form that caused it is still on
screen. Deleting is refused while sessions still run there, and says which
ones rather than cascading.
Every mutation goes through one `update`: clone, apply, save, then commit,
so a failed write leaves the previous state intact and reports why.
Verified against a running server, including a real ssh machine (this VM,
via a throwaway loopback key since removed): probing here found echo and
claude-cli; probing over ssh found claude-cli and correctly no echo, which
runs in-process and exists only where this server does; an unreachable
machine came back with ssh's own words ("connect to host ... Connection
timed out"); adding derived the id `loopback-vm` from "loopback vm";
renaming kept the id; deleting was refused while a session used it, naming
it, and succeeded once nothing did.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017xn8nHw1tw1R6PtiY1eEtw
|
||
|
|
ef019a7aea |
Mark the setups model as built, and say what is left
The three questions the plan left open are answered by having built it: where echo lives (seeded into the local setup, not implicit), what happens to an old config (refused with instructions, because the silent version loses everything), and that editing setups from the phone is still missing. That last one is the honest gap: GET /setups exists, writing them does not, so adding a machine is still a hand edit on the backend. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017xn8nHw1tw1R6PtiY1eEtw |
||
|
|
ecac404fd4 |
A setup is a machine, and it carries what that machine can run
Providers and hosts were two independent lists, and a session named one of each. They were never independent: a provider only exists on a machine where that program is installed, so the spawn screen offered the whole cross-product, including "the Claude CLI on the box that hasn't got it". The picker could not know, because nothing in the model said. Now a setup is a machine -- optional ssh, plus the providers it has -- and spawning is two choices in order: pick a setup, then one of its providers. The impossible pairs stop being expressible rather than being validated against. Provider names are unique within a setup and only within one, so two machines can each have a `claude-cli`, which was previously either a name collision or two entries called things like "claude" and "claude on the vm". It also settles the "Run on" problem properly. That control was offered for every provider but honoured only by the Claude driver -- an echo session sent to a host ran locally and said otherwise. There is no such control now: the machine is chosen first, and echo is a provider of the setup with no ssh, where it belongs, since it runs in-process and has no transport to cross. The built-in echo provider is gone as a concept. It used to be conjured at read time and never written to the file, which meant a provider nobody could see or edit; it is now seeded into the config on first run alongside claude-cli. What the file says is what there is, and deleting it is a choice rather than a state to be repaired. A config in the old shape is refused with instructions rather than loaded. `Config` defaults unknown fields away, so `providers:` and `hosts:` would otherwise have vanished into an empty config that was then seeded over -- a migration nobody would notice until their setups were gone. Verified against a running server and on the emulator: a fresh install seeds "this machine" with echo and claude-cli and the file reads cleanly; a two-setup config lists both with their own providers; spawning on a setup works and the session row names it; asking for a provider a setup lacks says which it offers, and an unknown setup says which exist. On the phone, selecting "dev vm" narrows the provider chips to that machine's one and shows its address. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017xn8nHw1tw1R6PtiY1eEtw |
||
|
|
7cf36005ae |
Browse, download and delete models from the phone
The point of the llama.cpp work was that models are managed from the app, not by editing the backend's filesystem, so this is the screen for it: search HuggingFace, expand a repository to see its GGUFs with sizes, download one and watch it, cancel it, delete what is no longer wanted. Everything shown is the server's state rather than the screen's. A download started here keeps going when the screen closes, is visible from any enrolled device, and its outcome outlives it -- demonstrated by accident while testing, when a 538 MB download finished during an app rebuild and was still there, complete, after reinstalling. Polled rather than streamed, at 1.5s. A download belongs to the machine rather than to any session, so it has no event stream of its own; this is the one screen in the app that asks repeatedly instead of being told. Three things the screenshots decided rather than the diff: - **The list header no longer squeezes its title.** Adding a fourth action to the row wrapped "AI Sessions" onto three lines. Title and actions now have a row each, so a fifth costs nothing and the title is never what gives. - **A repository's files render inside its own card**, not as a section after the list -- drawn after every card they read as belonging to whichever was last. - **A file already downloading says so** and is disabled, rather than offering a Download button whose effect nobody can see. The progress bar is determinate only when the server reported a size, and says "total size unknown" otherwise rather than inventing a position. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017xn8nHw1tw1R6PtiY1eEtw |
||
|
|
3be25c2f64 |
Write down what the llama.cpp work actually does
Phase 4 is no longer deferred and phase 5 is exercised, so the status section says so. The two decisions worth not undoing by accident get named: the conversation lives in the transcript rather than the driver, and a llama session is refused on an ssh host rather than half-working. Also the local testing recipe, including the trap that cost me twenty minutes -- a 2-bit quant produces fluent nonsense that reads exactly like a broken driver, and llama-cli on the same file is how to tell the two apart. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017xn8nHw1tw1R6PtiY1eEtw |
||
|
|
3deeffd1e7 |
Run GGUF models through llama-server, and stop orphaning them
The second half of the llama.cpp work: a session can now name a downloaded model and talk to it. `llama-server` is spawned through the same transport as any other driver, polled until the model is loaded, then driven over its OpenAI-compatible streaming endpoint and translated into the same events the Claude driver emits -- so the transcript, the SSE stream and the phone need to know nothing new. **The conversation is rebuilt from the transcript, not held in the driver.** llama-server is stateless between requests, so the whole history goes with every one, and the obvious place to keep it is a Vec in the driver. That fails the requirement: memory in a driver is invisible to a second device and gone on restart, and this app is meant to work across devices. Reading it back also means the model is prompted with exactly what the phone was shown -- including a reply that was interrupted half way, which is in the transcript because the deltas were already emitted. That leaves the Claude driver as the odd one out rather than this one: the CLI's memory of a conversation is a cache in front of the same transcript, not a second truth. Said so at the top of llama.rs, because it is the sort of inconsistency that gets "fixed" in the wrong direction. Session settings arrive as a driver-interpreted `params` map rather than new typed fields, so the shared schema does not grow one dialect's vocabulary. Context size, gpu layers and threads become server flags; temperature and the rest ride on each request, so changing them need not reload a model. **Also fixes an orphan this feature would have created.** Drivers set kill_on_drop, which covers a session being deleted -- but nothing drops on the way out of a SIGTERM, so signalling the server left its children running. For the Claude CLI that is untidy; for a llama-server holding a model it is gigabytes belonging to nobody. The server now stops its sessions on SIGTERM and SIGINT. Found by killing a test server and noticing two 600 MB processes still resident. Remote llama sessions are refused rather than half-working: the model is reached over HTTP, and forwarding that port to an ssh host is the "reach this port" operation the transport does not have yet. Verified end to end against a real model: downloaded Qwen3-0.6B Q8_0 through the app's own download route, spawned a session on it, and held a two-turn conversation -- "my favourite colour is teal" then "what is my favourite colour?", answered "teal", which is the transcript replay doing its job. Token counts arrive. An earlier attempt with the IQ2_XXS quant produced fluent nonsense, which turned out to be the quantisation rather than the pipeline: llama-cli produces the same from that file directly. Four unit tests cover the fold and the path guard. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017xn8nHw1tw1R6PtiY1eEtw |
||
|
|
e50d1a2bbf |
Take the resumed total from Content-Range, not Content-Length
On a 206 those two headers answer different questions: Content-Length is the length of the range, so a resume at 162 MB reports 72 MB and a bar drawn from it fills at a third of the model. The arithmetic that was here (`have + length`) happened to be right, but only because the range always starts exactly at what is on disk -- it was correct by coincidence of two things agreeing rather than by asking for the number wanted. Content-Range carries the whole size as its last field and does not care where the range began. Confirmed against HuggingFace: `content-range: bytes 162000000-234074815/234074816` beside `content-length: 72074816`, and a resumed download now reports 234.1 MB rather than 72. Two other things checked rather than assumed, both fine as they stood. Downloads are already single-flight per file -- the check and the insert happen under one lock, keyed by the model, so a second client asking for the same file joins the running download instead of starting a second writer onto the same partial. And HuggingFace's ETag is stable across requests, with no weak prefix or per-edge variation, so the identity check will not discard good partials and refetch gigabytes for nothing. One hypothesis worth recording as false: HF's ETag is not the content sha256 for these files (`db6593d0…` against a published `55e0d0b8…`), so the published-hash check cannot collapse into the identity check. Both earn their place. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017xn8nHw1tw1R6PtiY1eEtw |
||
|
|
6f149398d0 |
Don't resume onto a partial from a different revision
A resume splices: it appends bytes from wherever the server is now onto whatever is already on disk. If the file changed upstream in between, the result is exactly the failure that survives every cheap check -- the right number of bytes, the wrong contents, and no error anywhere. HuggingFace files do get updated, so this is a real path rather than a theoretical one. `If-Range` is the header for this and would have been the tidy answer, but HuggingFace's CDN ignores it: probed today, a deliberately stale validator still answers 206 with the ranged bytes rather than 200 with the whole body. So the check is done here instead. A partial now has an identity file beside it holding the ETag it was written against, written before the body so an interrupted download still knows what it is a piece of. On resume, the response's ETag is compared against it, and a mismatch throws the partial away and asks again from zero. A partial with no identity at all is not resumed either -- it could be a fragment of anything. The sha256 HuggingFace publishes is now also checked before the file gets its real name, so a bad one is never offered to be run. That is belt-and-braces after the above rather than the primary defence, which is the right order: detecting corruption after downloading gigabytes is worth far less than not creating it. Verified by planting one: a 60 MB partial of random bytes with an identity file naming a revision that does not exist. The server logged "changed upstream since the partial was written -- starting again", restarted from zero rather than appending, and the finished file's sha256 matches the published one. Repeated the honest resume too -- cancel at 145 MB, restart, resume at 162 MB, correct hash. Thanks to dev-updater's session for the If-Range idea and for saying to confirm the CDN honours it rather than assume, which is exactly what it turned out not to do. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017xn8nHw1tw1R6PtiY1eEtw |
||
|
|
9d29776f02 |
Download GGUF models from HuggingFace, watchably
The first half of the llama.cpp work Bryan asked for: browse HuggingFace, fetch a model, and see how far it has got from any device. The design is dev-updater's build-progress shape with the four changes its author recommended after living with it, since a model download is an hour where a build is two minutes: - **A run has an id.** Without one "not downloading" means three different things -- finished, never started, or someone else's run ended while you were away -- and over an hour that ambiguity is certain rather than theoretical. A device compares the run it was watching to the run reported now. - **Outcomes outlive their run**, so a phone that was asleep at the moment of completion can still find out what happened. - **Cancel exists.** Retrofitting cancellation into a blocking loop is miserable, and several gigabytes over someone's data plan is not something to have no answer for. - **Progress is bytes, not a parsed marker.** We own the loop, so it counts directly; `total` is whatever Content-Length said and nothing else, and stays absent when the server sends none rather than becoming a bar drawn from a guess. The download owns its own thread rather than the blocking pool, which exists for short work. It resumes through HTTP Range, and trusts the 206 rather than the request -- a server that ignores Range answers 200 with the whole file, and appending to that would corrupt it. `truncate(false)` on the open is load-bearing for the same reason and says so. Searching is proxied through the server rather than done from the phone, because the app trusts exactly one certificate -- this one -- and the machine that must do the downloading is also the one whose view of what exists matters. Verified against the real HuggingFace, not a mock: searched, listed a repository's GGUFs, downloaded 234 MB with live byte progress, cancelled mid-flight, confirmed the partial survived, restarted and watched it resume at 162 MB rather than 0, and let it finish. The result's sha256 matches the one HuggingFace publishes for that file, so the resume is byte-correct and not merely the right length. llama.cpp then loaded it and ran inference. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017xn8nHw1tw1R6PtiY1eEtw |
||
|
|
4004183acf |
Answer the "how do we test SSH here" question by testing it
The status section had carried an open question since phase 3: SSH could not be exercised in this VM because there is no second machine and no key in ~/.ssh. There is a second machine, though -- this one. Ssh it to itself with a throwaway key and a host of bob@127.0.0.1, point the provider's command at /bin/echo rather than claude, and the whole path runs: connection, remote exec, and the process's death arriving as `status: exited` in the transcript. It costs no tokens and touches nothing real, and the key comes back out afterwards. Done that way just now against the new transport, so the technique is written down as something that worked rather than something that should. Also recorded: the login shell in this VM is fish. `cd '…' && exec '…'` is valid there and the POSIX single-quote escaping happens to mean the same thing, but both are luck, and a non-POSIX remote shell is the first thing to suspect if an argument is ever mangled on the way over. Phase 4 is no longer deferred -- Bryan asked for llama.cpp today -- so the status paragraph stops saying it is. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017xn8nHw1tw1R6PtiY1eEtw |
||
|
|
4cdcbd204a |
Put the transport above the drivers instead of inside one
ClaudeDriver::spawn called ssh::command itself, so a translator whose job is a wire format also knew how sessions reach other machines, and every future driver would have had to remember the same. It now emits a `Launch` -- program, arguments, working directory -- and hands it to a `Transport` the manager chose from the session's host. This is the inversion Bryan asked for, and it pays for itself immediately in a place I had reported as a UI bug: "Run on" is offered for every provider but only the Claude driver honoured it, so choosing a host for an echo session silently ran it locally. With the transport above the driver that cannot be written -- EchoDriver builds no Launch, so there is nothing to wrap and nothing to misreport. The picker still needs to stop offering it, but the code no longer lies underneath. crate::ssh keeps the quoting, the forced options and the remote script, with its tests; transport.rs only decides which of the two it is. The two failure messages move with it, since they are transport-specific -- a missing ssh client here is a different thing to check than a program missing from a remote PATH. Noted in transport.rs rather than built, because nothing needs it yet: a remote llama-server is spawned as a process but spoken to over HTTP, so a transport eventually needs "reach this port" as well as "run this". Verified: cargo test (35), clippy, fmt. Nothing outside transport.rs mentions ssh now. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017xn8nHw1tw1R6PtiY1eEtw |
||
|
|
ba6c770be3 |
Record the setups model, and put the transport above the drivers
Two decisions from Bryan today, written down before they are built, since this file is where the reasoning is supposed to live rather than in a conversation. A setup is a machine carrying the providers that machine has, and spawning picks a setup then one of its providers. The independent providers × hosts model it replaces is left in place below it, because its reasoning is worth keeping and the code still implements it. What that model got wrong is that the axes are not independent: it offers combinations that cannot work, and "Run on" is already a control that does nothing for the echo driver, which takes no host at all. The transport wraps the driver rather than the driver reaching for the transport. ClaudeDriver::spawn calls ssh::command itself today, which puts transport knowledge inside a translator whose job is a wire format, and obliges every future driver to remember the same. Inverted, a driver that emits no command has nothing to wrap, which is the same "Run on" problem solved structurally instead of by a special case. Also noted: llama.cpp is the case where "wrap a command" is not enough on its own, since a managed llama-server is spawned but then spoken to over HTTP, so a transport is "run this" plus "reach this port". And this section's claim that remote attachments need scp was never true of the code -- images are base64 inside the stream-json message in both directions, so nothing has to exist on the remote filesystem. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017xn8nHw1tw1R6PtiY1eEtw |
||
|
|
8bf99f4b76 |
Write down the UI polish noticed while cleaning up
One item so far: the session header's status sits tight against the right edge while Back looks roomier, because the row's 8dp padding is measured against a TextButton whose touch target is wider than its text. Not a defect and not urgent, but the kind of thing that gets re-noticed and re-diagnosed every few months unless it is written down once. PLAN.md rather than a new file, since that is where this project's decisions live, and this is a decision to defer rather than a bug to track. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017xn8nHw1tw1R6PtiY1eEtw |
||
|
|
2cefede450 |
Compose Multiplatform 1.12.0
The last four lint warnings were all this one thing, and they were right: 1.11.1 with 1.12.0 out. Nothing else here is behind -- material3 1.9.0 and activity-compose 1.13.0 are both current, checked against Maven Central and Google Maven today, which is also why the header comment's date moves. Android Lint now reports no issues at all, from 11 errors and 8 warnings this morning. Verified by running it rather than by the build succeeding, since a Compose minor can change how things draw: the session list, the usage screen and a session transcript with a message sent through it all render as before. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017xn8nHw1tw1R6PtiY1eEtw |