Files
dev-updater/server/src/service.rs
T
irisandClaude Opus 5 b67c70fa2f A component can name the files its build reads elsewhere
Freshness compared the commits under a component's own cwd, which is
right about where its files are and wrong about what its build reads.
tdep-survey's two clients each source one shared scripts/android-sdk-env.sh:
a fix committed there changed what every build does, moved no component's
subtree head, rebuilt nothing, and left every card correctly reporting
"current" while answering a narrower question than the reader was asking.

alsoWatch names the rest. The paths join the component's own in the same
subtree-head and dirty comparisons -- git takes several pathspecs, so it
stays one call each -- and it is part of the acceptance gate like any
other declared field. Declared rather than inferred: which files a build
reads is not knowable from here, and both wrong guesses are expensive.

Raised by the tdep-survey session, which traced why its dioxus client
could never become stale on the serving host.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-01 11:11:07 -04:00

543 lines
21 KiB
Rust

//! Driving a project's long-running server through the script it carries.
//!
//! This server knows nothing about systemd or OpenRC and deliberately
//! never will: which init system is present, and how a unit gets written
//! into it, is knowledge that belongs where the service does. A project
//! declares one script and this runs `<script> <subcommand>` -- `install`,
//! `uninstall`, `start`, `stop`, `restart`, `status`, `logs`.
//!
//! `status` is the one with a contract, because five answers have to be
//! told apart and three of them are not failures: see [`ServiceState`].
//!
//! Nothing here ever runs on the manifest path. Asking a service manager
//! costs a process spawn, and the manifest is fetched on every open,
//! resume and Refresh -- so the answer is fetched in the background and
//! read from [`ServiceChecks`], exactly as `git::RemoteChecks` does for
//! remotes.
use std::os::unix::process::CommandExt;
use std::path::{Path, PathBuf};
use std::sync::Arc;
use serde::Serialize;
use crate::config::{Command, Component, Service};
/// What a service script says about its service.
///
/// Four states rather than a boolean, because each is a different thing
/// to offer and a different thing to say. "Not installed" gets an Install
/// button and no controls; "stopped" gets Start; "failed" gets Start too,
/// but must not be *described* as stopped.
///
/// `Failed` exists because its absence made the card lie. A service that
/// fell over is not stopped -- stopped is a state somebody chose, and
/// reading "stopped" about a crash sends you looking for who stopped it.
/// The only alternative a script had was to exit non-zero, which reads as
/// "couldn't check" and is equally wrong: it found out perfectly well, it
/// just had no word for the answer.
///
/// A fifth answer, "the script could not tell us", is still the `Err`
/// case rather than a variant: that one is not a state the service is in,
/// it is this server failing to find out.
#[derive(Clone, Copy, PartialEq, Eq, Debug, Serialize)]
#[serde(rename_all = "camelCase")]
pub enum ServiceState {
Running,
Stopped,
/// Installed, not running, and not because anybody asked -- it exited
/// non-zero or was killed.
Failed,
NotInstalled,
}
impl ServiceState {
/// The word a script prints on stdout. Deliberately words rather than
/// exit codes: a script that returns 3 has to be read against a table
/// nobody remembers, and the exit status is wanted for the separate
/// question of whether the script worked at all.
fn parse(output: &str) -> Option<Self> {
match output.trim() {
"running" => Some(Self::Running),
"stopped" => Some(Self::Stopped),
"failed" => Some(Self::Failed),
"not-installed" => Some(Self::NotInstalled),
_ => None,
}
}
}
/// The command that drives this component's service, whichever way it
/// declared one.
///
/// The single place the two variants of [`Service`] become the same thing.
/// Everything downstream takes a command and runs `<command> <subcommand>`,
/// so nothing but this function knows that a built-in script exists -- and
/// a component that switches from its own script to the built-in one
/// changes nothing anywhere else.
///
/// `None` covers both a component that is not a server and one that
/// declares no service, which are the same thing to every caller: there
/// is nothing to drive.
pub fn driver(key: &str, component: &Component) -> Option<Command> {
match component.service()? {
Service::Script(script) if !script.is_empty() => Some(script.clone()),
Service::Managed(run) if !run.is_empty() => Some(Command::from_words(vec![
crate::shipped::service_default()
.to_string_lossy()
.into_owned(),
"--name".to_string(),
unit_name(key, component.name()),
"--exec".to_string(),
run.to_line(),
])),
// Declared empty, which says the same as not declaring it.
Service::Script(_) | Service::Managed(_) => None,
}
}
/// What a managed service is called to its service manager.
///
/// The project key and the component, because a service manager's names
/// are one flat namespace across every project on the machine and
/// "backend" is not a name two projects can share. A project that wants to
/// choose its own name carries its own script, which is one of the things
/// that is for.
fn unit_name(key: &str, component: &str) -> String {
format!("{key}-{component}")
}
/// How long a service command is given before it is given up on.
///
/// Same reasoning as the git remote timeout: a script that hangs -- on a
/// password prompt it should never have shown, most likely -- must not
/// hold a slot forever. Longer than a status check needs, because
/// `install` may be writing units and enabling them.
const TIMEOUT: std::time::Duration = std::time::Duration::from_secs(60);
/// Runs one subcommand of `script`, returning its stdout.
///
/// Stdin is closed rather than inherited. A script that prompts for a
/// password gets end-of-file and fails, instead of hanging until the
/// timeout with a card stuck on "installing" -- the same bargain git's
/// `BatchMode` makes.
pub fn run(
script: &Command,
project: &Path,
cwd: Option<&Path>,
subcommand: &str,
) -> Result<String, String> {
let mut child = script
.to_process(project, cwd, &[subcommand])?
.stdin(std::process::Stdio::null())
.stdout(std::process::Stdio::piped())
.stderr(std::process::Stdio::piped())
.spawn()
.map_err(|err| format!("failed to run the service script: {err}"))?;
let deadline = std::time::Instant::now() + TIMEOUT;
loop {
match child.try_wait() {
Ok(Some(_)) => {
let output = child
.wait_with_output()
.map_err(|err| format!("reading the service script's output: {err}"))?;
return if output.status.success() {
Ok(String::from_utf8_lossy(&output.stdout).trim().to_string())
} else {
// The script's own words, which is what the card shows
// -- it knows why it could not do the thing and this
// does not.
let stderr = String::from_utf8_lossy(&output.stderr).trim().to_string();
Err(if stderr.is_empty() {
format!("{subcommand} failed ({})", output.status)
} else {
first_line(&stderr)
})
};
}
Ok(None) if std::time::Instant::now() >= deadline => {
let _ = child.kill();
let _ = child.wait();
return Err(format!(
"{subcommand} took longer than {}s and was stopped -- a service script must \
never wait for input",
TIMEOUT.as_secs()
));
}
Ok(None) => std::thread::sleep(std::time::Duration::from_millis(50)),
Err(err) => return Err(format!("waiting on the service script: {err}")),
}
}
}
/// Runs one subcommand and does **not** wait for it, in its own process
/// group.
///
/// For the one case where the caller is the target: this server asking its
/// own manager to restart it. Both halves matter.
///
/// *Not waiting*, because the thing being waited for is a restart of this
/// process -- there is no result to collect, and blocking a thread on it
/// only creates something for the stop to interrupt.
///
/// *Its own process group*, because on OpenRC `restart` is a shell script
/// doing stop-then-start, run as a child of the very process it is
/// stopping. Waiting on it deadlocks: this server's shutdown waits for its
/// children, and the restart's stop phase waits for this server to exit.
/// Neither moves, `start-stop-daemon` gives up with "1 process refused to
/// stop", and the restart aborts *before* the start ever runs -- which is
/// what left one service parked in `stopping` with the phone unable to
/// reach anything.
///
/// Reproduced on OpenRC 0.63.3, and the reproduction is worth knowing:
/// **it only bites a server that shuts down gracefully.** A test process
/// that dies instantly on SIGTERM passes -- the restart child is orphaned
/// and finishes the job -- so the toy version of this test says everything
/// is fine. Anything that drains its work first, which is every real
/// server, deadlocks.
///
/// systemd does not have the problem at all, because there a restart is a
/// job the daemon owns and the client is free to die. That difference is
/// precisely why this was not caught here.
pub fn spawn_detached(
script: &Command,
project: &Path,
cwd: Option<&Path>,
subcommand: &str,
) -> Result<(), String> {
script
.to_process(project, cwd, &[subcommand])?
.stdin(std::process::Stdio::null())
// Inherited, so whatever the script says about a failed restart
// lands in this server's log -- the one place it can still be read
// afterwards.
.process_group(0)
.spawn()
.map(|_| ())
.map_err(|err| format!("failed to run the service script: {err}"))
}
/// Asks `script` what state its service is in.
pub fn status(
script: &Command,
project: &Path,
cwd: Option<&Path>,
) -> Result<ServiceState, String> {
let output = run(script, project, cwd, "status")?;
ServiceState::parse(&output).ok_or_else(|| {
format!("status printed {output:?}, not running, stopped, failed or not-installed")
})
}
/// Where a component's own log files are, newest first.
///
/// The script both arranges the logging and reports it: neither service
/// manager writes a file by default -- systemd goes to the journal,
/// OpenRC's backgrounded output goes nowhere -- so a unit has to be
/// written to redirect, and only the script knows how. That is the whole
/// point of asking it: this server needs no service-manager-specific code
/// at all, it just gets paths.
///
/// **Not supporting logs is a first-class answer.** A script that prints
/// nothing, or exits non-zero, means "no logs from me" and the card
/// simply offers no button. An older script that has never heard of
/// `logs` falls through to its usage case and exits non-zero, which lands
/// in the same place -- so this is additive and no script has to change
/// before the server does.
///
/// The paths are reported rather than the contents. What to do with a log
/// -- how much of it, which generation -- is this server's business and a
/// phone's, not the script's.
pub fn logs(script: &Command, project: &Path, cwd: Option<&Path>) -> Vec<PathBuf> {
let Ok(output) = run(script, project, cwd, "logs") else {
return Vec::new();
};
output
.lines()
.map(str::trim)
.filter(|line| !line.is_empty())
.map(|line| {
// Relative paths resolve against the project, the same rule
// the script's own command follows.
let path = Path::new(line);
if path.is_absolute() {
path.to_path_buf()
} else {
project.join(path)
}
})
.collect()
}
/// A script's failures run to a paragraph; a card gets a line. The rest is
/// in this server's log.
fn first_line(text: &str) -> String {
text.lines()
.map(str::trim)
.find(|line| !line.is_empty())
.unwrap_or(text)
.to_string()
}
/// What each server component was last found to be doing, refreshed off
/// the request path.
///
/// The typed face of [`crate::checks`], which is where the mechanism
/// lives: the list must not wait on a process spawn, so the answer lands
/// on the next look and the card says it is still being worked out until
/// it does. Keyed by project key and component name, because the actions
/// are per component.
#[derive(Default, Clone)]
pub struct ServiceChecks(Arc<crate::checks::Checks<(String, String), Answer>>);
/// What is known about one component, as one value.
///
/// Both parts come from the same trip, because both cost the same process
/// spawn and the manifest must not pay it per request. They are kept
/// together rather than as two checks so that a status that failed cannot
/// leave the logs looking like they belong to a different moment.
#[derive(Default, Clone)]
pub struct Answer {
state: Option<ServiceState>,
/// The log files this component's script reported, newest first. The
/// card only needs to know *whether* there are logs; reading one is an
/// explicit tap and can pay for its own ask.
logs: Vec<PathBuf>,
}
/// One component to ask about: which project it belongs to, where that
/// project is, and the script.
pub struct Target {
pub key: String,
pub component: String,
pub project: PathBuf,
pub cwd: Option<PathBuf>,
pub script: Command,
}
impl ServiceChecks {
pub fn state(&self, key: &str, component: &str) -> Option<ServiceState> {
self.0.answer(&id(key, component)).and_then(|it| it.state)
}
pub fn is_checking(&self, key: &str, component: &str) -> bool {
self.0.is_checking(&id(key, component))
}
/// The log files last reported for this component. Empty before
/// anything has been asked, and for a script that offers none.
pub fn logs(&self, key: &str, component: &str) -> Vec<PathBuf> {
self.0
.answer(&id(key, component))
.map(|it| it.logs)
.unwrap_or_default()
}
pub fn error(&self, key: &str, component: &str) -> Option<String> {
self.0.error(&id(key, component))
}
/// Asks every target that isn't already being asked. Returns at once.
pub fn refresh(&self, targets: Vec<Target>) {
for target in targets {
let id = id(&target.key, &target.component);
self.0.start(id, move |previous| {
let state = status(&target.script, &target.project, target.cwd.as_deref());
// Same trip, because it is the same spawn cost.
let logs = logs(&target.script, &target.project, target.cwd.as_deref());
if let Err(err) = &state {
tracing::warn!("asking {}'s {} failed: {err}", target.key, target.component);
}
// A failed status keeps the previous one and records why,
// so the card can qualify what it shows rather than
// passing a stale state off as current -- while the logs,
// which were found, are kept either way.
crate::checks::Report {
answer: Some(Answer {
state: state
.as_ref()
.ok()
.copied()
.or_else(|| previous.and_then(|previous: Answer| previous.state)),
logs,
}),
error: state.err(),
}
});
}
}
/// Records a state this server just caused, so the card reflects an
/// action without waiting for the next background check. Leaves the
/// logs alone: they are still the ones that were found.
pub fn mark(&self, key: &str, component: &str, state: ServiceState) {
self.0
.update(id(key, component), |answer| answer.state = Some(state));
}
/// Forgets a project's components, for one being removed.
pub fn forget(&self, key: &str) {
self.0.retain(|(project, _)| project != key);
}
}
fn id(key: &str, component: &str) -> (String, String) {
(key.to_string(), component.to_string())
}
#[cfg(test)]
mod tests {
use std::os::unix::fs::PermissionsExt;
use super::*;
fn server(service: Option<Service>) -> Component {
Component::Server {
name: "backend".to_string(),
build: Command::default(),
cwd: None,
stale_when: None,
also_watch: Vec::new(),
service,
built_from: None,
}
}
/// The one place the two variants become the same thing, so this is
/// where it is worth pinning down what each turns into.
#[test]
fn both_variants_resolve_to_one_command_to_run() {
// A project's own script is passed through untouched -- it is
// already the thing the contract describes.
let own = Command::from_line("server/service");
assert_eq!(
driver("app", &server(Some(Service::Script(own.clone())))),
Some(own)
);
// Managed becomes the built-in script with the component's
// identity and command as arguments, so the subcommand the caller
// appends still lands last.
let managed = driver(
"app",
&server(Some(Service::Managed(Command::from_line(
"target/release/ai-server --port 8080",
)))),
)
.expect("a managed component has a driver");
let (program, arguments) = managed.split_first().expect("a program");
assert_eq!(
program,
&crate::shipped::service_default()
.to_string_lossy()
.into_owned()
);
assert_eq!(
arguments,
[
"--name",
// The key as well as the component: a service manager's
// names are one namespace across every project here.
"app-backend",
"--exec",
"target/release/ai-server --port 8080",
]
);
}
/// A server with no service, and one whose declaration is empty, are
/// the same answer to every caller: there is nothing to drive.
#[test]
fn nothing_to_drive_is_none_however_it_was_said() {
assert_eq!(driver("app", &server(None)), None);
assert_eq!(
driver("app", &server(Some(Service::Script(Command::default())))),
None
);
assert_eq!(
driver("app", &server(Some(Service::Managed(Command::default())))),
None
);
}
/// The property the OpenRC failure turned on: the script this server
/// asks to restart it must not be something this server then waits on,
/// or the two wait for each other and the start never happens. Its own
/// process group is what makes that structurally true rather than
/// remembered.
///
/// Tested by having the child report its own group, because that is
/// the thing that decides whether a group-directed signal reaches it.
/// Nothing here can test OpenRC itself -- this machine has systemd --
/// so this pins the mechanism rather than the outcome.
#[test]
fn a_detached_child_is_outside_this_process_group() {
let dir = tempfile::tempdir().expect("tempdir");
let script = dir.path().join("report");
std::fs::write(
&script,
"#!/bin/sh\nps -o pgid= -p $$ | tr -d ' ' > \"$(dirname \"$0\")/pgid\"\n",
)
.expect("write");
std::fs::set_permissions(&script, std::fs::Permissions::from_mode(0o700)).expect("chmod");
spawn_detached(
&Command::from_line(&script.to_string_lossy()),
dir.path(),
None,
"restart",
)
.expect("spawn");
let reported = dir.path().join("pgid");
let mut waited = 0;
while !reported.is_file() && waited < 100 {
std::thread::sleep(std::time::Duration::from_millis(50));
waited += 1;
}
let child_group: i32 = std::fs::read_to_string(&reported)
.expect("the detached child should have reported its group")
.trim()
.parse()
.expect("a process group id");
// Read the same way the child reported its own, so the two
// numbers are comparable and nothing new is depended on for it.
let ours: i32 = String::from_utf8_lossy(
&std::process::Command::new("ps")
.args(["-o", "pgid=", "-p", &std::process::id().to_string()])
.output()
.expect("ps")
.stdout,
)
.trim()
.parse()
.expect("our own process group id");
assert_ne!(
child_group, ours,
"a child in our own group is one the stop would kill with us"
);
}
/// The three words and nothing else. A script printing something else
/// has a bug, and saying so beats picking whichever state is nearest.
#[test]
fn only_the_three_words_are_states() {
assert_eq!(ServiceState::parse("running"), Some(ServiceState::Running));
assert_eq!(
ServiceState::parse(" stopped\n"),
Some(ServiceState::Stopped)
);
assert_eq!(ServiceState::parse("failed"), Some(ServiceState::Failed));
assert_eq!(
ServiceState::parse("not-installed"),
Some(ServiceState::NotInstalled)
);
assert_eq!(ServiceState::parse("active"), None);
assert_eq!(ServiceState::parse("Running"), None);
assert_eq!(ServiceState::parse(""), None);
}
}