A llama.cpp session had its own `llama-server`: two sessions on one model held two copies of it in memory, a model change bought a load only that session benefited from, and the process was a session's to end. A machine's models are now served by one `llama-server` in **router mode** -- no `-m`, a preset file naming models and their flags, a child server per model asked for, and each request routed by its `model` field. So one server per model with that model's own settings is what a machine runs, while this backend has one process, one port and one record per machine to keep track of. The record is the mechanism every other driver already uses, so a restart adopts it; a session records the same pid in its own directory as `Detail::Shared`, and `process::signal` refuses to signal one of those -- which is what keeps stopping, deleting or cleaning up after one session from unloading a model every other session is using. Nothing stops a router on its own. That is deliberate (a loaded model is minutes of disk) and it is why the machines tab now has a card per provider that opens its own screen: how each model is loaded, how many stay in memory, Unload, and Stop. How a model is *loaded* therefore belongs to the model on its machine rather than to a session -- context size, GPU layers, threads, slots, speculative decoding -- written into the preset as llama-server's own argument names. Saving them re-reads that file, which unloads the model; that is the change taking effect, and the dialog says so before you save. What stays a session's is everything that rides on a request, including which tools it offers: the router hosts one set for the machine and the choice is a filter applied here, so it costs no reload (2,181 tokens of prompt with all seven, 698 with none). Verified end to end against the scratch backend and the emulator: two sessions sharing one loaded model with one child process, a second session joining it with a 26ms prefill, a backend restart adopting the router and answering with the prompt cache intact, the same over ssh to this VM, a model's settings reaching the running server, Unload, and Stop leaving every session `exited` with no error line.
565 lines
23 KiB
Rust
565 lines
23 KiB
Rust
//! What a session's process is, and how far this server has read it --
|
|
//! written down so a *later* run of this server can find the same process
|
|
//! rather than start a second one.
|
|
//!
|
|
//! Stopping the backend must not kill a turn that is in flight, so session
|
|
//! processes are left running and adopted again on the way back up. That only
|
|
//! works if "is this still mine?" has an answer, which is what this module is.
|
|
//!
|
|
//! **A pid is not an identity.** Pids are reused, so adopting one by number
|
|
//! alone eventually means treating a stranger's process as a session -- never
|
|
//! resuming the real conversation, and signalling something unrelated when the
|
|
//! session is deleted. The kernel's start time for that pid is recorded beside
|
|
//! it; the pair is unique for as long as the machine has been up.
|
|
//!
|
|
//! **How to reach it again belongs here too**, in the same record and the same
|
|
//! write, because it answers the other half of the same question. Splitting
|
|
//! them would be two files that can disagree about one process. What it takes
|
|
//! differs by driver, so it is a typed [`Detail`] rather than a union of every
|
|
//! driver's fields.
|
|
//!
|
|
//! The record is rewritten in place as reading advances. A crash during that
|
|
//! write leaves a record that does not parse, which is read as "no live
|
|
//! process" -- so the failure is the old behaviour rather than a wrong
|
|
//! adoption.
|
|
|
|
use std::os::unix::fs::OpenOptionsExt;
|
|
use std::path::{Path, PathBuf};
|
|
|
|
use anyhow::{Context, Result};
|
|
use serde::{Deserialize, Serialize};
|
|
|
|
const RECORD_FILE: &str = "process.json";
|
|
const STOP_REQUEST_FILE: &str = "stop-requested";
|
|
|
|
/// A process this server started and expects to outlive it.
|
|
#[derive(Debug, Clone, PartialEq, Serialize, Deserialize)]
|
|
pub struct Record {
|
|
pub pid: u32,
|
|
/// The kernel's start time for `pid`, in clock ticks since boot. See the
|
|
/// module comment: this is what makes the pid an identity.
|
|
pub started: u64,
|
|
/// What the driver needs in order to pick this process back up.
|
|
#[serde(flatten)]
|
|
pub detail: Detail,
|
|
}
|
|
|
|
/// How a reattaching driver reaches a process it did not start.
|
|
#[derive(Debug, Clone, PartialEq, Serialize, Deserialize)]
|
|
#[serde(tag = "kind", rename_all = "snake_case")]
|
|
pub enum Detail {
|
|
/// Spoken to over stdio, which outlives the server as files in the session
|
|
/// directory. `stdout_read` is how many bytes of the stdout log have
|
|
/// already become events: everything after it is what a reattaching server
|
|
/// owes the conversation.
|
|
Stdio { stdout_read: u64 },
|
|
/// Spoken to over HTTP on a loopback port, which is all it takes to find
|
|
/// it again -- there is no stream to be partway through.
|
|
Http { port: u16 },
|
|
/// The same, for a process this session reaches but does not own: the
|
|
/// llama.cpp router serving every session on its machine.
|
|
///
|
|
/// A variant rather than a flag because of what it forbids. Liveness is
|
|
/// the identical question -- a session whose router has gone has no model
|
|
/// -- but ending it is not this session's to ask, and [`signal`] is where
|
|
/// that is enforced: stopping, deleting or cleaning up after a session
|
|
/// must not take a model out of memory for every other session on that
|
|
/// machine.
|
|
Shared { port: u16 },
|
|
}
|
|
|
|
/// Whether a recorded process is still there.
|
|
///
|
|
/// Three answers rather than a boolean, because "I could not find out" is a
|
|
/// real one and is not the same as "no". Treating it as "no" is what would
|
|
/// start a second process against a conversation that already has one.
|
|
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
|
|
pub enum Liveness {
|
|
Alive,
|
|
Dead,
|
|
Unknown,
|
|
}
|
|
|
|
impl Record {
|
|
/// The record for a process this server just started, or `None` when the
|
|
/// kernel will not say when it started -- which is the same answer as "do
|
|
/// not adopt this later", and the safe one.
|
|
pub fn of(pid: u32, detail: Detail) -> Option<Self> {
|
|
Some(Self {
|
|
pid,
|
|
started: stat_of(pid).ok().flatten()?.started,
|
|
detail,
|
|
})
|
|
}
|
|
|
|
/// Whether this server may end that process.
|
|
///
|
|
/// False for the one it shares -- see [`Detail::Shared`]. Liveness is the
|
|
/// identical question for both, which is why this is separate from it:
|
|
/// "is it there?" and "is it mine to end?" are asked in different places.
|
|
pub fn ours(&self) -> bool {
|
|
!matches!(self.detail, Detail::Shared { .. })
|
|
}
|
|
|
|
pub fn liveness(&self) -> Liveness {
|
|
match stat_of(self.pid) {
|
|
// A different start time is a reused pid, so definitely not ours.
|
|
Ok(Some(stat)) if stat.started == self.started => {
|
|
if stat.exited {
|
|
Liveness::Dead
|
|
} else {
|
|
Liveness::Alive
|
|
}
|
|
}
|
|
Ok(Some(_)) | Ok(None) => Liveness::Dead,
|
|
Err(_) => Liveness::Unknown,
|
|
}
|
|
}
|
|
}
|
|
|
|
fn path(session_dir: &Path) -> PathBuf {
|
|
session_dir.join(RECORD_FILE)
|
|
}
|
|
|
|
fn stop_request_path(session_dir: &Path) -> PathBuf {
|
|
session_dir.join(STOP_REQUEST_FILE)
|
|
}
|
|
|
|
/// Marks the process as one the server deliberately asked to end, so its watcher reports an exit
|
|
/// without presenting ordinary stderr from the process's lifetime as the cause.
|
|
pub fn mark_stopping(session_dir: &Path) -> Result<()> {
|
|
std::fs::OpenOptions::new()
|
|
.create(true)
|
|
.write(true)
|
|
.truncate(true)
|
|
.mode(0o600)
|
|
.open(stop_request_path(session_dir))
|
|
.with_context(|| format!("marking {} as stopping", session_dir.display()))?;
|
|
Ok(())
|
|
}
|
|
|
|
/// Whether this process's exit was deliberately requested.
|
|
pub fn stopping(session_dir: &Path) -> bool {
|
|
stop_request_path(session_dir).is_file()
|
|
}
|
|
|
|
/// The recorded process and whether it is still there, or `None` when nothing
|
|
/// usable is recorded. A record that does not parse reads as no record: the
|
|
/// only way to get one is a crash partway through writing it, and the safe
|
|
/// reading is that this server has no claim on anything.
|
|
pub fn recorded(session_dir: &Path) -> Option<(Record, Liveness)> {
|
|
let text = std::fs::read_to_string(path(session_dir)).ok()?;
|
|
let record: Record = serde_json::from_str(text.trim_end()).ok()?;
|
|
let liveness = record.liveness();
|
|
Some((record, liveness))
|
|
}
|
|
|
|
/// The recorded process if it is definitely still running. One function rather
|
|
/// than a read plus a liveness check at each caller: the caller that forgets
|
|
/// the second half is the one that starts a duplicate.
|
|
pub fn live(session_dir: &Path) -> Option<Record> {
|
|
match recorded(session_dir) {
|
|
Some((record, Liveness::Alive)) => Some(record),
|
|
_ => None,
|
|
}
|
|
}
|
|
|
|
/// Writes `record` where [`live`] will find it, atomically -- to a neighbouring
|
|
/// file, renamed over the real name, so a reader sees either the whole old
|
|
/// record or the whole new one.
|
|
///
|
|
/// Writing in place would not be, and the consequence is severe rather than
|
|
/// untidy. `fs::write` truncates before it fills, so a crash inside that window
|
|
/// leaves no readable record -- and a missing record reads as "nothing is
|
|
/// running", which is the single answer that makes the next launch start a
|
|
/// *second* process against a conversation that already has one. The window is
|
|
/// not rare: this runs on every read that makes progress, so many times a
|
|
/// second while a turn is producing output.
|
|
///
|
|
/// Errors are logged rather than returned: this runs on the reading path, and a
|
|
/// session that cannot save its position is still worth having.
|
|
pub fn write(session_dir: &Path, record: &Record) {
|
|
let path = path(session_dir);
|
|
let text = match serde_json::to_string(record) {
|
|
Ok(text) => text,
|
|
Err(err) => {
|
|
tracing::error!("couldn't serialize the process record: {err}");
|
|
return;
|
|
}
|
|
};
|
|
// Beside the real file so the rename stays within one filesystem, which is
|
|
// what makes it atomic.
|
|
let temp = path.with_extension("json.new");
|
|
let written = std::fs::OpenOptions::new()
|
|
.create(true)
|
|
.write(true)
|
|
.truncate(true)
|
|
// Owner-only, like everything else in a session directory.
|
|
.mode(0o600)
|
|
.open(&temp)
|
|
.and_then(|mut file| {
|
|
use std::io::Write;
|
|
file.write_all(text.as_bytes())?;
|
|
file.write_all(b"\n")
|
|
})
|
|
.and_then(|()| std::fs::rename(&temp, &path));
|
|
if let Err(err) = written {
|
|
tracing::error!(
|
|
"couldn't record the session process in {}: {err}",
|
|
path.display()
|
|
);
|
|
let _ = std::fs::remove_file(&temp);
|
|
}
|
|
}
|
|
|
|
/// How many bytes `path` holds, or 0 if it is not there. Exists so a caller
|
|
/// wanting only the length does not have to read the file to find it.
|
|
pub fn size_of(path: &Path) -> u64 {
|
|
std::fs::metadata(path).map(|meta| meta.len()).unwrap_or(0)
|
|
}
|
|
|
|
/// Forgets the recorded process -- for one confirmed dead, or a session
|
|
/// being deleted. The path out for [`write`].
|
|
pub fn clear(session_dir: &Path) {
|
|
for path in [path(session_dir), stop_request_path(session_dir)] {
|
|
if let Err(err) = std::fs::remove_file(&path)
|
|
&& err.kind() != std::io::ErrorKind::NotFound
|
|
{
|
|
tracing::warn!("couldn't remove {}: {err}", path.display());
|
|
}
|
|
}
|
|
}
|
|
|
|
/// Creates a session stdin fifo and opens it read-write for the child.
|
|
///
|
|
/// The child holding the write end is what keeps a detached JSON server from
|
|
/// reading EOF when ai-server restarts and temporarily closes its own writer.
|
|
pub fn make_fifo(path: &Path) -> Result<std::fs::File> {
|
|
if !path.exists() {
|
|
let c_path = std::ffi::CString::new(path.as_os_str().as_encoded_bytes())
|
|
.with_context(|| format!("{} is not a usable path", path.display()))?;
|
|
// SAFETY: `c_path` is nul-terminated and this call only reads it.
|
|
let made = unsafe { libc::mkfifo(c_path.as_ptr(), 0o600) };
|
|
if made != 0 {
|
|
return Err(std::io::Error::last_os_error())
|
|
.with_context(|| format!("creating the fifo {}", path.display()));
|
|
}
|
|
}
|
|
std::fs::OpenOptions::new()
|
|
.read(true)
|
|
.write(true)
|
|
.open(path)
|
|
.with_context(|| format!("opening the fifo {}", path.display()))
|
|
}
|
|
|
|
/// A fresh owner-only log for a detached session process.
|
|
pub fn create_log(path: &Path) -> Result<std::fs::File> {
|
|
std::fs::OpenOptions::new()
|
|
.create(true)
|
|
.write(true)
|
|
.truncate(true)
|
|
.mode(0o600)
|
|
.open(path)
|
|
.with_context(|| format!("creating {}", path.display()))
|
|
}
|
|
|
|
/// Grace period between asking a session's process to stop and killing it.
|
|
/// Here rather than beside each caller: two drivers plus the manager had
|
|
/// written the same five seconds down separately.
|
|
pub const STOP_GRACE: std::time::Duration = std::time::Duration::from_secs(5);
|
|
|
|
/// Asks it to stop, then makes sure. Used where a leaked process must actually
|
|
/// end: a deleted session, or one being replaced.
|
|
///
|
|
/// SIGTERM first because the CLI writes its own session file on the way out and
|
|
/// a SIGKILL would cost whatever it had not flushed; SIGKILL after the grace
|
|
/// period because a session the phone has deleted must not still be running.
|
|
pub fn stop(record: &Record, grace: std::time::Duration) {
|
|
if !record.ours() || record.liveness() != Liveness::Alive {
|
|
return;
|
|
}
|
|
signal(record, libc::SIGTERM);
|
|
let record = record.clone();
|
|
tokio::spawn(async move {
|
|
tokio::time::sleep(grace).await;
|
|
kill_if_still_there(&record, grace);
|
|
});
|
|
}
|
|
|
|
/// Waits for processes already asked to stop, and kills whichever have not, for
|
|
/// a caller that is about to exit.
|
|
///
|
|
/// The waiting cannot be [`stop`]'s here, and that is the whole reason this
|
|
/// exists: the kill it leaves behind is a timer inside the tokio runtime, and a
|
|
/// runtime that is shutting down never runs it. That is how the original
|
|
/// `shutdown_all` leaked the processes it had just asked to stop -- it reported
|
|
/// them stopped, too, which is worse than not asking.
|
|
///
|
|
/// One deadline for all of them rather than one each: they were signalled
|
|
/// together, so waiting is bounded by the grace period however many there are.
|
|
pub fn wait_gone(records: &[Record], grace: std::time::Duration) {
|
|
/// How often to look. Short enough that a process that goes at once costs
|
|
/// nothing noticeable, and long enough not to spin.
|
|
const LOOK: std::time::Duration = std::time::Duration::from_millis(20);
|
|
|
|
let deadline = std::time::Instant::now() + grace;
|
|
for record in records {
|
|
// Never asked to stop, so there is nothing to wait out.
|
|
if !record.ours() {
|
|
continue;
|
|
}
|
|
while record.liveness() == Liveness::Alive && std::time::Instant::now() < deadline {
|
|
std::thread::sleep(LOOK);
|
|
}
|
|
kill_if_still_there(record, grace);
|
|
}
|
|
}
|
|
|
|
/// The end of both paths above: a process that was asked to stop and did not is
|
|
/// killed. Written once because the two callers differ only in how they wait,
|
|
/// and a grace period meaning one thing in one and something else in the other
|
|
/// is exactly the drift `STOP_GRACE` was gathered here to prevent.
|
|
fn kill_if_still_there(record: &Record, grace: std::time::Duration) {
|
|
if record.liveness() == Liveness::Alive {
|
|
tracing::warn!(
|
|
"session process {} did not stop within {:?}; killing it",
|
|
record.pid,
|
|
grace
|
|
);
|
|
signal(record, libc::SIGKILL);
|
|
}
|
|
}
|
|
|
|
/// The one place a session's process is signalled, which is why the refusal to
|
|
/// signal a shared one lives here rather than at each caller: every path out of
|
|
/// a session -- stopped, deleted, cleaned up on the way down -- ends in this
|
|
/// function, and the one that forgot would be a model unloaded under somebody
|
|
/// else's turn.
|
|
fn signal(record: &Record, signal: libc::c_int) {
|
|
if !record.ours() {
|
|
return;
|
|
}
|
|
// SAFETY: `kill` with a positive pid touches only that process, and the pid
|
|
// came from a record whose start time was just confirmed to match -- so it
|
|
// is still the process this server started, not a reused number. A failure
|
|
// (already gone) is nothing to act on.
|
|
unsafe {
|
|
libc::kill(record.pid as libc::pid_t, signal);
|
|
}
|
|
}
|
|
|
|
/// What `/proc` says about a pid.
|
|
struct Stat {
|
|
/// The kernel's start time in clock ticks since boot -- see
|
|
/// [`Record::started`].
|
|
started: u64,
|
|
/// State `Z`: the process has ended, and the kernel is keeping its entry
|
|
/// only until somebody collects the exit status.
|
|
///
|
|
/// Read rather than ignored, because that entry has the same pid *and* the
|
|
/// same start time, so a finished process goes on answering "still there"
|
|
/// for as long as nothing reaps it -- which makes `Exited` unsayable: the
|
|
/// session shows `unknown`, its Start button never appears, and stopping it
|
|
/// says there is nothing to stop.
|
|
exited: bool,
|
|
}
|
|
|
|
/// The kernel's start time for `pid`, in clock ticks since boot.
|
|
///
|
|
/// Field 22 of `/proc/<pid>/stat`, counted from the closing parenthesis of
|
|
/// field 2 rather than from the start of the line: a process's name is field 2,
|
|
/// it is wrapped in parentheses, and it may itself contain spaces and
|
|
/// parentheses. Splitting the whole line on whitespace reads the wrong field
|
|
/// for anything with a space in its name.
|
|
///
|
|
/// Three outcomes, and they are not the same: `Ok(None)` is "no such process",
|
|
/// `Err` is "could not find out". Collapsing the second into the first is what
|
|
/// would let a machine without a readable `/proc` look like a machine with
|
|
/// nothing running on it. Linux-specific, like `import`'s use of GNU `stat`.
|
|
fn stat_of(pid: u32) -> std::io::Result<Option<Stat>> {
|
|
let stat = match std::fs::read_to_string(format!("/proc/{pid}/stat")) {
|
|
Ok(stat) => stat,
|
|
Err(err) if err.kind() == std::io::ErrorKind::NotFound => return Ok(None),
|
|
Err(err) => return Err(err),
|
|
};
|
|
// A `/proc` entry that exists but does not have the shape this reads is not
|
|
// a process that has gone away; it is a reading this code cannot make.
|
|
let unreadable =
|
|
|| std::io::Error::new(std::io::ErrorKind::InvalidData, "unreadable /proc stat");
|
|
let after_name = stat.rsplit_once(')').ok_or_else(unreadable)?.1;
|
|
// Field 3 is the first after the name, so the state is the first here and
|
|
// field 22 is the 20th.
|
|
let mut fields = after_name.split_whitespace();
|
|
let exited = fields.next().ok_or_else(unreadable)? == "Z";
|
|
let started = fields
|
|
.nth(18)
|
|
.ok_or_else(unreadable)?
|
|
.parse()
|
|
.map_err(|_| unreadable())?;
|
|
Ok(Some(Stat { started, exited }))
|
|
}
|
|
|
|
/// Reads `path` from `from`, returning what is there and where reading reached.
|
|
/// A file truncated or replaced under us reads from the start, since the offset
|
|
/// no longer means anything in it.
|
|
pub fn read_from(path: &Path, from: u64) -> Result<(Vec<u8>, u64)> {
|
|
use std::io::{Read, Seek, SeekFrom};
|
|
let mut file = match std::fs::File::open(path) {
|
|
Ok(file) => file,
|
|
Err(err) if err.kind() == std::io::ErrorKind::NotFound => return Ok((Vec::new(), from)),
|
|
Err(err) => return Err(err).with_context(|| format!("open {}", path.display())),
|
|
};
|
|
let len = file
|
|
.metadata()
|
|
.with_context(|| format!("stat {}", path.display()))?
|
|
.len();
|
|
let from = if from > len { 0 } else { from };
|
|
file.seek(SeekFrom::Start(from))
|
|
.with_context(|| format!("seek {}", path.display()))?;
|
|
let mut bytes = Vec::new();
|
|
file.read_to_end(&mut bytes)
|
|
.with_context(|| format!("read {}", path.display()))?;
|
|
let read = from + bytes.len() as u64;
|
|
Ok((bytes, read))
|
|
}
|
|
|
|
#[cfg(test)]
|
|
mod tests {
|
|
use super::*;
|
|
|
|
#[test]
|
|
fn this_process_is_alive_and_a_wrong_start_time_is_not() {
|
|
let mine = Record::of(std::process::id(), Detail::Stdio { stdout_read: 0 })
|
|
.expect("this process has a start time");
|
|
assert_eq!(mine.liveness(), Liveness::Alive);
|
|
|
|
// The same pid with a different start time is a different process --
|
|
// which is the whole reason the start time is recorded.
|
|
let recycled = Record {
|
|
started: mine.started + 1,
|
|
..mine.clone()
|
|
};
|
|
assert_eq!(recycled.liveness(), Liveness::Dead);
|
|
}
|
|
|
|
#[test]
|
|
fn a_record_round_trips_through_the_padded_file() {
|
|
let dir = tempfile::tempdir().expect("tempdir");
|
|
let mut record = Record::of(std::process::id(), Detail::Stdio { stdout_read: 4096 })
|
|
.expect("start time");
|
|
write(dir.path(), &record);
|
|
assert_eq!(live(dir.path()), Some(record.clone()));
|
|
|
|
// A shorter value must not leave a readable tail of the longer one.
|
|
record.detail = Detail::Stdio { stdout_read: 1 };
|
|
write(dir.path(), &record);
|
|
assert_eq!(live(dir.path()), Some(record.clone()));
|
|
|
|
// And the other shapes round trip through the same file.
|
|
for detail in [Detail::Http { port: 8080 }, Detail::Shared { port: 8080 }] {
|
|
record.detail = detail;
|
|
write(dir.path(), &record);
|
|
assert_eq!(live(dir.path()), Some(record.clone()));
|
|
}
|
|
|
|
mark_stopping(dir.path()).expect("mark stopping");
|
|
assert!(stopping(dir.path()));
|
|
clear(dir.path());
|
|
assert_eq!(live(dir.path()), None);
|
|
assert!(!stopping(dir.path()));
|
|
}
|
|
|
|
#[test]
|
|
fn writing_leaves_no_temporary_behind_and_stays_readable() {
|
|
let dir = tempfile::tempdir().expect("tempdir");
|
|
let mut record =
|
|
Record::of(std::process::id(), Detail::Stdio { stdout_read: 0 }).expect("start time");
|
|
// Rewritten the way the reader rewrites it: constantly, as the position
|
|
// advances. Each one must land whole.
|
|
for read in [1u64, 4096, 2, 999_999] {
|
|
record.detail = Detail::Stdio { stdout_read: read };
|
|
write(dir.path(), &record);
|
|
assert_eq!(
|
|
live(dir.path()),
|
|
Some(record.clone()),
|
|
"after offset {read}"
|
|
);
|
|
}
|
|
// The rename is what makes it atomic; a leftover neighbour would mean it
|
|
// had not happened.
|
|
let stray: Vec<_> = std::fs::read_dir(dir.path())
|
|
.expect("read dir")
|
|
.filter_map(Result::ok)
|
|
.map(|e| e.file_name().to_string_lossy().into_owned())
|
|
.filter(|name| name != RECORD_FILE)
|
|
.collect();
|
|
assert!(stray.is_empty(), "left behind {stray:?}");
|
|
}
|
|
|
|
/// The whole of what [`Detail::Shared`] is for: a session ending must not
|
|
/// take the machine's llama.cpp router with it.
|
|
#[tokio::test]
|
|
async fn a_shared_process_is_not_stopped_with_the_session_that_reached_it() {
|
|
let mut child = std::process::Command::new("sleep")
|
|
.arg("30")
|
|
.spawn()
|
|
.expect("spawn sleep");
|
|
let shared = Record::of(child.id(), Detail::Shared { port: 1 }).expect("start time");
|
|
stop(&shared, std::time::Duration::from_millis(50));
|
|
std::thread::sleep(std::time::Duration::from_millis(200));
|
|
assert_eq!(shared.liveness(), Liveness::Alive, "the router was killed");
|
|
|
|
// The same process, recorded as one this session owns, does stop.
|
|
let owned = Record {
|
|
detail: Detail::Http { port: 1 },
|
|
..shared
|
|
};
|
|
stop(&owned, std::time::Duration::from_millis(50));
|
|
let _ = child.wait();
|
|
assert_eq!(owned.liveness(), Liveness::Dead);
|
|
}
|
|
|
|
#[test]
|
|
fn a_dead_or_unreadable_record_is_not_live() {
|
|
let dir = tempfile::tempdir().expect("tempdir");
|
|
assert_eq!(live(dir.path()), None);
|
|
|
|
// Pid 0 is never a process we started.
|
|
write(
|
|
dir.path(),
|
|
&Record {
|
|
pid: 0,
|
|
started: 1,
|
|
detail: Detail::Stdio { stdout_read: 0 },
|
|
},
|
|
);
|
|
assert_eq!(live(dir.path()), None);
|
|
|
|
std::fs::write(dir.path().join(RECORD_FILE), "not json").expect("write");
|
|
assert_eq!(live(dir.path()), None);
|
|
}
|
|
|
|
#[test]
|
|
fn reading_resumes_from_an_offset_and_restarts_on_truncation() {
|
|
let dir = tempfile::tempdir().expect("tempdir");
|
|
let path = dir.path().join("stdout.log");
|
|
std::fs::write(&path, b"hello world").expect("write");
|
|
|
|
let (bytes, read) = read_from(&path, 6).expect("read");
|
|
assert_eq!(bytes, b"world");
|
|
assert_eq!(read, 11);
|
|
|
|
// An offset past the end means the file was replaced, so the offset
|
|
// describes a file that no longer exists.
|
|
std::fs::write(&path, b"new").expect("truncate");
|
|
let (bytes, read) = read_from(&path, 11).expect("read");
|
|
assert_eq!(bytes, b"new");
|
|
assert_eq!(read, 3);
|
|
|
|
// A missing file is not an error: the process has said nothing.
|
|
let (bytes, read) = read_from(&dir.path().join("nope"), 7).expect("read");
|
|
assert!(bytes.is_empty());
|
|
assert_eq!(read, 7);
|
|
}
|
|
}
|