docs/ had 26 developer documents flat beside the manual, and the two audiences are very differently sized: most readers want the manual and the gesture reference, a few want the register, the designs and the measurements. The manual and gestures.md stay at the top; everything for someone changing the code moves to docs/dev/, and the two documents that name their own successors — the v0.1 milestone and the UI-refinement plan — go to docs/dev/archive/ rather than being deleted, since both are still cited. docs/README.md is the index, users first. Every reference follows: code comments, Cargo manifests, the workflows, the pre-commit hook, the bench and traceability tools (which locate the repo root by docs/dev/requirements.md now), packaging, the Docker READMEs, CLAUDE.md, CONTRIBUTING.md and the README. The matrix links one level deeper and is regenerated. Links out of the moved documents into the tree gain a level; a link checker over every Markdown file finds none broken.
301 lines
12 KiB
Rust
301 lines
12 KiB
Rust
//! The committed numbers, and what counts as a regression against them.
|
|
//!
|
|
//! `docs/dev/frame-budget.md` commits its measurements by hand and says why: *"a
|
|
//! regression should be a diff rather than somebody's memory."* This is the
|
|
//! same idea in a form a program can read, because §8 asks for more than a
|
|
//! record — *"a regression beyond stated tolerance fails the build"*.
|
|
//!
|
|
//! # Two gates, and they are not the same gate
|
|
//!
|
|
//! **The budget** is the requirement's own number: 2 s to open a catalog, 100
|
|
//! images a second through the preview path. It does not move. A build that
|
|
//! violates it has violated a requirement, and no amount of "but it was always
|
|
//! like that" changes it.
|
|
//!
|
|
//! **The baseline** is what this machine last measured. It moves — deliberately
|
|
//! and by hand, through `dr-bench record` — and its job is to catch the change
|
|
//! that is still inside the budget but has halved the headroom. Most real
|
|
//! performance rot arrives that way: never over the line, always a little
|
|
//! worse, until one day the line is crossed by a change that was not the cause.
|
|
//!
|
|
//! # Why a machine-sensitive metric skips its budget off the reference desktop
|
|
//!
|
|
//! §8 names *"the reference desktop"*, not CI, and it is right to. A container
|
|
//! with two cores cannot speak to a throughput target written for a
|
|
//! twenty-four-thread machine, and asserting it there would produce exactly
|
|
//! what `core/dr-gpu/tests/frame_budget.rs` refused to produce: *"a red suite
|
|
//! that everyone learns to ignore"*. So a metric declares whether its budget
|
|
//! is machine-sensitive. Those budgets are asserted under `--reference` and
|
|
//! reported everywhere else; the ones with orders of magnitude of headroom —
|
|
//! catalog open against two seconds — are asserted everywhere, because a
|
|
//! failure there is a real failure on any machine.
|
|
//!
|
|
//! # Why the committed file starts with no numbers in it
|
|
//!
|
|
//! Because nobody had run it yet. Writing plausible-looking figures into a
|
|
//! baseline is the one thing that would make the whole suite worthless: every
|
|
//! later comparison would be against a guess, and the first genuine regression
|
|
//! would be invisible or, worse, a fabricated improvement. `recorded` is
|
|
//! therefore `null` until somebody runs `dr-bench record --reference` on the
|
|
//! reference desktop and commits the diff. Until then the budget gate works
|
|
//! and the regression gate says so rather than pretending.
|
|
|
|
use std::collections::BTreeMap;
|
|
use std::path::{Path, PathBuf};
|
|
|
|
use anyhow::{Context, Result};
|
|
use serde::{Deserialize, Serialize};
|
|
|
|
use crate::fixture::Stamp;
|
|
|
|
/// Which way is better.
|
|
#[derive(Debug, Clone, Copy, PartialEq, Eq, Serialize, Deserialize)]
|
|
#[serde(rename_all = "snake_case")]
|
|
pub enum Direction {
|
|
LowerIsBetter,
|
|
HigherIsBetter,
|
|
}
|
|
|
|
/// One measured quantity: what it is, what it must be, and what it was.
|
|
#[derive(Debug, Clone, Serialize, Deserialize)]
|
|
pub struct Metric {
|
|
/// The requirement ID this speaks to, or a note saying it speaks to only
|
|
/// part of one. Free text on purpose: several of these describe a fraction
|
|
/// of a requirement, and a bare ID here would read as the whole of it.
|
|
pub requirement: String,
|
|
/// One sentence a reader of the JSON can understand without the code.
|
|
pub what: String,
|
|
pub unit: String,
|
|
pub direction: Direction,
|
|
/// Whether the budget below is a statement about a machine as much as
|
|
/// about the code. See this module's header.
|
|
pub machine_sensitive: bool,
|
|
/// The requirement's own threshold, where it has one this can check.
|
|
pub budget: Option<f64>,
|
|
/// What the last `record` measured. `null` until one has been taken.
|
|
pub recorded: Option<f64>,
|
|
}
|
|
|
|
/// The committed file.
|
|
#[derive(Debug, Clone, Serialize, Deserialize)]
|
|
pub struct Baseline {
|
|
/// Prose for whoever opens the JSON first. Kept in the file rather than
|
|
/// only in `docs/dev/benchmarks.md`, because the person who finds this in a
|
|
/// failing CI log is not reading the docs directory at that moment.
|
|
#[serde(rename = "_readme")]
|
|
pub readme: Vec<String>,
|
|
/// How much worse than [`Metric::recorded`] a metric may get before the
|
|
/// build fails, as a fraction. 0.15 is fifteen per cent.
|
|
pub tolerance: f64,
|
|
/// The machine the recorded figures came from, as [`machine_id`] spells
|
|
/// it. Compared before a regression is judged: drift against a different
|
|
/// machine's numbers is not a regression, it is a different machine.
|
|
pub recorded_on: Option<String>,
|
|
/// Unix seconds. An integer rather than a formatted date because this
|
|
/// workspace has no date library and adding one for a comment would be a
|
|
/// poor trade.
|
|
pub recorded_at_unix: Option<i64>,
|
|
/// The fixture the recorded figures describe. A number measured against a
|
|
/// different workload is not comparable, and this is what says so.
|
|
pub fixture: Option<Stamp>,
|
|
pub metrics: BTreeMap<String, Metric>,
|
|
}
|
|
|
|
/// How a measurement compares.
|
|
#[derive(Debug, Clone, Copy)]
|
|
pub struct Judgement {
|
|
/// Past the requirement's own threshold. A build failure.
|
|
pub over_budget: bool,
|
|
/// Worse than the recorded baseline by more than the tolerance. A build
|
|
/// failure.
|
|
pub regressed: bool,
|
|
/// Fractional change against the baseline, positive meaning worse.
|
|
pub drift: Option<f64>,
|
|
/// A budget exists but was not asserted, because it is machine-sensitive
|
|
/// and this is not the reference desktop.
|
|
pub budget_deferred: bool,
|
|
}
|
|
|
|
/// What a judgement is made in the light of.
|
|
#[derive(Debug, Clone, Copy)]
|
|
pub struct Judging {
|
|
/// This run declares itself the reference desktop.
|
|
pub reference: bool,
|
|
/// This run is on the machine the baseline was recorded on.
|
|
pub same_machine: bool,
|
|
pub tolerance: f64,
|
|
}
|
|
|
|
impl Metric {
|
|
/// Judge `measured` against the budget and the baseline.
|
|
pub fn judge(&self, measured: f64, cx: Judging) -> Judgement {
|
|
let violates = |threshold: f64| match self.direction {
|
|
Direction::LowerIsBetter => measured > threshold,
|
|
Direction::HigherIsBetter => measured < threshold,
|
|
};
|
|
let assert_budget = self.budget.is_some() && (cx.reference || !self.machine_sensitive);
|
|
|
|
let drift = match self.recorded {
|
|
// A recorded zero would divide by nothing, and a recorded figure
|
|
// of zero is a broken record rather than a very fast one.
|
|
Some(was) if was > 0.0 => Some(match self.direction {
|
|
Direction::LowerIsBetter => (measured - was) / was,
|
|
Direction::HigherIsBetter => (was - measured) / was,
|
|
}),
|
|
_ => None,
|
|
};
|
|
|
|
Judgement {
|
|
over_budget: assert_budget && self.budget.is_some_and(violates),
|
|
regressed: cx.same_machine && drift.is_some_and(|d| d > cx.tolerance),
|
|
drift,
|
|
budget_deferred: self.budget.is_some() && !assert_budget,
|
|
}
|
|
}
|
|
}
|
|
|
|
impl Baseline {
|
|
pub fn load(path: &Path) -> Result<Baseline> {
|
|
let text = std::fs::read_to_string(path)
|
|
.with_context(|| format!("reading the baseline at {}", path.display()))?;
|
|
serde_json::from_str(&text)
|
|
.with_context(|| format!("parsing the baseline at {}", path.display()))
|
|
}
|
|
|
|
pub fn save(&self, path: &Path) -> Result<()> {
|
|
let mut text = serde_json::to_string_pretty(self)?;
|
|
// A trailing newline, so the file is a well-behaved text file and a
|
|
// `record` that changed nothing produces an empty diff.
|
|
text.push('\n');
|
|
std::fs::write(path, text)
|
|
.with_context(|| format!("writing the baseline to {}", path.display()))
|
|
}
|
|
|
|
/// Where the committed baseline lives, found the way `tools/traceability`
|
|
/// finds the repo root: by walking up from this crate's manifest until
|
|
/// `docs/dev/requirements.md` appears.
|
|
pub fn default_path() -> Result<PathBuf> {
|
|
let mut dir = PathBuf::from(env!("CARGO_MANIFEST_DIR"));
|
|
while !dir.join("docs/dev/requirements.md").exists() {
|
|
if !dir.pop() {
|
|
anyhow::bail!("could not locate the repo root above this crate");
|
|
}
|
|
}
|
|
Ok(dir.join("docs/dev/bench-baseline.json"))
|
|
}
|
|
}
|
|
|
|
/// How this machine is named in the baseline.
|
|
///
|
|
/// Host name plus thread count. Not a hardware inventory — it exists to answer
|
|
/// one question, "are these numbers from here?", and to answer it the same way
|
|
/// twice on the same box. A container whose hostname changes per run therefore
|
|
/// never matches, which is the correct answer for CI: its drift is information,
|
|
/// not a verdict.
|
|
pub fn machine_id() -> String {
|
|
let host = std::fs::read_to_string("/proc/sys/kernel/hostname")
|
|
.map(|s| s.trim().to_string())
|
|
.unwrap_or_else(|_| "unknown-host".to_string());
|
|
let threads = std::thread::available_parallelism()
|
|
.map(|n| n.get())
|
|
.unwrap_or(0);
|
|
format!("{host} ({threads} threads)")
|
|
}
|
|
|
|
/// Unix seconds now, or 0 if the clock is before 1970, which it is not.
|
|
pub fn now_unix() -> i64 {
|
|
std::time::SystemTime::now()
|
|
.duration_since(std::time::UNIX_EPOCH)
|
|
.map(|d| d.as_secs() as i64)
|
|
.unwrap_or(0)
|
|
}
|
|
|
|
#[cfg(test)]
|
|
mod tests {
|
|
use super::*;
|
|
|
|
fn metric(direction: Direction, budget: Option<f64>, recorded: Option<f64>) -> Metric {
|
|
Metric {
|
|
requirement: "NFR-TEST".into(),
|
|
what: "a test metric".into(),
|
|
unit: "ms".into(),
|
|
direction,
|
|
machine_sensitive: false,
|
|
budget,
|
|
recorded,
|
|
}
|
|
}
|
|
|
|
fn cx(same_machine: bool) -> Judging {
|
|
Judging {
|
|
reference: true,
|
|
same_machine,
|
|
tolerance: 0.15,
|
|
}
|
|
}
|
|
|
|
#[test]
|
|
fn a_budget_is_directional() {
|
|
// The bug this exists to prevent: judging a throughput target the way
|
|
// a latency target is judged, so a suite that got twice as slow passes
|
|
// and one that got twice as fast fails.
|
|
let latency = metric(Direction::LowerIsBetter, Some(100.0), None);
|
|
assert!(latency.judge(101.0, cx(true)).over_budget);
|
|
assert!(!latency.judge(99.0, cx(true)).over_budget);
|
|
|
|
let throughput = metric(Direction::HigherIsBetter, Some(100.0), None);
|
|
assert!(throughput.judge(99.0, cx(true)).over_budget);
|
|
assert!(!throughput.judge(101.0, cx(true)).over_budget);
|
|
}
|
|
|
|
#[test]
|
|
fn drift_is_positive_when_things_got_worse_whichever_way_that_is() {
|
|
let latency = metric(Direction::LowerIsBetter, None, Some(100.0));
|
|
assert!(latency.judge(120.0, cx(true)).drift.unwrap() > 0.0);
|
|
let throughput = metric(Direction::HigherIsBetter, None, Some(100.0));
|
|
assert!(throughput.judge(80.0, cx(true)).drift.unwrap() > 0.0);
|
|
}
|
|
|
|
#[test]
|
|
fn drift_against_another_machine_is_reported_but_never_a_failure() {
|
|
// CI is not the reference desktop. Its numbers are worth printing and
|
|
// are not a verdict on anybody's commit.
|
|
let m = metric(Direction::LowerIsBetter, None, Some(100.0));
|
|
let elsewhere = m.judge(400.0, cx(false));
|
|
assert!(elsewhere.drift.unwrap() > 0.15);
|
|
assert!(!elsewhere.regressed);
|
|
assert!(m.judge(400.0, cx(true)).regressed);
|
|
}
|
|
|
|
#[test]
|
|
fn an_unrecorded_metric_cannot_regress() {
|
|
// The state the committed file ships in. It must gate on the budget
|
|
// and stay silent about drift, rather than treating null as zero and
|
|
// declaring an infinite regression.
|
|
let m = metric(Direction::LowerIsBetter, Some(100.0), None);
|
|
let j = m.judge(50.0, cx(true));
|
|
assert!(j.drift.is_none());
|
|
assert!(!j.regressed);
|
|
assert!(!j.over_budget);
|
|
}
|
|
|
|
#[test]
|
|
fn a_machine_sensitive_budget_defers_off_the_reference_desktop() {
|
|
let mut m = metric(Direction::HigherIsBetter, Some(100.0), None);
|
|
m.machine_sensitive = true;
|
|
let on_ci = Judging {
|
|
reference: false,
|
|
same_machine: false,
|
|
tolerance: 0.15,
|
|
};
|
|
let j = m.judge(10.0, on_ci);
|
|
// A two-core runner must not fail a target written for twenty-four.
|
|
assert!(!j.over_budget);
|
|
// And the report has to say the budget was not applied, rather than
|
|
// letting a deferred budget read as a passed one.
|
|
assert!(j.budget_deferred);
|
|
// On the reference desktop the same figure is judged.
|
|
assert!(m.judge(10.0, cx(false)).over_budget);
|
|
}
|
|
}
|