Files
DarkRoom/tools/bench/src/stats.rs
T
dtourolle e8e96eed40 Measure the performance targets §8 has been promising, and fail on a regression
docs/requirements.md §8 has said since it was written that performance is
verified by "an automated benchmark suite against a synthetic 50k catalog, run
per-commit … A regression beyond stated tolerance fails the build." There was
none. No benches/, no [[bench]], no criterion, no synthetic catalog, and three
CI workflows that between them measured nothing. Ten performance requirements
could therefore be neither passed nor failed, and five of them carried a
TRACES: tag regardless.

tools/bench is the half of that promise that can be kept honestly on a runner
with no GPU and no display.

# The fixture

Rows are cheap and pixels are not, so it builds fifty thousand catalog rows
over a pool of a dozen real files, each referenced by several thousand of them.
Everything the catalog half touches is rows and is exact at full scale;
everything the pixel half touches is one file at a time and does not care how
many rows point at it. Fourteen megabytes on disk instead of two terabytes, and
neither half is flattered by the trade. It is reproducible from a seed, and a
stamp beside it — seed, row count, source size, dr-catalog's schema version —
rebuilds it rather than letting a run be compared against a baseline that
describes a different library.

# What it can now pass or fail

NFR-P1, and R2's second sentence with it: Catalog::open plus the count, first
window and timeline the grid cannot paint without. The interesting part turned
out to be the open itself — schema::backfill runs three passes over the images
table on every open, which is O(library) work on a path whose budget is stated
in absolute seconds. Tagged TRACES: NFR-P1, on a gate that fails if it breaks.

NFR-P3: thumbnail throughput on the embedded preview path, through the same
per-image work spawn_thumbnail_sweep does and in the same shape — chunks of 96,
lanes owning disjoint slices, the single thread that owns the store writing the
finished chunk. Mirrored rather than called, because that function takes a
RemoteBackend and would measure somebody's network. Tagged TRACES: NFR-P3.

# What it deliberately does not claim

NFR-P7 is the whole chain, and only the encode half of it runs without an
adapter. So the export row is a one-sided gate — over two seconds in the encode
alone violates the requirement; under it proves nothing — and there is no
TRACES: NFR-P7 anywhere. NFR-P8 is about the application at idle, and the probe
is a process holding the catalog and nothing else, so it records the catalog
layer's share and carries no budget until somebody decides what that share
should be. No tag there either. CONTRIBUTING.md asks that a requirement be
closed by a test that would fail if the behaviour were removed, and two more
plumbing tags is what this repository already has too many of.

NFR-P8 also gets the answer §4.1 demands: RSS is exclusive of device-local GPU
allocations and cannot be made otherwise, because such an allocation never
enters the process's address space. The requirement should be restated as two
figures, and docs/benchmarks.md says so.

# Two gates, and why one of them steps aside off the reference desktop

The budget is the requirement's own number and never moves. The baseline is
what the reference desktop last measured, and drifting 15% past it fails the
build even while still inside the budget — which is how performance rot
actually arrives, never over the line, always a little worse.

A budget written for twenty-four threads cannot be asserted on a two-core
container. §8 names the reference desktop, not CI, so each metric declares
whether its budget is machine-sensitive; those are asserted under --reference
and reported everywhere else. Catalog open is not one of them: two seconds
against an expected figure two orders of magnitude smaller is a threshold any
machine can be held to. This is the trap core/dr-gpu/tests/frame_budget.rs
already refuses — a red gate everybody learns to ignore.

# The baseline ships with no numbers in it

Every recorded field is null, because nobody has run it yet. Writing
plausible-looking figures would make every later comparison a comparison
against a guess, and the first real regression would be invisible. Run
`dr-bench record --reference` on the reference desktop and commit the diff;
until then the budget gate works and the report says the other one cannot.

# CI

.gitea/workflows/benchmark.yml, and its own workflow rather than a step in
build-and-test.yml: a red "Build and test" says the code is wrong, a red
"Benchmarks" says it got slower, and the second must not be reachable by
retrying a flaky compile. The cpu job runs on every push and builds -p dr-bench
alone — which is why that crate depends on no GPU and no UI crate. The gpu job
is the frame budget that already exists and already skips without an adapter,
on workflow_dispatch, because building wgpu on every commit to rediscover that
the runner has no device is not a use of anybody's minutes.
2026-08-30 10:40:10 +02:00

123 lines
4.5 KiB
Rust

//! Ranking a set of samples, the way `dr-gpu`'s frame budget ranks them.
//!
//! Copied in spirit rather than shared, because the two live in different
//! dependency worlds — `core/dr-gpu/examples/frame_budget.rs` is an example
//! inside a crate this one deliberately does not depend on (see `Cargo.toml`).
//! The arithmetic is identical on purpose: two percentile definitions in one
//! repository is how two benchmarks come to disagree about the same machine.
//!
//! # Nearest-rank, not an interpolating definition
//!
//! The samples *are* the population. There is no distribution being estimated
//! here, only a set of catalog opens or thumbnail encodes that either happened
//! inside the target or did not. At 100 samples the 99th percentile is the
//! second-worst, which is the honest reading of "one bad one in a hundred is
//! one too many" without letting a single scheduler hiccup on an unrelated
//! process decide the verdict.
use std::time::Duration;
/// Nearest-rank percentiles over a set of samples, in the caller's unit.
#[derive(Debug, Clone, Copy)]
pub struct Percentiles {
pub p50: f64,
pub p99: f64,
pub max: f64,
}
impl Percentiles {
/// Rank `samples`. Panics on an empty set, which is a harness bug rather
/// than a measurement: a row with nothing in it must not print a zero that
/// reads like a very fast result.
pub fn of(mut samples: Vec<f64>) -> Self {
assert!(
!samples.is_empty(),
"percentiles of an empty sample set — the measurement produced nothing"
);
samples.sort_by(f64::total_cmp);
let rank = |p: f64| {
let n = samples.len();
let i = ((p * n as f64).ceil() as usize).clamp(1, n) - 1;
samples[i]
};
Percentiles {
p50: rank(0.50),
p99: rank(0.99),
max: samples[samples.len() - 1],
}
}
}
/// A duration in milliseconds, which is the unit every timing here is stated
/// in. One spelling, so no row is accidentally in seconds.
pub fn ms(d: Duration) -> f64 {
d.as_secs_f64() * 1e3
}
/// A deterministic generator, so a fixture is reproducible from its seed.
///
/// SplitMix64. Chosen because it is eight lines, has no dependency, and passes
/// the only test that matters here — that the same seed produces the same
/// catalog on the reference desktop and on the CI runner, so a number measured
/// in one place describes the same workload as a number measured in the other.
/// Nothing cryptographic depends on it.
pub struct Rng(u64);
impl Rng {
pub fn new(seed: u64) -> Self {
Rng(seed)
}
pub fn next_u64(&mut self) -> u64 {
self.0 = self.0.wrapping_add(0x9E37_79B9_7F4A_7C15);
let mut z = self.0;
z = (z ^ (z >> 30)).wrapping_mul(0xBF58_476D_1CE4_E5B9);
z = (z ^ (z >> 27)).wrapping_mul(0x94D0_49BB_1331_11EB);
z ^ (z >> 31)
}
/// A value in `0..n`. Modulo-biased, which does not matter for a fixture:
/// nothing here is a statistical test, only a spread of plausible values.
pub fn below(&mut self, n: u64) -> u64 {
self.next_u64() % n.max(1)
}
}
#[cfg(test)]
mod tests {
use super::*;
#[test]
fn the_ninety_ninth_of_a_hundred_is_the_second_worst() {
// The property the whole suite's verdict rests on. Off by one here and
// every threshold is judged against the worst sample instead.
let samples: Vec<f64> = (1..=100).map(|n| n as f64).collect();
let p = Percentiles::of(samples);
assert_eq!(p.p99, 99.0);
assert_eq!(p.max, 100.0);
assert_eq!(p.p50, 50.0);
}
#[test]
fn a_single_sample_ranks_as_itself() {
// A row measured once — a cold catalog open — must not divide by zero
// or index off the end.
let p = Percentiles::of(vec![7.5]);
assert_eq!((p.p50, p.p99, p.max), (7.5, 7.5, 7.5));
}
#[test]
fn the_same_seed_gives_the_same_sequence() {
// Reproducibility from a seed is what makes a committed baseline mean
// anything: two runs must describe the same catalog.
let mut a = Rng::new(20_260_829);
let mut b = Rng::new(20_260_829);
let mut c = Rng::new(20_260_830);
let first: Vec<u64> = (0..8).map(|_| a.next_u64()).collect();
let same: Vec<u64> = (0..8).map(|_| b.next_u64()).collect();
let other: Vec<u64> = (0..8).map(|_| c.next_u64()).collect();
assert_eq!(first, same);
assert_ne!(first, other);
}
}