Measure the performance targets §8 has been promising, and fail on a regression

docs/requirements.md §8 has said since it was written that performance is
verified by "an automated benchmark suite against a synthetic 50k catalog, run
per-commit … A regression beyond stated tolerance fails the build." There was
none. No benches/, no [[bench]], no criterion, no synthetic catalog, and three
CI workflows that between them measured nothing. Ten performance requirements
could therefore be neither passed nor failed, and five of them carried a
TRACES: tag regardless.

tools/bench is the half of that promise that can be kept honestly on a runner
with no GPU and no display.

# The fixture

Rows are cheap and pixels are not, so it builds fifty thousand catalog rows
over a pool of a dozen real files, each referenced by several thousand of them.
Everything the catalog half touches is rows and is exact at full scale;
everything the pixel half touches is one file at a time and does not care how
many rows point at it. Fourteen megabytes on disk instead of two terabytes, and
neither half is flattered by the trade. It is reproducible from a seed, and a
stamp beside it — seed, row count, source size, dr-catalog's schema version —
rebuilds it rather than letting a run be compared against a baseline that
describes a different library.

# What it can now pass or fail

NFR-P1, and R2's second sentence with it: Catalog::open plus the count, first
window and timeline the grid cannot paint without. The interesting part turned
out to be the open itself — schema::backfill runs three passes over the images
table on every open, which is O(library) work on a path whose budget is stated
in absolute seconds. Tagged TRACES: NFR-P1, on a gate that fails if it breaks.

NFR-P3: thumbnail throughput on the embedded preview path, through the same
per-image work spawn_thumbnail_sweep does and in the same shape — chunks of 96,
lanes owning disjoint slices, the single thread that owns the store writing the
finished chunk. Mirrored rather than called, because that function takes a
RemoteBackend and would measure somebody's network. Tagged TRACES: NFR-P3.

# What it deliberately does not claim

NFR-P7 is the whole chain, and only the encode half of it runs without an
adapter. So the export row is a one-sided gate — over two seconds in the encode
alone violates the requirement; under it proves nothing — and there is no
TRACES: NFR-P7 anywhere. NFR-P8 is about the application at idle, and the probe
is a process holding the catalog and nothing else, so it records the catalog
layer's share and carries no budget until somebody decides what that share
should be. No tag there either. CONTRIBUTING.md asks that a requirement be
closed by a test that would fail if the behaviour were removed, and two more
plumbing tags is what this repository already has too many of.

NFR-P8 also gets the answer §4.1 demands: RSS is exclusive of device-local GPU
allocations and cannot be made otherwise, because such an allocation never
enters the process's address space. The requirement should be restated as two
figures, and docs/benchmarks.md says so.

# Two gates, and why one of them steps aside off the reference desktop

The budget is the requirement's own number and never moves. The baseline is
what the reference desktop last measured, and drifting 15% past it fails the
build even while still inside the budget — which is how performance rot
actually arrives, never over the line, always a little worse.

A budget written for twenty-four threads cannot be asserted on a two-core
container. §8 names the reference desktop, not CI, so each metric declares
whether its budget is machine-sensitive; those are asserted under --reference
and reported everywhere else. Catalog open is not one of them: two seconds
against an expected figure two orders of magnitude smaller is a threshold any
machine can be held to. This is the trap core/dr-gpu/tests/frame_budget.rs
already refuses — a red gate everybody learns to ignore.

# The baseline ships with no numbers in it

Every recorded field is null, because nobody has run it yet. Writing
plausible-looking figures would make every later comparison a comparison
against a guess, and the first real regression would be invisible. Run
`dr-bench record --reference` on the reference desktop and commit the diff;
until then the budget gate works and the report says the other one cannot.

# CI

.gitea/workflows/benchmark.yml, and its own workflow rather than a step in
build-and-test.yml: a red "Build and test" says the code is wrong, a red
"Benchmarks" says it got slower, and the second must not be reachable by
retrying a flaky compile. The cpu job runs on every push and builds -p dr-bench
alone — which is why that crate depends on no GPU and no UI crate. The gpu job
is the frame budget that already exists and already skips without an adapter,
on workflow_dispatch, because building wgpu on every commit to rediscover that
the runner has no device is not a use of anybody's minutes.
This commit is contained in:
2026-08-30 10:40:10 +02:00
parent ef07e6ca3e
commit e8e96eed40
16 changed files with 2821 additions and 21 deletions
+179
View File
@@ -0,0 +1,179 @@
//! Idle memory with a 50k catalog open — and the question NFR-P8 leaves open.
//!
//! # The question §4.1 asks, answered
//!
//! §4.1 says of NFR-P8: *"must state whether it measures RSS inclusive or
//! exclusive of GPU allocations, and whether it holds after SQLite's page cache
//! warms on a 50k catalog."* Both halves have an answer, and neither is
//! flattering.
//!
//! **On GPU memory: what this reports is RSS, and RSS is exclusive of
//! device-local GPU allocations.** A Vulkan allocation in device-local heap
//! never enters the process's address space, so no counter under
//! `/proc/self/status` can see it; what *does* land in RSS is the host-visible
//! side — staging buffers, mapped upload rings, the read-back `AdjustPass`
//! performs on export — and the driver's own resident pages. So "RSS < 500 MB"
//! is not one budget, it is two questions wearing one number, and a build that
//! kept RSS at 400 MB while holding 3 GB of textures would pass it.
//!
//! The recommendation this measurement exists to support: **NFR-P8 should be
//! restated as two figures** — host RSS exclusive of device-local memory, and
//! a separate VRAM ceiling read from the adapter — because the second is the
//! one that decides whether the application survives beside a browser on an
//! 8 GB card, and nothing in this repository currently measures it.
//!
//! **On the page cache: warm.** The probe runs the queries before it reads the
//! counter, so SQLite's page cache holds the b-tree pages a grid scroll
//! touches. That is the right side to err on — a figure taken before the cache
//! warms would understate a steady-state library — and it is why the probe
//! scrolls rather than opening and stopping.
//!
//! # Why this is a subprocess
//!
//! RSS is a high-water-influenced property of a *process*, not of a function.
//! Building a 50k fixture allocates hundreds of megabytes; decoding thumbnails
//! allocates more; the allocator returns some of it to the OS and keeps the
//! rest. Measuring after any of that would report the harness's history rather
//! than the catalog's cost. So the probe is a fresh process that opens the
//! catalog, does the grid's work, reads its own counters and exits.
//!
//! # What this cannot certify, said plainly
//!
//! Not NFR-P8. The requirement is about the *application* at idle — Slint, the
//! wgpu device, the font stack, the decode cache and the catalog together — and
//! this process contains only the last of those. No requirement tag in this
//! crate names NFR-P8, for that reason — and see `exporting.rs` for why that
//! sentence avoids spelling the tag out.
//!
//! What it is, is the catalog layer's share, measured rather than guessed. The
//! decision NFR-P8 actually needs — how much of the 500 MB belongs to the
//! catalog and how much to everything above it — is a decision somebody has to
//! take, and taking it against a recorded number is better than taking it
//! against an estimate. That is what this records. Until it is taken, the
//! metric carries no budget and gates only against its own baseline.
use std::path::Path;
use anyhow::Result;
use dr_catalog::{Catalog, Granularity, Query};
/// Rows fetched per window while the probe scrolls. The same 400
/// [`crate::catalog_open`] uses, for the same reason.
const WINDOW: usize = 400;
/// Windows the probe pages through before reading the counter.
///
/// Twenty-five is ten thousand rows: enough that SQLite's page cache holds a
/// realistic working set and that any per-window leak would be visible, and
/// far short of the whole library, which FR-CAT-4 forbids holding anyway.
const WINDOWS: usize = 25;
/// A fixed clock, for the reason `catalog_open`'s `NOW` gives: nothing here
/// should depend on the day it runs.
const NOW: i64 = 2_000_000_000;
/// Resident memory, in kilobytes, as Linux reports it.
#[derive(Debug, Clone, Copy)]
pub struct Rss {
/// `VmRSS`: resident now.
pub now_kb: u64,
/// `VmHWM`: the peak this process reached. Reported alongside because a
/// process that touched 900 MB and gave it back is not idling at 200 MB in
/// any sense a user would recognise — the pages came from somewhere.
pub peak_kb: u64,
}
impl Rss {
pub fn now_mb(&self) -> f64 {
self.now_kb as f64 / 1024.0
}
pub fn peak_mb(&self) -> f64 {
self.peak_kb as f64 / 1024.0
}
}
/// Read this process's own counters.
///
/// `None` anywhere without a Linux-shaped `/proc` — including Android, where
/// the file exists but a benchmark does not run, and macOS, where it does not.
/// Returning `None` rather than zero is deliberate: a memory figure of zero
/// would be reported as an excellent result.
pub fn of_this_process() -> Option<Rss> {
let status = std::fs::read_to_string("/proc/self/status").ok()?;
let mut now = None;
let mut peak = None;
for line in status.lines() {
if let Some(rest) = line.strip_prefix("VmRSS:") {
now = rest.split_whitespace().next()?.parse::<u64>().ok();
} else if let Some(rest) = line.strip_prefix("VmHWM:") {
peak = rest.split_whitespace().next()?.parse::<u64>().ok();
}
}
Some(Rss {
now_kb: now?,
peak_kb: peak?,
})
}
/// The probe: open the catalog, do what the grid does, print the counters.
///
/// Stdout is one line of `key=value` pairs rather than JSON, because the only
/// reader is [`in_a_fresh_process`] and a format a human can read in a log is
/// worth more here than one a parser prefers.
pub fn probe(catalog_path: &Path) -> Result<()> {
let catalog = Catalog::open(catalog_path)
.map_err(|e| anyhow::anyhow!("opening {} : {e}", catalog_path.display()))?;
let q = Query::default();
let images = catalog.count(&q, NOW)?;
let buckets = catalog.timeline(&q, Granularity::Month, NOW)?.len();
// Scroll, keeping only the window in hand — which is what the grid does,
// and what FR-CAT-4 requires it to do. If this ever starts costing memory
// proportional to how far the user scrolled, that is the bug this figure
// exists to catch.
let mut rows = 0usize;
let span = images.saturating_sub(WINDOW).max(1);
for i in 0..WINDOWS {
let start = (i * span) / WINDOWS.max(1);
rows = catalog.window(&q, start..start + WINDOW, NOW)?.len();
}
let Some(rss) = of_this_process() else {
anyhow::bail!("no /proc/self/status on this platform; RSS cannot be read");
};
println!(
"rss_kb={} peak_kb={} images={images} buckets={buckets} last_window={rows}",
rss.now_kb, rss.peak_kb
);
Ok(())
}
/// Run [`probe`] in a fresh copy of this executable and read back its counters.
pub fn in_a_fresh_process(catalog_path: &Path) -> Result<Rss> {
let exe = std::env::current_exe()?;
let output = std::process::Command::new(&exe)
.arg("memory-probe")
.arg(catalog_path)
.output()?;
if !output.status.success() {
anyhow::bail!(
"the memory probe exited with {}: {}",
output.status,
String::from_utf8_lossy(&output.stderr).trim()
);
}
let text = String::from_utf8_lossy(&output.stdout);
let field = |key: &str| -> Option<u64> {
text.split_whitespace()
.find_map(|pair| pair.strip_prefix(key))
.and_then(|v| v.parse::<u64>().ok())
};
let (Some(now_kb), Some(peak_kb)) = (field("rss_kb="), field("peak_kb=")) else {
anyhow::bail!("the memory probe printed something unreadable: {}", text.trim());
};
Ok(Rss { now_kb, peak_kb })
}