docs/requirements.md §8 has said since it was written that performance is verified by "an automated benchmark suite against a synthetic 50k catalog, run per-commit … A regression beyond stated tolerance fails the build." There was none. No benches/, no [[bench]], no criterion, no synthetic catalog, and three CI workflows that between them measured nothing. Ten performance requirements could therefore be neither passed nor failed, and five of them carried a TRACES: tag regardless. tools/bench is the half of that promise that can be kept honestly on a runner with no GPU and no display. # The fixture Rows are cheap and pixels are not, so it builds fifty thousand catalog rows over a pool of a dozen real files, each referenced by several thousand of them. Everything the catalog half touches is rows and is exact at full scale; everything the pixel half touches is one file at a time and does not care how many rows point at it. Fourteen megabytes on disk instead of two terabytes, and neither half is flattered by the trade. It is reproducible from a seed, and a stamp beside it — seed, row count, source size, dr-catalog's schema version — rebuilds it rather than letting a run be compared against a baseline that describes a different library. # What it can now pass or fail NFR-P1, and R2's second sentence with it: Catalog::open plus the count, first window and timeline the grid cannot paint without. The interesting part turned out to be the open itself — schema::backfill runs three passes over the images table on every open, which is O(library) work on a path whose budget is stated in absolute seconds. Tagged TRACES: NFR-P1, on a gate that fails if it breaks. NFR-P3: thumbnail throughput on the embedded preview path, through the same per-image work spawn_thumbnail_sweep does and in the same shape — chunks of 96, lanes owning disjoint slices, the single thread that owns the store writing the finished chunk. Mirrored rather than called, because that function takes a RemoteBackend and would measure somebody's network. Tagged TRACES: NFR-P3. # What it deliberately does not claim NFR-P7 is the whole chain, and only the encode half of it runs without an adapter. So the export row is a one-sided gate — over two seconds in the encode alone violates the requirement; under it proves nothing — and there is no TRACES: NFR-P7 anywhere. NFR-P8 is about the application at idle, and the probe is a process holding the catalog and nothing else, so it records the catalog layer's share and carries no budget until somebody decides what that share should be. No tag there either. CONTRIBUTING.md asks that a requirement be closed by a test that would fail if the behaviour were removed, and two more plumbing tags is what this repository already has too many of. NFR-P8 also gets the answer §4.1 demands: RSS is exclusive of device-local GPU allocations and cannot be made otherwise, because such an allocation never enters the process's address space. The requirement should be restated as two figures, and docs/benchmarks.md says so. # Two gates, and why one of them steps aside off the reference desktop The budget is the requirement's own number and never moves. The baseline is what the reference desktop last measured, and drifting 15% past it fails the build even while still inside the budget — which is how performance rot actually arrives, never over the line, always a little worse. A budget written for twenty-four threads cannot be asserted on a two-core container. §8 names the reference desktop, not CI, so each metric declares whether its budget is machine-sensitive; those are asserted under --reference and reported everywhere else. Catalog open is not one of them: two seconds against an expected figure two orders of magnitude smaller is a threshold any machine can be held to. This is the trap core/dr-gpu/tests/frame_budget.rs already refuses — a red gate everybody learns to ignore. # The baseline ships with no numbers in it Every recorded field is null, because nobody has run it yet. Writing plausible-looking figures would make every later comparison a comparison against a guess, and the first real regression would be invisible. Run `dr-bench record --reference` on the reference desktop and commit the diff; until then the budget gate works and the report says the other one cannot. # CI .gitea/workflows/benchmark.yml, and its own workflow rather than a step in build-and-test.yml: a red "Build and test" says the code is wrong, a red "Benchmarks" says it got slower, and the second must not be reachable by retrying a flaky compile. The cpu job runs on every push and builds -p dr-bench alone — which is why that crate depends on no GPU and no UI crate. The gpu job is the frame budget that already exists and already skips without an adapter, on workflow_dispatch, because building wgpu on every commit to rediscover that the runner has no device is not a use of anybody's minutes.
179 lines
7.6 KiB
Rust
179 lines
7.6 KiB
Rust
//! TRACES: NFR-P1
|
|
//! Opening a fifty-thousand-image catalog, and what that actually involves.
|
|
//!
|
|
//! NFR-P1 says under two seconds on the reference desktop. Until this file
|
|
//! existed the number had never been measured, which made it a wish — and
|
|
//! `dr-catalog`'s `lib.rs` has carried an `NFR-P1` tag the whole time on code
|
|
//! that describes the target rather than checking it. This is the check.
|
|
//!
|
|
//! # What counts as "open"
|
|
//!
|
|
//! Not `Catalog::open` alone. That call returns before anything is on screen,
|
|
//! and a user's "the catalog opened" is the moment the grid has cells in it.
|
|
//! So the measured span is the four things the library view cannot paint
|
|
//! without:
|
|
//!
|
|
//! 1. [`Catalog::open`] — connect, migrate if needed, and **backfill**. The
|
|
//! backfill is the interesting one: `schema::backfill` runs on every open
|
|
//! and is three passes over the images table, so it is O(library) work on a
|
|
//! path whose budget is stated in absolute seconds.
|
|
//! 2. [`Catalog::count`] — the total, which is what sizes the scrollbar.
|
|
//! 3. [`Catalog::window`] — the first screenful of rows.
|
|
//! 4. [`Catalog::timeline`] — the scrubber's buckets, drawn beside the grid
|
|
//! from the first frame.
|
|
//!
|
|
//! `open` alone is reported separately anyway, because if the two ever diverge
|
|
//! sharply the fix is in a different place.
|
|
//!
|
|
//! # What this does *not* measure, said out loud
|
|
//!
|
|
//! `ui/dr-ui/src/library.rs` does not call [`Catalog::count`] or
|
|
//! [`Catalog::window`]. It issues its own SQL — `total_images_scoped`,
|
|
//! `total_images_filtered` and friends — against the same tables, with a
|
|
//! `VISIBLE` predicate and a burst-folding clause that this crate cannot see
|
|
//! without depending on the UI, which would drag Slint into a benchmark job
|
|
//! that has no display. So the number here is the **catalog crate's** open
|
|
//! path, and the application's is that plus whatever those queries cost.
|
|
//!
|
|
//! That gap is a real limit on what this file can certify, and it has a
|
|
//! falsifiable end: when the grid's queries move down into `dr-catalog` —
|
|
//! which is where SQL over catalog tables belongs — this measurement becomes
|
|
//! the whole of the application's open, and the caveat can be deleted rather
|
|
//! than argued about.
|
|
//!
|
|
//! # Cold and warm
|
|
//!
|
|
//! Both are reported. The first open in a process pays for SQLite's page cache
|
|
//! being empty and for the schema being read; the second pays for neither, and
|
|
//! is what a user gets when they close and reopen a library in the same
|
|
//! session. §4.1 asks the question directly for NFR-P8 and it is worth having
|
|
//! the answer here too. Neither figure is a genuinely cold *disk*: the fixture
|
|
//! was written by this same suite or by an earlier run of it, so the file is
|
|
//! in the OS page cache. On the reference desktop's NVMe a truly cold read of
|
|
//! a ~14 MB file is a few tens of milliseconds; on a spinning disk it is not.
|
|
|
|
use std::path::Path;
|
|
use std::time::Instant;
|
|
|
|
use anyhow::Result;
|
|
use dr_catalog::{Catalog, Granularity, Query};
|
|
use dr_types::Selector;
|
|
|
|
use crate::stats::{ms, Percentiles, Rng};
|
|
|
|
/// Rows fetched for the first screenful.
|
|
///
|
|
/// A dense grid on a 4K display is around three hundred cells; four hundred is
|
|
/// that plus the prefetch margin R2 asks for. Not the whole library, because
|
|
/// FR-CAT-4 is explicit that memory must not scale with it — a benchmark that
|
|
/// asked for 50,000 rows would be measuring the requirement's violation.
|
|
const WINDOW: usize = 400;
|
|
|
|
/// The clock the query compiler is handed.
|
|
///
|
|
/// Fixed rather than read from the system, so a rolling date filter would
|
|
/// compile to the same SQL on every run. The unfiltered query does not consult
|
|
/// it at all; this is here so that adding a dated row later does not silently
|
|
/// make the suite time-dependent.
|
|
const NOW: i64 = 2_000_000_000;
|
|
|
|
/// What one measured open produced.
|
|
pub struct Open {
|
|
/// Open, count, first window, timeline — the whole span, first time.
|
|
pub cold_ms: f64,
|
|
/// The same four calls on a second connection in the same process.
|
|
pub warm_ms: f64,
|
|
/// [`Catalog::open`] on its own, out of the cold span.
|
|
pub open_only_ms: f64,
|
|
/// How many images the count found. Reported so a fixture that failed to
|
|
/// populate cannot masquerade as a very fast open.
|
|
pub images: usize,
|
|
/// Timeline buckets at monthly granularity.
|
|
pub buckets: usize,
|
|
/// One window fetched at a random offset — the scroll, minus the drawing.
|
|
pub window_ms: Percentiles,
|
|
/// Count plus first window under a rating filter, which compiles to a
|
|
/// correlated subquery over `versions` (see `query::default_version_scalar`).
|
|
pub filtered_ms: f64,
|
|
}
|
|
|
|
/// Measure an open of the catalog at `path`, then `windows` random windows.
|
|
pub fn measure(path: &Path, windows: usize) -> Result<Open> {
|
|
let q = Query::default();
|
|
|
|
let started = Instant::now();
|
|
let catalog = open(path)?;
|
|
let open_only_ms = ms(started.elapsed());
|
|
let images = catalog.count(&q, NOW)?;
|
|
let rows = catalog.window(&q, 0..WINDOW, NOW)?;
|
|
let buckets = catalog.timeline(&q, Granularity::Month, NOW)?.len();
|
|
let cold_ms = ms(started.elapsed());
|
|
|
|
// A catalog that returned nothing would post an excellent time. Checked
|
|
// rather than trusted, because the failure mode is a *fast* wrong answer.
|
|
anyhow::ensure!(
|
|
!rows.is_empty() && images > 0 && buckets > 0,
|
|
"the fixture catalog answered with {images} images, {} rows and {buckets} buckets — \
|
|
the measurement below would be meaningless",
|
|
rows.len()
|
|
);
|
|
drop(catalog);
|
|
|
|
let started = Instant::now();
|
|
let catalog = open(path)?;
|
|
let _ = catalog.count(&q, NOW)?;
|
|
let _ = catalog.window(&q, 0..WINDOW, NOW)?;
|
|
let _ = catalog.timeline(&q, Granularity::Month, NOW)?;
|
|
let warm_ms = ms(started.elapsed());
|
|
|
|
// The scroll. Offsets are drawn from a fixed seed rather than swept in
|
|
// order, because a sequential sweep would be answered increasingly out of
|
|
// SQLite's own cache and would flatter the deep end of the library — which
|
|
// is exactly the end a person reaches by dragging the scrollbar.
|
|
let mut rng = Rng::new(0x5C_20_11);
|
|
let span = images.saturating_sub(WINDOW).max(1) as u64;
|
|
// Discarded: the first window of a new connection compiles the statement
|
|
// and faults in the b-tree's upper levels, and neither recurs while
|
|
// scrolling.
|
|
for _ in 0..4 {
|
|
let start = rng.below(span) as usize;
|
|
let _ = catalog.window(&q, start..start + WINDOW, NOW)?;
|
|
}
|
|
let mut samples = Vec::with_capacity(windows);
|
|
for _ in 0..windows {
|
|
let start = rng.below(span) as usize;
|
|
let t = Instant::now();
|
|
let rows = catalog.window(&q, start..start + WINDOW, NOW)?;
|
|
samples.push(ms(t.elapsed()));
|
|
debug_assert!(!rows.is_empty());
|
|
}
|
|
|
|
// A filter that has to reach the default version for every candidate row.
|
|
// Cheap to add and the one query shape in the grid that is not a scan of
|
|
// `images` alone, so a regression in it would otherwise show up first as a
|
|
// user complaint.
|
|
let rated = Query {
|
|
filter: Selector::Rating { min: 2 },
|
|
..Query::default()
|
|
};
|
|
let t = Instant::now();
|
|
let _ = catalog.count(&rated, NOW)?;
|
|
let _ = catalog.window(&rated, 0..WINDOW, NOW)?;
|
|
let filtered_ms = ms(t.elapsed());
|
|
|
|
Ok(Open {
|
|
cold_ms,
|
|
warm_ms,
|
|
open_only_ms,
|
|
images,
|
|
buckets,
|
|
window_ms: Percentiles::of(samples),
|
|
filtered_ms,
|
|
})
|
|
}
|
|
|
|
fn open(path: &Path) -> Result<Catalog> {
|
|
Catalog::open(path)
|
|
.map_err(|e| anyhow::anyhow!("opening the fixture catalog at {}: {e}", path.display()))
|
|
}
|