Files
dtourolle 1b8d0e740f Say the benchmark's second open no longer pays the backfill
benchmarks.md and catalog_open.rs both said schema::backfill runs on
every Catalog::open. Since ffdd640 it runs on the first open of a path
in a process and is skipped while the stamp matches, so in dr-bench
catalog_open_ms still includes it and catalog_open_warm_ms, the second
open in the same process, no longer does. That is what a library
reopened in one session costs, and both now say which figure is which.
The comment keeps its line count, so no tag below it moves.
2026-09-26 14:58:20 -04:00

12 KiB
Raw Permalink Blame History

The benchmark suite

Status: Built, not yet recorded · 2026-08-30 Companion to: requirements.md §4.1 (performance targets) · §8 (verification) Instrument: tools/bench — cargo run --release -p dr-bench -- check Committed numbers: bench-baseline.json GPU half: core/dr-gpu/tests/frame_budget.rs · frame-budget.md

§8 has said since it was written that performance is verified by "an automated benchmark suite against a synthetic 50k catalog, run per-commit … A regression beyond stated tolerance fails the build." Until this suite there was none. No benches/, no [[bench]], no criterion, no fixture — and ten performance requirements that could therefore be neither passed nor failed, five of them carrying a TRACES: tag regardless.

This file is what the suite covers, what it deliberately does not, and how to read a failure.


The state of it, first

No numbers have been recorded yet. Every recorded field in bench-baseline.json is null, on purpose: writing plausible-looking figures into a baseline would make every later comparison a comparison against a guess, and the first real regression would be invisible.

To record them, on the reference desktop:

cargo run --release -p dr-bench -- record --reference

and commit the diff. Until that happens the budget gate works — a catalog that takes three seconds to open fails the build today — and the regression gate reports that it has nothing to compare against, rather than pretending.


What it measures

Metric Requirement Gated?
catalog_open_ms NFR-P1, and R2's second sentence Yes, everywhere — budget 2000 ms
catalog_open_warm_ms NFR-P1, page cache warm Yes, everywhere — budget 2000 ms
catalog_window_p99_ms FR-CAT-4 Regression only
catalog_filtered_ms FR-CAT-6 Regression only
thumbnail_throughput_ips NFR-P3 Budget 100 img/s, on the reference desktop
thumbnail_per_image_p99_ms NFR-P3 Regression only
export_24mp_original_ms NFR-P7, encode half only One-sided: can fail it, cannot pass it
export_24mp_long_edge_2048_ms FR-EXP-3 Regression only
catalog_idle_rss_mb NFR-P8, catalog layer only Regression only — see below

Two of those rows carry a qualifier, and the qualifiers are the point.

Requirements this can now pass or fail

NFR-P1 — catalog open under 2 s. The measured span is the four things the library view cannot paint without: Catalog::open (which connects, migrates and backfills, and the backfill is passes over the whole images table), count, the first 400-row window, and the monthly timeline. Since 0.17.0 the backfill runs on the first open of a catalog in a process and is skipped by later ones while nothing has changed (catalog.md §2), so catalog_open_ms includes it and catalog_open_warm_ms, the second open in the same process, does not — which is what a library reopened in the same session costs. Tagged TRACES: NFR-P1 in tools/bench/src/catalog_open.rs, because a build that breaks it fails this gate.

NFR-P3 — ≥ 100 images per second on the embedded preview path. The per-image work is exactly what spawn_thumbnail_sweep does — decode_jpeg, Preview::downscale_to, Preview::apply_orientation, encode_rgba, ThumbStore::put — arranged in the same shape: chunks of 96, lanes owning disjoint slices, and the single thread that owns the store writing the finished chunk. Tagged TRACES: NFR-P3 in tools/bench/src/thumbnails.rs.

Requirements this can only half-answer, and is not tagged for

NFR-P7 — 24 MP export under 2 s, full chain. The full chain is decode, demosaic, a full-resolution GPU render, a read-back, then resize, sharpen and encode. Only the last three run without an adapter. So the figure here is a lower bound on the requirement: exceeding 2 s in the encode alone violates NFR-P7 no matter how fast the render is, and coming in under it proves nothing. The budget is gated on that basis and there is no TRACES: NFR-P7 anywhere in tools/bench.

NFR-P8 — idle memory under 500 MB. The probe is a fresh process holding the catalog and nothing else: no Slint, no wgpu device, no font stack, no decode cache. Its RSS is the catalog layer's share of that 500 MB, not the figure the requirement is about. It carries no budget for a reason given below.

Requirements out of scope, listed so their absence reads as a decision

NFR-P2 (grid scroll at 60 fps), P4 (open in develop), P5 (slider to visible), P6 (pan/zoom), P9 (UI-executor blocking), P10 (touch response), P11 (layout transition), P12 (warm shader setup), P13 (next image in culling), P14 (focus peaking), P15 (drawn mask stroke). Every one of them needs a frame-timing probe inside a running Slint application, a GPU adapter, or both. None is faked here.

The GPU half of the story that does exist is frame-budget.md and its guard test, which asserts FR-DSP-3 and skips itself where there is no adapter. .gitea/workflows/benchmark.yml runs it as its own job for exactly that reason.


The fixture

Fifty thousand rows over a pool of twelve real image files. Rows are cheap and pixels are not: everything the catalog half touches is rows and is therefore exact at full scale, and everything the pixel half touches is one file at a time and does not care how many rows point at it. The result is ~14 MB on disk instead of ~2 TB, and neither half is flattered by that.

Rows 50,000 images, 50,000 default versions, 400 folders, one root
Capture times Twelve years from a fixed epoch, so the timeline has ~144 monthly buckets
Sources 12 synthesised JPEGs at 1620 × 1080 — the size dr-decode records a CR2 carrying in IFD2
Seed 20260829, in tools/bench/src/main.rs
Location $DR_BENCH_DIR, else the system temporary directory

It is reproducible from the seed, and a stamp.json beside it records what it was built from — seed, row count, source count, preview size, and dr-catalog's schema version. A mismatch rebuilds rather than silently measuring a different workload than the baseline describes.

Two honest limits on it:

  • The page cache is warm. The fixture was written by this suite or by an earlier run of it, so neither the catalog open nor the thumbnail sweep pays for a cold disk. On the reference desktop's NVMe a genuinely cold read of a 14 MB catalog is tens of milliseconds; on spinning rust it is not.
  • The sources are synthetic. A coarse gradient with a fine dither, which is what frame_budget.rs synthesises for the same reason — a flat frame lets the memory system serve every sample from one cache line and flatters a box filter, and pure noise defeats the entropy coder in the other direction.

Two gates, and how to read a failure

Budget. The requirement's own threshold. It does not move. Failing it means a requirement is violated.

Regression. More than 15% worse than the last recorded figure on the same machine, against the same fixture. Failing it means the code got slower while still inside the requirement — which is how most performance rot actually arrives, never over the line, always a little worse, until one day the line is crossed by a change that was not the cause.

A metric declares whether its budget is machine_sensitive. Those are asserted only under --reference, and reported everywhere else. §8 names "the reference desktop", not CI, and it is right to: a container with two cores cannot speak to a throughput target written for twenty-four threads, and asserting one there would produce exactly what core/dr-gpu/tests/frame_budget.rs refused to produce — "a red suite that everyone learns to ignore". Catalog open is not machine-sensitive: 2 s against an expected figure two orders of magnitude smaller is a threshold any machine can be held to.

Exit codes: 0 everything passed, 1 a gate failed, 2 the harness itself could not run. Distinguished so a CI log that says "failed" does not leave anyone guessing whether the code got slower or the fixture would not build.

Release, always. The workspace builds its own crates at opt-level = 0 in dev, and every figure here is dominated by this workspace's own code — the JPEG decode, the box filter, the resample, the sharpen. A debug run measures rustc's shadow. The report says which profile it was built in on its second line.


NFR-P8, and the question §4.1 asks

§4.1 says NFR-P8 "must state whether it measures RSS inclusive or exclusive of GPU allocations, and whether it holds after SQLite's page cache warms on a 50k catalog." Both halves have an answer.

On the page cache: warm. The probe runs the count, the timeline and twenty-five windows before it reads its counters, so SQLite's cache holds the b-tree pages a scroll touches. That is the right side to err on — a figure taken before the cache warms would understate a steady-state library.

On GPU memory: RSS is exclusive of device-local allocations, and cannot be made otherwise. A Vulkan allocation in a device-local heap never enters the process's address space, so nothing under /proc/self/status can see it. What does land in RSS is the host-visible side — staging buffers, mapped upload rings, the read-back AdjustPass performs on export — plus the driver's own resident pages.

So "idle memory < 500 MB" is two questions wearing one number, and a build holding 400 MB of RSS and 3 GB of textures would pass it.

Recommendation: NFR-P8 should be restated as two figures — host RSS exclusive of device-local memory, and a separate VRAM ceiling read from the adapter — because the second is the one that decides whether the application survives beside a browser on an 8 GB card, and nothing in this repository measures it today.

And a decision is outstanding. catalog_idle_rss_mb carries no budget because nobody has decided how much of the 500 MB belongs to the catalog layer and how much to everything above it. The suite records the number so that decision can be taken against a measurement rather than an estimate. When it is taken, put the figure in budget and the metric becomes a gate.


What is not measured, and would be worth adding

  • The UI's own open. ui/dr-ui/src/library.rs does not call Catalog::count or Catalog::window; it issues its own SQL against the same tables, with a VISIBLE predicate and a burst-folding clause. dr-bench cannot see those without depending on dr-ui, which would drag Slint into a job that has no display. Falsifiable end: when the grid's queries move down into dr-catalog — which is where SQL over catalog tables belongs — catalog_open_ms becomes the whole of the application's open and this caveat can be deleted rather than argued about.
  • The remote sweep. spawn_thumbnail_sweep's wall clock against a real server is latency, not CPU, and is what FR-NC-3's design is judged by. It needs a server and belongs in a different kind of test.
  • A cold disk. See the fixture's limits above.
  • Android. §4.1 states a second column of targets and §8 asks for "periodically on the named reference Android devices". Nothing here runs on a device. Spike S10 is the piece of work that would start it.
  • Everything with a frame in it. See the out-of-scope list above.

Running it

# Measure and print. Judges nothing.
cargo run --release -p dr-bench -- run

# Measure and gate. What CI runs.
cargo run --release -p dr-bench -- check

# The same, with machine-sensitive budgets asserted too.
cargo run --release -p dr-bench -- check --reference

# Rewrite bench-baseline.json from this run, and commit the diff.
cargo run --release -p dr-bench -- record --reference

Useful flags: --fixture <dir> (or $DR_BENCH_DIR) to put the synthetic catalog somewhere specific, --lanes <n> to pin the sweep's parallelism, and --thumbnails <n> to lengthen or shorten the throughput row.