docs/requirements.md §8 has said since it was written that performance is verified by "an automated benchmark suite against a synthetic 50k catalog, run per-commit … A regression beyond stated tolerance fails the build." There was none. No benches/, no [[bench]], no criterion, no synthetic catalog, and three CI workflows that between them measured nothing. Ten performance requirements could therefore be neither passed nor failed, and five of them carried a TRACES: tag regardless. tools/bench is the half of that promise that can be kept honestly on a runner with no GPU and no display. # The fixture Rows are cheap and pixels are not, so it builds fifty thousand catalog rows over a pool of a dozen real files, each referenced by several thousand of them. Everything the catalog half touches is rows and is exact at full scale; everything the pixel half touches is one file at a time and does not care how many rows point at it. Fourteen megabytes on disk instead of two terabytes, and neither half is flattered by the trade. It is reproducible from a seed, and a stamp beside it — seed, row count, source size, dr-catalog's schema version — rebuilds it rather than letting a run be compared against a baseline that describes a different library. # What it can now pass or fail NFR-P1, and R2's second sentence with it: Catalog::open plus the count, first window and timeline the grid cannot paint without. The interesting part turned out to be the open itself — schema::backfill runs three passes over the images table on every open, which is O(library) work on a path whose budget is stated in absolute seconds. Tagged TRACES: NFR-P1, on a gate that fails if it breaks. NFR-P3: thumbnail throughput on the embedded preview path, through the same per-image work spawn_thumbnail_sweep does and in the same shape — chunks of 96, lanes owning disjoint slices, the single thread that owns the store writing the finished chunk. Mirrored rather than called, because that function takes a RemoteBackend and would measure somebody's network. Tagged TRACES: NFR-P3. # What it deliberately does not claim NFR-P7 is the whole chain, and only the encode half of it runs without an adapter. So the export row is a one-sided gate — over two seconds in the encode alone violates the requirement; under it proves nothing — and there is no TRACES: NFR-P7 anywhere. NFR-P8 is about the application at idle, and the probe is a process holding the catalog and nothing else, so it records the catalog layer's share and carries no budget until somebody decides what that share should be. No tag there either. CONTRIBUTING.md asks that a requirement be closed by a test that would fail if the behaviour were removed, and two more plumbing tags is what this repository already has too many of. NFR-P8 also gets the answer §4.1 demands: RSS is exclusive of device-local GPU allocations and cannot be made otherwise, because such an allocation never enters the process's address space. The requirement should be restated as two figures, and docs/benchmarks.md says so. # Two gates, and why one of them steps aside off the reference desktop The budget is the requirement's own number and never moves. The baseline is what the reference desktop last measured, and drifting 15% past it fails the build even while still inside the budget — which is how performance rot actually arrives, never over the line, always a little worse. A budget written for twenty-four threads cannot be asserted on a two-core container. §8 names the reference desktop, not CI, so each metric declares whether its budget is machine-sensitive; those are asserted under --reference and reported everywhere else. Catalog open is not one of them: two seconds against an expected figure two orders of magnitude smaller is a threshold any machine can be held to. This is the trap core/dr-gpu/tests/frame_budget.rs already refuses — a red gate everybody learns to ignore. # The baseline ships with no numbers in it Every recorded field is null, because nobody has run it yet. Writing plausible-looking figures would make every later comparison a comparison against a guess, and the first real regression would be invisible. Run `dr-bench record --reference` on the reference desktop and commit the diff; until then the budget gate works and the report says the other one cannot. # CI .gitea/workflows/benchmark.yml, and its own workflow rather than a step in build-and-test.yml: a red "Build and test" says the code is wrong, a red "Benchmarks" says it got slower, and the second must not be reachable by retrying a flaky compile. The cpu job runs on every push and builds -p dr-bench alone — which is why that crate depends on no GPU and no UI crate. The gpu job is the frame budget that already exists and already skips without an adapter, on workflow_dispatch, because building wgpu on every commit to rediscover that the runner has no device is not a use of anybody's minutes.
244 lines
12 KiB
Markdown
244 lines
12 KiB
Markdown
# The benchmark suite
|
||
|
||
**Status:** Built, not yet recorded · 2026-08-30
|
||
**Companion to:** [requirements.md](requirements.md) §4.1 (performance targets) · §8 (verification)
|
||
**Instrument:** [`tools/bench`](../tools/bench) — `cargo run --release -p dr-bench -- check`
|
||
**Committed numbers:** [`bench-baseline.json`](bench-baseline.json)
|
||
**GPU half:** [`core/dr-gpu/tests/frame_budget.rs`](../core/dr-gpu/tests/frame_budget.rs) ·
|
||
[frame-budget.md](frame-budget.md)
|
||
|
||
§8 has said since it was written that performance is verified by *"an automated
|
||
benchmark suite against a synthetic 50k catalog, run per-commit … A regression
|
||
beyond stated tolerance fails the build."* Until this suite there was none. No
|
||
`benches/`, no `[[bench]]`, no criterion, no fixture — and ten performance
|
||
requirements that could therefore be neither passed nor failed, five of them
|
||
carrying a `TRACES:` tag regardless.
|
||
|
||
This file is what the suite covers, what it deliberately does not, and how to
|
||
read a failure.
|
||
|
||
---
|
||
|
||
## The state of it, first
|
||
|
||
**No numbers have been recorded yet.** Every `recorded` field in
|
||
[`bench-baseline.json`](bench-baseline.json) is `null`, on purpose: writing
|
||
plausible-looking figures into a baseline would make every later comparison a
|
||
comparison against a guess, and the first real regression would be invisible.
|
||
|
||
To record them, on the reference desktop:
|
||
|
||
```sh
|
||
cargo run --release -p dr-bench -- record --reference
|
||
```
|
||
|
||
and commit the diff. Until that happens the **budget** gate works — a catalog
|
||
that takes three seconds to open fails the build today — and the **regression**
|
||
gate reports that it has nothing to compare against, rather than pretending.
|
||
|
||
---
|
||
|
||
## What it measures
|
||
|
||
| Metric | Requirement | Gated? |
|
||
|---|---|---|
|
||
| `catalog_open_ms` | **NFR-P1**, and R2's second sentence | Yes, everywhere — budget 2000 ms |
|
||
| `catalog_open_warm_ms` | NFR-P1, page cache warm | Yes, everywhere — budget 2000 ms |
|
||
| `catalog_window_p99_ms` | FR-CAT-4 | Regression only |
|
||
| `catalog_filtered_ms` | FR-CAT-6 | Regression only |
|
||
| `thumbnail_throughput_ips` | **NFR-P3** | Budget 100 img/s, on the reference desktop |
|
||
| `thumbnail_per_image_p99_ms` | NFR-P3 | Regression only |
|
||
| `export_24mp_original_ms` | NFR-P7, **encode half only** | One-sided: can fail it, cannot pass it |
|
||
| `export_24mp_long_edge_2048_ms` | FR-EXP-3 | Regression only |
|
||
| `catalog_idle_rss_mb` | NFR-P8, **catalog layer only** | Regression only — see below |
|
||
|
||
Two of those rows carry a qualifier, and the qualifiers are the point.
|
||
|
||
### Requirements this can now pass *or* fail
|
||
|
||
**NFR-P1 — catalog open under 2 s.** The measured span is the four things the
|
||
library view cannot paint without: `Catalog::open` (which connects, migrates and
|
||
**backfills**, and the backfill is three passes over the images table on every
|
||
open), `count`, the first 400-row `window`, and the monthly `timeline`. Tagged
|
||
`TRACES: NFR-P1` in [`tools/bench/src/catalog_open.rs`](../tools/bench/src/catalog_open.rs),
|
||
because a build that breaks it fails this gate.
|
||
|
||
**NFR-P3 — ≥ 100 images per second on the embedded preview path.** The
|
||
per-image work is exactly what `spawn_thumbnail_sweep` does — `decode_jpeg`,
|
||
`Preview::downscale_to`, `Preview::apply_orientation`, `encode_rgba`,
|
||
`ThumbStore::put` — arranged in the same shape: chunks of 96, lanes owning
|
||
disjoint slices, and the single thread that owns the store writing the finished
|
||
chunk. Tagged `TRACES: NFR-P3` in
|
||
[`tools/bench/src/thumbnails.rs`](../tools/bench/src/thumbnails.rs).
|
||
|
||
### Requirements this can only half-answer, and is not tagged for
|
||
|
||
**NFR-P7 — 24 MP export under 2 s, full chain.** The full chain is decode,
|
||
demosaic, a full-resolution GPU render, a read-back, then resize, sharpen and
|
||
encode. Only the last three run without an adapter. So the figure here is a
|
||
**lower bound** on the requirement: exceeding 2 s in the encode alone violates
|
||
NFR-P7 no matter how fast the render is, and coming in under it proves nothing.
|
||
The budget is gated on that basis and there is no `TRACES: NFR-P7` anywhere in
|
||
`tools/bench`.
|
||
|
||
**NFR-P8 — idle memory under 500 MB.** The probe is a fresh process holding the
|
||
catalog and nothing else: no Slint, no wgpu device, no font stack, no decode
|
||
cache. Its RSS is the catalog layer's *share* of that 500 MB, not the figure the
|
||
requirement is about. It carries no budget for a reason given below.
|
||
|
||
### Requirements out of scope, listed so their absence reads as a decision
|
||
|
||
NFR-P2 (grid scroll at 60 fps), P4 (open in develop), P5 (slider to visible),
|
||
P6 (pan/zoom), P9 (UI-executor blocking), P10 (touch response), P11 (layout
|
||
transition), P12 (warm shader setup), P13 (next image in culling), P14 (focus
|
||
peaking), P15 (drawn mask stroke). Every one of them needs a frame-timing probe
|
||
inside a running Slint application, a GPU adapter, or both. None is faked here.
|
||
|
||
The GPU half of the story that *does* exist is
|
||
[frame-budget.md](frame-budget.md) and its guard test, which asserts FR-DSP-3
|
||
and skips itself where there is no adapter. `.gitea/workflows/benchmark.yml`
|
||
runs it as its own job for exactly that reason.
|
||
|
||
---
|
||
|
||
## The fixture
|
||
|
||
Fifty thousand rows over a pool of twelve real image files. Rows are cheap and
|
||
pixels are not: everything the catalog half touches is rows and is therefore
|
||
exact at full scale, and everything the pixel half touches is one file at a time
|
||
and does not care how many rows point at it. The result is ~14 MB on disk
|
||
instead of ~2 TB, and neither half is flattered by that.
|
||
|
||
| | |
|
||
|---|---|
|
||
| Rows | 50,000 images, 50,000 default versions, 400 folders, one root |
|
||
| Capture times | Twelve years from a fixed epoch, so the timeline has ~144 monthly buckets |
|
||
| Sources | 12 synthesised JPEGs at 1620 × 1080 — the size `dr-decode` records a CR2 carrying in IFD2 |
|
||
| Seed | 20260829, in [`tools/bench/src/main.rs`](../tools/bench/src/main.rs) |
|
||
| Location | `$DR_BENCH_DIR`, else the system temporary directory |
|
||
|
||
It is reproducible from the seed, and a `stamp.json` beside it records what it
|
||
was built from — seed, row count, source count, preview size, and `dr-catalog`'s
|
||
schema version. A mismatch rebuilds rather than silently measuring a different
|
||
workload than the baseline describes.
|
||
|
||
Two honest limits on it:
|
||
|
||
- **The page cache is warm.** The fixture was written by this suite or by an
|
||
earlier run of it, so neither the catalog open nor the thumbnail sweep pays
|
||
for a cold disk. On the reference desktop's NVMe a genuinely cold read of a
|
||
14 MB catalog is tens of milliseconds; on spinning rust it is not.
|
||
- **The sources are synthetic.** A coarse gradient with a fine dither, which is
|
||
what `frame_budget.rs` synthesises for the same reason — a flat frame lets the
|
||
memory system serve every sample from one cache line and flatters a box
|
||
filter, and pure noise defeats the entropy coder in the other direction.
|
||
|
||
---
|
||
|
||
## Two gates, and how to read a failure
|
||
|
||
**Budget.** The requirement's own threshold. It does not move. Failing it means
|
||
a requirement is violated.
|
||
|
||
**Regression.** More than 15% worse than the last recorded figure *on the same
|
||
machine, against the same fixture*. Failing it means the code got slower while
|
||
still inside the requirement — which is how most performance rot actually
|
||
arrives, never over the line, always a little worse, until one day the line is
|
||
crossed by a change that was not the cause.
|
||
|
||
A metric declares whether its budget is `machine_sensitive`. Those are asserted
|
||
only under `--reference`, and reported everywhere else. §8 names *"the reference
|
||
desktop"*, not CI, and it is right to: a container with two cores cannot speak
|
||
to a throughput target written for twenty-four threads, and asserting one there
|
||
would produce exactly what `core/dr-gpu/tests/frame_budget.rs` refused to
|
||
produce — *"a red suite that everyone learns to ignore"*. Catalog open is not
|
||
machine-sensitive: 2 s against an expected figure two orders of magnitude
|
||
smaller is a threshold any machine can be held to.
|
||
|
||
Exit codes: `0` everything passed, `1` a gate failed, `2` the harness itself
|
||
could not run. Distinguished so a CI log that says "failed" does not leave
|
||
anyone guessing whether the code got slower or the fixture would not build.
|
||
|
||
**Release, always.** The workspace builds its own crates at `opt-level = 0` in
|
||
dev, and every figure here is dominated by this workspace's own code — the JPEG
|
||
decode, the box filter, the resample, the sharpen. A debug run measures rustc's
|
||
shadow. The report says which profile it was built in on its second line.
|
||
|
||
---
|
||
|
||
## NFR-P8, and the question §4.1 asks
|
||
|
||
§4.1 says NFR-P8 *"must state whether it measures RSS inclusive or exclusive of
|
||
GPU allocations, and whether it holds after SQLite's page cache warms on a 50k
|
||
catalog."* Both halves have an answer.
|
||
|
||
**On the page cache: warm.** The probe runs the count, the timeline and
|
||
twenty-five windows before it reads its counters, so SQLite's cache holds the
|
||
b-tree pages a scroll touches. That is the right side to err on — a figure taken
|
||
before the cache warms would understate a steady-state library.
|
||
|
||
**On GPU memory: RSS is exclusive of device-local allocations, and cannot be
|
||
made otherwise.** A Vulkan allocation in a device-local heap never enters the
|
||
process's address space, so nothing under `/proc/self/status` can see it. What
|
||
*does* land in RSS is the host-visible side — staging buffers, mapped upload
|
||
rings, the read-back `AdjustPass` performs on export — plus the driver's own
|
||
resident pages.
|
||
|
||
So "idle memory < 500 MB" is two questions wearing one number, and a build
|
||
holding 400 MB of RSS and 3 GB of textures would pass it.
|
||
|
||
**Recommendation: NFR-P8 should be restated as two figures** — host RSS
|
||
exclusive of device-local memory, and a separate VRAM ceiling read from the
|
||
adapter — because the second is the one that decides whether the application
|
||
survives beside a browser on an 8 GB card, and nothing in this repository
|
||
measures it today.
|
||
|
||
**And a decision is outstanding.** `catalog_idle_rss_mb` carries no budget
|
||
because nobody has decided how much of the 500 MB belongs to the catalog layer
|
||
and how much to everything above it. The suite records the number so that
|
||
decision can be taken against a measurement rather than an estimate. When it is
|
||
taken, put the figure in `budget` and the metric becomes a gate.
|
||
|
||
---
|
||
|
||
## What is not measured, and would be worth adding
|
||
|
||
- **The UI's own open.** `ui/dr-ui/src/library.rs` does not call
|
||
`Catalog::count` or `Catalog::window`; it issues its own SQL against the same
|
||
tables, with a `VISIBLE` predicate and a burst-folding clause. `dr-bench`
|
||
cannot see those without depending on `dr-ui`, which would drag Slint into a
|
||
job that has no display. **Falsifiable end:** when the grid's queries move
|
||
down into `dr-catalog` — which is where SQL over catalog tables belongs —
|
||
`catalog_open_ms` becomes the whole of the application's open and this caveat
|
||
can be deleted rather than argued about.
|
||
- **The remote sweep.** `spawn_thumbnail_sweep`'s wall clock against a real
|
||
server is latency, not CPU, and is what FR-NC-3's design is judged by. It
|
||
needs a server and belongs in a different kind of test.
|
||
- **A cold disk.** See the fixture's limits above.
|
||
- **Android.** §4.1 states a second column of targets and §8 asks for
|
||
"periodically on the named reference Android devices". Nothing here runs on a
|
||
device. Spike S10 is the piece of work that would start it.
|
||
- **Everything with a frame in it.** See the out-of-scope list above.
|
||
|
||
---
|
||
|
||
## Running it
|
||
|
||
```sh
|
||
# Measure and print. Judges nothing.
|
||
cargo run --release -p dr-bench -- run
|
||
|
||
# Measure and gate. What CI runs.
|
||
cargo run --release -p dr-bench -- check
|
||
|
||
# The same, with machine-sensitive budgets asserted too.
|
||
cargo run --release -p dr-bench -- check --reference
|
||
|
||
# Rewrite bench-baseline.json from this run, and commit the diff.
|
||
cargo run --release -p dr-bench -- record --reference
|
||
```
|
||
|
||
Useful flags: `--fixture <dir>` (or `$DR_BENCH_DIR`) to put the synthetic
|
||
catalog somewhere specific, `--lanes <n>` to pin the sweep's parallelism, and
|
||
`--thumbnails <n>` to lengthen or shorten the throughput row.
|