docs/ had 26 developer documents flat beside the manual, and the two audiences are very differently sized: most readers want the manual and the gesture reference, a few want the register, the designs and the measurements. The manual and gestures.md stay at the top; everything for someone changing the code moves to docs/dev/, and the two documents that name their own successors — the v0.1 milestone and the UI-refinement plan — go to docs/dev/archive/ rather than being deleted, since both are still cited. docs/README.md is the index, users first. Every reference follows: code comments, Cargo manifests, the workflows, the pre-commit hook, the bench and traceability tools (which locate the repo root by docs/dev/requirements.md now), packaging, the Docker READMEs, CLAUDE.md, CONTRIBUTING.md and the README. The matrix links one level deeper and is regenerated. Links out of the moved documents into the tree gain a level; a link checker over every Markdown file finds none broken.
244 lines
12 KiB
Markdown
244 lines
12 KiB
Markdown
# The benchmark suite
|
||
|
||
**Status:** Built, not yet recorded · 2026-08-30
|
||
**Companion to:** [requirements.md](requirements.md) §4.1 (performance targets) · §8 (verification)
|
||
**Instrument:** [`tools/bench`](../../tools/bench) — `cargo run --release -p dr-bench -- check`
|
||
**Committed numbers:** [`bench-baseline.json`](bench-baseline.json)
|
||
**GPU half:** [`core/dr-gpu/tests/frame_budget.rs`](../../core/dr-gpu/tests/frame_budget.rs) ·
|
||
[frame-budget.md](frame-budget.md)
|
||
|
||
§8 has said since it was written that performance is verified by *"an automated
|
||
benchmark suite against a synthetic 50k catalog, run per-commit … A regression
|
||
beyond stated tolerance fails the build."* Until this suite there was none. No
|
||
`benches/`, no `[[bench]]`, no criterion, no fixture — and ten performance
|
||
requirements that could therefore be neither passed nor failed, five of them
|
||
carrying a `TRACES:` tag regardless.
|
||
|
||
This file is what the suite covers, what it deliberately does not, and how to
|
||
read a failure.
|
||
|
||
---
|
||
|
||
## The state of it, first
|
||
|
||
**No numbers have been recorded yet.** Every `recorded` field in
|
||
[`bench-baseline.json`](bench-baseline.json) is `null`, on purpose: writing
|
||
plausible-looking figures into a baseline would make every later comparison a
|
||
comparison against a guess, and the first real regression would be invisible.
|
||
|
||
To record them, on the reference desktop:
|
||
|
||
```sh
|
||
cargo run --release -p dr-bench -- record --reference
|
||
```
|
||
|
||
and commit the diff. Until that happens the **budget** gate works — a catalog
|
||
that takes three seconds to open fails the build today — and the **regression**
|
||
gate reports that it has nothing to compare against, rather than pretending.
|
||
|
||
---
|
||
|
||
## What it measures
|
||
|
||
| Metric | Requirement | Gated? |
|
||
|---|---|---|
|
||
| `catalog_open_ms` | **NFR-P1**, and R2's second sentence | Yes, everywhere — budget 2000 ms |
|
||
| `catalog_open_warm_ms` | NFR-P1, page cache warm | Yes, everywhere — budget 2000 ms |
|
||
| `catalog_window_p99_ms` | FR-CAT-4 | Regression only |
|
||
| `catalog_filtered_ms` | FR-CAT-6 | Regression only |
|
||
| `thumbnail_throughput_ips` | **NFR-P3** | Budget 100 img/s, on the reference desktop |
|
||
| `thumbnail_per_image_p99_ms` | NFR-P3 | Regression only |
|
||
| `export_24mp_original_ms` | NFR-P7, **encode half only** | One-sided: can fail it, cannot pass it |
|
||
| `export_24mp_long_edge_2048_ms` | FR-EXP-3 | Regression only |
|
||
| `catalog_idle_rss_mb` | NFR-P8, **catalog layer only** | Regression only — see below |
|
||
|
||
Two of those rows carry a qualifier, and the qualifiers are the point.
|
||
|
||
### Requirements this can now pass *or* fail
|
||
|
||
**NFR-P1 — catalog open under 2 s.** The measured span is the four things the
|
||
library view cannot paint without: `Catalog::open` (which connects, migrates and
|
||
**backfills**, and the backfill is three passes over the images table on every
|
||
open), `count`, the first 400-row `window`, and the monthly `timeline`. Tagged
|
||
`TRACES: NFR-P1` in [`tools/bench/src/catalog_open.rs`](../../tools/bench/src/catalog_open.rs),
|
||
because a build that breaks it fails this gate.
|
||
|
||
**NFR-P3 — ≥ 100 images per second on the embedded preview path.** The
|
||
per-image work is exactly what `spawn_thumbnail_sweep` does — `decode_jpeg`,
|
||
`Preview::downscale_to`, `Preview::apply_orientation`, `encode_rgba`,
|
||
`ThumbStore::put` — arranged in the same shape: chunks of 96, lanes owning
|
||
disjoint slices, and the single thread that owns the store writing the finished
|
||
chunk. Tagged `TRACES: NFR-P3` in
|
||
[`tools/bench/src/thumbnails.rs`](../../tools/bench/src/thumbnails.rs).
|
||
|
||
### Requirements this can only half-answer, and is not tagged for
|
||
|
||
**NFR-P7 — 24 MP export under 2 s, full chain.** The full chain is decode,
|
||
demosaic, a full-resolution GPU render, a read-back, then resize, sharpen and
|
||
encode. Only the last three run without an adapter. So the figure here is a
|
||
**lower bound** on the requirement: exceeding 2 s in the encode alone violates
|
||
NFR-P7 no matter how fast the render is, and coming in under it proves nothing.
|
||
The budget is gated on that basis and there is no `TRACES: NFR-P7` anywhere in
|
||
`tools/bench`.
|
||
|
||
**NFR-P8 — idle memory under 500 MB.** The probe is a fresh process holding the
|
||
catalog and nothing else: no Slint, no wgpu device, no font stack, no decode
|
||
cache. Its RSS is the catalog layer's *share* of that 500 MB, not the figure the
|
||
requirement is about. It carries no budget for a reason given below.
|
||
|
||
### Requirements out of scope, listed so their absence reads as a decision
|
||
|
||
NFR-P2 (grid scroll at 60 fps), P4 (open in develop), P5 (slider to visible),
|
||
P6 (pan/zoom), P9 (UI-executor blocking), P10 (touch response), P11 (layout
|
||
transition), P12 (warm shader setup), P13 (next image in culling), P14 (focus
|
||
peaking), P15 (drawn mask stroke). Every one of them needs a frame-timing probe
|
||
inside a running Slint application, a GPU adapter, or both. None is faked here.
|
||
|
||
The GPU half of the story that *does* exist is
|
||
[frame-budget.md](frame-budget.md) and its guard test, which asserts FR-DSP-3
|
||
and skips itself where there is no adapter. `.gitea/workflows/benchmark.yml`
|
||
runs it as its own job for exactly that reason.
|
||
|
||
---
|
||
|
||
## The fixture
|
||
|
||
Fifty thousand rows over a pool of twelve real image files. Rows are cheap and
|
||
pixels are not: everything the catalog half touches is rows and is therefore
|
||
exact at full scale, and everything the pixel half touches is one file at a time
|
||
and does not care how many rows point at it. The result is ~14 MB on disk
|
||
instead of ~2 TB, and neither half is flattered by that.
|
||
|
||
| | |
|
||
|---|---|
|
||
| Rows | 50,000 images, 50,000 default versions, 400 folders, one root |
|
||
| Capture times | Twelve years from a fixed epoch, so the timeline has ~144 monthly buckets |
|
||
| Sources | 12 synthesised JPEGs at 1620 × 1080 — the size `dr-decode` records a CR2 carrying in IFD2 |
|
||
| Seed | 20260829, in [`tools/bench/src/main.rs`](../../tools/bench/src/main.rs) |
|
||
| Location | `$DR_BENCH_DIR`, else the system temporary directory |
|
||
|
||
It is reproducible from the seed, and a `stamp.json` beside it records what it
|
||
was built from — seed, row count, source count, preview size, and `dr-catalog`'s
|
||
schema version. A mismatch rebuilds rather than silently measuring a different
|
||
workload than the baseline describes.
|
||
|
||
Two honest limits on it:
|
||
|
||
- **The page cache is warm.** The fixture was written by this suite or by an
|
||
earlier run of it, so neither the catalog open nor the thumbnail sweep pays
|
||
for a cold disk. On the reference desktop's NVMe a genuinely cold read of a
|
||
14 MB catalog is tens of milliseconds; on spinning rust it is not.
|
||
- **The sources are synthetic.** A coarse gradient with a fine dither, which is
|
||
what `frame_budget.rs` synthesises for the same reason — a flat frame lets the
|
||
memory system serve every sample from one cache line and flatters a box
|
||
filter, and pure noise defeats the entropy coder in the other direction.
|
||
|
||
---
|
||
|
||
## Two gates, and how to read a failure
|
||
|
||
**Budget.** The requirement's own threshold. It does not move. Failing it means
|
||
a requirement is violated.
|
||
|
||
**Regression.** More than 15% worse than the last recorded figure *on the same
|
||
machine, against the same fixture*. Failing it means the code got slower while
|
||
still inside the requirement — which is how most performance rot actually
|
||
arrives, never over the line, always a little worse, until one day the line is
|
||
crossed by a change that was not the cause.
|
||
|
||
A metric declares whether its budget is `machine_sensitive`. Those are asserted
|
||
only under `--reference`, and reported everywhere else. §8 names *"the reference
|
||
desktop"*, not CI, and it is right to: a container with two cores cannot speak
|
||
to a throughput target written for twenty-four threads, and asserting one there
|
||
would produce exactly what `core/dr-gpu/tests/frame_budget.rs` refused to
|
||
produce — *"a red suite that everyone learns to ignore"*. Catalog open is not
|
||
machine-sensitive: 2 s against an expected figure two orders of magnitude
|
||
smaller is a threshold any machine can be held to.
|
||
|
||
Exit codes: `0` everything passed, `1` a gate failed, `2` the harness itself
|
||
could not run. Distinguished so a CI log that says "failed" does not leave
|
||
anyone guessing whether the code got slower or the fixture would not build.
|
||
|
||
**Release, always.** The workspace builds its own crates at `opt-level = 0` in
|
||
dev, and every figure here is dominated by this workspace's own code — the JPEG
|
||
decode, the box filter, the resample, the sharpen. A debug run measures rustc's
|
||
shadow. The report says which profile it was built in on its second line.
|
||
|
||
---
|
||
|
||
## NFR-P8, and the question §4.1 asks
|
||
|
||
§4.1 says NFR-P8 *"must state whether it measures RSS inclusive or exclusive of
|
||
GPU allocations, and whether it holds after SQLite's page cache warms on a 50k
|
||
catalog."* Both halves have an answer.
|
||
|
||
**On the page cache: warm.** The probe runs the count, the timeline and
|
||
twenty-five windows before it reads its counters, so SQLite's cache holds the
|
||
b-tree pages a scroll touches. That is the right side to err on — a figure taken
|
||
before the cache warms would understate a steady-state library.
|
||
|
||
**On GPU memory: RSS is exclusive of device-local allocations, and cannot be
|
||
made otherwise.** A Vulkan allocation in a device-local heap never enters the
|
||
process's address space, so nothing under `/proc/self/status` can see it. What
|
||
*does* land in RSS is the host-visible side — staging buffers, mapped upload
|
||
rings, the read-back `AdjustPass` performs on export — plus the driver's own
|
||
resident pages.
|
||
|
||
So "idle memory < 500 MB" is two questions wearing one number, and a build
|
||
holding 400 MB of RSS and 3 GB of textures would pass it.
|
||
|
||
**Recommendation: NFR-P8 should be restated as two figures** — host RSS
|
||
exclusive of device-local memory, and a separate VRAM ceiling read from the
|
||
adapter — because the second is the one that decides whether the application
|
||
survives beside a browser on an 8 GB card, and nothing in this repository
|
||
measures it today.
|
||
|
||
**And a decision is outstanding.** `catalog_idle_rss_mb` carries no budget
|
||
because nobody has decided how much of the 500 MB belongs to the catalog layer
|
||
and how much to everything above it. The suite records the number so that
|
||
decision can be taken against a measurement rather than an estimate. When it is
|
||
taken, put the figure in `budget` and the metric becomes a gate.
|
||
|
||
---
|
||
|
||
## What is not measured, and would be worth adding
|
||
|
||
- **The UI's own open.** `ui/dr-ui/src/library.rs` does not call
|
||
`Catalog::count` or `Catalog::window`; it issues its own SQL against the same
|
||
tables, with a `VISIBLE` predicate and a burst-folding clause. `dr-bench`
|
||
cannot see those without depending on `dr-ui`, which would drag Slint into a
|
||
job that has no display. **Falsifiable end:** when the grid's queries move
|
||
down into `dr-catalog` — which is where SQL over catalog tables belongs —
|
||
`catalog_open_ms` becomes the whole of the application's open and this caveat
|
||
can be deleted rather than argued about.
|
||
- **The remote sweep.** `spawn_thumbnail_sweep`'s wall clock against a real
|
||
server is latency, not CPU, and is what FR-NC-3's design is judged by. It
|
||
needs a server and belongs in a different kind of test.
|
||
- **A cold disk.** See the fixture's limits above.
|
||
- **Android.** §4.1 states a second column of targets and §8 asks for
|
||
"periodically on the named reference Android devices". Nothing here runs on a
|
||
device. Spike S10 is the piece of work that would start it.
|
||
- **Everything with a frame in it.** See the out-of-scope list above.
|
||
|
||
---
|
||
|
||
## Running it
|
||
|
||
```sh
|
||
# Measure and print. Judges nothing.
|
||
cargo run --release -p dr-bench -- run
|
||
|
||
# Measure and gate. What CI runs.
|
||
cargo run --release -p dr-bench -- check
|
||
|
||
# The same, with machine-sensitive budgets asserted too.
|
||
cargo run --release -p dr-bench -- check --reference
|
||
|
||
# Rewrite bench-baseline.json from this run, and commit the diff.
|
||
cargo run --release -p dr-bench -- record --reference
|
||
```
|
||
|
||
Useful flags: `--fixture <dir>` (or `$DR_BENCH_DIR`) to put the synthetic
|
||
catalog somewhere specific, `--lanes <n>` to pin the sweep's parallelism, and
|
||
`--thumbnails <n>` to lengthen or shorten the throughput row.
|