Merge: measure the catalog, so §8's promise stops being a promise

There was no benchmark harness of any kind -- no benches/, no criterion,
no synthetic fixture -- while §8 promised a suite run per commit that
fails the build on regression. Ten performance requirements could be
neither passed nor failed.

tools/bench builds a deterministic 50,000-row catalog over a pool of
twelve generated JPEGs, about 14 MB, reproducible from a seed, with a
stamp so it rebuilds rather than silently comparing against a different
workload. It depends on nothing GPU or UI, which is what makes the CI job
affordable.

NFR-P1 and NFR-P3 are gated and tagged. NFR-P7, NFR-P8 and R2 are
measured but deliberately untagged: the export gate is one-sided, the
memory figure is the catalog layer's share rather than the whole, and
R2's first sentence is a 60 fps scroll a catalog benchmark cannot claim.

Every recorded value in the baseline is null. Nobody has run this on the
reference desktop, and a fabricated figure would make every later
comparison a comparison against a guess.

First run on this machine: catalog opens in 70 ms against a 2 s budget,
and thumbnail throughput measures 37 img/s against a target of 100 --
reported rather than asserted here, and the first evidence that the
target may not hold.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
2026-08-30 13:13:57 +02:00
co-authored by Claude Opus 5
16 changed files with 2825 additions and 21 deletions
+113
View File
@@ -0,0 +1,113 @@
{
"_readme": [
"The committed numbers for DarkRoom's benchmark suite (docs/requirements.md §8).",
"Produced and checked by `cargo run --release -p dr-bench`; docs/benchmarks.md explains each metric.",
"",
"Two gates, and they are not the same gate. `budget` is the requirement's own threshold and never moves.",
"`recorded` is what the reference desktop last measured, and a run that drifts past `tolerance` beyond it",
"fails the build even while still inside the budget — which is how most performance rot actually arrives.",
"",
"`recorded` is null on every metric because nobody has run the suite yet. That is deliberate: writing",
"plausible-looking figures here would make every later comparison a comparison against a guess. Run",
"`cargo run --release -p dr-bench -- record --reference` on the reference desktop and commit the diff.",
"Until then the budget gate works and the regression gate says, in the report, that it cannot.",
"",
"`machine_sensitive` says whether a budget is a statement about a machine as much as about the code.",
"Those budgets are asserted only under --reference: §8 names the reference desktop, and a two-core CI",
"container cannot speak to a target written for twenty-four threads. Asserting one there would produce a",
"red gate everybody learns to ignore, which is the trap core/dr-gpu/tests/frame_budget.rs already avoids.",
"",
"Several metrics carry a requirement ID with a qualifier. Read those literally. NFR-P7's budget here is",
"checked against the encode half of an export only — no GPU render is in the figure — so it can fail the",
"requirement and cannot pass it, and no TRACES tag claims otherwise. NFR-P8 has no budget at all yet,",
"because nobody has decided how much of its 500 MB belongs to the catalog layer; this records the number",
"that decision needs."
],
"tolerance": 0.15,
"recorded_on": null,
"recorded_at_unix": null,
"fixture": null,
"metrics": {
"catalog_filtered_ms": {
"requirement": "FR-CAT-6",
"what": "Count plus first window under a rating filter, which compiles to a correlated subquery over versions.",
"unit": "ms",
"direction": "lower_is_better",
"machine_sensitive": true,
"budget": null,
"recorded": null
},
"catalog_idle_rss_mb": {
"requirement": "NFR-P8 (the catalog layer's share only — no toolkit, no adapter, no decode cache)",
"what": "Resident memory of a process that opened the 50k catalog and scrolled ten thousand rows.",
"unit": "MB",
"direction": "lower_is_better",
"machine_sensitive": false,
"budget": null,
"recorded": null
},
"catalog_open_ms": {
"requirement": "NFR-P1",
"what": "Catalog::open plus the count, first window and timeline the grid cannot paint without.",
"unit": "ms",
"direction": "lower_is_better",
"machine_sensitive": false,
"budget": 2000.0,
"recorded": null
},
"catalog_open_warm_ms": {
"requirement": "NFR-P1",
"what": "The same four calls on a second connection, with SQLite's page cache already warm.",
"unit": "ms",
"direction": "lower_is_better",
"machine_sensitive": false,
"budget": 2000.0,
"recorded": null
},
"catalog_window_p99_ms": {
"requirement": "FR-CAT-4",
"what": "One 400-row grid window at a random offset, p99 of one hundred.",
"unit": "ms",
"direction": "lower_is_better",
"machine_sensitive": true,
"budget": null,
"recorded": null
},
"export_24mp_long_edge_2048_ms": {
"requirement": "FR-EXP-3",
"what": "The web export: resample a 24 MP frame to a 2048 px long edge, sharpen, encode. p99 of five.",
"unit": "ms",
"direction": "lower_is_better",
"machine_sensitive": true,
"budget": null,
"recorded": null
},
"export_24mp_original_ms": {
"requirement": "NFR-P7 (the encode half only — the GPU render is not in this figure)",
"what": "Resample, output-sharpen and JPEG-encode a 24 MP frame at source size. p99 of five.",
"unit": "ms",
"direction": "lower_is_better",
"machine_sensitive": true,
"budget": 2000.0,
"recorded": null
},
"thumbnail_per_image_p99_ms": {
"requirement": "NFR-P3",
"what": "One thumbnail on its own lane: decode the preview, downscale, orient, encode. p99.",
"unit": "ms",
"direction": "lower_is_better",
"machine_sensitive": true,
"budget": null,
"recorded": null
},
"thumbnail_throughput_ips": {
"requirement": "NFR-P3",
"what": "Whole-sweep throughput: 1200 thumbnails through the sweep's chunk-and-lane shape, wall clock.",
"unit": "img/s",
"direction": "higher_is_better",
"machine_sensitive": true,
"budget": 100.0,
"recorded": null
}
}
}
+243
View File
@@ -0,0 +1,243 @@
# The benchmark suite
**Status:** Built, not yet recorded · 2026-08-30
**Companion to:** [requirements.md](requirements.md) §4.1 (performance targets) · §8 (verification)
**Instrument:** [`tools/bench`](../tools/bench) — `cargo run --release -p dr-bench -- check`
**Committed numbers:** [`bench-baseline.json`](bench-baseline.json)
**GPU half:** [`core/dr-gpu/tests/frame_budget.rs`](../core/dr-gpu/tests/frame_budget.rs) ·
[frame-budget.md](frame-budget.md)
§8 has said since it was written that performance is verified by *"an automated
benchmark suite against a synthetic 50k catalog, run per-commit … A regression
beyond stated tolerance fails the build."* Until this suite there was none. No
`benches/`, no `[[bench]]`, no criterion, no fixture — and ten performance
requirements that could therefore be neither passed nor failed, five of them
carrying a `TRACES:` tag regardless.
This file is what the suite covers, what it deliberately does not, and how to
read a failure.
---
## The state of it, first
**No numbers have been recorded yet.** Every `recorded` field in
[`bench-baseline.json`](bench-baseline.json) is `null`, on purpose: writing
plausible-looking figures into a baseline would make every later comparison a
comparison against a guess, and the first real regression would be invisible.
To record them, on the reference desktop:
```sh
cargo run --release -p dr-bench -- record --reference
```
and commit the diff. Until that happens the **budget** gate works — a catalog
that takes three seconds to open fails the build today — and the **regression**
gate reports that it has nothing to compare against, rather than pretending.
---
## What it measures
| Metric | Requirement | Gated? |
|---|---|---|
| `catalog_open_ms` | **NFR-P1**, and R2's second sentence | Yes, everywhere — budget 2000 ms |
| `catalog_open_warm_ms` | NFR-P1, page cache warm | Yes, everywhere — budget 2000 ms |
| `catalog_window_p99_ms` | FR-CAT-4 | Regression only |
| `catalog_filtered_ms` | FR-CAT-6 | Regression only |
| `thumbnail_throughput_ips` | **NFR-P3** | Budget 100 img/s, on the reference desktop |
| `thumbnail_per_image_p99_ms` | NFR-P3 | Regression only |
| `export_24mp_original_ms` | NFR-P7, **encode half only** | One-sided: can fail it, cannot pass it |
| `export_24mp_long_edge_2048_ms` | FR-EXP-3 | Regression only |
| `catalog_idle_rss_mb` | NFR-P8, **catalog layer only** | Regression only — see below |
Two of those rows carry a qualifier, and the qualifiers are the point.
### Requirements this can now pass *or* fail
**NFR-P1 — catalog open under 2 s.** The measured span is the four things the
library view cannot paint without: `Catalog::open` (which connects, migrates and
**backfills**, and the backfill is three passes over the images table on every
open), `count`, the first 400-row `window`, and the monthly `timeline`. Tagged
`TRACES: NFR-P1` in [`tools/bench/src/catalog_open.rs`](../tools/bench/src/catalog_open.rs),
because a build that breaks it fails this gate.
**NFR-P3 — ≥ 100 images per second on the embedded preview path.** The
per-image work is exactly what `spawn_thumbnail_sweep` does — `decode_jpeg`,
`Preview::downscale_to`, `Preview::apply_orientation`, `encode_rgba`,
`ThumbStore::put` — arranged in the same shape: chunks of 96, lanes owning
disjoint slices, and the single thread that owns the store writing the finished
chunk. Tagged `TRACES: NFR-P3` in
[`tools/bench/src/thumbnails.rs`](../tools/bench/src/thumbnails.rs).
### Requirements this can only half-answer, and is not tagged for
**NFR-P7 — 24 MP export under 2 s, full chain.** The full chain is decode,
demosaic, a full-resolution GPU render, a read-back, then resize, sharpen and
encode. Only the last three run without an adapter. So the figure here is a
**lower bound** on the requirement: exceeding 2 s in the encode alone violates
NFR-P7 no matter how fast the render is, and coming in under it proves nothing.
The budget is gated on that basis and there is no `TRACES: NFR-P7` anywhere in
`tools/bench`.
**NFR-P8 — idle memory under 500 MB.** The probe is a fresh process holding the
catalog and nothing else: no Slint, no wgpu device, no font stack, no decode
cache. Its RSS is the catalog layer's *share* of that 500 MB, not the figure the
requirement is about. It carries no budget for a reason given below.
### Requirements out of scope, listed so their absence reads as a decision
NFR-P2 (grid scroll at 60 fps), P4 (open in develop), P5 (slider to visible),
P6 (pan/zoom), P9 (UI-executor blocking), P10 (touch response), P11 (layout
transition), P12 (warm shader setup), P13 (next image in culling), P14 (focus
peaking), P15 (drawn mask stroke). Every one of them needs a frame-timing probe
inside a running Slint application, a GPU adapter, or both. None is faked here.
The GPU half of the story that *does* exist is
[frame-budget.md](frame-budget.md) and its guard test, which asserts FR-DSP-3
and skips itself where there is no adapter. `.gitea/workflows/benchmark.yml`
runs it as its own job for exactly that reason.
---
## The fixture
Fifty thousand rows over a pool of twelve real image files. Rows are cheap and
pixels are not: everything the catalog half touches is rows and is therefore
exact at full scale, and everything the pixel half touches is one file at a time
and does not care how many rows point at it. The result is ~14 MB on disk
instead of ~2 TB, and neither half is flattered by that.
| | |
|---|---|
| Rows | 50,000 images, 50,000 default versions, 400 folders, one root |
| Capture times | Twelve years from a fixed epoch, so the timeline has ~144 monthly buckets |
| Sources | 12 synthesised JPEGs at 1620 × 1080 — the size `dr-decode` records a CR2 carrying in IFD2 |
| Seed | 20260829, in [`tools/bench/src/main.rs`](../tools/bench/src/main.rs) |
| Location | `$DR_BENCH_DIR`, else the system temporary directory |
It is reproducible from the seed, and a `stamp.json` beside it records what it
was built from — seed, row count, source count, preview size, and `dr-catalog`'s
schema version. A mismatch rebuilds rather than silently measuring a different
workload than the baseline describes.
Two honest limits on it:
- **The page cache is warm.** The fixture was written by this suite or by an
earlier run of it, so neither the catalog open nor the thumbnail sweep pays
for a cold disk. On the reference desktop's NVMe a genuinely cold read of a
14 MB catalog is tens of milliseconds; on spinning rust it is not.
- **The sources are synthetic.** A coarse gradient with a fine dither, which is
what `frame_budget.rs` synthesises for the same reason — a flat frame lets the
memory system serve every sample from one cache line and flatters a box
filter, and pure noise defeats the entropy coder in the other direction.
---
## Two gates, and how to read a failure
**Budget.** The requirement's own threshold. It does not move. Failing it means
a requirement is violated.
**Regression.** More than 15% worse than the last recorded figure *on the same
machine, against the same fixture*. Failing it means the code got slower while
still inside the requirement — which is how most performance rot actually
arrives, never over the line, always a little worse, until one day the line is
crossed by a change that was not the cause.
A metric declares whether its budget is `machine_sensitive`. Those are asserted
only under `--reference`, and reported everywhere else. §8 names *"the reference
desktop"*, not CI, and it is right to: a container with two cores cannot speak
to a throughput target written for twenty-four threads, and asserting one there
would produce exactly what `core/dr-gpu/tests/frame_budget.rs` refused to
produce — *"a red suite that everyone learns to ignore"*. Catalog open is not
machine-sensitive: 2 s against an expected figure two orders of magnitude
smaller is a threshold any machine can be held to.
Exit codes: `0` everything passed, `1` a gate failed, `2` the harness itself
could not run. Distinguished so a CI log that says "failed" does not leave
anyone guessing whether the code got slower or the fixture would not build.
**Release, always.** The workspace builds its own crates at `opt-level = 0` in
dev, and every figure here is dominated by this workspace's own code — the JPEG
decode, the box filter, the resample, the sharpen. A debug run measures rustc's
shadow. The report says which profile it was built in on its second line.
---
## NFR-P8, and the question §4.1 asks
§4.1 says NFR-P8 *"must state whether it measures RSS inclusive or exclusive of
GPU allocations, and whether it holds after SQLite's page cache warms on a 50k
catalog."* Both halves have an answer.
**On the page cache: warm.** The probe runs the count, the timeline and
twenty-five windows before it reads its counters, so SQLite's cache holds the
b-tree pages a scroll touches. That is the right side to err on — a figure taken
before the cache warms would understate a steady-state library.
**On GPU memory: RSS is exclusive of device-local allocations, and cannot be
made otherwise.** A Vulkan allocation in a device-local heap never enters the
process's address space, so nothing under `/proc/self/status` can see it. What
*does* land in RSS is the host-visible side — staging buffers, mapped upload
rings, the read-back `AdjustPass` performs on export — plus the driver's own
resident pages.
So "idle memory < 500 MB" is two questions wearing one number, and a build
holding 400 MB of RSS and 3 GB of textures would pass it.
**Recommendation: NFR-P8 should be restated as two figures** — host RSS
exclusive of device-local memory, and a separate VRAM ceiling read from the
adapter — because the second is the one that decides whether the application
survives beside a browser on an 8 GB card, and nothing in this repository
measures it today.
**And a decision is outstanding.** `catalog_idle_rss_mb` carries no budget
because nobody has decided how much of the 500 MB belongs to the catalog layer
and how much to everything above it. The suite records the number so that
decision can be taken against a measurement rather than an estimate. When it is
taken, put the figure in `budget` and the metric becomes a gate.
---
## What is not measured, and would be worth adding
- **The UI's own open.** `ui/dr-ui/src/library.rs` does not call
`Catalog::count` or `Catalog::window`; it issues its own SQL against the same
tables, with a `VISIBLE` predicate and a burst-folding clause. `dr-bench`
cannot see those without depending on `dr-ui`, which would drag Slint into a
job that has no display. **Falsifiable end:** when the grid's queries move
down into `dr-catalog` — which is where SQL over catalog tables belongs —
`catalog_open_ms` becomes the whole of the application's open and this caveat
can be deleted rather than argued about.
- **The remote sweep.** `spawn_thumbnail_sweep`'s wall clock against a real
server is latency, not CPU, and is what FR-NC-3's design is judged by. It
needs a server and belongs in a different kind of test.
- **A cold disk.** See the fixture's limits above.
- **Android.** §4.1 states a second column of targets and §8 asks for
"periodically on the named reference Android devices". Nothing here runs on a
device. Spike S10 is the piece of work that would start it.
- **Everything with a frame in it.** See the out-of-scope list above.
---
## Running it
```sh
# Measure and print. Judges nothing.
cargo run --release -p dr-bench -- run
# Measure and gate. What CI runs.
cargo run --release -p dr-bench -- check
# The same, with machine-sensitive budgets asserted too.
cargo run --release -p dr-bench -- check --reference
# Rewrite bench-baseline.json from this run, and commit the diff.
cargo run --release -p dr-bench -- record --reference
```
Useful flags: `--fixture <dir>` (or `$DR_BENCH_DIR`) to put the synthetic
catalog somewhere specific, `--lanes <n>` to pin the sweep's parallelism, and
`--thumbnails <n>` to lengthen or shorten the throughput row.
+33 -21
View File
@@ -316,31 +316,43 @@ for interoperating with the editors FR-CAT-14 imports from.
---
## 8. The performance targets are unverified, not unmet
## 8. The performance targets are half-verified, and the half that is left is the hard one
Eleven of the fifteen §4.1 targets carry no tag: NFR-P2, -P3, -P4, -P6, -P7, -P8, -P10, -P11, -P12,
-P14, -P15. That is the uninteresting part of this section.
§8 and §4.1 both require the same thing in the same words: an automated benchmark suite against a
synthetic 50k catalog, run per commit, where **"a regression beyond a stated tolerance is a build
failure, not a notification."** For most of this project's life it did not exist — no `benches/`, no
criterion, no synthetic catalog, and three CI workflows that between them measured nothing.
The interesting part is that §8 and §4.1 both require the same thing, in the same words, and it does
not exist: an automated benchmark suite against a synthetic 50k catalog, run per commit, where **"a
regression beyond a stated tolerance is a build failure, not a notification."** There is no
`benches/` directory in the workspace, no criterion dependency, and no synthetic catalog. The three
CI workflows run `cargo fmt --check`, clippy, `cargo test --workspace`, a release build, an Android
cross-build and a layering check. None of them measures anything, so there is no baseline to
regress against and no tolerance to exceed.
**It exists now, for everything that does not need a frame.** [`tools/bench`](../tools/bench) builds
a deterministic 50,000-row catalog over a pool of a dozen real files, measures against it, and fails
the build on a violated budget or a drift past tolerance;
[`.gitea/workflows/benchmark.yml`](../.gitea/workflows/benchmark.yml) runs it on every push, and
[benchmarks.md](benchmarks.md) is the account of what it does and does not cover. **NFR-P1** and
**NFR-P3** are now genuinely gated, and R2's "catalog opens in under 2s" clause with them.
What does exist is narrower and genuinely good: `dr-gpu/examples/frame_budget` is a real instrument,
its results are committed in [frame-budget.md](frame-budget.md) with the machine and profile named,
and TD-4's before-and-after was measured with it. But it is run by hand — frame-budget.md's own
instruction is "rerun and diff this file" — and the guard version that does live in CI skips itself
where there is no GPU adapter, which the workflow notes is the normal case on a runner, while
asserting its CPU half only when `debug_assertions` is off, which a dev-profile `cargo test` is not.
In CI it therefore asserts approximately nothing.
Three qualifications, all of them stated in the harness itself rather than only here:
**The claim to take from this is precise.** Nothing here says the performance targets are missed.
Several are plausibly met. It says that if one were broken tomorrow, nobody would find out — which
is the failure mode §8 was written to prevent, and the reason it belongs in this document rather
than in a backlog.
- **The numbers have not been recorded yet.** Every `recorded` field in
[bench-baseline.json](bench-baseline.json) is `null`, deliberately: a fabricated baseline is worse
than none. Until `dr-bench record --reference` is run on the reference desktop and committed, the
budget gate works and the regression gate does not.
- **NFR-P7 and NFR-P8 are half-measured and are not tagged.** The export row covers the encode half
of the chain and no GPU render, so it can fail the requirement and cannot pass it. The memory row
covers a process holding the catalog and nothing else — no toolkit, no adapter — so it is the
catalog layer's share of the 500 MB rather than the figure NFR-P8 is about. Neither carries a
`TRACES:` tag, which is the point.
- **NFR-P8 needs a decision, not more code.** How much of its 500 MB belongs below the UI is
unstated, and until somebody says, the metric can record but not judge. [benchmarks.md](benchmarks.md)
also answers the question §4.1 raises about GPU memory — RSS cannot see device-local allocations
at all — and recommends restating the requirement as two figures.
**What is left is the frame-timing half, and it is the hard one.** NFR-P2, -P4, -P5, -P6, -P9, -P10,
-P11, -P12, -P13, -P14 and -P15 all need a probe inside a running Slint application, a GPU adapter,
or both. `dr-gpu/examples/frame_budget` is a real instrument for the GPU part and its results are
committed in [frame-budget.md](frame-budget.md) with the machine and profile named — but it is run
by hand, and the guard version in CI skips itself where there is no adapter, which is the normal
case on a runner. So the claim to take from this section is now narrower than it was, and still
true: **a scroll that dropped to 30 fps tomorrow would reach a user before it reached CI.**
---