Merge: measure the catalog, so §8's promise stops being a promise

There was no benchmark harness of any kind -- no benches/, no criterion,
no synthetic fixture -- while §8 promised a suite run per commit that
fails the build on regression. Ten performance requirements could be
neither passed nor failed.

tools/bench builds a deterministic 50,000-row catalog over a pool of
twelve generated JPEGs, about 14 MB, reproducible from a seed, with a
stamp so it rebuilds rather than silently comparing against a different
workload. It depends on nothing GPU or UI, which is what makes the CI job
affordable.

NFR-P1 and NFR-P3 are gated and tagged. NFR-P7, NFR-P8 and R2 are
measured but deliberately untagged: the export gate is one-sided, the
memory figure is the catalog layer's share rather than the whole, and
R2's first sentence is a 60 fps scroll a catalog benchmark cannot claim.

Every recorded value in the baseline is null. Nobody has run this on the
reference desktop, and a fabricated figure would make every later
comparison a comparison against a guess.

First run on this machine: catalog opens in 70 ms against a 2 s budget,
and thumbnail throughput measures 37 img/s against a target of 100 --
reported rather than asserted here, and the first evidence that the
target may not hold.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
2026-08-30 13:13:57 +02:00
co-authored by Claude Opus 5
16 changed files with 2825 additions and 21 deletions
+33 -21
View File
@@ -316,31 +316,43 @@ for interoperating with the editors FR-CAT-14 imports from.
---
## 8. The performance targets are unverified, not unmet
## 8. The performance targets are half-verified, and the half that is left is the hard one
Eleven of the fifteen §4.1 targets carry no tag: NFR-P2, -P3, -P4, -P6, -P7, -P8, -P10, -P11, -P12,
-P14, -P15. That is the uninteresting part of this section.
§8 and §4.1 both require the same thing in the same words: an automated benchmark suite against a
synthetic 50k catalog, run per commit, where **"a regression beyond a stated tolerance is a build
failure, not a notification."** For most of this project's life it did not exist — no `benches/`, no
criterion, no synthetic catalog, and three CI workflows that between them measured nothing.
The interesting part is that §8 and §4.1 both require the same thing, in the same words, and it does
not exist: an automated benchmark suite against a synthetic 50k catalog, run per commit, where **"a
regression beyond a stated tolerance is a build failure, not a notification."** There is no
`benches/` directory in the workspace, no criterion dependency, and no synthetic catalog. The three
CI workflows run `cargo fmt --check`, clippy, `cargo test --workspace`, a release build, an Android
cross-build and a layering check. None of them measures anything, so there is no baseline to
regress against and no tolerance to exceed.
**It exists now, for everything that does not need a frame.** [`tools/bench`](../tools/bench) builds
a deterministic 50,000-row catalog over a pool of a dozen real files, measures against it, and fails
the build on a violated budget or a drift past tolerance;
[`.gitea/workflows/benchmark.yml`](../.gitea/workflows/benchmark.yml) runs it on every push, and
[benchmarks.md](benchmarks.md) is the account of what it does and does not cover. **NFR-P1** and
**NFR-P3** are now genuinely gated, and R2's "catalog opens in under 2s" clause with them.
What does exist is narrower and genuinely good: `dr-gpu/examples/frame_budget` is a real instrument,
its results are committed in [frame-budget.md](frame-budget.md) with the machine and profile named,
and TD-4's before-and-after was measured with it. But it is run by hand — frame-budget.md's own
instruction is "rerun and diff this file" — and the guard version that does live in CI skips itself
where there is no GPU adapter, which the workflow notes is the normal case on a runner, while
asserting its CPU half only when `debug_assertions` is off, which a dev-profile `cargo test` is not.
In CI it therefore asserts approximately nothing.
Three qualifications, all of them stated in the harness itself rather than only here:
**The claim to take from this is precise.** Nothing here says the performance targets are missed.
Several are plausibly met. It says that if one were broken tomorrow, nobody would find out — which
is the failure mode §8 was written to prevent, and the reason it belongs in this document rather
than in a backlog.
- **The numbers have not been recorded yet.** Every `recorded` field in
[bench-baseline.json](bench-baseline.json) is `null`, deliberately: a fabricated baseline is worse
than none. Until `dr-bench record --reference` is run on the reference desktop and committed, the
budget gate works and the regression gate does not.
- **NFR-P7 and NFR-P8 are half-measured and are not tagged.** The export row covers the encode half
of the chain and no GPU render, so it can fail the requirement and cannot pass it. The memory row
covers a process holding the catalog and nothing else — no toolkit, no adapter — so it is the
catalog layer's share of the 500 MB rather than the figure NFR-P8 is about. Neither carries a
`TRACES:` tag, which is the point.
- **NFR-P8 needs a decision, not more code.** How much of its 500 MB belongs below the UI is
unstated, and until somebody says, the metric can record but not judge. [benchmarks.md](benchmarks.md)
also answers the question §4.1 raises about GPU memory — RSS cannot see device-local allocations
at all — and recommends restating the requirement as two figures.
**What is left is the frame-timing half, and it is the hard one.** NFR-P2, -P4, -P5, -P6, -P9, -P10,
-P11, -P12, -P13, -P14 and -P15 all need a probe inside a running Slint application, a GPU adapter,
or both. `dr-gpu/examples/frame_budget` is a real instrument for the GPU part and its results are
committed in [frame-budget.md](frame-budget.md) with the machine and profile named — but it is run
by hand, and the guard version in CI skips itself where there is no adapter, which is the normal
case on a runner. So the claim to take from this section is now narrower than it was, and still
true: **a scroll that dropped to 30 fps tomorrow would reach a user before it reached CI.**
---