Merge: measure the catalog, so §8's promise stops being a promise
There was no benchmark harness of any kind -- no benches/, no criterion, no synthetic fixture -- while §8 promised a suite run per commit that fails the build on regression. Ten performance requirements could be neither passed nor failed. tools/bench builds a deterministic 50,000-row catalog over a pool of twelve generated JPEGs, about 14 MB, reproducible from a seed, with a stamp so it rebuilds rather than silently comparing against a different workload. It depends on nothing GPU or UI, which is what makes the CI job affordable. NFR-P1 and NFR-P3 are gated and tagged. NFR-P7, NFR-P8 and R2 are measured but deliberately untagged: the export gate is one-sided, the memory figure is the catalog layer's share rather than the whole, and R2's first sentence is a 60 fps scroll a catalog benchmark cannot claim. Every recorded value in the baseline is null. Nobody has run this on the reference desktop, and a fabricated figure would make every later comparison a comparison against a guess. First run on this machine: catalog opens in 70 ms against a 2 s budget, and thumbnail throughput measures 37 img/s against a target of 100 -- reported rather than asserted here, and the first evidence that the target may not hold. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
+33
-21
@@ -316,31 +316,43 @@ for interoperating with the editors FR-CAT-14 imports from.
|
||||
|
||||
---
|
||||
|
||||
## 8. The performance targets are unverified, not unmet
|
||||
## 8. The performance targets are half-verified, and the half that is left is the hard one
|
||||
|
||||
Eleven of the fifteen §4.1 targets carry no tag: NFR-P2, -P3, -P4, -P6, -P7, -P8, -P10, -P11, -P12,
|
||||
-P14, -P15. That is the uninteresting part of this section.
|
||||
§8 and §4.1 both require the same thing in the same words: an automated benchmark suite against a
|
||||
synthetic 50k catalog, run per commit, where **"a regression beyond a stated tolerance is a build
|
||||
failure, not a notification."** For most of this project's life it did not exist — no `benches/`, no
|
||||
criterion, no synthetic catalog, and three CI workflows that between them measured nothing.
|
||||
|
||||
The interesting part is that §8 and §4.1 both require the same thing, in the same words, and it does
|
||||
not exist: an automated benchmark suite against a synthetic 50k catalog, run per commit, where **"a
|
||||
regression beyond a stated tolerance is a build failure, not a notification."** There is no
|
||||
`benches/` directory in the workspace, no criterion dependency, and no synthetic catalog. The three
|
||||
CI workflows run `cargo fmt --check`, clippy, `cargo test --workspace`, a release build, an Android
|
||||
cross-build and a layering check. None of them measures anything, so there is no baseline to
|
||||
regress against and no tolerance to exceed.
|
||||
**It exists now, for everything that does not need a frame.** [`tools/bench`](../tools/bench) builds
|
||||
a deterministic 50,000-row catalog over a pool of a dozen real files, measures against it, and fails
|
||||
the build on a violated budget or a drift past tolerance;
|
||||
[`.gitea/workflows/benchmark.yml`](../.gitea/workflows/benchmark.yml) runs it on every push, and
|
||||
[benchmarks.md](benchmarks.md) is the account of what it does and does not cover. **NFR-P1** and
|
||||
**NFR-P3** are now genuinely gated, and R2's "catalog opens in under 2s" clause with them.
|
||||
|
||||
What does exist is narrower and genuinely good: `dr-gpu/examples/frame_budget` is a real instrument,
|
||||
its results are committed in [frame-budget.md](frame-budget.md) with the machine and profile named,
|
||||
and TD-4's before-and-after was measured with it. But it is run by hand — frame-budget.md's own
|
||||
instruction is "rerun and diff this file" — and the guard version that does live in CI skips itself
|
||||
where there is no GPU adapter, which the workflow notes is the normal case on a runner, while
|
||||
asserting its CPU half only when `debug_assertions` is off, which a dev-profile `cargo test` is not.
|
||||
In CI it therefore asserts approximately nothing.
|
||||
Three qualifications, all of them stated in the harness itself rather than only here:
|
||||
|
||||
**The claim to take from this is precise.** Nothing here says the performance targets are missed.
|
||||
Several are plausibly met. It says that if one were broken tomorrow, nobody would find out — which
|
||||
is the failure mode §8 was written to prevent, and the reason it belongs in this document rather
|
||||
than in a backlog.
|
||||
- **The numbers have not been recorded yet.** Every `recorded` field in
|
||||
[bench-baseline.json](bench-baseline.json) is `null`, deliberately: a fabricated baseline is worse
|
||||
than none. Until `dr-bench record --reference` is run on the reference desktop and committed, the
|
||||
budget gate works and the regression gate does not.
|
||||
- **NFR-P7 and NFR-P8 are half-measured and are not tagged.** The export row covers the encode half
|
||||
of the chain and no GPU render, so it can fail the requirement and cannot pass it. The memory row
|
||||
covers a process holding the catalog and nothing else — no toolkit, no adapter — so it is the
|
||||
catalog layer's share of the 500 MB rather than the figure NFR-P8 is about. Neither carries a
|
||||
`TRACES:` tag, which is the point.
|
||||
- **NFR-P8 needs a decision, not more code.** How much of its 500 MB belongs below the UI is
|
||||
unstated, and until somebody says, the metric can record but not judge. [benchmarks.md](benchmarks.md)
|
||||
also answers the question §4.1 raises about GPU memory — RSS cannot see device-local allocations
|
||||
at all — and recommends restating the requirement as two figures.
|
||||
|
||||
**What is left is the frame-timing half, and it is the hard one.** NFR-P2, -P4, -P5, -P6, -P9, -P10,
|
||||
-P11, -P12, -P13, -P14 and -P15 all need a probe inside a running Slint application, a GPU adapter,
|
||||
or both. `dr-gpu/examples/frame_budget` is a real instrument for the GPU part and its results are
|
||||
committed in [frame-budget.md](frame-budget.md) with the machine and profile named — but it is run
|
||||
by hand, and the guard version in CI skips itself where there is no adapter, which is the normal
|
||||
case on a runner. So the claim to take from this section is now narrower than it was, and still
|
||||
true: **a scroll that dropped to 30 fps tomorrow would reach a user before it reached CI.**
|
||||
|
||||
---
|
||||
|
||||
|
||||
Reference in New Issue
Block a user