Put the developer docs under docs/dev and index the folder for users first
docs/ had 26 developer documents flat beside the manual, and the two audiences are very differently sized: most readers want the manual and the gesture reference, a few want the register, the designs and the measurements. The manual and gestures.md stay at the top; everything for someone changing the code moves to docs/dev/, and the two documents that name their own successors — the v0.1 milestone and the UI-refinement plan — go to docs/dev/archive/ rather than being deleted, since both are still cited. docs/README.md is the index, users first. Every reference follows: code comments, Cargo manifests, the workflows, the pre-commit hook, the bench and traceability tools (which locate the repo root by docs/dev/requirements.md now), packaging, the Docker READMEs, CLAUDE.md, CONTRIBUTING.md and the README. The matrix links one level deeper and is regenerated. Links out of the moved documents into the tree gain a level; a link checker over every Markdown file finds none broken.
This commit is contained in:
@@ -0,0 +1,243 @@
|
||||
# The benchmark suite
|
||||
|
||||
**Status:** Built, not yet recorded · 2026-08-30
|
||||
**Companion to:** [requirements.md](requirements.md) §4.1 (performance targets) · §8 (verification)
|
||||
**Instrument:** [`tools/bench`](../../tools/bench) — `cargo run --release -p dr-bench -- check`
|
||||
**Committed numbers:** [`bench-baseline.json`](bench-baseline.json)
|
||||
**GPU half:** [`core/dr-gpu/tests/frame_budget.rs`](../../core/dr-gpu/tests/frame_budget.rs) ·
|
||||
[frame-budget.md](frame-budget.md)
|
||||
|
||||
§8 has said since it was written that performance is verified by *"an automated
|
||||
benchmark suite against a synthetic 50k catalog, run per-commit … A regression
|
||||
beyond stated tolerance fails the build."* Until this suite there was none. No
|
||||
`benches/`, no `[[bench]]`, no criterion, no fixture — and ten performance
|
||||
requirements that could therefore be neither passed nor failed, five of them
|
||||
carrying a `TRACES:` tag regardless.
|
||||
|
||||
This file is what the suite covers, what it deliberately does not, and how to
|
||||
read a failure.
|
||||
|
||||
---
|
||||
|
||||
## The state of it, first
|
||||
|
||||
**No numbers have been recorded yet.** Every `recorded` field in
|
||||
[`bench-baseline.json`](bench-baseline.json) is `null`, on purpose: writing
|
||||
plausible-looking figures into a baseline would make every later comparison a
|
||||
comparison against a guess, and the first real regression would be invisible.
|
||||
|
||||
To record them, on the reference desktop:
|
||||
|
||||
```sh
|
||||
cargo run --release -p dr-bench -- record --reference
|
||||
```
|
||||
|
||||
and commit the diff. Until that happens the **budget** gate works — a catalog
|
||||
that takes three seconds to open fails the build today — and the **regression**
|
||||
gate reports that it has nothing to compare against, rather than pretending.
|
||||
|
||||
---
|
||||
|
||||
## What it measures
|
||||
|
||||
| Metric | Requirement | Gated? |
|
||||
|---|---|---|
|
||||
| `catalog_open_ms` | **NFR-P1**, and R2's second sentence | Yes, everywhere — budget 2000 ms |
|
||||
| `catalog_open_warm_ms` | NFR-P1, page cache warm | Yes, everywhere — budget 2000 ms |
|
||||
| `catalog_window_p99_ms` | FR-CAT-4 | Regression only |
|
||||
| `catalog_filtered_ms` | FR-CAT-6 | Regression only |
|
||||
| `thumbnail_throughput_ips` | **NFR-P3** | Budget 100 img/s, on the reference desktop |
|
||||
| `thumbnail_per_image_p99_ms` | NFR-P3 | Regression only |
|
||||
| `export_24mp_original_ms` | NFR-P7, **encode half only** | One-sided: can fail it, cannot pass it |
|
||||
| `export_24mp_long_edge_2048_ms` | FR-EXP-3 | Regression only |
|
||||
| `catalog_idle_rss_mb` | NFR-P8, **catalog layer only** | Regression only — see below |
|
||||
|
||||
Two of those rows carry a qualifier, and the qualifiers are the point.
|
||||
|
||||
### Requirements this can now pass *or* fail
|
||||
|
||||
**NFR-P1 — catalog open under 2 s.** The measured span is the four things the
|
||||
library view cannot paint without: `Catalog::open` (which connects, migrates and
|
||||
**backfills**, and the backfill is three passes over the images table on every
|
||||
open), `count`, the first 400-row `window`, and the monthly `timeline`. Tagged
|
||||
`TRACES: NFR-P1` in [`tools/bench/src/catalog_open.rs`](../../tools/bench/src/catalog_open.rs),
|
||||
because a build that breaks it fails this gate.
|
||||
|
||||
**NFR-P3 — ≥ 100 images per second on the embedded preview path.** The
|
||||
per-image work is exactly what `spawn_thumbnail_sweep` does — `decode_jpeg`,
|
||||
`Preview::downscale_to`, `Preview::apply_orientation`, `encode_rgba`,
|
||||
`ThumbStore::put` — arranged in the same shape: chunks of 96, lanes owning
|
||||
disjoint slices, and the single thread that owns the store writing the finished
|
||||
chunk. Tagged `TRACES: NFR-P3` in
|
||||
[`tools/bench/src/thumbnails.rs`](../../tools/bench/src/thumbnails.rs).
|
||||
|
||||
### Requirements this can only half-answer, and is not tagged for
|
||||
|
||||
**NFR-P7 — 24 MP export under 2 s, full chain.** The full chain is decode,
|
||||
demosaic, a full-resolution GPU render, a read-back, then resize, sharpen and
|
||||
encode. Only the last three run without an adapter. So the figure here is a
|
||||
**lower bound** on the requirement: exceeding 2 s in the encode alone violates
|
||||
NFR-P7 no matter how fast the render is, and coming in under it proves nothing.
|
||||
The budget is gated on that basis and there is no `TRACES: NFR-P7` anywhere in
|
||||
`tools/bench`.
|
||||
|
||||
**NFR-P8 — idle memory under 500 MB.** The probe is a fresh process holding the
|
||||
catalog and nothing else: no Slint, no wgpu device, no font stack, no decode
|
||||
cache. Its RSS is the catalog layer's *share* of that 500 MB, not the figure the
|
||||
requirement is about. It carries no budget for a reason given below.
|
||||
|
||||
### Requirements out of scope, listed so their absence reads as a decision
|
||||
|
||||
NFR-P2 (grid scroll at 60 fps), P4 (open in develop), P5 (slider to visible),
|
||||
P6 (pan/zoom), P9 (UI-executor blocking), P10 (touch response), P11 (layout
|
||||
transition), P12 (warm shader setup), P13 (next image in culling), P14 (focus
|
||||
peaking), P15 (drawn mask stroke). Every one of them needs a frame-timing probe
|
||||
inside a running Slint application, a GPU adapter, or both. None is faked here.
|
||||
|
||||
The GPU half of the story that *does* exist is
|
||||
[frame-budget.md](frame-budget.md) and its guard test, which asserts FR-DSP-3
|
||||
and skips itself where there is no adapter. `.gitea/workflows/benchmark.yml`
|
||||
runs it as its own job for exactly that reason.
|
||||
|
||||
---
|
||||
|
||||
## The fixture
|
||||
|
||||
Fifty thousand rows over a pool of twelve real image files. Rows are cheap and
|
||||
pixels are not: everything the catalog half touches is rows and is therefore
|
||||
exact at full scale, and everything the pixel half touches is one file at a time
|
||||
and does not care how many rows point at it. The result is ~14 MB on disk
|
||||
instead of ~2 TB, and neither half is flattered by that.
|
||||
|
||||
| | |
|
||||
|---|---|
|
||||
| Rows | 50,000 images, 50,000 default versions, 400 folders, one root |
|
||||
| Capture times | Twelve years from a fixed epoch, so the timeline has ~144 monthly buckets |
|
||||
| Sources | 12 synthesised JPEGs at 1620 × 1080 — the size `dr-decode` records a CR2 carrying in IFD2 |
|
||||
| Seed | 20260829, in [`tools/bench/src/main.rs`](../../tools/bench/src/main.rs) |
|
||||
| Location | `$DR_BENCH_DIR`, else the system temporary directory |
|
||||
|
||||
It is reproducible from the seed, and a `stamp.json` beside it records what it
|
||||
was built from — seed, row count, source count, preview size, and `dr-catalog`'s
|
||||
schema version. A mismatch rebuilds rather than silently measuring a different
|
||||
workload than the baseline describes.
|
||||
|
||||
Two honest limits on it:
|
||||
|
||||
- **The page cache is warm.** The fixture was written by this suite or by an
|
||||
earlier run of it, so neither the catalog open nor the thumbnail sweep pays
|
||||
for a cold disk. On the reference desktop's NVMe a genuinely cold read of a
|
||||
14 MB catalog is tens of milliseconds; on spinning rust it is not.
|
||||
- **The sources are synthetic.** A coarse gradient with a fine dither, which is
|
||||
what `frame_budget.rs` synthesises for the same reason — a flat frame lets the
|
||||
memory system serve every sample from one cache line and flatters a box
|
||||
filter, and pure noise defeats the entropy coder in the other direction.
|
||||
|
||||
---
|
||||
|
||||
## Two gates, and how to read a failure
|
||||
|
||||
**Budget.** The requirement's own threshold. It does not move. Failing it means
|
||||
a requirement is violated.
|
||||
|
||||
**Regression.** More than 15% worse than the last recorded figure *on the same
|
||||
machine, against the same fixture*. Failing it means the code got slower while
|
||||
still inside the requirement — which is how most performance rot actually
|
||||
arrives, never over the line, always a little worse, until one day the line is
|
||||
crossed by a change that was not the cause.
|
||||
|
||||
A metric declares whether its budget is `machine_sensitive`. Those are asserted
|
||||
only under `--reference`, and reported everywhere else. §8 names *"the reference
|
||||
desktop"*, not CI, and it is right to: a container with two cores cannot speak
|
||||
to a throughput target written for twenty-four threads, and asserting one there
|
||||
would produce exactly what `core/dr-gpu/tests/frame_budget.rs` refused to
|
||||
produce — *"a red suite that everyone learns to ignore"*. Catalog open is not
|
||||
machine-sensitive: 2 s against an expected figure two orders of magnitude
|
||||
smaller is a threshold any machine can be held to.
|
||||
|
||||
Exit codes: `0` everything passed, `1` a gate failed, `2` the harness itself
|
||||
could not run. Distinguished so a CI log that says "failed" does not leave
|
||||
anyone guessing whether the code got slower or the fixture would not build.
|
||||
|
||||
**Release, always.** The workspace builds its own crates at `opt-level = 0` in
|
||||
dev, and every figure here is dominated by this workspace's own code — the JPEG
|
||||
decode, the box filter, the resample, the sharpen. A debug run measures rustc's
|
||||
shadow. The report says which profile it was built in on its second line.
|
||||
|
||||
---
|
||||
|
||||
## NFR-P8, and the question §4.1 asks
|
||||
|
||||
§4.1 says NFR-P8 *"must state whether it measures RSS inclusive or exclusive of
|
||||
GPU allocations, and whether it holds after SQLite's page cache warms on a 50k
|
||||
catalog."* Both halves have an answer.
|
||||
|
||||
**On the page cache: warm.** The probe runs the count, the timeline and
|
||||
twenty-five windows before it reads its counters, so SQLite's cache holds the
|
||||
b-tree pages a scroll touches. That is the right side to err on — a figure taken
|
||||
before the cache warms would understate a steady-state library.
|
||||
|
||||
**On GPU memory: RSS is exclusive of device-local allocations, and cannot be
|
||||
made otherwise.** A Vulkan allocation in a device-local heap never enters the
|
||||
process's address space, so nothing under `/proc/self/status` can see it. What
|
||||
*does* land in RSS is the host-visible side — staging buffers, mapped upload
|
||||
rings, the read-back `AdjustPass` performs on export — plus the driver's own
|
||||
resident pages.
|
||||
|
||||
So "idle memory < 500 MB" is two questions wearing one number, and a build
|
||||
holding 400 MB of RSS and 3 GB of textures would pass it.
|
||||
|
||||
**Recommendation: NFR-P8 should be restated as two figures** — host RSS
|
||||
exclusive of device-local memory, and a separate VRAM ceiling read from the
|
||||
adapter — because the second is the one that decides whether the application
|
||||
survives beside a browser on an 8 GB card, and nothing in this repository
|
||||
measures it today.
|
||||
|
||||
**And a decision is outstanding.** `catalog_idle_rss_mb` carries no budget
|
||||
because nobody has decided how much of the 500 MB belongs to the catalog layer
|
||||
and how much to everything above it. The suite records the number so that
|
||||
decision can be taken against a measurement rather than an estimate. When it is
|
||||
taken, put the figure in `budget` and the metric becomes a gate.
|
||||
|
||||
---
|
||||
|
||||
## What is not measured, and would be worth adding
|
||||
|
||||
- **The UI's own open.** `ui/dr-ui/src/library.rs` does not call
|
||||
`Catalog::count` or `Catalog::window`; it issues its own SQL against the same
|
||||
tables, with a `VISIBLE` predicate and a burst-folding clause. `dr-bench`
|
||||
cannot see those without depending on `dr-ui`, which would drag Slint into a
|
||||
job that has no display. **Falsifiable end:** when the grid's queries move
|
||||
down into `dr-catalog` — which is where SQL over catalog tables belongs —
|
||||
`catalog_open_ms` becomes the whole of the application's open and this caveat
|
||||
can be deleted rather than argued about.
|
||||
- **The remote sweep.** `spawn_thumbnail_sweep`'s wall clock against a real
|
||||
server is latency, not CPU, and is what FR-NC-3's design is judged by. It
|
||||
needs a server and belongs in a different kind of test.
|
||||
- **A cold disk.** See the fixture's limits above.
|
||||
- **Android.** §4.1 states a second column of targets and §8 asks for
|
||||
"periodically on the named reference Android devices". Nothing here runs on a
|
||||
device. Spike S10 is the piece of work that would start it.
|
||||
- **Everything with a frame in it.** See the out-of-scope list above.
|
||||
|
||||
---
|
||||
|
||||
## Running it
|
||||
|
||||
```sh
|
||||
# Measure and print. Judges nothing.
|
||||
cargo run --release -p dr-bench -- run
|
||||
|
||||
# Measure and gate. What CI runs.
|
||||
cargo run --release -p dr-bench -- check
|
||||
|
||||
# The same, with machine-sensitive budgets asserted too.
|
||||
cargo run --release -p dr-bench -- check --reference
|
||||
|
||||
# Rewrite bench-baseline.json from this run, and commit the diff.
|
||||
cargo run --release -p dr-bench -- record --reference
|
||||
```
|
||||
|
||||
Useful flags: `--fixture <dir>` (or `$DR_BENCH_DIR`) to put the synthetic
|
||||
catalog somewhere specific, `--lanes <n>` to pin the sweep's parallelism, and
|
||||
`--thumbnails <n>` to lengthen or shorten the throughput row.
|
||||
Reference in New Issue
Block a user