Files
DarkRoom/.gitea/workflows/benchmark.yml
T
dtourolle e8e96eed40 Measure the performance targets §8 has been promising, and fail on a regression
docs/requirements.md §8 has said since it was written that performance is
verified by "an automated benchmark suite against a synthetic 50k catalog, run
per-commit … A regression beyond stated tolerance fails the build." There was
none. No benches/, no [[bench]], no criterion, no synthetic catalog, and three
CI workflows that between them measured nothing. Ten performance requirements
could therefore be neither passed nor failed, and five of them carried a
TRACES: tag regardless.

tools/bench is the half of that promise that can be kept honestly on a runner
with no GPU and no display.

# The fixture

Rows are cheap and pixels are not, so it builds fifty thousand catalog rows
over a pool of a dozen real files, each referenced by several thousand of them.
Everything the catalog half touches is rows and is exact at full scale;
everything the pixel half touches is one file at a time and does not care how
many rows point at it. Fourteen megabytes on disk instead of two terabytes, and
neither half is flattered by the trade. It is reproducible from a seed, and a
stamp beside it — seed, row count, source size, dr-catalog's schema version —
rebuilds it rather than letting a run be compared against a baseline that
describes a different library.

# What it can now pass or fail

NFR-P1, and R2's second sentence with it: Catalog::open plus the count, first
window and timeline the grid cannot paint without. The interesting part turned
out to be the open itself — schema::backfill runs three passes over the images
table on every open, which is O(library) work on a path whose budget is stated
in absolute seconds. Tagged TRACES: NFR-P1, on a gate that fails if it breaks.

NFR-P3: thumbnail throughput on the embedded preview path, through the same
per-image work spawn_thumbnail_sweep does and in the same shape — chunks of 96,
lanes owning disjoint slices, the single thread that owns the store writing the
finished chunk. Mirrored rather than called, because that function takes a
RemoteBackend and would measure somebody's network. Tagged TRACES: NFR-P3.

# What it deliberately does not claim

NFR-P7 is the whole chain, and only the encode half of it runs without an
adapter. So the export row is a one-sided gate — over two seconds in the encode
alone violates the requirement; under it proves nothing — and there is no
TRACES: NFR-P7 anywhere. NFR-P8 is about the application at idle, and the probe
is a process holding the catalog and nothing else, so it records the catalog
layer's share and carries no budget until somebody decides what that share
should be. No tag there either. CONTRIBUTING.md asks that a requirement be
closed by a test that would fail if the behaviour were removed, and two more
plumbing tags is what this repository already has too many of.

NFR-P8 also gets the answer §4.1 demands: RSS is exclusive of device-local GPU
allocations and cannot be made otherwise, because such an allocation never
enters the process's address space. The requirement should be restated as two
figures, and docs/benchmarks.md says so.

# Two gates, and why one of them steps aside off the reference desktop

The budget is the requirement's own number and never moves. The baseline is
what the reference desktop last measured, and drifting 15% past it fails the
build even while still inside the budget — which is how performance rot
actually arrives, never over the line, always a little worse.

A budget written for twenty-four threads cannot be asserted on a two-core
container. §8 names the reference desktop, not CI, so each metric declares
whether its budget is machine-sensitive; those are asserted under --reference
and reported everywhere else. Catalog open is not one of them: two seconds
against an expected figure two orders of magnitude smaller is a threshold any
machine can be held to. This is the trap core/dr-gpu/tests/frame_budget.rs
already refuses — a red gate everybody learns to ignore.

# The baseline ships with no numbers in it

Every recorded field is null, because nobody has run it yet. Writing
plausible-looking figures would make every later comparison a comparison
against a guess, and the first real regression would be invisible. Run
`dr-bench record --reference` on the reference desktop and commit the diff;
until then the budget gate works and the report says the other one cannot.

# CI

.gitea/workflows/benchmark.yml, and its own workflow rather than a step in
build-and-test.yml: a red "Build and test" says the code is wrong, a red
"Benchmarks" says it got slower, and the second must not be reachable by
retrying a flaky compile. The cpu job runs on every push and builds -p dr-bench
alone — which is why that crate depends on no GPU and no UI crate. The gpu job
is the frame budget that already exists and already skips without an adapter,
on workflow_dispatch, because building wgpu on every commit to rediscover that
the runner has no device is not a use of anybody's minutes.
2026-08-30 10:40:10 +02:00

194 lines
8.4 KiB
YAML

name: Benchmarks
# The suite docs/requirements.md §8 has been promising since it was written:
# "an automated benchmark suite against a synthetic 50k catalog, run per-commit
# … A regression beyond stated tolerance fails the build."
#
# Its own workflow rather than a step inside build-and-test.yml, and the reason
# is what a failure here means. A red `Build and test` says the code is wrong; a
# red `Benchmarks` says the code is slower than it was, which is a different
# conversation, is read by different people, and must not be reachable by
# retrying a flaky compile.
#
# # Why this is split in two
#
# §8 names "the reference desktop", not CI, and it is right to. So:
#
# cpu — runs on every push. It needs no adapter and no display, and the
# budgets it asserts (a 50k catalog opening inside two seconds) have
# two orders of magnitude of headroom, so a modest runner can be held
# to them honestly. Machine-sensitive budgets — throughput targets
# written for a 24-thread desktop — are reported here rather than
# asserted; `dr-bench` decides that per metric and says so in its
# report. Asserting them on a two-core container would produce exactly
# what core/dr-gpu/tests/frame_budget.rs refused to produce: a red gate
# everybody learns to ignore.
#
# gpu — the frame budget, which already exists and already skips itself where
# there is no adapter. Not on push: it would build wgpu and naga on
# every commit to establish, every time, that this runner has no GPU. It
# runs on demand (Actions → Run workflow) so that a runner that *does*
# have one can be pointed at it, and the numbers it produces belong in
# docs/frame-budget.md by hand, as they already are.
on:
push:
branches: [main, master, develop]
pull_request:
branches: [main, master, develop]
workflow_dispatch:
jobs:
cpu:
runs-on: linux/amd64
name: CPU and I/O (per commit)
# Node for actions/checkout and actions/cache, which the bare runner image
# cannot execute. Rust is installed below.
container:
image: catthehacker/ubuntu:act-latest
env:
# Same reasoning as the desktop job in build-and-test.yml: incremental
# state exists to make the *second* build in a working tree fast, which is
# not a thing a fresh checkout has, and it fills the runner's disk.
CARGO_INCREMENTAL: 0
# The fixture, out of the workspace so actions/cache never picks it up.
# A 14 MB synthetic catalog is two seconds to regenerate and would
# otherwise be uploaded and downloaded on every push to save them.
DR_BENCH_DIR: /tmp/darkroom-bench
steps:
- name: Checkout
uses: actions/checkout@v4
# No `git lfs pull` here, deliberately. `dr-bench` depends on the catalog,
# the decoder, the thumbnail store and the encoder, and on nothing that
# reaches `dr-segment` — so the model this repository keeps in LFS is not
# part of this job's dependency graph and fetching it would be a minute
# spent on a file nothing opens.
- name: Cache cargo
uses: actions/cache@v4
with:
path: |
~/.cargo/registry
~/.cargo/git
target
key: bench-${{ runner.os }}-${{ hashFiles('**/Cargo.lock') }}
# Pinned to the workspace rust-version, as every other job here is: a
# floating toolchain turns an unrelated push into a mystery failure, and
# for a benchmark it would turn one into a mystery *regression*.
#
# rust-analyzer is named for the reason build-and-test.yml gives: rustup
# reconciles rust-toolchain.toml on the first cargo call whatever this
# step asks for, so naming it keeps the download inside the step that says
# it is installing things.
- name: Install Rust 1.92.0
run: |
set -e
curl -fsSL https://sh.rustup.rs | sh -s -- \
-y --no-modify-path --profile minimal --default-toolchain 1.92.0 \
--component rust-analyzer
echo "$HOME/.cargo/bin" >> "$GITHUB_PATH"
# `-p dr-bench`, not `--workspace`. The whole point of that crate having
# no GPU and no UI dependency is that this job resolves the catalog, the
# decoder and the encoders and stops there — a few minutes rather than the
# release build of Slint and wgpu the desktop job pays for.
#
# Release, and it is not optional: the workspace builds its own crates at
# opt-level = 0 in dev, and every figure this produces is dominated by
# this workspace's own code. A debug run would measure rustc.
- name: Build the suite
run: cargo build --release -p dr-bench
# Exit 1 is a violated budget or a regression past tolerance; exit 2 is
# the harness failing to run at all. Both fail the job, and the report
# above the failure says which.
- name: Measure, and gate
run: cargo run --release -p dr-bench -- check
- name: Disk after
if: always()
run: df -h /workspace 2>/dev/null || df -h .
gpu:
# On demand only — see the header. A runner with a Vulkan device can be
# pointed at this; one without will skip the measurement and say so, which
# is the same posture the rest of this repository's device tests take.
if: github.event_name == 'workflow_dispatch'
runs-on: linux/amd64
name: Frame budget (on demand)
container:
image: catthehacker/ubuntu:act-latest
env:
CARGO_INCREMENTAL: 0
steps:
- name: Checkout
uses: actions/checkout@v4
# `dr-gpu` depends on `dr-segment` for the watershed's pixel passes. Its
# default features are off, so no weights are compiled in — but the fetch
# is cheap insurance and its failure is not fatal. The header of the same
# step in build-and-test.yml explains why the extraheader is stripped
# rather than reused: two Authorization headers is a 400 from Gitea, one
# step after the batch call that had just succeeded.
- name: Fetch the segmentation model
continue-on-error: true
env:
LFS_TOKEN: ${{ secrets.GITEA_TOKEN || github.token }}
run: |
set -e
git lfs install --local
git config --local --get-regexp '^http\..*extraheader$' \
| cut -d' ' -f1 | sort -u \
| while read -r key; do git config --local --unset-all "$key"; done || true
git config --local lfs.url \
"https://x-access-token:${LFS_TOKEN}@gitea.tourolle.paris/dtourolle/DarkRoom.git/info/lfs"
git lfs pull
- name: Cache cargo
uses: actions/cache@v4
with:
path: |
~/.cargo/registry
~/.cargo/git
target
key: bench-gpu-${{ runner.os }}-${{ hashFiles('**/Cargo.lock') }}
- name: Build dependencies
run: |
apt-get update -qq
apt-get install -y -qq pkg-config libfontconfig1-dev libxkbcommon-dev
- name: Install Rust 1.92.0
run: |
set -e
curl -fsSL https://sh.rustup.rs | sh -s -- \
-y --no-modify-path --profile minimal --default-toolchain 1.92.0 \
--component rust-analyzer
echo "$HOME/.cargo/bin" >> "$GITHUB_PATH"
# The guard, in release. Its own module documentation is explicit that a
# release run checks strictly more than a dev one: the CPU half of a frame
# is shader-string assembly, which is several times slower unoptimised, so
# it is folded into the assertion only when debug_assertions is off.
#
# With no adapter this prints "skipping: no GPU adapter" and passes. A
# test that cannot run is not evidence either way, and turning that into a
# failure would make the job useless on the runner it usually lands on.
- name: Frame budget (FR-DSP-3)
run: cargo test --release -p dr-gpu --test frame_budget -- --nocapture
# The instrument behind docs/frame-budget.md. It exits non-zero with no
# adapter, which is right for a tool a person runs deliberately and wrong
# for a job that usually has none — hence continue-on-error. Its table is
# in the log for whoever asked for this run; the committed numbers are
# still updated by hand, as that file says.
- name: Frame budget table
continue-on-error: true
run: cargo run --release -p dr-gpu --example frame_budget