Record the benchmark baseline, so the regression gate turns on #42

Open
opened 2026-09-05 16:20:20 +00:00 by dtourolle · 1 comment
Owner

Every recorded field in docs/bench-baseline.json is null. Until the reference numbers are committed, the budget gate works and the regression gate does not.

Why this is the cheapest high-value ticket in the tracker

tools/bench is built: it constructs a deterministic 50,000-row catalog over a pool of a dozen real files, measures against it, and fails the build on a violated budget or a drift past tolerance. .gitea/workflows/benchmark.yml runs it on every push. NFR-P1 and NFR-P3 are genuinely gated, and R2's "catalog opens in under 2s" clause with them.

Half of that machinery is idle for want of one command being run once on the reference desktop.

The nulls are deliberate — a fabricated baseline is worse than none — so this is not an oversight to fix, it is a step to take.

Acceptance

  • dr-bench record --reference run on the reference desktop.
  • The machine, the profile and the date recorded alongside, as frame-budget.md already does for the GPU measurements.
  • bench-baseline.json committed with real figures.
  • A deliberate regression confirmed to fail CI — the gate is untested until something has failed it on purpose.

See docs/outstanding.md §8 and docs/benchmarks.md.

**Every `recorded` field in `docs/bench-baseline.json` is `null`.** Until the reference numbers are committed, the budget gate works and **the regression gate does not**. ## Why this is the cheapest high-value ticket in the tracker `tools/bench` is built: it constructs a deterministic 50,000-row catalog over a pool of a dozen real files, measures against it, and fails the build on a violated budget or a drift past tolerance. `.gitea/workflows/benchmark.yml` runs it on every push. NFR-P1 and NFR-P3 are genuinely gated, and R2's "catalog opens in under 2s" clause with them. Half of that machinery is idle for want of one command being run once on the reference desktop. The nulls are **deliberate** — a fabricated baseline is worse than none — so this is not an oversight to fix, it is a step to take. ## Acceptance - [ ] `dr-bench record --reference` run on the reference desktop. - [ ] The machine, the profile and the date recorded alongside, as `frame-budget.md` already does for the GPU measurements. - [ ] `bench-baseline.json` committed with real figures. - [ ] A deliberate regression confirmed to fail CI — the gate is untested until something has failed it on purpose. See `docs/outstanding.md` §8 and `docs/benchmarks.md`.
dtourolle added the size:Sperformance labels 2026-09-05 16:20:20 +00:00
Author
Owner

Turns on the gate that #43 and #46 both depend on. Cheapest high-value item in the tracker: one command, run once, on the reference desktop.

**Turns on the gate that** #43 and #46 both depend on. Cheapest high-value item in the tracker: one command, run once, on the reference desktop.
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: dtourolle/DarkRoom#42