Commit Graph
2 Commits
Author SHA1 Message Date
dtourolleandClaude Opus 5 8b251b84a0 Record what the reduced base actually bought
TD-4 asked for the measurement as well as the change, and this is it: 25.05 ms
to 4.17 ms at 3840 x 2160, six times faster, with clarity no longer dominating
the neighbourhood stage it used to be 97% of.

Measured before and after on the same machine and the same adapter minutes
apart, baseline at the branch's merge-base, so the only variable is the change.
That adapter is not the RTX 3050 the rest of this document was measured on, so
the new table says to read it on its own rather than against the ones above —
the before/after is comparable, the absolute figures are not, and quietly
replacing the existing tables would have changed the instrument.

Also recorded: the declared halo is now quantised to multiples of the output
scale, because the 2-sigma truncation rounds on the reduced grid. 29 px becomes
28 at 1920x1200 and 38 becomes 40 at 2560x1600. It is inside what the cross-form
test holds — 0.03 stops of peak, 2% of reach — but it is a change in reach and
not only in cost, and a tile scheduler would be handed it. Better written down
now than found later as a seam.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-29 13:18:37 +02:00
dtourolleandClaude Opus 5 13deaa2fbb Assert the frame budget, and commit the numbers behind the FR-DSP-2 verdict
FR-DSP-3 states a latency requirement and nothing checked it, which makes it a
wish. This adds the check and the measurements it guards.

`docs/frame-budget.md` is the bench's output with the reading of §2's decision
rule attached. The short version: every point-operation chain at every viewport
size, fit and at 1:1, is inside 16 ms at the 99th percentile — the widest is
4.5 ms of GPU at 4K — so FR-DSP-2 should be rewritten rather than implemented.
The measurement did find a stage that misses the budget, and it is the one §2
predicted: clarity's 52-pixel separable kernel costs 34 ms at 4K. Tiles make
that worse rather than better, since a tiled convolution reads a halo per tile;
the fix `local_contrast` already names for itself is a base computed at reduced
resolution.

The test guards the fused path and says so, at length, rather than quietly
excluding the expensive stage and letting the tag imply otherwise (§7). What it
asserts is exactly the claim the recommendation rests on: one dispatch over a
viewport-sized target, at a full chain, is comfortably inside a frame.

Two things the numbers forced:

- The two cases are one `#[test]`. As two they ran on a thread each, contended
  for the same device, and took the 1:1 case from 2.5 ms to 14.9 ms — a
  measurement of the harness that would have flickered either side of the
  budget forever.

- The CPU half of the frame is judged only in an optimised build. Composition
  is real per-frame work on the UI thread and belongs in the budget, but the
  workspace builds its own crates at `opt-level = 0` in dev and `cargo test` is
  a dev build, so measuring it there measures rustc. The GPU half is asserted
  either way.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-27 19:00:45 +02:00