Record what the reduced base actually bought

TD-4 asked for the measurement as well as the change, and this is it: 25.05 ms
to 4.17 ms at 3840 x 2160, six times faster, with clarity no longer dominating
the neighbourhood stage it used to be 97% of.

Measured before and after on the same machine and the same adapter minutes
apart, baseline at the branch's merge-base, so the only variable is the change.
That adapter is not the RTX 3050 the rest of this document was measured on, so
the new table says to read it on its own rather than against the ones above —
the before/after is comparable, the absolute figures are not, and quietly
replacing the existing tables would have changed the instrument.

Also recorded: the declared halo is now quantised to multiples of the output
scale, because the 2-sigma truncation rounds on the reduced grid. 29 px becomes
28 at 1920x1200 and 38 becomes 40 at 2560x1600. It is inside what the cross-form
test holds — 0.03 stops of peak, 2% of reach — but it is a change in reach and
not only in cost, and a tile scheduler would be handed it. Better written down
now than found later as a seam.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
2026-08-29 13:18:37 +02:00
co-authored by Claude Opus 5
parent bff95e25ad
commit 8b251b84a0
2 changed files with 106 additions and 2 deletions
+31 -1
View File
@@ -142,7 +142,7 @@ in well under a second.
---
## TD-4 — The local-contrast base is computed at full render resolution
## TD-4 — The local-contrast base is computed at full render resolution ✅ PAID OFF
**Where:** `dr_pipeline::ops::local_contrast::LocalContrast::passes` — the `base` and `combine`
passes, and the stage that dispatches them, `dr_pipeline::detail`.
@@ -201,6 +201,36 @@ resolution; the scale therefore belongs on the `DetailPass`, not on the stage.
`tests/local_contrast.rs` still pass unchanged, and `examples/frame_budget`'s M3 table in
[frame-budget.md](frame-budget.md) has been rerun and committed.
### Paid off
A `DetailPass` now declares `output_scale`, and clarity's base is computed on a grid a quarter the
size on each axis. Measured before and after on the same machine, same adapter, same build profile,
with only the change between them — see [frame-budget.md](frame-budget.md) §"The reduced base,
measured":
| viewport, fit | before | after | |
|---|---:|---:|---:|
| 1920 × 1200 | 3.94 ms | 1.97 ms | 2.0× |
| 2560 × 1600 | 10.40 ms | 2.37 ms | 4.4× |
| 3840 × 2160 | **25.05 ms** | **4.17 ms** | **6.0×** |
Clarity is no longer the stage that misses the budget, and no longer dominates the neighbourhood
stage: at 4K it is 4.17 ms against 4.61 ms for all four neighbourhood operations together, where it
was 97% of that total at every size.
**Two things worth recording, because neither is visible in the table.**
The declared halo is now quantised to multiples of `output_scale`. The kernel truncates at 2σ and
that rounding now happens on the reduced grid, so 1920 × 1200 reports 28 render pixels where it
reported 29, and 2560 × 1600 reports 40 where it reported 38. At 2σ the Gaussian is already down to
`e⁻²` of its peak, and the cross-form test holds the difference to 0.03 stops of peak excursion and
2% of frame reach — but it is a real change in reach, not a pure speed-up, and a tile scheduler
would see it.
The measurement was taken on an AMD RX 5700 XT, not the RTX 3050 the M1/M2/M3 tables above were
measured on, so the *absolute* figures are not comparable with those. The before/after is, because
both halves of it were measured on the same card minutes apart.
---
## TD-5 — The fused shader is reassembled from strings on every frame