From 8b251b84a06c64600467703a08e29a5e1c31ec61 Mon Sep 17 00:00:00 2001 From: Duncan Tourolle Date: Sat, 29 Aug 2026 13:18:37 +0200 Subject: [PATCH] Record what the reduced base actually bought MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit TD-4 asked for the measurement as well as the change, and this is it: 25.05 ms to 4.17 ms at 3840 x 2160, six times faster, with clarity no longer dominating the neighbourhood stage it used to be 97% of. Measured before and after on the same machine and the same adapter minutes apart, baseline at the branch's merge-base, so the only variable is the change. That adapter is not the RTX 3050 the rest of this document was measured on, so the new table says to read it on its own rather than against the ones above — the before/after is comparable, the absolute figures are not, and quietly replacing the existing tables would have changed the instrument. Also recorded: the declared halo is now quantised to multiples of the output scale, because the 2-sigma truncation rounds on the reduced grid. 29 px becomes 28 at 1920x1200 and 38 becomes 40 at 2560x1600. It is inside what the cross-form test holds — 0.03 stops of peak, 2% of reach — but it is a change in reach and not only in cost, and a tile scheduler would be handed it. Better written down now than found later as a seam. Co-Authored-By: Claude Opus 5 (1M context) --- docs/frame-budget.md | 76 +++++++++++++++++++++++++++++++++++++++++- docs/technical-debt.md | 32 +++++++++++++++++- 2 files changed, 106 insertions(+), 2 deletions(-) diff --git a/docs/frame-budget.md b/docs/frame-budget.md index 7b09057..9b74187 100644 --- a/docs/frame-budget.md +++ b/docs/frame-budget.md @@ -215,7 +215,14 @@ So FR-DSP-2 becomes what §2 said it actually is for this architecture: a scheduling concern for export and thumbnailing, both of which already run off the frame path. The interactive path does not tile. -### The stage that does need work — and it is not tiling +### The stage that did need work — and it is not tiling + +**Resolved.** The fix described below landed; the measurement is in +§[The reduced base, measured](#the-reduced-base-measured) at the foot of this +file, and `docs/technical-debt.md` TD-4 is closed. What follows is the +reasoning as it stood, kept because it is what the numbers above argue for and +because the tiling half of it is still live. + The measurement's real product is naming the stage. It is `local_contrast`, and the fix is stated in that module's own documentation: @@ -295,3 +302,70 @@ handful of floats per operation. So caching the source string against the structure hash and rebuilding only the uniforms would take these milliseconds to approximately nothing, on the path that needs them most. Worth its own entry in [technical-debt.md](technical-debt.md). + +--- + +## The reduced base, measured + +**Status:** Measured · 2026-08-29 · closes TD-4 + +`DetailPass` gained an `output_scale`, and clarity's base is now computed on a +grid a quarter the size on each axis — the change §M3 argued for above. + +**Read this table on its own, not against the ones above.** It was taken on a +different adapter, so the absolute figures are not comparable with the RTX 3050 +measurements this document is otherwise built from. What *is* comparable is the +before and the after, which were measured on the same machine, same card, same +release profile, minutes apart, with nothing between them but the change — the +baseline at `0407fb8` and the result at `bff95e2`. + +| | | +|---|---| +| Adapter | AMD Radeon RX 5700 XT (RADV NAVI10) (Vulkan) | +| Source | 9504 × 6336 (60.2 MP, 482 MB as `rgba16f`) | +| Baseline | `0407fb8`, the branch's merge-base | +| Result | `bff95e2` | +| Date | 2026-08-29 | + +### M3 — the neighbourhood stage alone, before and after + +| size | stage | view | before p99 | after p99 | | +|------------:|---------:|:-----|-----------:|----------:|-----:| +| 1920 × 1200 | clarity | fit | 3.94ms | 1.97ms | 2.0× | +| 1920 × 1200 | all four | fit | 2.57ms | 1.81ms | 1.4× | +| 1920 × 1200 | clarity | 1:1 | 3.07ms | 2.41ms | 1.3× | +| 2560 × 1600 | clarity | fit | 10.40ms | 2.37ms | 4.4× | +| 2560 × 1600 | all four | fit | 5.50ms | 2.64ms | 2.1× | +| 2560 × 1600 | clarity | 1:1 | 5.78ms | 3.85ms | 1.5× | +| 3840 × 2160 | clarity | fit | **25.05ms**|**4.17ms** | 6.0× | +| 3840 × 2160 | all four | fit | 25.48ms | 4.61ms | 5.5× | +| 3840 × 2160 | clarity | 1:1 | 27.60ms | 6.93ms | 4.0× | +| 3840 × 2160 | all four | 1:1 | 15.88ms | 7.74ms | 2.1× | + +The passes went from 2 to 4 and got faster, which is the point: three of the +four now run on a grid a sixteenth the area, and the fourth — the combine — is a +single bilinear read where it used to be a 105-tap convolution. + +**Clarity is no longer the stage that dominates.** At 4K fit it is 4.17 ms +against 4.61 ms for all four neighbourhood operations together; it was 97% of +that total at every size before. Every M1 and M2 row now sits inside the 16 ms +budget on this adapter, including the six `all` rows that carried `OVER`. + +### The one thing that is not a pure speed-up + +**The declared halo is now quantised to multiples of `output_scale`.** The +kernel truncates at 2σ, and that rounding now happens on the reduced grid +before being multiplied back up: + +| viewport | before | after | +|---|---:|---:| +| 1920 × 1200 | 29 px | 28 px | +| 2560 × 1600 | 38 px | 40 px | +| 3840 × 2160 | 52 px | 52 px | + +At 2σ the Gaussian is already down to `e⁻²` of its peak, and +`crossing_the_reduction_threshold_does_not_change_the_picture` holds the +difference between a quarter-scale and a half-scale base to 0.03 stops of peak +excursion and 2% of frame reach. But it is a change in reach rather than only +in cost, it is what a tile scheduler would be handed, and it is worth knowing +that the number moved rather than discovering it later as a seam. diff --git a/docs/technical-debt.md b/docs/technical-debt.md index 0c6dda8..406aadd 100644 --- a/docs/technical-debt.md +++ b/docs/technical-debt.md @@ -142,7 +142,7 @@ in well under a second. --- -## TD-4 — The local-contrast base is computed at full render resolution +## TD-4 — The local-contrast base is computed at full render resolution ✅ PAID OFF **Where:** `dr_pipeline::ops::local_contrast::LocalContrast::passes` — the `base` and `combine` passes, and the stage that dispatches them, `dr_pipeline::detail`. @@ -201,6 +201,36 @@ resolution; the scale therefore belongs on the `DetailPass`, not on the stage. `tests/local_contrast.rs` still pass unchanged, and `examples/frame_budget`'s M3 table in [frame-budget.md](frame-budget.md) has been rerun and committed. +### Paid off + +A `DetailPass` now declares `output_scale`, and clarity's base is computed on a grid a quarter the +size on each axis. Measured before and after on the same machine, same adapter, same build profile, +with only the change between them — see [frame-budget.md](frame-budget.md) §"The reduced base, +measured": + +| viewport, fit | before | after | | +|---|---:|---:|---:| +| 1920 × 1200 | 3.94 ms | 1.97 ms | 2.0× | +| 2560 × 1600 | 10.40 ms | 2.37 ms | 4.4× | +| 3840 × 2160 | **25.05 ms** | **4.17 ms** | **6.0×** | + +Clarity is no longer the stage that misses the budget, and no longer dominates the neighbourhood +stage: at 4K it is 4.17 ms against 4.61 ms for all four neighbourhood operations together, where it +was 97% of that total at every size. + +**Two things worth recording, because neither is visible in the table.** + +The declared halo is now quantised to multiples of `output_scale`. The kernel truncates at 2σ and +that rounding now happens on the reduced grid, so 1920 × 1200 reports 28 render pixels where it +reported 29, and 2560 × 1600 reports 40 where it reported 38. At 2σ the Gaussian is already down to +`e⁻²` of its peak, and the cross-form test holds the difference to 0.03 stops of peak excursion and +2% of frame reach — but it is a real change in reach, not a pure speed-up, and a tile scheduler +would see it. + +The measurement was taken on an AMD RX 5700 XT, not the RTX 3050 the M1/M2/M3 tables above were +measured on, so the *absolute* figures are not comparable with those. The before/after is, because +both halves of it were measured on the same card minutes apart. + --- ## TD-5 — The fused shader is reassembled from strings on every frame