Tile the neighbourhood kernels: NR, sharpen and dehaze dominate the frames that use them #74

Closed
opened 2026-09-26 14:43:22 +00:00 by dtourolle · 1 comment
Owner

Neighbourhood kernels are now most of a develop frame that uses them. After v0.16.0's fit-view gather cache and dropping the empty detail pass: dehaze costs 26–34 ms, and every operation together ~67 ms, at 2560×1600 fit on the power-capped reference RTX 3050 (420/810 MHz). Each separable or bilateral tap reads through tap() straight from the texture.

Why

NFR-P5's slider latency fails first on the chains that include noise reduction, sharpening, texture or dehaze. The 2026-09-25 pass estimated these wins, but did not do them:

  • Shared-memory tiling of the taps in tap() codegen, which keeps the arithmetic order and so the output bit-identical: probably 1.5–2× on NR, sharpen and dehaze.
  • Dehaze's five passes sharing the erosion's x/y halves in workgroup memory: a similar size of win.
  • Fusing texture's and capture sharpening's x-pass with their combine: one render-sized rgba16f round trip saved per operation, ~2–4 ms, with care needed to stay bit-identical.

Workgroup shapes (16×8, 16×16, 32×8) were measured and gave nothing. Hoisting source_origin and textureDimensions out of tap() gave 0% on NVIDIA and is unmeasured on Adreno and Mali.

Deliverable

Tiled neighbourhood taps in the generated WGSL, within the limits the device already requests (Limits::default(), no features), including workgroup storage.

Acceptance

  • rgba8 output hashes identical before and after, per scene, for the frame-budget probe scenes
  • Before and after medians on the same GPU, held at the same clocks, run back to back
  • Measured on the reference tablet as well (see #23), or the Android path left unchanged
**Neighbourhood kernels are now most of a develop frame that uses them.** After v0.16.0's fit-view gather cache and dropping the empty detail pass: dehaze costs 26–34 ms, and every operation together ~67 ms, at 2560×1600 fit on the power-capped reference RTX 3050 (420/810 MHz). Each separable or bilateral tap reads through `tap()` straight from the texture. ## Why NFR-P5's slider latency fails first on the chains that include noise reduction, sharpening, texture or dehaze. The 2026-09-25 pass estimated these wins, but did not do them: - Shared-memory tiling of the taps in `tap()` codegen, which keeps the arithmetic order and so the output bit-identical: probably 1.5–2× on NR, sharpen and dehaze. - Dehaze's five passes sharing the erosion's x/y halves in workgroup memory: a similar size of win. - Fusing texture's and capture sharpening's x-pass with their combine: one render-sized rgba16f round trip saved per operation, ~2–4 ms, with care needed to stay bit-identical. Workgroup shapes (16×8, 16×16, 32×8) were measured and gave nothing. Hoisting `source_origin` and `textureDimensions` out of `tap()` gave 0% on NVIDIA and is unmeasured on Adreno and Mali. ## Deliverable Tiled neighbourhood taps in the generated WGSL, within the limits the device already requests (`Limits::default()`, no features), including workgroup storage. ## Acceptance - [ ] rgba8 output hashes identical before and after, per scene, for the frame-budget probe scenes - [ ] Before and after medians on the same GPU, held at the same clocks, run back to back - [ ] Measured on the reference tablet as well (see #23), or the Android path left unchanged
dtourolle added the pipelinesize:Mgpuperformance labels 2026-09-26 14:43:22 +00:00
Author
Owner

Closed with a measured outcome, released in v0.17.0 (bee5c58). Shared-memory tiling of tap() was built, bit-identical, and gave no gain on the desktop RTX 3050: within noise for rows, ~2 ms slower per pass for columns, and 14% slower for NR at 1:1 with squares. A detail pass there is dominated by its render-sized read and write (4.0 ms for an empty pass against ~4.6 ms for 17 taps), so removing passes is what pays. Dehaze now runs 2 passes instead of 5, bit-identical across 64 scene hashes: 2560×1600 fit 22.9 → 9.1 ms, and every op plus film 57.6 → 44.2 ms. Each pass now reads more texels, which is untested on the tablet; worth watching on Android. Fusing texture's and capture sharpening's x-pass with their combine (~4 ms each) was left out, because bit-identical f16 rounding in a shader cannot be guaranteed across Android drivers. The tiling diff is kept outside the repo.

Closed with a measured outcome, released in v0.17.0 (bee5c58). Shared-memory tiling of `tap()` was built, bit-identical, and gave no gain on the desktop RTX 3050: within noise for rows, ~2 ms slower per pass for columns, and 14% slower for NR at 1:1 with squares. A detail pass there is dominated by its render-sized read and write (4.0 ms for an empty pass against ~4.6 ms for 17 taps), so removing passes is what pays. Dehaze now runs 2 passes instead of 5, bit-identical across 64 scene hashes: 2560×1600 fit 22.9 → 9.1 ms, and every op plus film 57.6 → 44.2 ms. Each pass now reads more texels, which is untested on the tablet; worth watching on Android. Fusing texture's and capture sharpening's x-pass with their combine (~4 ms each) was left out, because bit-identical f16 rounding in a shader cannot be guaranteed across Android drivers. The tiling diff is kept outside the repo.
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: dtourolle/DarkRoom#74