Neighbourhood kernels are now most of a develop frame that uses them. After v0.16.0's fit-view gather cache and dropping the empty detail pass: dehaze costs 26–34 ms, and every operation together ~67 ms, at 2560×1600 fit on the power-capped reference RTX 3050 (420/810 MHz). Each separable or bilateral tap reads through tap() straight from the texture.
Why
NFR-P5's slider latency fails first on the chains that include noise reduction, sharpening, texture or dehaze. The 2026-09-25 pass estimated these wins, but did not do them:
Shared-memory tiling of the taps in tap() codegen, which keeps the arithmetic order and so the output bit-identical: probably 1.5–2× on NR, sharpen and dehaze.
Dehaze's five passes sharing the erosion's x/y halves in workgroup memory: a similar size of win.
Fusing texture's and capture sharpening's x-pass with their combine: one render-sized rgba16f round trip saved per operation, ~2–4 ms, with care needed to stay bit-identical.
Workgroup shapes (16×8, 16×16, 32×8) were measured and gave nothing. Hoisting source_origin and textureDimensions out of tap() gave 0% on NVIDIA and is unmeasured on Adreno and Mali.
Deliverable
Tiled neighbourhood taps in the generated WGSL, within the limits the device already requests (Limits::default(), no features), including workgroup storage.
Acceptance
rgba8 output hashes identical before and after, per scene, for the frame-budget probe scenes
Before and after medians on the same GPU, held at the same clocks, run back to back
Measured on the reference tablet as well (see #23), or the Android path left unchanged
**Neighbourhood kernels are now most of a develop frame that uses them.** After v0.16.0's fit-view gather cache and dropping the empty detail pass: dehaze costs 26–34 ms, and every operation together ~67 ms, at 2560×1600 fit on the power-capped reference RTX 3050 (420/810 MHz). Each separable or bilateral tap reads through `tap()` straight from the texture.
## Why
NFR-P5's slider latency fails first on the chains that include noise reduction, sharpening, texture or dehaze. The 2026-09-25 pass estimated these wins, but did not do them:
- Shared-memory tiling of the taps in `tap()` codegen, which keeps the arithmetic order and so the output bit-identical: probably 1.5–2× on NR, sharpen and dehaze.
- Dehaze's five passes sharing the erosion's x/y halves in workgroup memory: a similar size of win.
- Fusing texture's and capture sharpening's x-pass with their combine: one render-sized rgba16f round trip saved per operation, ~2–4 ms, with care needed to stay bit-identical.
Workgroup shapes (16×8, 16×16, 32×8) were measured and gave nothing. Hoisting `source_origin` and `textureDimensions` out of `tap()` gave 0% on NVIDIA and is unmeasured on Adreno and Mali.
## Deliverable
Tiled neighbourhood taps in the generated WGSL, within the limits the device already requests (`Limits::default()`, no features), including workgroup storage.
## Acceptance
- [ ] rgba8 output hashes identical before and after, per scene, for the frame-budget probe scenes
- [ ] Before and after medians on the same GPU, held at the same clocks, run back to back
- [ ] Measured on the reference tablet as well (see #23), or the Android path left unchanged
Closed with a measured outcome, released in v0.17.0 (bee5c58). Shared-memory tiling of tap() was built, bit-identical, and gave no gain on the desktop RTX 3050: within noise for rows, ~2 ms slower per pass for columns, and 14% slower for NR at 1:1 with squares. A detail pass there is dominated by its render-sized read and write (4.0 ms for an empty pass against ~4.6 ms for 17 taps), so removing passes is what pays. Dehaze now runs 2 passes instead of 5, bit-identical across 64 scene hashes: 2560×1600 fit 22.9 → 9.1 ms, and every op plus film 57.6 → 44.2 ms. Each pass now reads more texels, which is untested on the tablet; worth watching on Android. Fusing texture's and capture sharpening's x-pass with their combine (~4 ms each) was left out, because bit-identical f16 rounding in a shader cannot be guaranteed across Android drivers. The tiling diff is kept outside the repo.
Closed with a measured outcome, released in v0.17.0 (bee5c58). Shared-memory tiling of `tap()` was built, bit-identical, and gave no gain on the desktop RTX 3050: within noise for rows, ~2 ms slower per pass for columns, and 14% slower for NR at 1:1 with squares. A detail pass there is dominated by its render-sized read and write (4.0 ms for an empty pass against ~4.6 ms for 17 taps), so removing passes is what pays. Dehaze now runs 2 passes instead of 5, bit-identical across 64 scene hashes: 2560×1600 fit 22.9 → 9.1 ms, and every op plus film 57.6 → 44.2 ms. Each pass now reads more texels, which is untested on the tablet; worth watching on Android. Fusing texture's and capture sharpening's x-pass with their combine (~4 ms each) was left out, because bit-identical f16 rounding in a shader cannot be guaranteed across Android drivers. The tiling diff is kept outside the repo.
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
Neighbourhood kernels are now most of a develop frame that uses them. After v0.16.0's fit-view gather cache and dropping the empty detail pass: dehaze costs 26–34 ms, and every operation together ~67 ms, at 2560×1600 fit on the power-capped reference RTX 3050 (420/810 MHz). Each separable or bilateral tap reads through
tap()straight from the texture.Why
NFR-P5's slider latency fails first on the chains that include noise reduction, sharpening, texture or dehaze. The 2026-09-25 pass estimated these wins, but did not do them:
tap()codegen, which keeps the arithmetic order and so the output bit-identical: probably 1.5–2× on NR, sharpen and dehaze.Workgroup shapes (16×8, 16×16, 32×8) were measured and gave nothing. Hoisting
source_originandtextureDimensionsout oftap()gave 0% on NVIDIA and is unmeasured on Adreno and Mali.Deliverable
Tiled neighbourhood taps in the generated WGSL, within the limits the device already requests (
Limits::default(), no features), including workgroup storage.Acceptance
Closed with a measured outcome, released in v0.17.0 (
bee5c58). Shared-memory tiling oftap()was built, bit-identical, and gave no gain on the desktop RTX 3050: within noise for rows, ~2 ms slower per pass for columns, and 14% slower for NR at 1:1 with squares. A detail pass there is dominated by its render-sized read and write (4.0 ms for an empty pass against ~4.6 ms for 17 taps), so removing passes is what pays. Dehaze now runs 2 passes instead of 5, bit-identical across 64 scene hashes: 2560×1600 fit 22.9 → 9.1 ms, and every op plus film 57.6 → 44.2 ms. Each pass now reads more texels, which is untested on the tablet; worth watching on Android. Fusing texture's and capture sharpening's x-pass with their combine (~4 ms each) was left out, because bit-identical f16 rounding in a shader cannot be guaranteed across Android drivers. The tiling diff is kept outside the repo.