Read the fit view's source gather once per framing, not once per frame

At fit, every output pixel of the fused pass loads one texel from a source
three or four times its width, on a stride. The memory system fetches the
texels it skips along with the one it wanted, so on a 60 MP rgba16float
source that gather was most of what the fused pass cost: 10.6 ms of a
2560x1600 frame against 3.8 ms for the same shader reading a contiguous
window (the 1:1 view). At 3840x2160 it was 21.1 ms. Those are the laptop
RTX 3050 with its clocks held at 420/810 MHz by the power cap; unthrottled
the same frames were about 2.0 and 3.2 ms, and the gather is the same
share of them.

Which texel an output pixel reads depends only on the framing prologue,
the framing and warp uniforms, the source and the render size. None of
those move during a slider drag, so the gather is the same work every
frame. The fused shader now takes a render-sized rgba16float cache of it
(bindings 6 and 7, declared in every generated shader like the masks) and
a pair of uniform flags: write what was gathered, or read it back at the
pixel's own coordinate. AdjustPass keeps the cache and decides per
dispatch. The composer supplies `ComposedShader::sample_key`, a hash of
the prologue and those uniforms, and AdjustPass adds the image and the
size; an image gets a process-unique id for this rather than being held
alive by the key.

The picture is bit-for-bit the same. The source is rgba16float and so is
the cache, so the stored texel is the texel, and only the path that reads
a texel whole takes part: an interpolated sample (straightening, lens
warps, CA) is a blend that f16 could not hold exactly, so the composer
gives it no key and it reads directly as before.

The cache is written on the second frame with a given key, not the first:
a crop or zoom drag changes the key every frame, and writing then would
add a render-sized write to exactly the gestures that can afford it least.
It is kept only up to 3840x2400, so an export never parks a full-frame
copy on the device, and `release_caches` drops it.

Measured with a scratch probe rendering the synthetic 60 MP frame from
examples/frame_budget.rs, forty frames per run after six warm-up, five
runs of each binary alternated, median of the per-run p50 (GPU idle apart
from the power cap):

  scene                    before     after
  neutral   2560x1600 fit  10.62 ms   3.88 ms
  exposure  2560x1600 fit  10.83 ms   3.87 ms
  nr chroma 2560x1600 fit  19.84 ms  12.69 ms
  neutral   3840x2160 fit  21.05 ms   7.11 ms
  exposure  3840x2160 fit  21.08 ms   6.94 ms
  clarity   3840x2160 fit  42.20 ms  27.88 ms
  neutral   2560x1600 1:1   3.83 ms   3.84 ms  (control: nothing to gain)

The rgba8 output of every scene hashed identically before and after, in
isolated runs and across all 38 scene/size/view combinations of the
probe. New tests walk a pass through direct, write and read frames, a
slider move, a neighbourhood operation and a framing change, and compare
every frame with a fresh pass that can only have read directly.
This commit is contained in:
2026-09-26 07:10:35 -04:00
parent 284fc4a456
commit 1dc7b45cfe
5 changed files with 521 additions and 28 deletions
File diff suppressed because one or more lines are too long