Read the fit view's source gather once per framing, not once per frame
At fit, every output pixel of the fused pass loads one texel from a source three or four times its width, on a stride. The memory system fetches the texels it skips along with the one it wanted, so on a 60 MP rgba16float source that gather was most of what the fused pass cost: 10.6 ms of a 2560x1600 frame against 3.8 ms for the same shader reading a contiguous window (the 1:1 view). At 3840x2160 it was 21.1 ms. Those are the laptop RTX 3050 with its clocks held at 420/810 MHz by the power cap; unthrottled the same frames were about 2.0 and 3.2 ms, and the gather is the same share of them. Which texel an output pixel reads depends only on the framing prologue, the framing and warp uniforms, the source and the render size. None of those move during a slider drag, so the gather is the same work every frame. The fused shader now takes a render-sized rgba16float cache of it (bindings 6 and 7, declared in every generated shader like the masks) and a pair of uniform flags: write what was gathered, or read it back at the pixel's own coordinate. AdjustPass keeps the cache and decides per dispatch. The composer supplies `ComposedShader::sample_key`, a hash of the prologue and those uniforms, and AdjustPass adds the image and the size; an image gets a process-unique id for this rather than being held alive by the key. The picture is bit-for-bit the same. The source is rgba16float and so is the cache, so the stored texel is the texel, and only the path that reads a texel whole takes part: an interpolated sample (straightening, lens warps, CA) is a blend that f16 could not hold exactly, so the composer gives it no key and it reads directly as before. The cache is written on the second frame with a given key, not the first: a crop or zoom drag changes the key every frame, and writing then would add a render-sized write to exactly the gestures that can afford it least. It is kept only up to 3840x2400, so an export never parks a full-frame copy on the device, and `release_caches` drops it. Measured with a scratch probe rendering the synthetic 60 MP frame from examples/frame_budget.rs, forty frames per run after six warm-up, five runs of each binary alternated, median of the per-run p50 (GPU idle apart from the power cap): scene before after neutral 2560x1600 fit 10.62 ms 3.88 ms exposure 2560x1600 fit 10.83 ms 3.87 ms nr chroma 2560x1600 fit 19.84 ms 12.69 ms neutral 3840x2160 fit 21.05 ms 7.11 ms exposure 3840x2160 fit 21.08 ms 6.94 ms clarity 3840x2160 fit 42.20 ms 27.88 ms neutral 2560x1600 1:1 3.83 ms 3.84 ms (control: nothing to gain) The rgba8 output of every scene hashed identically before and after, in isolated runs and across all 38 scene/size/view combinations of the probe. New tests walk a pass through direct, write and read frames, a slider move, a neighbourhood operation and a framing change, and compare every frame with a fresh pass that can only have read directly.
This commit is contained in:
@@ -67,7 +67,7 @@ pub use lens::{compose_warps, ComposedWarp, LensProfile, Tca, Warp};
|
||||
pub use operation::{
|
||||
compose, compose_with_framing, Affects, ComposedShader, Helper, Invalidation, Operation,
|
||||
OutputMode, Uniform, BASE_CURVE_POINTS, BASE_CURVE_UNIFORM_OFFSET, CLIP_ONSET,
|
||||
RESERVED_UNIFORM_FIELDS,
|
||||
RESERVED_UNIFORM_FIELDS, SAMPLE_CACHE_UNIFORM_OFFSET,
|
||||
};
|
||||
pub use preset::{LibraryParseError, NameError, Preset, PresetLibrary, Scope};
|
||||
pub use sidecar::{Sidecar, Version};
|
||||
|
||||
Reference in New Issue
Block a user