Read the fit view's source gather once per framing, not once per frame
At fit, every output pixel of the fused pass loads one texel from a source three or four times its width, on a stride. The memory system fetches the texels it skips along with the one it wanted, so on a 60 MP rgba16float source that gather was most of what the fused pass cost: 10.6 ms of a 2560x1600 frame against 3.8 ms for the same shader reading a contiguous window (the 1:1 view). At 3840x2160 it was 21.1 ms. Those are the laptop RTX 3050 with its clocks held at 420/810 MHz by the power cap; unthrottled the same frames were about 2.0 and 3.2 ms, and the gather is the same share of them. Which texel an output pixel reads depends only on the framing prologue, the framing and warp uniforms, the source and the render size. None of those move during a slider drag, so the gather is the same work every frame. The fused shader now takes a render-sized rgba16float cache of it (bindings 6 and 7, declared in every generated shader like the masks) and a pair of uniform flags: write what was gathered, or read it back at the pixel's own coordinate. AdjustPass keeps the cache and decides per dispatch. The composer supplies `ComposedShader::sample_key`, a hash of the prologue and those uniforms, and AdjustPass adds the image and the size; an image gets a process-unique id for this rather than being held alive by the key. The picture is bit-for-bit the same. The source is rgba16float and so is the cache, so the stored texel is the texel, and only the path that reads a texel whole takes part: an interpolated sample (straightening, lens warps, CA) is a blend that f16 could not hold exactly, so the composer gives it no key and it reads directly as before. The cache is written on the second frame with a given key, not the first: a crop or zoom drag changes the key every frame, and writing then would add a render-sized write to exactly the gestures that can afford it least. It is kept only up to 3840x2400, so an export never parks a full-frame copy on the device, and `release_caches` drops it. Measured with a scratch probe rendering the synthetic 60 MP frame from examples/frame_budget.rs, forty frames per run after six warm-up, five runs of each binary alternated, median of the per-run p50 (GPU idle apart from the power cap): scene before after neutral 2560x1600 fit 10.62 ms 3.88 ms exposure 2560x1600 fit 10.83 ms 3.87 ms nr chroma 2560x1600 fit 19.84 ms 12.69 ms neutral 3840x2160 fit 21.05 ms 7.11 ms exposure 3840x2160 fit 21.08 ms 6.94 ms clarity 3840x2160 fit 42.20 ms 27.88 ms neutral 2560x1600 1:1 3.83 ms 3.84 ms (control: nothing to gain) The rgba8 output of every scene hashed identically before and after, in isolated runs and across all 38 scene/size/view combinations of the probe. New tests walk a pass through direct, write and read frames, a slider move, a neighbourhood operation and a framing change, and compare every frame with a fresh pass that can only have read directly.
This commit is contained in:
+406
-5
@@ -124,6 +124,9 @@ pub struct AdjustPass {
|
||||
/// Cleared by any render that does not write it, so a stale intermediate
|
||||
/// cannot survive a change of image and be handed to a later detail chain.
|
||||
colour_key: Option<(u64, u32, u32)>,
|
||||
/// The source texels the fused pass read on an earlier frame, kept so a
|
||||
/// slider drag at fit reads them contiguously. See [`SampleCache`].
|
||||
sample: SampleCache,
|
||||
/// Fused dispatches actually encoded. Exposed so a test can see the reuse
|
||||
/// above happening rather than take it on trust.
|
||||
colour_dispatches: usize,
|
||||
@@ -147,6 +150,205 @@ struct Target {
|
||||
const FILM_FORMAT: wgpu::TextureFormat = wgpu::TextureFormat::Rgba32Float;
|
||||
|
||||
/// TRACES: FR-DEV-3f
|
||||
/// What a cached sample is valid for: the composer's
|
||||
/// [`ComposedShader::sample_key`], the source image, and the render size.
|
||||
type SampleKey = (u64, u64, u32, u32);
|
||||
|
||||
/// The largest render the sample cache is kept for, in pixels.
|
||||
///
|
||||
/// 4K and a little over, which is every develop view there is. An export
|
||||
/// renders a whole sensor once and gains nothing from a cache it will not
|
||||
/// read again; without a ceiling, two exports of the same frame in a row
|
||||
/// would park a full-resolution copy of it on the device.
|
||||
const SAMPLE_CACHE_MAX_PIXELS: u64 = 3840 * 2400;
|
||||
|
||||
/// How the fused dispatch gets its source colour this frame.
|
||||
#[derive(Clone, Copy, PartialEq, Eq, Debug)]
|
||||
enum SampleUse {
|
||||
/// From the source, as it always did.
|
||||
Direct,
|
||||
/// From the source, and stored in the cache for the frames after.
|
||||
Write,
|
||||
/// From the cache.
|
||||
Read,
|
||||
}
|
||||
|
||||
/// TRACES: NFR-P5
|
||||
/// The fused pass's gather from the source, remembered across frames.
|
||||
///
|
||||
/// At fit, each output pixel of the fused pass reads one texel of a source
|
||||
/// three or four times its width, on a stride, and the memory system fetches
|
||||
/// the neighbours it skips along with it. On a 60 MP source that gather was
|
||||
/// most of the fused pass: 10.9 ms of a 2560 x 1600 frame against 3.8 ms for
|
||||
/// the same work reading contiguously (RTX 3050, clocks held down). But which
|
||||
/// texel a pixel reads depends on the framing and nothing else, and a slider
|
||||
/// drag does not move the framing. So the shader writes what it gathered to a
|
||||
/// render-sized texture on one frame and reads it back contiguously on every
|
||||
/// frame after, until the framing, the image or the size changes.
|
||||
///
|
||||
/// Bit-for-bit the same picture: the source is `rgba16float` and so is the
|
||||
/// cache, so the stored texel is the texel. Only the whole-texel sampling path
|
||||
/// takes part — [`ComposedShader::sample_key`] is `None` when the sample is
|
||||
/// interpolated.
|
||||
///
|
||||
/// **Written on the second frame with a key, not the first.** A drag of the
|
||||
/// crop or of a zoom changes the key on every frame, and a cache written then
|
||||
/// is never read — it would add a write per frame to exactly the gestures that
|
||||
/// can least afford one. Waiting for the key to repeat once costs a slider drag
|
||||
/// one uncached frame and costs a crop drag nothing.
|
||||
struct SampleCache {
|
||||
/// The cache itself, `rgba16float`, sampled and written as storage.
|
||||
target: Option<Target>,
|
||||
/// What `target` holds, once a dispatch has written it.
|
||||
holds: Option<SampleKey>,
|
||||
/// The key the previous fused dispatch had, cached or not.
|
||||
last: Option<SampleKey>,
|
||||
/// What the most recent plan decided. Read by the tests, which have no
|
||||
/// other way to tell a cached frame from an uncached one — the point being
|
||||
/// that the pictures are identical.
|
||||
last_use: SampleUse,
|
||||
/// Bound at `@binding(6)` when the cache is not being read.
|
||||
no_sampled: wgpu::TextureView,
|
||||
/// Bound at `@binding(7)` when the cache is not being written.
|
||||
no_sample_out: wgpu::TextureView,
|
||||
}
|
||||
|
||||
impl SampleCache {
|
||||
const FORMAT: wgpu::TextureFormat = DemosaicedImage::FORMAT;
|
||||
|
||||
fn new(ctx: &GpuContext) -> Self {
|
||||
let placeholder = |label, usage| {
|
||||
ctx.device
|
||||
.create_texture(&wgpu::TextureDescriptor {
|
||||
label: Some(label),
|
||||
size: wgpu::Extent3d {
|
||||
width: 1,
|
||||
height: 1,
|
||||
depth_or_array_layers: 1,
|
||||
},
|
||||
mip_level_count: 1,
|
||||
sample_count: 1,
|
||||
dimension: wgpu::TextureDimension::D2,
|
||||
format: Self::FORMAT,
|
||||
usage,
|
||||
view_formats: &[],
|
||||
})
|
||||
.create_view(&Default::default())
|
||||
};
|
||||
Self {
|
||||
target: None,
|
||||
holds: None,
|
||||
last: None,
|
||||
last_use: SampleUse::Direct,
|
||||
no_sampled: placeholder("adjust-no-sampled", wgpu::TextureUsages::TEXTURE_BINDING),
|
||||
no_sample_out: placeholder(
|
||||
"adjust-no-sample-out",
|
||||
wgpu::TextureUsages::STORAGE_BINDING,
|
||||
),
|
||||
}
|
||||
}
|
||||
|
||||
/// Decide how this dispatch samples, on the assumption that it will be
|
||||
/// submitted — call it after anything that can still fail.
|
||||
fn plan(
|
||||
&mut self,
|
||||
ctx: &GpuContext,
|
||||
source: &DemosaicedImage,
|
||||
shader: &ComposedShader,
|
||||
width: u32,
|
||||
height: u32,
|
||||
) -> SampleUse {
|
||||
self.last_use = self.decide(ctx, source, shader, width, height);
|
||||
self.last_use
|
||||
}
|
||||
|
||||
fn decide(
|
||||
&mut self,
|
||||
ctx: &GpuContext,
|
||||
source: &DemosaicedImage,
|
||||
shader: &ComposedShader,
|
||||
width: u32,
|
||||
height: u32,
|
||||
) -> SampleUse {
|
||||
let key = shader
|
||||
.sample_key
|
||||
.filter(|_| u64::from(width) * u64::from(height) <= SAMPLE_CACHE_MAX_PIXELS)
|
||||
.map(|k| (k, source.id(), width, height));
|
||||
let last = std::mem::replace(&mut self.last, key);
|
||||
let Some(key) = key else {
|
||||
return SampleUse::Direct;
|
||||
};
|
||||
if self.holds == Some(key) {
|
||||
return SampleUse::Read;
|
||||
}
|
||||
if last != Some(key) {
|
||||
return SampleUse::Direct;
|
||||
}
|
||||
|
||||
if !self
|
||||
.target
|
||||
.as_ref()
|
||||
.is_some_and(|t| t.width == width && t.height == height)
|
||||
{
|
||||
let texture = ctx.device.create_texture(&wgpu::TextureDescriptor {
|
||||
label: Some("adjust-sample-cache"),
|
||||
size: wgpu::Extent3d {
|
||||
width,
|
||||
height,
|
||||
depth_or_array_layers: 1,
|
||||
},
|
||||
mip_level_count: 1,
|
||||
sample_count: 1,
|
||||
dimension: wgpu::TextureDimension::D2,
|
||||
format: Self::FORMAT,
|
||||
usage: wgpu::TextureUsages::STORAGE_BINDING | wgpu::TextureUsages::TEXTURE_BINDING,
|
||||
view_formats: &[],
|
||||
});
|
||||
let view = texture.create_view(&Default::default());
|
||||
self.target = Some(Target {
|
||||
texture,
|
||||
view,
|
||||
width,
|
||||
height,
|
||||
});
|
||||
}
|
||||
// Marked as held now: the dispatch that writes it is submitted before
|
||||
// any that could read it, and one queue orders the two.
|
||||
self.holds = Some(key);
|
||||
SampleUse::Write
|
||||
}
|
||||
|
||||
/// Write the flags for `usage` into a fused uniform block.
|
||||
fn flag(usage: SampleUse, uniforms: &mut [f32]) {
|
||||
let o = dr_pipeline::SAMPLE_CACHE_UNIFORM_OFFSET;
|
||||
uniforms[o] = if usage == SampleUse::Read { 1.0 } else { 0.0 };
|
||||
uniforms[o + 1] = if usage == SampleUse::Write { 1.0 } else { 0.0 };
|
||||
}
|
||||
|
||||
/// The views for `@binding(6)` and `@binding(7)`.
|
||||
fn views(&self, usage: SampleUse) -> (wgpu::TextureView, wgpu::TextureView) {
|
||||
let cache = || {
|
||||
self.target
|
||||
.as_ref()
|
||||
.expect("planned with a target")
|
||||
.view
|
||||
.clone()
|
||||
};
|
||||
match usage {
|
||||
SampleUse::Direct => (self.no_sampled.clone(), self.no_sample_out.clone()),
|
||||
SampleUse::Write => (self.no_sampled.clone(), cache()),
|
||||
SampleUse::Read => (cache(), self.no_sample_out.clone()),
|
||||
}
|
||||
}
|
||||
|
||||
/// Forget everything; the next render starts over.
|
||||
fn release(&mut self) {
|
||||
self.target = None;
|
||||
self.holds = None;
|
||||
self.last = None;
|
||||
}
|
||||
}
|
||||
|
||||
/// A baked film stock, resident on the GPU.
|
||||
struct FilmTextures {
|
||||
curves: wgpu::TextureView,
|
||||
@@ -440,6 +642,7 @@ impl AdjustPass {
|
||||
camera_pipeline_layout,
|
||||
camera_target: None,
|
||||
colour_key: None,
|
||||
sample: SampleCache::new(ctx),
|
||||
colour_dispatches: 0,
|
||||
detail_dispatches: 0,
|
||||
}
|
||||
@@ -537,6 +740,30 @@ impl AdjustPass {
|
||||
},
|
||||
count: None,
|
||||
},
|
||||
// The sample cache, read and written. Present in every
|
||||
// layout for the reason the masks are, and bound to
|
||||
// placeholders whenever the flags leave it alone. See
|
||||
// `SampleCache`.
|
||||
wgpu::BindGroupLayoutEntry {
|
||||
binding: 6,
|
||||
visibility: wgpu::ShaderStages::COMPUTE,
|
||||
ty: wgpu::BindingType::Texture {
|
||||
sample_type: wgpu::TextureSampleType::Float { filterable: false },
|
||||
view_dimension: wgpu::TextureViewDimension::D2,
|
||||
multisampled: false,
|
||||
},
|
||||
count: None,
|
||||
},
|
||||
wgpu::BindGroupLayoutEntry {
|
||||
binding: 7,
|
||||
visibility: wgpu::ShaderStages::COMPUTE,
|
||||
ty: wgpu::BindingType::StorageTexture {
|
||||
access: wgpu::StorageTextureAccess::WriteOnly,
|
||||
format: SampleCache::FORMAT,
|
||||
view_dimension: wgpu::TextureViewDimension::D2,
|
||||
},
|
||||
count: None,
|
||||
},
|
||||
],
|
||||
})
|
||||
}
|
||||
@@ -716,7 +943,15 @@ impl AdjustPass {
|
||||
let (width, height) = (width.max(1), height.max(1));
|
||||
self.ensure_target(width, height);
|
||||
|
||||
let uniforms = Self::fused_uniforms(source, shader);
|
||||
// Compile first: `pipeline` takes &mut self, and the sample cache's
|
||||
// plan below assumes the dispatch it plans for is submitted, so nothing
|
||||
// after it may fail.
|
||||
let _ = self.pipeline(shader)?;
|
||||
|
||||
let mut uniforms = Self::fused_uniforms(source, shader);
|
||||
let sampling = self.sample.plan(&self.ctx, source, shader, width, height);
|
||||
SampleCache::flag(sampling, &mut uniforms);
|
||||
let (sampled, sample_out) = self.sample.views(sampling);
|
||||
|
||||
let params_buf = self
|
||||
.ctx
|
||||
@@ -727,8 +962,6 @@ impl AdjustPass {
|
||||
usage: wgpu::BufferUsages::UNIFORM,
|
||||
});
|
||||
|
||||
// Borrow order: compile first, since `pipeline` takes &mut self.
|
||||
let _ = self.pipeline(shader)?;
|
||||
let pipeline = self
|
||||
.cache
|
||||
.get(&shader.structure_hash)
|
||||
@@ -768,6 +1001,14 @@ impl AdjustPass {
|
||||
binding: 5,
|
||||
resource: wgpu::BindingResource::TextureView(self.film_lut_view()),
|
||||
},
|
||||
wgpu::BindGroupEntry {
|
||||
binding: 6,
|
||||
resource: wgpu::BindingResource::TextureView(&sampled),
|
||||
},
|
||||
wgpu::BindGroupEntry {
|
||||
binding: 7,
|
||||
resource: wgpu::BindingResource::TextureView(&sample_out),
|
||||
},
|
||||
],
|
||||
});
|
||||
|
||||
@@ -885,7 +1126,12 @@ impl AdjustPass {
|
||||
label: Some("adjust-detail-encoder"),
|
||||
});
|
||||
|
||||
let mut sampling = SampleUse::Direct;
|
||||
if !reuse {
|
||||
let mut uniforms = uniforms;
|
||||
sampling = self.sample.plan(&self.ctx, source, shader, width, height);
|
||||
SampleCache::flag(sampling, &mut uniforms);
|
||||
let (sampled, sample_out) = self.sample.views(sampling);
|
||||
let params_buf =
|
||||
self.ctx
|
||||
.device
|
||||
@@ -927,6 +1173,14 @@ impl AdjustPass {
|
||||
binding: 5,
|
||||
resource: wgpu::BindingResource::TextureView(&film_lut),
|
||||
},
|
||||
wgpu::BindGroupEntry {
|
||||
binding: 6,
|
||||
resource: wgpu::BindingResource::TextureView(&sampled),
|
||||
},
|
||||
wgpu::BindGroupEntry {
|
||||
binding: 7,
|
||||
resource: wgpu::BindingResource::TextureView(&sample_out),
|
||||
},
|
||||
],
|
||||
});
|
||||
let pipeline = self
|
||||
@@ -954,9 +1208,20 @@ impl AdjustPass {
|
||||
.expect("ensured above")
|
||||
.view
|
||||
.clone();
|
||||
let ran = self
|
||||
let ran = match self
|
||||
.detail
|
||||
.encode(&mut enc, detail, &target_view, width, height)?;
|
||||
.encode(&mut enc, detail, &target_view, width, height)
|
||||
{
|
||||
Ok(ran) => ran,
|
||||
Err(e) => {
|
||||
// Nothing is submitted, so a cache this frame was to write
|
||||
// holds nothing, and must not be read as though it did.
|
||||
if sampling == SampleUse::Write {
|
||||
self.sample.release();
|
||||
}
|
||||
return Err(e);
|
||||
}
|
||||
};
|
||||
self.ctx.queue.submit(Some(enc.finish()));
|
||||
self.detail_dispatches += ran;
|
||||
self.colour_key = Some((key, width, height));
|
||||
@@ -1090,6 +1355,7 @@ impl AdjustPass {
|
||||
self.detail.release_caches();
|
||||
self.targets = [None, None];
|
||||
self.colour_key = None;
|
||||
self.sample.release();
|
||||
}
|
||||
|
||||
/// How many distinct pipelines are compiled. Exposed for tests asserting
|
||||
@@ -1244,6 +1510,16 @@ impl AdjustPass {
|
||||
binding: 5,
|
||||
resource: wgpu::BindingResource::TextureView(self.film_lut_view()),
|
||||
},
|
||||
// The sample cache is the display's; a camera-space tap
|
||||
// reads its source directly (its flags are zero).
|
||||
wgpu::BindGroupEntry {
|
||||
binding: 6,
|
||||
resource: wgpu::BindingResource::TextureView(&self.sample.no_sampled),
|
||||
},
|
||||
wgpu::BindGroupEntry {
|
||||
binding: 7,
|
||||
resource: wgpu::BindingResource::TextureView(&self.sample.no_sample_out),
|
||||
},
|
||||
],
|
||||
});
|
||||
|
||||
@@ -2624,6 +2900,131 @@ mod tests {
|
||||
|
||||
/// A flat RGBA8 image on the JPEG path — already gamma-encoded, as a
|
||||
/// decoded JPEG is.
|
||||
/// A frame with a different value at every pixel, several times the size
|
||||
/// of the renders below, so that a fit view reads it on a stride and a
|
||||
/// texel read from the wrong place cannot go unnoticed.
|
||||
fn busy_image(ctx: &GpuContext) -> DemosaicedImage {
|
||||
let (w, h) = (97u32, 61u32);
|
||||
let mut data = Vec::with_capacity((w * h * 4) as usize);
|
||||
for y in 0..h {
|
||||
for x in 0..w {
|
||||
let n = (x.wrapping_mul(2_654_435_761) ^ y.wrapping_mul(1_640_531_527)) >> 7;
|
||||
data.extend_from_slice(&[n as u8, (n >> 8) as u8, (x * 2 + y) as u8, 255]);
|
||||
}
|
||||
}
|
||||
DemosaicedImage::from_rgba8(ctx, &data, w, h).expect("upload")
|
||||
}
|
||||
|
||||
/// Render `g` with a pass that has never seen it, which reads the source
|
||||
/// directly by construction: the reference a cached frame must equal.
|
||||
fn fresh(ctx: &GpuContext, img: &DemosaicedImage, g: &EditGraph) -> Vec<u8> {
|
||||
let mut pass = AdjustPass::new(ctx);
|
||||
let detail = g.compose_detail(img.size(), (23, 15));
|
||||
pass.render_detailed(img, &g.compose(), 23, 15, None, &detail, 1)
|
||||
.expect("render");
|
||||
assert_eq!(pass.sample.last_use, SampleUse::Direct);
|
||||
pass.export_pixels().expect("read").0
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn a_cached_sample_is_the_same_picture() {
|
||||
// TRACES: NFR-P5
|
||||
// The sample cache's whole claim: a frame that reads the source through
|
||||
// it is bit-for-bit the frame that reads the source directly. Walked
|
||||
// through each state — direct, writing, reading, reading after a
|
||||
// slider moved, and direct again once the framing moves — against a
|
||||
// fresh pass each time.
|
||||
let Some(ctx) = ctx() else { return };
|
||||
let img = busy_image(&ctx);
|
||||
let mut pass = AdjustPass::new(&ctx);
|
||||
let mut g = EditGraph::default_chain();
|
||||
|
||||
let frame = |pass: &mut AdjustPass, g: &EditGraph, key: u64| {
|
||||
let detail = g.compose_detail(img.size(), (23, 15));
|
||||
pass.render_detailed(&img, &g.compose(), 23, 15, None, &detail, key)
|
||||
.expect("render");
|
||||
(pass.sample.last_use, pass.export_pixels().expect("read").0)
|
||||
};
|
||||
|
||||
for (i, (expected, value)) in [
|
||||
(SampleUse::Direct, 0.3),
|
||||
(SampleUse::Write, 0.4),
|
||||
(SampleUse::Read, 0.5),
|
||||
(SampleUse::Read, -0.7),
|
||||
]
|
||||
.into_iter()
|
||||
.enumerate()
|
||||
{
|
||||
g.set_param(exposure::ID, exposure::EXPOSURE, value);
|
||||
let (used, pixels) = frame(&mut pass, &g, i as u64);
|
||||
assert_eq!(used, expected, "frame {i}");
|
||||
assert_eq!(pixels, fresh(&ctx, &img, &g), "frame {i} ({used:?})");
|
||||
}
|
||||
|
||||
// A neighbourhood operation: the fused pass writes the linear
|
||||
// intermediate instead, through the same sampling.
|
||||
g.set_param(
|
||||
dr_pipeline::ops::noise_reduction::ID,
|
||||
dr_pipeline::ops::noise_reduction::CHROMA,
|
||||
60.0,
|
||||
);
|
||||
for i in 10..13 {
|
||||
g.set_param(exposure::ID, exposure::EXPOSURE, i as f32 * 0.01);
|
||||
let (used, pixels) = frame(&mut pass, &g, i);
|
||||
assert_eq!(used, SampleUse::Read, "detail frame {i}");
|
||||
assert_eq!(pixels, fresh(&ctx, &img, &g), "detail frame {i}");
|
||||
}
|
||||
|
||||
// The framing moves: what was cached is for the old framing.
|
||||
g.set_param(dr_pipeline::framing::ID, dr_pipeline::framing::CROP_W, 0.6);
|
||||
let (used, pixels) = frame(&mut pass, &g, 20);
|
||||
assert_eq!(used, SampleUse::Direct, "a new framing reads directly");
|
||||
assert_eq!(pixels, fresh(&ctx, &img, &g));
|
||||
let (used, _) = frame(&mut pass, &g, 21);
|
||||
assert_eq!(used, SampleUse::Write, "and caches once it holds still");
|
||||
let (used, pixels) = frame(&mut pass, &g, 22);
|
||||
assert_eq!(used, SampleUse::Read);
|
||||
assert_eq!(pixels, fresh(&ctx, &img, &g));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn a_cache_is_not_read_for_another_image_or_size() {
|
||||
// The key is the composer's half plus the two things only this side
|
||||
// knows. A second photograph with the same edit and the same framing
|
||||
// must not be shown the first one's texels.
|
||||
let Some(ctx) = ctx() else { return };
|
||||
let (a, b) = (busy_image(&ctx), grey_image(&ctx, 8000));
|
||||
let mut pass = AdjustPass::new(&ctx);
|
||||
let shader = EditGraph::default_chain().compose();
|
||||
for _ in 0..3 {
|
||||
pass.render(&a, &shader, 23, 15).expect("render");
|
||||
}
|
||||
assert_eq!(pass.sample.last_use, SampleUse::Read);
|
||||
|
||||
pass.render(&b, &shader, 23, 15).expect("render");
|
||||
assert_eq!(pass.sample.last_use, SampleUse::Direct, "another image");
|
||||
pass.render(&b, &shader, 23, 15).expect("render");
|
||||
pass.render(&b, &shader, 24, 15).expect("render");
|
||||
assert_eq!(pass.sample.last_use, SampleUse::Direct, "another size");
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn a_straightened_frame_samples_directly() {
|
||||
// Interpolated: the sample is a blend of four texels, which the cache's
|
||||
// format could not hold exactly, so the composer offers no key.
|
||||
let Some(ctx) = ctx() else { return };
|
||||
let img = busy_image(&ctx);
|
||||
let mut pass = AdjustPass::new(&ctx);
|
||||
let mut g = EditGraph::default_chain();
|
||||
g.set_param(dr_pipeline::framing::ID, dr_pipeline::framing::ANGLE, 3.0);
|
||||
let shader = g.compose();
|
||||
assert!(shader.sample_key.is_none());
|
||||
for _ in 0..3 {
|
||||
pass.render(&img, &shader, 23, 15).expect("render");
|
||||
assert_eq!(pass.sample.last_use, SampleUse::Direct);
|
||||
}
|
||||
}
|
||||
|
||||
fn jpeg_image(ctx: &GpuContext, rgb: [u8; 3]) -> DemosaicedImage {
|
||||
let size = 16u32;
|
||||
let mut data = Vec::with_capacity((size * size) as usize * 4);
|
||||
|
||||
@@ -95,6 +95,15 @@ pub struct DemosaicedImage {
|
||||
base_curve: BaseCurve,
|
||||
/// Whether the texture holds gamma-encoded rather than linear values.
|
||||
non_linear: bool,
|
||||
/// Which upload this is, unique for the life of the process. See
|
||||
/// [`Self::id`].
|
||||
id: u64,
|
||||
}
|
||||
|
||||
/// The next [`DemosaicedImage::id`].
|
||||
fn next_image_id() -> u64 {
|
||||
static NEXT: std::sync::atomic::AtomicU64 = std::sync::atomic::AtomicU64::new(1);
|
||||
NEXT.fetch_add(1, std::sync::atomic::Ordering::Relaxed)
|
||||
}
|
||||
|
||||
impl DemosaicedImage {
|
||||
@@ -112,6 +121,18 @@ impl DemosaicedImage {
|
||||
(self.width, self.height)
|
||||
}
|
||||
|
||||
/// Which texture this is, as a number that is never reused.
|
||||
///
|
||||
/// For a cache that has to know it is still looking at the same pixels
|
||||
/// (`AdjustPass`'s sample cache) without holding the texture alive to find
|
||||
/// out: keeping a handle would keep half a gigabyte of a closed photograph
|
||||
/// on the device, and comparing addresses would mistake a new upload for an
|
||||
/// old one the moment the allocator reused the slot. A texture here is
|
||||
/// never written after it is built, so the same id is the same pixels.
|
||||
pub(crate) fn id(&self) -> u64 {
|
||||
self.id
|
||||
}
|
||||
|
||||
/// Camera RGB → linear sRGB, row-major. Identity where the body is
|
||||
/// uncalibrated, so the image renders uncalibrated rather than black.
|
||||
pub fn color_matrix(&self) -> [f32; 9] {
|
||||
@@ -246,6 +267,7 @@ impl DemosaicedImage {
|
||||
// the highlights of an image that was already finished.
|
||||
base_curve: BaseCurve::IDENTITY,
|
||||
non_linear: true,
|
||||
id: next_image_id(),
|
||||
})
|
||||
}
|
||||
}
|
||||
@@ -322,6 +344,7 @@ impl DemosaicedImage {
|
||||
as_shot_wb: [raw.wb_coeffs[0], raw.wb_coeffs[1], raw.wb_coeffs[2]],
|
||||
base_curve: raw.base_curve,
|
||||
non_linear: false,
|
||||
id: next_image_id(),
|
||||
})
|
||||
}
|
||||
}
|
||||
@@ -660,6 +683,7 @@ impl Demosaicer {
|
||||
// normalises against black and white levels and applies no
|
||||
// transfer function.
|
||||
non_linear: false,
|
||||
id: next_image_id(),
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
@@ -67,7 +67,7 @@ pub use lens::{compose_warps, ComposedWarp, LensProfile, Tca, Warp};
|
||||
pub use operation::{
|
||||
compose, compose_with_framing, Affects, ComposedShader, Helper, Invalidation, Operation,
|
||||
OutputMode, Uniform, BASE_CURVE_POINTS, BASE_CURVE_UNIFORM_OFFSET, CLIP_ONSET,
|
||||
RESERVED_UNIFORM_FIELDS,
|
||||
RESERVED_UNIFORM_FIELDS, SAMPLE_CACHE_UNIFORM_OFFSET,
|
||||
};
|
||||
pub use preset::{LibraryParseError, NameError, Preset, PresetLibrary, Scope};
|
||||
pub use sidecar::{Sidecar, Version};
|
||||
|
||||
@@ -456,6 +456,25 @@ pub struct ComposedShader {
|
||||
pub structure_hash: u64,
|
||||
/// What this shader writes. See [`OutputMode`].
|
||||
pub output_mode: OutputMode,
|
||||
/// What decides which source texel each output pixel reads, when that
|
||||
/// texel is read whole — `None` when it is interpolated.
|
||||
///
|
||||
/// A fit view reads one texel in every three or four of a 60 MP source,
|
||||
/// on a stride, and that gather is most of what the fused pass costs
|
||||
/// there: the texel it wants shares a cache line with neighbours nobody
|
||||
/// reads. But the gather depends on the framing and nothing else, so it
|
||||
/// is the same on every frame of a slider drag. The shader can therefore
|
||||
/// write what it gathered to a viewport-sized texture once and read it
|
||||
/// back contiguously thereafter; the flags in the uniform block at
|
||||
/// [`SAMPLE_CACHE_UNIFORM_OFFSET`] say which, and `dr-gpu` decides.
|
||||
///
|
||||
/// This key is the half of that decision only the composer can make: a
|
||||
/// hash of the generated prologue and the framing and warp uniforms, which
|
||||
/// together are everything that maps an output pixel to a source texel.
|
||||
/// The caller mixes in the source image and the render size. `None` for
|
||||
/// the interpolating paths, whose sample is a blend of four texels and
|
||||
/// not representable exactly in the source's own format.
|
||||
pub sample_key: Option<u64>,
|
||||
}
|
||||
|
||||
/// Fields the generated uniform struct always carries, before op uniforms.
|
||||
@@ -466,7 +485,19 @@ pub struct ComposedShader {
|
||||
/// Twelve of the twenty-eight are the camera profile's base curve
|
||||
/// ([`BASE_CURVE_UNIFORM_FIELDS`]); the rest are the matrix, the as-shot
|
||||
/// balance and framing's own block.
|
||||
const BASE_UNIFORM_FIELDS: usize = 16 + BASE_CURVE_UNIFORM_FIELDS;
|
||||
const BASE_UNIFORM_FIELDS: usize = 16 + SAMPLE_CACHE_UNIFORM_FIELDS + BASE_CURVE_UNIFORM_FIELDS;
|
||||
|
||||
/// Slots the sample cache's two flags occupy: read, write, and two spare to
|
||||
/// keep the block a whole `vec4`. See [`ComposedShader::sample_key`].
|
||||
const SAMPLE_CACHE_UNIFORM_FIELDS: usize = 4;
|
||||
|
||||
/// Where the sample cache's flags sit in the generated uniform block: `x` says
|
||||
/// read the source colour from the cache, `y` says write it there.
|
||||
///
|
||||
/// Exported for the reason [`BASE_CURVE_UNIFORM_OFFSET`] is — `dr-gpu` writes
|
||||
/// these by index — and zero in every block the composer hands out, so a
|
||||
/// caller that never heard of the cache gets the direct read it always had.
|
||||
pub const SAMPLE_CACHE_UNIFORM_OFFSET: usize = 16;
|
||||
|
||||
/// TRACES: FR-DEV-3e
|
||||
/// Slots the base curve occupies: five `(x, y)` points and an active flag.
|
||||
@@ -484,7 +515,8 @@ const BASE_CURVE_UNIFORM_FIELDS: usize = 12;
|
||||
/// Exported for the same reason [`RESERVED_UNIFORM_FIELDS`] is: `dr-gpu`
|
||||
/// writes these by index, and an offset computed independently at both ends is
|
||||
/// an offset that will eventually disagree with itself.
|
||||
pub const BASE_CURVE_UNIFORM_OFFSET: usize = 16;
|
||||
pub const BASE_CURVE_UNIFORM_OFFSET: usize =
|
||||
SAMPLE_CACHE_UNIFORM_OFFSET + SAMPLE_CACHE_UNIFORM_FIELDS;
|
||||
|
||||
/// How many control points a base curve carries.
|
||||
///
|
||||
@@ -756,6 +788,10 @@ fn compose_inner(
|
||||
\x20 // gamma-encoded JPEG, 0.0 for demosaiced sensor data), which the\n\
|
||||
\x20 // prologue reads to decide whether to linearise.\n\
|
||||
\x20 as_shot_wb: vec4<f32>,\n\
|
||||
\x20 // The sample cache (see `ComposedShader::sample_key`): `.x` reads\n\
|
||||
\x20 // the source colour from `sampled`, `.y` writes it to\n\
|
||||
\x20 // `sample_out`. Zero for both is the direct read.\n\
|
||||
\x20 sample_cache: vec4<f32>,\n\
|
||||
\x20 // The camera profile's base curve (FR-DEV-3e): five points on a\n\
|
||||
\x20 // monotone spline, packed as x0..x3, y0..y3, then (x4, y4, on).\n\
|
||||
\x20 // `.z` of the last is the flag, not padding — it is 0 for a\n\
|
||||
@@ -936,6 +972,19 @@ fn compose_inner(
|
||||
);
|
||||
let sampler_helper = if interpolate { BILINEAR_HELPER } else { "" };
|
||||
|
||||
// Everything that decides which texel an output pixel reads: the code that
|
||||
// computes `coord`, and the uniforms that code reads. Only on the path that
|
||||
// reads a texel whole — see `ComposedShader::sample_key`.
|
||||
let sample_key = (!interpolate && !warp.splits_channels).then(|| {
|
||||
framing
|
||||
.uniforms()
|
||||
.iter()
|
||||
.chain(&warp.uniforms)
|
||||
.fold(hash_source(&prologue), |h, v| {
|
||||
mix(h, u64::from(v.to_bits()))
|
||||
})
|
||||
});
|
||||
|
||||
// The tail, and it is the whole of the difference between the two output
|
||||
// modes. Everything above — the prologue, the fragments, the mask layers,
|
||||
// the camera matrix — is emitted identically either way, so an operation
|
||||
@@ -1087,6 +1136,12 @@ struct Params {{
|
||||
// stock is loaded, which costs eight bytes and no branch.
|
||||
@group(0) @binding(4) var film_curves: texture_2d<f32>;
|
||||
@group(0) @binding(5) var film_lut_texture: texture_3d<f32>;
|
||||
// The sample cache: the source texel each output pixel read on an earlier
|
||||
// frame with this framing, and where this frame writes it when asked. See
|
||||
// `ComposedShader::sample_key`. Declared unconditionally, like the masks, and
|
||||
// bound to 1x1 placeholders whenever the flags say not to touch them.
|
||||
@group(0) @binding(6) var sampled: texture_2d<f32>;
|
||||
@group(0) @binding(7) var sample_out: texture_storage_2d<rgba16float, write>;
|
||||
|
||||
{sampler_helper}{helper_src}{encode_output}
|
||||
// Display-encoded sRGB back to linear, for sources that arrive that way.
|
||||
@@ -1196,6 +1251,7 @@ fn main(@builtin(global_invocation_id) gid: vec3<u32>) {{
|
||||
uniforms: uniform_values,
|
||||
structure_hash,
|
||||
output_mode,
|
||||
sample_key,
|
||||
}
|
||||
}
|
||||
|
||||
@@ -1415,7 +1471,19 @@ pub(crate) fn sample_source(interpolate: bool, splits_channels: bool) -> &'stati
|
||||
// amount that changes with the aspect ratio. It reads as a correction that
|
||||
// is simply too weak, which is indistinguishable from a bad profile.
|
||||
let radius = length(p) / (0.5 * length(aspect));
|
||||
var c = textureLoad(source, coord, 0).rgb;
|
||||
// The texel itself, from the source or from the cache of it an earlier
|
||||
// frame wrote (see `ComposedShader::sample_key`). Both branches yield the
|
||||
// same bits: the source is `rgba16float` and so is the cache. The flags
|
||||
// are uniforms, so the whole dispatch takes one branch.
|
||||
var c: vec3<f32>;
|
||||
if (u.sample_cache.x > 0.5) {
|
||||
c = textureLoad(sampled, vec2<i32>(gid.xy), 0).rgb;
|
||||
} else {
|
||||
c = textureLoad(source, coord, 0).rgb;
|
||||
if (u.sample_cache.y > 0.5) {
|
||||
textureStore(sample_out, vec2<i32>(gid.xy), vec4<f32>(c, 1.0));
|
||||
}
|
||||
}
|
||||
"
|
||||
}
|
||||
}
|
||||
|
||||
+19
-19
File diff suppressed because one or more lines are too long
Reference in New Issue
Block a user