Let an operation read the pixel next to it, and settle where sharpening belongs
The fused pass hands a fragment a colour and no coordinate. That is what buys one dispatch for a whole edit, and it is also a wall: sharpening, noise reduction, clarity, texture, dehaze and spot removal are each defined by what the neighbours are doing, and FR-DEV-3 and FR-DEV-8 ask for all six. None of them could be written at any price. So there is now a detail stage. An operation implements `Operation` for its parameters exactly as before — the panel, the sidecar, the history and the presets all work unchanged — and additionally returns `Affects::Detail` and a `DetailStage` yielding one pass per dispatch. `Affects` grows the third variant `docs/requirements.md:250` designed and nothing had cut. Where the stage sits is a colour-science decision, not an arrangement of convenience. It runs after every point operation and every mask layer, so an amount chosen against a tone curve survives the curve moving; in linear sRGB after the camera matrix, because camera RGB has no luminance to sharpen against; and before the output transform and the clip, because FR-DEV-2 allows one quantisation and a highlight clipped before a convolution grows a dark ring. The fused pass therefore ends one of two ways, and when a detail stage follows it hands on unclipped f16 and the last detail pass encodes. At render resolution rather than on the source, which is the whole of FR-DSP-1: a pass before the framing prologue would cost 24 MP to draw a 2 MP preview. `RenderScale` is what makes that survivable — a radius is stored as a fraction of the frame's shorter edge, exactly as a mask feather already is, or as a count of source pixels, and converted per render. It also reports when a radius is smaller than a proxy pixel rather than drawing a plausible lie; zooming to 1:1 makes the preview exact with no second path. `Invalidation` gives FR-DEV-3d something to mean. Moving a detail parameter leaves the colour key alone, so `AdjustPass` keeps the linear intermediate and skips the fused dispatch: dragging a sharpening slider costs a convolution. Moving exposure does re-run the detail passes, because they read what the colour pass wrote, and there is no arrangement of keys that avoids it while keeping sharpening after tone. Validated by a separable box blur that is not a develop operation, behind the `detail-probe` feature and absent from a shipping build. An abstraction with no consumer is a guess; a box blur's answer is known in closed form, so the tests assert every byte of the ramp rather than that the edge got softer. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
+424
-74
@@ -16,9 +16,11 @@
|
||||
|
||||
use std::collections::HashMap;
|
||||
|
||||
use dr_pipeline::ComposedShader;
|
||||
use dr_pipeline::detail::ComposedDetail;
|
||||
use dr_pipeline::{ComposedShader, OutputMode};
|
||||
use wgpu::util::DeviceExt;
|
||||
|
||||
use crate::detail::DetailRunner;
|
||||
use crate::readback::await_mapping;
|
||||
use crate::{DemosaicedImage, GpuContext, GpuError};
|
||||
|
||||
@@ -62,6 +64,44 @@ pub struct AdjustPass {
|
||||
current: usize,
|
||||
/// Bound at `@binding(3)` when the edit carries no mask layers.
|
||||
empty_masks: wgpu::TextureView,
|
||||
/// TRACES: FR-DEV-3 | FR-DEV-3d
|
||||
/// The neighbourhood stage — sharpening, noise reduction, clarity and the
|
||||
/// rest of FR-DEV-3's detail set, which cannot be fused into the shader
|
||||
/// above because they read pixels they are not writing.
|
||||
///
|
||||
/// It lives here rather than beside this pass because the two are one
|
||||
/// render: when a detail chain is present the fused pass writes a linear
|
||||
/// intermediate the runner owns, and the runner's last pass writes
|
||||
/// [`Self::targets`]. Kept as separate objects, a caller could hold a
|
||||
/// stale intermediate against a fresh colour result with nothing to tell
|
||||
/// it apart.
|
||||
detail: DetailRunner,
|
||||
/// The bind group layout for a fused pass writing a linear intermediate.
|
||||
///
|
||||
/// A second layout rather than a second pass: the only difference is the
|
||||
/// storage texture's format, which is part of the layout and cannot be
|
||||
/// varied per bind group. Built once here, so a detail operation being
|
||||
/// switched on does not build a pipeline layout mid-frame.
|
||||
linear_bind_group_layout: wgpu::BindGroupLayout,
|
||||
linear_pipeline_layout: wgpu::PipelineLayout,
|
||||
/// TRACES: FR-DEV-3d
|
||||
/// What the linear intermediate currently holds, and at what size.
|
||||
///
|
||||
/// **This is where `Affects::Detail` stops being bookkeeping.** The key is
|
||||
/// everything the fused dispatch depends on — the caller's
|
||||
/// `Invalidation::through(Affects::Colour)`, the compiled structure, the
|
||||
/// uniform values and the output size. When it matches, the colour pass is
|
||||
/// skipped and only the detail passes run, so dragging a sharpening slider
|
||||
/// costs a convolution and not a re-render of the whole chain (FR-DEV-3d).
|
||||
///
|
||||
/// Cleared by any render that does not write it, so a stale intermediate
|
||||
/// cannot survive a change of image and be handed to a later detail chain.
|
||||
colour_key: Option<(u64, u32, u32)>,
|
||||
/// Fused dispatches actually encoded. Exposed so a test can see the reuse
|
||||
/// above happening rather than take it on trust.
|
||||
colour_dispatches: usize,
|
||||
/// Detail dispatches encoded.
|
||||
detail_dispatches: usize,
|
||||
}
|
||||
|
||||
struct Target {
|
||||
@@ -75,58 +115,7 @@ impl AdjustPass {
|
||||
pub const FORMAT: wgpu::TextureFormat = wgpu::TextureFormat::Rgba8Unorm;
|
||||
|
||||
pub fn new(ctx: &GpuContext) -> Self {
|
||||
let bind_group_layout =
|
||||
ctx.device
|
||||
.create_bind_group_layout(&wgpu::BindGroupLayoutDescriptor {
|
||||
label: Some("adjust-bgl"),
|
||||
entries: &[
|
||||
// The demosaiced source.
|
||||
wgpu::BindGroupLayoutEntry {
|
||||
binding: 0,
|
||||
visibility: wgpu::ShaderStages::COMPUTE,
|
||||
ty: wgpu::BindingType::Texture {
|
||||
sample_type: wgpu::TextureSampleType::Float { filterable: true },
|
||||
view_dimension: wgpu::TextureViewDimension::D2,
|
||||
multisampled: false,
|
||||
},
|
||||
count: None,
|
||||
},
|
||||
wgpu::BindGroupLayoutEntry {
|
||||
binding: 1,
|
||||
visibility: wgpu::ShaderStages::COMPUTE,
|
||||
ty: wgpu::BindingType::Buffer {
|
||||
ty: wgpu::BufferBindingType::Uniform,
|
||||
has_dynamic_offset: false,
|
||||
min_binding_size: None,
|
||||
},
|
||||
count: None,
|
||||
},
|
||||
wgpu::BindGroupLayoutEntry {
|
||||
binding: 2,
|
||||
visibility: wgpu::ShaderStages::COMPUTE,
|
||||
ty: wgpu::BindingType::StorageTexture {
|
||||
access: wgpu::StorageTextureAccess::WriteOnly,
|
||||
format: Self::FORMAT,
|
||||
view_dimension: wgpu::TextureViewDimension::D2,
|
||||
},
|
||||
count: None,
|
||||
},
|
||||
// The local-adjustment masks. Present in every layout
|
||||
// whether or not the edit has any, because the layout
|
||||
// is built once here and the generated shader declares
|
||||
// the binding unconditionally for exactly that reason.
|
||||
wgpu::BindGroupLayoutEntry {
|
||||
binding: 3,
|
||||
visibility: wgpu::ShaderStages::COMPUTE,
|
||||
ty: wgpu::BindingType::Texture {
|
||||
sample_type: wgpu::TextureSampleType::Float { filterable: true },
|
||||
view_dimension: wgpu::TextureViewDimension::D2Array,
|
||||
multisampled: false,
|
||||
},
|
||||
count: None,
|
||||
},
|
||||
],
|
||||
});
|
||||
let bind_group_layout = Self::layout_writing(ctx, Self::FORMAT, "adjust-bgl");
|
||||
|
||||
let pipeline_layout = ctx
|
||||
.device
|
||||
@@ -136,6 +125,23 @@ impl AdjustPass {
|
||||
immediate_size: 0,
|
||||
});
|
||||
|
||||
// The same layout with an `Rgba16Float` storage texture, for the fused
|
||||
// pass when a detail stage follows it and it hands on linear working
|
||||
// values instead of encoding (see `dr_pipeline::OutputMode`). The
|
||||
// format is part of a bind group layout and cannot be varied per bind
|
||||
// group, so this is a second layout rather than a second binding —
|
||||
// built here, once, so that switching sharpening on does not construct
|
||||
// a pipeline layout in the middle of a frame.
|
||||
let linear_bind_group_layout =
|
||||
Self::layout_writing(ctx, crate::detail::INTERMEDIATE_FORMAT, "adjust-linear-bgl");
|
||||
let linear_pipeline_layout =
|
||||
ctx.device
|
||||
.create_pipeline_layout(&wgpu::PipelineLayoutDescriptor {
|
||||
label: Some("adjust-linear-layout"),
|
||||
bind_group_layouts: &[Some(&linear_bind_group_layout)],
|
||||
immediate_size: 0,
|
||||
});
|
||||
|
||||
// A 1x1 single-layer mask, bound when the edit has no local
|
||||
// adjustments. The generated shader never samples it — no layer block
|
||||
// is emitted — but a bind group must still satisfy the layout.
|
||||
@@ -167,9 +173,81 @@ impl AdjustPass {
|
||||
targets: [None, None],
|
||||
current: 0,
|
||||
empty_masks,
|
||||
detail: DetailRunner::new(ctx),
|
||||
linear_bind_group_layout,
|
||||
linear_pipeline_layout,
|
||||
colour_key: None,
|
||||
colour_dispatches: 0,
|
||||
detail_dispatches: 0,
|
||||
}
|
||||
}
|
||||
|
||||
/// The fused pass's bind group layout, for a given storage format.
|
||||
///
|
||||
/// Two of these exist — one writing `Rgba8Unorm` and one writing
|
||||
/// `Rgba16Float` — and they differ in exactly one field. Written once and
|
||||
/// parameterised rather than copied, because two copies of a four-entry
|
||||
/// layout is how the mask binding comes to be present in one and absent
|
||||
/// from the other, and a bind group that satisfies neither is a validation
|
||||
/// error a long way from its cause.
|
||||
fn layout_writing(
|
||||
ctx: &GpuContext,
|
||||
format: wgpu::TextureFormat,
|
||||
label: &str,
|
||||
) -> wgpu::BindGroupLayout {
|
||||
ctx.device
|
||||
.create_bind_group_layout(&wgpu::BindGroupLayoutDescriptor {
|
||||
label: Some(label),
|
||||
entries: &[
|
||||
// The demosaiced source.
|
||||
wgpu::BindGroupLayoutEntry {
|
||||
binding: 0,
|
||||
visibility: wgpu::ShaderStages::COMPUTE,
|
||||
ty: wgpu::BindingType::Texture {
|
||||
sample_type: wgpu::TextureSampleType::Float { filterable: true },
|
||||
view_dimension: wgpu::TextureViewDimension::D2,
|
||||
multisampled: false,
|
||||
},
|
||||
count: None,
|
||||
},
|
||||
wgpu::BindGroupLayoutEntry {
|
||||
binding: 1,
|
||||
visibility: wgpu::ShaderStages::COMPUTE,
|
||||
ty: wgpu::BindingType::Buffer {
|
||||
ty: wgpu::BufferBindingType::Uniform,
|
||||
has_dynamic_offset: false,
|
||||
min_binding_size: None,
|
||||
},
|
||||
count: None,
|
||||
},
|
||||
wgpu::BindGroupLayoutEntry {
|
||||
binding: 2,
|
||||
visibility: wgpu::ShaderStages::COMPUTE,
|
||||
ty: wgpu::BindingType::StorageTexture {
|
||||
access: wgpu::StorageTextureAccess::WriteOnly,
|
||||
format,
|
||||
view_dimension: wgpu::TextureViewDimension::D2,
|
||||
},
|
||||
count: None,
|
||||
},
|
||||
// The local-adjustment masks. Present in every layout
|
||||
// whether or not the edit has any, because the layout is
|
||||
// built once here and the generated shader declares the
|
||||
// binding unconditionally for exactly that reason.
|
||||
wgpu::BindGroupLayoutEntry {
|
||||
binding: 3,
|
||||
visibility: wgpu::ShaderStages::COMPUTE,
|
||||
ty: wgpu::BindingType::Texture {
|
||||
sample_type: wgpu::TextureSampleType::Float { filterable: true },
|
||||
view_dimension: wgpu::TextureViewDimension::D2Array,
|
||||
multisampled: false,
|
||||
},
|
||||
count: None,
|
||||
},
|
||||
],
|
||||
})
|
||||
}
|
||||
|
||||
/// Compile a composed shader, or return the cached pipeline.
|
||||
///
|
||||
/// Compilation errors carry the generated source, since a stray line
|
||||
@@ -197,12 +275,22 @@ impl AdjustPass {
|
||||
source: wgpu::ShaderSource::Wgsl(shader.source.as_str().into()),
|
||||
});
|
||||
|
||||
// The layout matching what this shader was composed to write. The
|
||||
// structure hash covers the generated source and the source
|
||||
// carries the storage format, so the two can never disagree — a
|
||||
// cached pipeline is always paired with the layout it was built
|
||||
// against.
|
||||
let layout = match shader.output_mode {
|
||||
OutputMode::Encoded => &self.pipeline_layout,
|
||||
OutputMode::LinearWorking => &self.linear_pipeline_layout,
|
||||
};
|
||||
|
||||
let pipeline =
|
||||
self.ctx
|
||||
.device
|
||||
.create_compute_pipeline(&wgpu::ComputePipelineDescriptor {
|
||||
label: Some("adjust-pipeline"),
|
||||
layout: Some(&self.pipeline_layout),
|
||||
layout: Some(layout),
|
||||
module: &module,
|
||||
entry_point: Some("main"),
|
||||
compilation_options: Default::default(),
|
||||
@@ -310,28 +398,31 @@ impl AdjustPass {
|
||||
height: u32,
|
||||
masks: Option<&crate::MaskArray>,
|
||||
) -> Result<&wgpu::Texture, GpuError> {
|
||||
if shader.output_mode != OutputMode::Encoded {
|
||||
// Composed for a detail stage and dispatched without one. The
|
||||
// shader writes `rgba16float` and this path binds an `rgba8unorm`
|
||||
// storage texture, which wgpu rejects — but well after the point
|
||||
// where the mistake is legible. Saying so here names the actual
|
||||
// error: the edit has a neighbourhood operation and needs
|
||||
// `render_detailed`.
|
||||
return Err(GpuError::ShaderCompilation(
|
||||
"this shader was composed with a detail stage and writes linear \
|
||||
working values; render it with `render_detailed` and the \
|
||||
matching chain from `EditGraph::compose_detail`"
|
||||
.into(),
|
||||
));
|
||||
}
|
||||
// Any render that does not write the linear intermediate leaves
|
||||
// whatever is in it belonging to some other edit — or some other
|
||||
// photograph. Forgetting this is how a detail chain comes to be run
|
||||
// over a stale colour result, so the key is dropped rather than
|
||||
// reasoned about.
|
||||
self.colour_key = None;
|
||||
|
||||
let (width, height) = (width.max(1), height.max(1));
|
||||
self.ensure_target(width, height);
|
||||
|
||||
// Base uniforms: the camera matrix and as-shot white balance, which
|
||||
// every generated shader reads regardless of which operations are
|
||||
// active. Framing's slots follow them and are filled by the composer,
|
||||
// which is why only the first sixteen are written here.
|
||||
let mut uniforms = shader.uniforms.clone();
|
||||
if uniforms.len() < RESERVED_FIELDS {
|
||||
uniforms.resize(RESERVED_FIELDS, 0.0);
|
||||
}
|
||||
let m = source.color_matrix();
|
||||
let wb = source.as_shot_wb();
|
||||
// Rows padded to vec4 for std140 alignment.
|
||||
uniforms[0..4].copy_from_slice(&[m[0], m[1], m[2], 0.0]);
|
||||
uniforms[4..8].copy_from_slice(&[m[3], m[4], m[5], 0.0]);
|
||||
uniforms[8..12].copy_from_slice(&[m[6], m[7], m[8], 0.0]);
|
||||
// The fourth slot is the non-linear flag, not padding: it tells the
|
||||
// shader whether to linearise the sampled texel before any operation
|
||||
// runs. See `DemosaicedImage::is_non_linear`.
|
||||
let non_linear = if source.is_non_linear() { 1.0 } else { 0.0 };
|
||||
uniforms[12..16].copy_from_slice(&[wb[0], wb[1], wb[2], non_linear]);
|
||||
let uniforms = Self::fused_uniforms(source, shader);
|
||||
|
||||
let params_buf = self
|
||||
.ctx
|
||||
@@ -394,6 +485,7 @@ impl AdjustPass {
|
||||
pass.dispatch_workgroups(width.div_ceil(8), height.div_ceil(8), 1);
|
||||
}
|
||||
self.ctx.queue.submit(Some(enc.finish()));
|
||||
self.colour_dispatches += 1;
|
||||
|
||||
Ok(&self.targets[self.current]
|
||||
.as_ref()
|
||||
@@ -401,12 +493,270 @@ impl AdjustPass {
|
||||
.texture)
|
||||
}
|
||||
|
||||
/// TRACES: FR-DEV-3 | FR-DEV-3d | FR-DEV-4 | FR-DSP-1
|
||||
/// Render one frame with a neighbourhood stage.
|
||||
///
|
||||
/// `shader` and `detail` must be the two halves of **one** composition —
|
||||
/// `EditGraph::compose_for` and `EditGraph::compose_detail_for` on the same
|
||||
/// graph, at the same output space. The fused pass stops at linear working
|
||||
/// values when a detail stage exists and the last detail pass performs the
|
||||
/// output transform, so a mismatched pair either encodes twice or not at
|
||||
/// all.
|
||||
///
|
||||
/// An empty `detail` falls through to [`Self::render_masked`], which is
|
||||
/// the honest thing to do rather than an optimisation: an edit with no
|
||||
/// active sharpening *is* an ordinary edit, and it should cost exactly
|
||||
/// what one costs.
|
||||
///
|
||||
/// # `colour_key`, and why the caller supplies it
|
||||
///
|
||||
/// It is `Invalidation::through(Affects::Colour)` for this edit, mixed
|
||||
/// with whatever names the photograph — a `VersionId`, typically. When it
|
||||
/// is unchanged, and the size and the composed shader and its uniforms are
|
||||
/// unchanged with it, the fused dispatch is **skipped** and the linear
|
||||
/// intermediate from the previous frame is convolved again. Dragging a
|
||||
/// sharpening slider then costs the detail passes alone, which is the
|
||||
/// reuse FR-DEV-3d asks for and the operational meaning of
|
||||
/// `Affects::Detail`.
|
||||
///
|
||||
/// The caller supplies it rather than this pass deriving it because only
|
||||
/// the caller knows which *image* is on screen. Everything else that goes
|
||||
/// into the fused dispatch — the shader's structure, its uniform values,
|
||||
/// the output size — is mixed in here, so a caller cannot make the reuse
|
||||
/// unsound by supplying a key that is merely coarse. It can only do so by
|
||||
/// supplying one that fails to distinguish two photographs, which is why
|
||||
/// the identity of the image is spelled out as its job.
|
||||
// Eight arguments, and every one of them is a distinct thing the render
|
||||
// depends on: the image, both halves of the composition, the size, the
|
||||
// masks and the cache key. Bundling them into a struct would move the
|
||||
// problem rather than solve it — the caller would fill in the same eight
|
||||
// fields — and would hide that composing the two halves apart is the one
|
||||
// mistake this signature exists to make visible.
|
||||
#[allow(clippy::too_many_arguments)]
|
||||
pub fn render_detailed(
|
||||
&mut self,
|
||||
source: &DemosaicedImage,
|
||||
shader: &ComposedShader,
|
||||
width: u32,
|
||||
height: u32,
|
||||
masks: Option<&crate::MaskArray>,
|
||||
detail: &ComposedDetail,
|
||||
colour_key: u64,
|
||||
) -> Result<&wgpu::Texture, GpuError> {
|
||||
if detail.is_empty() {
|
||||
return self.render_masked(source, shader, width, height, masks);
|
||||
}
|
||||
if shader.output_mode != OutputMode::LinearWorking {
|
||||
return Err(GpuError::ShaderCompilation(
|
||||
"this detail chain expects a fused pass composed to hand on \
|
||||
linear working values, but the shader given encodes its own \
|
||||
output; compose both halves from the same graph"
|
||||
.into(),
|
||||
));
|
||||
}
|
||||
|
||||
let (width, height) = (width.max(1), height.max(1));
|
||||
self.ensure_target(width, height);
|
||||
|
||||
let uniforms = Self::fused_uniforms(source, shader);
|
||||
let key = Self::colour_signature(colour_key, shader, &uniforms, masks);
|
||||
let reuse = self.colour_key == Some((key, width, height));
|
||||
|
||||
// Compile before borrowing anything: `pipeline` and `colour_target`
|
||||
// both want `&mut self`, and the second holds its borrow across the
|
||||
// encode below.
|
||||
self.pipeline(shader)?;
|
||||
let colour_view = self
|
||||
.detail
|
||||
.colour_target(detail.len(), width, height)
|
||||
.clone();
|
||||
|
||||
let mut enc = self
|
||||
.ctx
|
||||
.device
|
||||
.create_command_encoder(&wgpu::CommandEncoderDescriptor {
|
||||
label: Some("adjust-detail-encoder"),
|
||||
});
|
||||
|
||||
if !reuse {
|
||||
let params_buf = self
|
||||
.ctx
|
||||
.device
|
||||
.create_buffer_init(&wgpu::util::BufferInitDescriptor {
|
||||
label: Some("adjust-params"),
|
||||
contents: bytemuck::cast_slice(&uniforms),
|
||||
usage: wgpu::BufferUsages::UNIFORM,
|
||||
});
|
||||
let bind_group = self
|
||||
.ctx
|
||||
.device
|
||||
.create_bind_group(&wgpu::BindGroupDescriptor {
|
||||
label: Some("adjust-linear-bg"),
|
||||
layout: &self.linear_bind_group_layout,
|
||||
entries: &[
|
||||
wgpu::BindGroupEntry {
|
||||
binding: 0,
|
||||
resource: wgpu::BindingResource::TextureView(source.view()),
|
||||
},
|
||||
wgpu::BindGroupEntry {
|
||||
binding: 1,
|
||||
resource: params_buf.as_entire_binding(),
|
||||
},
|
||||
wgpu::BindGroupEntry {
|
||||
binding: 2,
|
||||
resource: wgpu::BindingResource::TextureView(&colour_view),
|
||||
},
|
||||
wgpu::BindGroupEntry {
|
||||
binding: 3,
|
||||
resource: wgpu::BindingResource::TextureView(
|
||||
masks.map_or(&self.empty_masks, |m| m.view()),
|
||||
),
|
||||
},
|
||||
],
|
||||
});
|
||||
let pipeline = self
|
||||
.cache
|
||||
.get(&shader.structure_hash)
|
||||
.expect("compiled above");
|
||||
|
||||
let mut pass = enc.begin_compute_pass(&wgpu::ComputePassDescriptor {
|
||||
label: Some("adjust-pass"),
|
||||
timestamp_writes: None,
|
||||
});
|
||||
pass.set_pipeline(pipeline);
|
||||
pass.set_bind_group(0, &bind_group, &[]);
|
||||
pass.dispatch_workgroups(width.div_ceil(8), height.div_ceil(8), 1);
|
||||
drop(pass);
|
||||
self.colour_dispatches += 1;
|
||||
}
|
||||
|
||||
// One encoder for the colour pass and every detail pass, submitted
|
||||
// once — the shape `MaskPass::render` established. Submission order is
|
||||
// the whole of the synchronisation: each pass reads what the previous
|
||||
// one wrote, through the same queue.
|
||||
let target_view = self.targets[self.current]
|
||||
.as_ref()
|
||||
.expect("ensured above")
|
||||
.view
|
||||
.clone();
|
||||
let ran = self
|
||||
.detail
|
||||
.encode(&mut enc, detail, &target_view, width, height)?;
|
||||
self.ctx.queue.submit(Some(enc.finish()));
|
||||
self.detail_dispatches += ran;
|
||||
self.colour_key = Some((key, width, height));
|
||||
|
||||
Ok(&self.targets[self.current]
|
||||
.as_ref()
|
||||
.expect("ensured above")
|
||||
.texture)
|
||||
}
|
||||
|
||||
/// The fused pass's uniform block, with the source's own values written in.
|
||||
///
|
||||
/// Split out because both render paths need exactly this and a second copy
|
||||
/// would eventually disagree about where the camera matrix goes — which is
|
||||
/// silent, and corrupts every operation's uniforms downstream of it.
|
||||
fn fused_uniforms(source: &DemosaicedImage, shader: &ComposedShader) -> Vec<f32> {
|
||||
// Base uniforms: the camera matrix and as-shot white balance, which
|
||||
// every generated shader reads regardless of which operations are
|
||||
// active. Framing's slots follow them and are filled by the composer,
|
||||
// which is why only the first sixteen are written here.
|
||||
let mut uniforms = shader.uniforms.clone();
|
||||
if uniforms.len() < RESERVED_FIELDS {
|
||||
uniforms.resize(RESERVED_FIELDS, 0.0);
|
||||
}
|
||||
let m = source.color_matrix();
|
||||
let wb = source.as_shot_wb();
|
||||
// Rows padded to vec4 for std140 alignment.
|
||||
uniforms[0..4].copy_from_slice(&[m[0], m[1], m[2], 0.0]);
|
||||
uniforms[4..8].copy_from_slice(&[m[3], m[4], m[5], 0.0]);
|
||||
uniforms[8..12].copy_from_slice(&[m[6], m[7], m[8], 0.0]);
|
||||
// The fourth slot is the non-linear flag, not padding: it tells the
|
||||
// shader whether to linearise the sampled texel before any operation
|
||||
// runs. See `DemosaicedImage::is_non_linear`.
|
||||
let non_linear = if source.is_non_linear() { 1.0 } else { 0.0 };
|
||||
uniforms[12..16].copy_from_slice(&[wb[0], wb[1], wb[2], non_linear]);
|
||||
uniforms
|
||||
}
|
||||
|
||||
/// TRACES: FR-DEV-3d
|
||||
/// Everything the fused dispatch depends on, in one integer.
|
||||
///
|
||||
/// The caller's edit key, plus the three things the caller does not know
|
||||
/// about: which pipeline was compiled, what was uploaded to it, and which
|
||||
/// mask array was bound. Hashing the uniforms rather than trusting the
|
||||
/// caller's key to cover them is what makes the reuse safe against a
|
||||
/// caller whose key is coarser than it should be — and the uniforms are
|
||||
/// parameters and matrix coefficients from the CPU, never rendered floats,
|
||||
/// so hashing their bit patterns satisfies ARCH §6.13.
|
||||
fn colour_signature(
|
||||
caller: u64,
|
||||
shader: &ComposedShader,
|
||||
uniforms: &[f32],
|
||||
masks: Option<&crate::MaskArray>,
|
||||
) -> u64 {
|
||||
let mut h: u64 = 0xcbf2_9ce4_8422_2325;
|
||||
let mut mix = |v: u64| {
|
||||
for byte in v.to_le_bytes() {
|
||||
h ^= u64::from(byte);
|
||||
h = h.wrapping_mul(0x100_0000_01b3);
|
||||
}
|
||||
};
|
||||
mix(caller);
|
||||
mix(shader.structure_hash);
|
||||
for v in uniforms {
|
||||
// Negative zero folded onto zero: the two render identically, and
|
||||
// a slider that reached zero from below must not miss the cache.
|
||||
mix(u64::from(if *v == 0.0 { 0 } else { v.to_bits() }));
|
||||
}
|
||||
match masks {
|
||||
None => mix(0),
|
||||
Some(m) => {
|
||||
let (w, h) = m.size();
|
||||
mix(1);
|
||||
mix(u64::from(w));
|
||||
mix(u64::from(h));
|
||||
mix(u64::from(m.layers()));
|
||||
}
|
||||
}
|
||||
h
|
||||
}
|
||||
|
||||
/// How many distinct pipelines are compiled. Exposed for tests asserting
|
||||
/// that slider movement does not recompile.
|
||||
pub fn cached_pipelines(&self) -> usize {
|
||||
self.cache.len()
|
||||
}
|
||||
|
||||
/// How many detail-pass pipelines are compiled. As above, for the stage
|
||||
/// that runs after this one.
|
||||
pub fn cached_detail_pipelines(&self) -> usize {
|
||||
self.detail.cached_pipelines()
|
||||
}
|
||||
|
||||
/// TRACES: FR-DEV-3d
|
||||
/// Fused colour dispatches encoded since this pass was created.
|
||||
///
|
||||
/// Exists to be asserted on. The saving `Affects::Detail` buys — a
|
||||
/// sharpening slider that does not re-run the colour chain — is invisible
|
||||
/// in the output by construction, since the picture is meant to be
|
||||
/// identical either way. A counter is the only thing that can see it.
|
||||
pub fn colour_dispatches(&self) -> usize {
|
||||
self.colour_dispatches
|
||||
}
|
||||
|
||||
/// Detail dispatches encoded since this pass was created.
|
||||
pub fn detail_dispatches(&self) -> usize {
|
||||
self.detail_dispatches
|
||||
}
|
||||
|
||||
/// How many linear intermediates have been allocated. For tests: see
|
||||
/// [`crate::MaskPass::allocations`] for the regression this catches.
|
||||
pub fn detail_allocations(&self) -> usize {
|
||||
self.detail.allocations()
|
||||
}
|
||||
|
||||
/// The texture the last render wrote, if there has been one.
|
||||
pub fn output(&self) -> Option<&wgpu::Texture> {
|
||||
self.targets[self.current].as_ref().map(|t| &t.texture)
|
||||
@@ -501,7 +851,7 @@ impl AdjustPass {
|
||||
}
|
||||
|
||||
/// Number the lines of generated source, so a compiler error can be located.
|
||||
fn numbered(src: &str) -> String {
|
||||
pub(crate) fn numbered(src: &str) -> String {
|
||||
src.lines()
|
||||
.enumerate()
|
||||
.map(|(i, l)| format!("{:>4} | {l}", i + 1))
|
||||
|
||||
@@ -0,0 +1,396 @@
|
||||
//! The detail stage — running `dr-pipeline`'s neighbourhood passes.
|
||||
//!
|
||||
//! Where [`crate::AdjustPass`] fuses every point operation into one dispatch,
|
||||
//! this runs the operations that cannot be fused because they read pixels they
|
||||
//! are not writing: sharpening, noise reduction, clarity, texture, dehaze,
|
||||
//! spot removal (FR-DEV-3, FR-DEV-8). `dr_pipeline::detail` decides *what* they
|
||||
//! are and generates their WGSL; this compiles it, finds it somewhere to
|
||||
//! write, and dispatches it.
|
||||
//!
|
||||
//! # Nothing round-trips
|
||||
//!
|
||||
//! Every intermediate here is a `wgpu::Texture` and none of them is ever
|
||||
//! mapped. The chain is `demosaiced -> fused -> f16 -> f16 -> ... -> rgba8`,
|
||||
//! all of it on the device, and the last write lands in the same texture the
|
||||
//! compositor was already being handed. ARCH §6.1 and FR-DEV-4 are satisfied
|
||||
//! by there being no code here that could violate them, which is the only
|
||||
//! guarantee worth having.
|
||||
//!
|
||||
//! # Following the mask pass rather than inventing a second pattern
|
||||
//!
|
||||
//! `mask.rs` established how multi-target work is done in this crate, and this
|
||||
//! copies it deliberately:
|
||||
//!
|
||||
//! - **One encoder for the whole chain.** The mask pass rasterises every layer
|
||||
//! into one command buffer and submits once; this does the same for every
|
||||
//! pass. Submission order is the only synchronisation either needs, because
|
||||
//! both write and then read through the same queue.
|
||||
//! - **Textures reallocated on size change, never per frame.** `ensure_array`
|
||||
//! there, [`Intermediates::ensure`] here. Steady-state rendering at one
|
||||
//! viewport size allocates nothing.
|
||||
//! - **An allocation counter that exists to be asserted on.** Reallocating per
|
||||
//! frame instead of per resize costs a great deal of bandwidth and shows up
|
||||
//! nowhere in the output, which is exactly the kind of regression that needs
|
||||
//! a test that can see it.
|
||||
//! - **Pipelines cached by structure hash**, as `AdjustPass` caches its own.
|
||||
//! Moving a slider re-uploads a uniform buffer; it does not recompile.
|
||||
//!
|
||||
//! # The ping-pong, and why there are at most three textures
|
||||
//!
|
||||
//! Slot 0 holds what the fused colour pass wrote. It is kept **across frames**,
|
||||
//! which is what makes [`dr_pipeline::Affects::Detail`] mean something: when
|
||||
//! only a detail parameter has moved, the colour key is unchanged, the fused
|
||||
//! dispatch is skipped, and dragging a sharpening slider costs the detail
|
||||
//! passes alone (FR-DEV-3d).
|
||||
//!
|
||||
//! The remaining passes alternate between slots 1 and 2, and the last one
|
||||
//! writes the display texture directly rather than an intermediate — so a
|
||||
//! chain of *N* passes costs *N* dispatches and not *N* + 1, and there is no
|
||||
//! resolve pass to pay for. That leaves the allocation at `1 + min(N-1, 2)`
|
||||
//! textures: one for a single-pass operation, two for a separable blur, three
|
||||
//! however long the chain gets after that.
|
||||
|
||||
use std::collections::HashMap;
|
||||
|
||||
use dr_pipeline::detail::{ComposedDetail, ComposedDetailPass};
|
||||
use wgpu::util::DeviceExt as _;
|
||||
|
||||
use crate::{GpuContext, GpuError};
|
||||
|
||||
/// The format every intermediate carries.
|
||||
///
|
||||
/// The same `Rgba16Float` the demosaicer produces and the same one ARCH §5.2
|
||||
/// names as the working precision (FR-DEV-2). It is not a free choice: the
|
||||
/// stage exists between the colour pass and the output transform precisely so
|
||||
/// that a kernel runs on linear values at full internal precision, and an
|
||||
/// 8-bit intermediate would quantise twice and convolve display-encoded
|
||||
/// numbers — which is how sharpening comes to band a clear sky.
|
||||
pub const INTERMEDIATE_FORMAT: wgpu::TextureFormat = wgpu::TextureFormat::Rgba16Float;
|
||||
|
||||
/// One linear working texture.
|
||||
struct Slot {
|
||||
#[allow(dead_code)]
|
||||
texture: wgpu::Texture,
|
||||
view: wgpu::TextureView,
|
||||
}
|
||||
|
||||
/// The pool of linear intermediates, sized to the chain and the viewport.
|
||||
struct Intermediates {
|
||||
slots: Vec<Slot>,
|
||||
width: u32,
|
||||
height: u32,
|
||||
allocations: usize,
|
||||
}
|
||||
|
||||
impl Intermediates {
|
||||
fn new() -> Self {
|
||||
Self {
|
||||
slots: Vec::new(),
|
||||
width: 0,
|
||||
height: 0,
|
||||
allocations: 0,
|
||||
}
|
||||
}
|
||||
|
||||
/// Make sure `count` textures of this size exist.
|
||||
///
|
||||
/// Grows but never shrinks within a size: an edit that briefly had a
|
||||
/// three-pass chain and then a one-pass one keeps the spare texture rather
|
||||
/// than freeing and reallocating it the next time the user turns the
|
||||
/// operation back on. A size change drops the lot, because none of them
|
||||
/// fits any more.
|
||||
fn ensure(&mut self, ctx: &GpuContext, count: usize, width: u32, height: u32) {
|
||||
if self.width != width || self.height != height {
|
||||
self.slots.clear();
|
||||
self.width = width;
|
||||
self.height = height;
|
||||
}
|
||||
while self.slots.len() < count {
|
||||
let texture = ctx.device.create_texture(&wgpu::TextureDescriptor {
|
||||
label: Some("detail-intermediate"),
|
||||
size: wgpu::Extent3d {
|
||||
width,
|
||||
height,
|
||||
depth_or_array_layers: 1,
|
||||
},
|
||||
mip_level_count: 1,
|
||||
sample_count: 1,
|
||||
dimension: wgpu::TextureDimension::D2,
|
||||
format: INTERMEDIATE_FORMAT,
|
||||
// STORAGE_BINDING to be written by a compute pass and
|
||||
// TEXTURE_BINDING to be read by the next one. Nothing else:
|
||||
// no RENDER_ATTACHMENT, because unlike the adjust pass's
|
||||
// output these are never handed to a compositor, and no
|
||||
// COPY_SRC, because nothing reads them back — that is the
|
||||
// point (ARCH §6.1).
|
||||
usage: wgpu::TextureUsages::STORAGE_BINDING
|
||||
| wgpu::TextureUsages::TEXTURE_BINDING,
|
||||
view_formats: &[],
|
||||
});
|
||||
let view = texture.create_view(&Default::default());
|
||||
self.slots.push(Slot { texture, view });
|
||||
self.allocations += 1;
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
/// Runs the detail stage.
|
||||
///
|
||||
/// Owned by [`crate::AdjustPass`] rather than standing alone, because the two
|
||||
/// halves are one render: the fused pass writes slot 0, this reads it, and the
|
||||
/// last pass writes the adjust pass's own output texture. Splitting them into
|
||||
/// two objects with two lifetimes would mean a caller could hold a stale
|
||||
/// intermediate against a fresh colour result and never be told.
|
||||
pub(crate) struct DetailRunner {
|
||||
ctx: GpuContext,
|
||||
/// Layout for a pass writing another linear intermediate.
|
||||
to_linear: Layout,
|
||||
/// Layout for the last pass, which writes the display texture.
|
||||
to_output: Layout,
|
||||
/// Compiled pipelines by pass structure hash.
|
||||
cache: HashMap<u64, wgpu::ComputePipeline>,
|
||||
pool: Intermediates,
|
||||
}
|
||||
|
||||
struct Layout {
|
||||
bind_group: wgpu::BindGroupLayout,
|
||||
pipeline: wgpu::PipelineLayout,
|
||||
}
|
||||
|
||||
impl DetailRunner {
|
||||
pub(crate) fn new(ctx: &GpuContext) -> Self {
|
||||
Self {
|
||||
ctx: ctx.clone(),
|
||||
to_linear: Layout::new(ctx, INTERMEDIATE_FORMAT, "detail-linear"),
|
||||
to_output: Layout::new(ctx, crate::AdjustPass::FORMAT, "detail-output"),
|
||||
cache: HashMap::new(),
|
||||
pool: Intermediates::new(),
|
||||
}
|
||||
}
|
||||
|
||||
/// The view the fused colour pass should write, given a chain of `passes`.
|
||||
///
|
||||
/// Slot 0, always — it is the one that survives between frames so that a
|
||||
/// detail-only change can skip the colour dispatch entirely.
|
||||
pub(crate) fn colour_target(
|
||||
&mut self,
|
||||
passes: usize,
|
||||
width: u32,
|
||||
height: u32,
|
||||
) -> &wgpu::TextureView {
|
||||
// One for the colour pass's result, then one per hand-off between
|
||||
// detail passes, capped at two because a ping-pong needs no more: the
|
||||
// last pass writes the display texture rather than an intermediate.
|
||||
let needed = 1 + passes.saturating_sub(1).min(2);
|
||||
self.pool.ensure(&self.ctx, needed, width, height);
|
||||
&self.pool.slots[0].view
|
||||
}
|
||||
|
||||
/// Encode every pass of `chain`, the last one writing `output`.
|
||||
///
|
||||
/// The caller must already have run the fused colour pass into
|
||||
/// [`Self::colour_target`] — or established that a previous frame's is
|
||||
/// still valid, which is the whole point of keeping slot 0.
|
||||
pub(crate) fn encode(
|
||||
&mut self,
|
||||
encoder: &mut wgpu::CommandEncoder,
|
||||
chain: &ComposedDetail,
|
||||
output: &wgpu::TextureView,
|
||||
width: u32,
|
||||
height: u32,
|
||||
) -> Result<usize, GpuError> {
|
||||
for pass in &chain.passes {
|
||||
self.compile(pass)?;
|
||||
}
|
||||
|
||||
for (index, pass) in chain.passes.iter().enumerate() {
|
||||
// Read what the previous pass wrote; write the next slot, or the
|
||||
// display texture if this is the last one. `index % 2` alternates
|
||||
// between slots 1 and 2, so a pass never reads the texture it is
|
||||
// writing — which on a compute pass is not an error the driver
|
||||
// reports, merely a picture that depends on scheduling.
|
||||
let source_slot = if index == 0 { 0 } else { 2 - (index % 2) };
|
||||
let source = &self.pool.slots[source_slot].view;
|
||||
let destination = if pass.writes_output {
|
||||
output
|
||||
} else {
|
||||
&self.pool.slots[1 + (index % 2)].view
|
||||
};
|
||||
let layout = if pass.writes_output {
|
||||
&self.to_output
|
||||
} else {
|
||||
&self.to_linear
|
||||
};
|
||||
|
||||
let params = self
|
||||
.ctx
|
||||
.device
|
||||
.create_buffer_init(&wgpu::util::BufferInitDescriptor {
|
||||
label: Some("detail-params"),
|
||||
contents: bytemuck::cast_slice(&pass.uniforms),
|
||||
usage: wgpu::BufferUsages::UNIFORM,
|
||||
});
|
||||
|
||||
let bind_group = self
|
||||
.ctx
|
||||
.device
|
||||
.create_bind_group(&wgpu::BindGroupDescriptor {
|
||||
label: Some("detail-bg"),
|
||||
layout: &layout.bind_group,
|
||||
entries: &[
|
||||
wgpu::BindGroupEntry {
|
||||
binding: 0,
|
||||
resource: wgpu::BindingResource::TextureView(source),
|
||||
},
|
||||
wgpu::BindGroupEntry {
|
||||
binding: 1,
|
||||
resource: params.as_entire_binding(),
|
||||
},
|
||||
wgpu::BindGroupEntry {
|
||||
binding: 2,
|
||||
resource: wgpu::BindingResource::TextureView(destination),
|
||||
},
|
||||
],
|
||||
});
|
||||
|
||||
let pipeline = self
|
||||
.cache
|
||||
.get(&pass.structure_hash)
|
||||
.expect("compiled above");
|
||||
|
||||
let mut compute = encoder.begin_compute_pass(&wgpu::ComputePassDescriptor {
|
||||
label: Some(pass.label.as_str()),
|
||||
timestamp_writes: None,
|
||||
});
|
||||
compute.set_pipeline(pipeline);
|
||||
compute.set_bind_group(0, &bind_group, &[]);
|
||||
compute.dispatch_workgroups(width.div_ceil(8), height.div_ceil(8), 1);
|
||||
}
|
||||
|
||||
Ok(chain.passes.len())
|
||||
}
|
||||
|
||||
/// Compile one pass, or leave the cached pipeline in place.
|
||||
///
|
||||
/// A validation error here is a codegen bug rather than anything the user
|
||||
/// did, so it is caught in an error scope and returned with the generated
|
||||
/// source and the pass's label attached — a line number against code
|
||||
/// nobody wrote, from one of several passes, is otherwise close to
|
||||
/// unactionable.
|
||||
fn compile(&mut self, pass: &ComposedDetailPass) -> Result<(), GpuError> {
|
||||
if self.cache.contains_key(&pass.structure_hash) {
|
||||
return Ok(());
|
||||
}
|
||||
|
||||
let scope = self
|
||||
.ctx
|
||||
.device
|
||||
.push_error_scope(wgpu::ErrorFilter::Validation);
|
||||
|
||||
let module = self
|
||||
.ctx
|
||||
.device
|
||||
.create_shader_module(wgpu::ShaderModuleDescriptor {
|
||||
label: Some(pass.label.as_str()),
|
||||
source: wgpu::ShaderSource::Wgsl(pass.source.as_str().into()),
|
||||
});
|
||||
|
||||
let layout = if pass.writes_output {
|
||||
&self.to_output
|
||||
} else {
|
||||
&self.to_linear
|
||||
};
|
||||
|
||||
let pipeline = self
|
||||
.ctx
|
||||
.device
|
||||
.create_compute_pipeline(&wgpu::ComputePipelineDescriptor {
|
||||
label: Some(pass.label.as_str()),
|
||||
layout: Some(&layout.pipeline),
|
||||
module: &module,
|
||||
entry_point: Some("main"),
|
||||
compilation_options: Default::default(),
|
||||
cache: None,
|
||||
});
|
||||
|
||||
if let Some(err) = pollster::block_on(scope.pop()) {
|
||||
return Err(GpuError::ShaderCompilation(format!(
|
||||
"detail pass {}: {err}\n\n--- generated source ---\n{}",
|
||||
pass.label,
|
||||
crate::adjust::numbered(&pass.source)
|
||||
)));
|
||||
}
|
||||
|
||||
self.cache.insert(pass.structure_hash, pipeline);
|
||||
Ok(())
|
||||
}
|
||||
|
||||
/// How many distinct detail pipelines are compiled. For tests asserting
|
||||
/// that slider movement does not recompile.
|
||||
pub(crate) fn cached_pipelines(&self) -> usize {
|
||||
self.cache.len()
|
||||
}
|
||||
|
||||
/// How many intermediate textures have been allocated since this pass was
|
||||
/// created. For tests — see [`crate::MaskPass::allocations`] for the
|
||||
/// regression this shape of counter exists to catch.
|
||||
pub(crate) fn allocations(&self) -> usize {
|
||||
self.pool.allocations
|
||||
}
|
||||
}
|
||||
|
||||
impl Layout {
|
||||
fn new(ctx: &GpuContext, format: wgpu::TextureFormat, label: &str) -> Self {
|
||||
let bind_group = ctx
|
||||
.device
|
||||
.create_bind_group_layout(&wgpu::BindGroupLayoutDescriptor {
|
||||
label: Some(label),
|
||||
entries: &[
|
||||
// The previous stage's result.
|
||||
wgpu::BindGroupLayoutEntry {
|
||||
binding: 0,
|
||||
visibility: wgpu::ShaderStages::COMPUTE,
|
||||
ty: wgpu::BindingType::Texture {
|
||||
sample_type: wgpu::TextureSampleType::Float { filterable: true },
|
||||
view_dimension: wgpu::TextureViewDimension::D2,
|
||||
multisampled: false,
|
||||
},
|
||||
count: None,
|
||||
},
|
||||
wgpu::BindGroupLayoutEntry {
|
||||
binding: 1,
|
||||
visibility: wgpu::ShaderStages::COMPUTE,
|
||||
ty: wgpu::BindingType::Buffer {
|
||||
ty: wgpu::BufferBindingType::Uniform,
|
||||
has_dynamic_offset: false,
|
||||
min_binding_size: None,
|
||||
},
|
||||
count: None,
|
||||
},
|
||||
wgpu::BindGroupLayoutEntry {
|
||||
binding: 2,
|
||||
visibility: wgpu::ShaderStages::COMPUTE,
|
||||
ty: wgpu::BindingType::StorageTexture {
|
||||
access: wgpu::StorageTextureAccess::WriteOnly,
|
||||
format,
|
||||
view_dimension: wgpu::TextureViewDimension::D2,
|
||||
},
|
||||
count: None,
|
||||
},
|
||||
],
|
||||
});
|
||||
|
||||
let pipeline = ctx
|
||||
.device
|
||||
.create_pipeline_layout(&wgpu::PipelineLayoutDescriptor {
|
||||
label: Some(label),
|
||||
bind_group_layouts: &[Some(&bind_group)],
|
||||
immediate_size: 0,
|
||||
});
|
||||
|
||||
Self {
|
||||
bind_group,
|
||||
pipeline,
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -19,12 +19,18 @@ use wgpu::util::DeviceExt;
|
||||
|
||||
mod adjust;
|
||||
mod demosaic;
|
||||
mod detail;
|
||||
mod error;
|
||||
mod histogram;
|
||||
mod mask;
|
||||
mod readback;
|
||||
mod segment;
|
||||
pub use adjust::AdjustPass;
|
||||
// The format the neighbourhood stage works in. Public because it is a promise
|
||||
// rather than an implementation detail: a detail pass is guaranteed linear,
|
||||
// unclipped, full internal precision (FR-DEV-2), and anyone reasoning about
|
||||
// VRAM at 24 MP needs to know what an intermediate costs.
|
||||
pub use detail::INTERMEDIATE_FORMAT as DETAIL_INTERMEDIATE_FORMAT;
|
||||
pub use demosaic::{DemosaicedImage, Demosaicer};
|
||||
pub use error::GpuError;
|
||||
// Renamed on the way out: `BINS` says enough inside `histogram`, and nothing
|
||||
|
||||
Reference in New Issue
Block a user