Let an operation read the pixel next to it, and settle where sharpening belongs

The fused pass hands a fragment a colour and no coordinate. That is what buys
one dispatch for a whole edit, and it is also a wall: sharpening, noise
reduction, clarity, texture, dehaze and spot removal are each defined by what
the neighbours are doing, and FR-DEV-3 and FR-DEV-8 ask for all six. None of
them could be written at any price.

So there is now a detail stage. An operation implements `Operation` for its
parameters exactly as before — the panel, the sidecar, the history and the
presets all work unchanged — and additionally returns `Affects::Detail` and a
`DetailStage` yielding one pass per dispatch. `Affects` grows the third variant
`docs/requirements.md:250` designed and nothing had cut.

Where the stage sits is a colour-science decision, not an arrangement of
convenience. It runs after every point operation and every mask layer, so an
amount chosen against a tone curve survives the curve moving; in linear sRGB
after the camera matrix, because camera RGB has no luminance to sharpen
against; and before the output transform and the clip, because FR-DEV-2 allows
one quantisation and a highlight clipped before a convolution grows a dark
ring. The fused pass therefore ends one of two ways, and when a detail stage
follows it hands on unclipped f16 and the last detail pass encodes.

At render resolution rather than on the source, which is the whole of FR-DSP-1:
a pass before the framing prologue would cost 24 MP to draw a 2 MP preview.
`RenderScale` is what makes that survivable — a radius is stored as a fraction
of the frame's shorter edge, exactly as a mask feather already is, or as a
count of source pixels, and converted per render. It also reports when a radius
is smaller than a proxy pixel rather than drawing a plausible lie; zooming to
1:1 makes the preview exact with no second path.

`Invalidation` gives FR-DEV-3d something to mean. Moving a detail parameter
leaves the colour key alone, so `AdjustPass` keeps the linear intermediate and
skips the fused dispatch: dragging a sharpening slider costs a convolution.
Moving exposure does re-run the detail passes, because they read what the
colour pass wrote, and there is no arrangement of keys that avoids it while
keeping sharpening after tone.

Validated by a separable box blur that is not a develop operation, behind the
`detail-probe` feature and absent from a shipping build. An abstraction with no
consumer is a guess; a box blur's answer is known in closed form, so the tests
assert every byte of the ramp rather than that the edge got softer.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
2026-08-22 14:39:35 +02:00
co-authored by Claude Opus 5
parent 586698db00
commit 7407a82aa7
14 changed files with 3105 additions and 138 deletions
+424 -74
View File
@@ -16,9 +16,11 @@
use std::collections::HashMap;
use dr_pipeline::ComposedShader;
use dr_pipeline::detail::ComposedDetail;
use dr_pipeline::{ComposedShader, OutputMode};
use wgpu::util::DeviceExt;
use crate::detail::DetailRunner;
use crate::readback::await_mapping;
use crate::{DemosaicedImage, GpuContext, GpuError};
@@ -62,6 +64,44 @@ pub struct AdjustPass {
current: usize,
/// Bound at `@binding(3)` when the edit carries no mask layers.
empty_masks: wgpu::TextureView,
/// TRACES: FR-DEV-3 | FR-DEV-3d
/// The neighbourhood stage — sharpening, noise reduction, clarity and the
/// rest of FR-DEV-3's detail set, which cannot be fused into the shader
/// above because they read pixels they are not writing.
///
/// It lives here rather than beside this pass because the two are one
/// render: when a detail chain is present the fused pass writes a linear
/// intermediate the runner owns, and the runner's last pass writes
/// [`Self::targets`]. Kept as separate objects, a caller could hold a
/// stale intermediate against a fresh colour result with nothing to tell
/// it apart.
detail: DetailRunner,
/// The bind group layout for a fused pass writing a linear intermediate.
///
/// A second layout rather than a second pass: the only difference is the
/// storage texture's format, which is part of the layout and cannot be
/// varied per bind group. Built once here, so a detail operation being
/// switched on does not build a pipeline layout mid-frame.
linear_bind_group_layout: wgpu::BindGroupLayout,
linear_pipeline_layout: wgpu::PipelineLayout,
/// TRACES: FR-DEV-3d
/// What the linear intermediate currently holds, and at what size.
///
/// **This is where `Affects::Detail` stops being bookkeeping.** The key is
/// everything the fused dispatch depends on — the caller's
/// `Invalidation::through(Affects::Colour)`, the compiled structure, the
/// uniform values and the output size. When it matches, the colour pass is
/// skipped and only the detail passes run, so dragging a sharpening slider
/// costs a convolution and not a re-render of the whole chain (FR-DEV-3d).
///
/// Cleared by any render that does not write it, so a stale intermediate
/// cannot survive a change of image and be handed to a later detail chain.
colour_key: Option<(u64, u32, u32)>,
/// Fused dispatches actually encoded. Exposed so a test can see the reuse
/// above happening rather than take it on trust.
colour_dispatches: usize,
/// Detail dispatches encoded.
detail_dispatches: usize,
}
struct Target {
@@ -75,58 +115,7 @@ impl AdjustPass {
pub const FORMAT: wgpu::TextureFormat = wgpu::TextureFormat::Rgba8Unorm;
pub fn new(ctx: &GpuContext) -> Self {
let bind_group_layout =
ctx.device
.create_bind_group_layout(&wgpu::BindGroupLayoutDescriptor {
label: Some("adjust-bgl"),
entries: &[
// The demosaiced source.
wgpu::BindGroupLayoutEntry {
binding: 0,
visibility: wgpu::ShaderStages::COMPUTE,
ty: wgpu::BindingType::Texture {
sample_type: wgpu::TextureSampleType::Float { filterable: true },
view_dimension: wgpu::TextureViewDimension::D2,
multisampled: false,
},
count: None,
},
wgpu::BindGroupLayoutEntry {
binding: 1,
visibility: wgpu::ShaderStages::COMPUTE,
ty: wgpu::BindingType::Buffer {
ty: wgpu::BufferBindingType::Uniform,
has_dynamic_offset: false,
min_binding_size: None,
},
count: None,
},
wgpu::BindGroupLayoutEntry {
binding: 2,
visibility: wgpu::ShaderStages::COMPUTE,
ty: wgpu::BindingType::StorageTexture {
access: wgpu::StorageTextureAccess::WriteOnly,
format: Self::FORMAT,
view_dimension: wgpu::TextureViewDimension::D2,
},
count: None,
},
// The local-adjustment masks. Present in every layout
// whether or not the edit has any, because the layout
// is built once here and the generated shader declares
// the binding unconditionally for exactly that reason.
wgpu::BindGroupLayoutEntry {
binding: 3,
visibility: wgpu::ShaderStages::COMPUTE,
ty: wgpu::BindingType::Texture {
sample_type: wgpu::TextureSampleType::Float { filterable: true },
view_dimension: wgpu::TextureViewDimension::D2Array,
multisampled: false,
},
count: None,
},
],
});
let bind_group_layout = Self::layout_writing(ctx, Self::FORMAT, "adjust-bgl");
let pipeline_layout = ctx
.device
@@ -136,6 +125,23 @@ impl AdjustPass {
immediate_size: 0,
});
// The same layout with an `Rgba16Float` storage texture, for the fused
// pass when a detail stage follows it and it hands on linear working
// values instead of encoding (see `dr_pipeline::OutputMode`). The
// format is part of a bind group layout and cannot be varied per bind
// group, so this is a second layout rather than a second binding —
// built here, once, so that switching sharpening on does not construct
// a pipeline layout in the middle of a frame.
let linear_bind_group_layout =
Self::layout_writing(ctx, crate::detail::INTERMEDIATE_FORMAT, "adjust-linear-bgl");
let linear_pipeline_layout =
ctx.device
.create_pipeline_layout(&wgpu::PipelineLayoutDescriptor {
label: Some("adjust-linear-layout"),
bind_group_layouts: &[Some(&linear_bind_group_layout)],
immediate_size: 0,
});
// A 1x1 single-layer mask, bound when the edit has no local
// adjustments. The generated shader never samples it — no layer block
// is emitted — but a bind group must still satisfy the layout.
@@ -167,9 +173,81 @@ impl AdjustPass {
targets: [None, None],
current: 0,
empty_masks,
detail: DetailRunner::new(ctx),
linear_bind_group_layout,
linear_pipeline_layout,
colour_key: None,
colour_dispatches: 0,
detail_dispatches: 0,
}
}
/// The fused pass's bind group layout, for a given storage format.
///
/// Two of these exist — one writing `Rgba8Unorm` and one writing
/// `Rgba16Float` — and they differ in exactly one field. Written once and
/// parameterised rather than copied, because two copies of a four-entry
/// layout is how the mask binding comes to be present in one and absent
/// from the other, and a bind group that satisfies neither is a validation
/// error a long way from its cause.
fn layout_writing(
ctx: &GpuContext,
format: wgpu::TextureFormat,
label: &str,
) -> wgpu::BindGroupLayout {
ctx.device
.create_bind_group_layout(&wgpu::BindGroupLayoutDescriptor {
label: Some(label),
entries: &[
// The demosaiced source.
wgpu::BindGroupLayoutEntry {
binding: 0,
visibility: wgpu::ShaderStages::COMPUTE,
ty: wgpu::BindingType::Texture {
sample_type: wgpu::TextureSampleType::Float { filterable: true },
view_dimension: wgpu::TextureViewDimension::D2,
multisampled: false,
},
count: None,
},
wgpu::BindGroupLayoutEntry {
binding: 1,
visibility: wgpu::ShaderStages::COMPUTE,
ty: wgpu::BindingType::Buffer {
ty: wgpu::BufferBindingType::Uniform,
has_dynamic_offset: false,
min_binding_size: None,
},
count: None,
},
wgpu::BindGroupLayoutEntry {
binding: 2,
visibility: wgpu::ShaderStages::COMPUTE,
ty: wgpu::BindingType::StorageTexture {
access: wgpu::StorageTextureAccess::WriteOnly,
format,
view_dimension: wgpu::TextureViewDimension::D2,
},
count: None,
},
// The local-adjustment masks. Present in every layout
// whether or not the edit has any, because the layout is
// built once here and the generated shader declares the
// binding unconditionally for exactly that reason.
wgpu::BindGroupLayoutEntry {
binding: 3,
visibility: wgpu::ShaderStages::COMPUTE,
ty: wgpu::BindingType::Texture {
sample_type: wgpu::TextureSampleType::Float { filterable: true },
view_dimension: wgpu::TextureViewDimension::D2Array,
multisampled: false,
},
count: None,
},
],
})
}
/// Compile a composed shader, or return the cached pipeline.
///
/// Compilation errors carry the generated source, since a stray line
@@ -197,12 +275,22 @@ impl AdjustPass {
source: wgpu::ShaderSource::Wgsl(shader.source.as_str().into()),
});
// The layout matching what this shader was composed to write. The
// structure hash covers the generated source and the source
// carries the storage format, so the two can never disagree — a
// cached pipeline is always paired with the layout it was built
// against.
let layout = match shader.output_mode {
OutputMode::Encoded => &self.pipeline_layout,
OutputMode::LinearWorking => &self.linear_pipeline_layout,
};
let pipeline =
self.ctx
.device
.create_compute_pipeline(&wgpu::ComputePipelineDescriptor {
label: Some("adjust-pipeline"),
layout: Some(&self.pipeline_layout),
layout: Some(layout),
module: &module,
entry_point: Some("main"),
compilation_options: Default::default(),
@@ -310,28 +398,31 @@ impl AdjustPass {
height: u32,
masks: Option<&crate::MaskArray>,
) -> Result<&wgpu::Texture, GpuError> {
if shader.output_mode != OutputMode::Encoded {
// Composed for a detail stage and dispatched without one. The
// shader writes `rgba16float` and this path binds an `rgba8unorm`
// storage texture, which wgpu rejects — but well after the point
// where the mistake is legible. Saying so here names the actual
// error: the edit has a neighbourhood operation and needs
// `render_detailed`.
return Err(GpuError::ShaderCompilation(
"this shader was composed with a detail stage and writes linear \
working values; render it with `render_detailed` and the \
matching chain from `EditGraph::compose_detail`"
.into(),
));
}
// Any render that does not write the linear intermediate leaves
// whatever is in it belonging to some other edit — or some other
// photograph. Forgetting this is how a detail chain comes to be run
// over a stale colour result, so the key is dropped rather than
// reasoned about.
self.colour_key = None;
let (width, height) = (width.max(1), height.max(1));
self.ensure_target(width, height);
// Base uniforms: the camera matrix and as-shot white balance, which
// every generated shader reads regardless of which operations are
// active. Framing's slots follow them and are filled by the composer,
// which is why only the first sixteen are written here.
let mut uniforms = shader.uniforms.clone();
if uniforms.len() < RESERVED_FIELDS {
uniforms.resize(RESERVED_FIELDS, 0.0);
}
let m = source.color_matrix();
let wb = source.as_shot_wb();
// Rows padded to vec4 for std140 alignment.
uniforms[0..4].copy_from_slice(&[m[0], m[1], m[2], 0.0]);
uniforms[4..8].copy_from_slice(&[m[3], m[4], m[5], 0.0]);
uniforms[8..12].copy_from_slice(&[m[6], m[7], m[8], 0.0]);
// The fourth slot is the non-linear flag, not padding: it tells the
// shader whether to linearise the sampled texel before any operation
// runs. See `DemosaicedImage::is_non_linear`.
let non_linear = if source.is_non_linear() { 1.0 } else { 0.0 };
uniforms[12..16].copy_from_slice(&[wb[0], wb[1], wb[2], non_linear]);
let uniforms = Self::fused_uniforms(source, shader);
let params_buf = self
.ctx
@@ -394,6 +485,7 @@ impl AdjustPass {
pass.dispatch_workgroups(width.div_ceil(8), height.div_ceil(8), 1);
}
self.ctx.queue.submit(Some(enc.finish()));
self.colour_dispatches += 1;
Ok(&self.targets[self.current]
.as_ref()
@@ -401,12 +493,270 @@ impl AdjustPass {
.texture)
}
/// TRACES: FR-DEV-3 | FR-DEV-3d | FR-DEV-4 | FR-DSP-1
/// Render one frame with a neighbourhood stage.
///
/// `shader` and `detail` must be the two halves of **one** composition —
/// `EditGraph::compose_for` and `EditGraph::compose_detail_for` on the same
/// graph, at the same output space. The fused pass stops at linear working
/// values when a detail stage exists and the last detail pass performs the
/// output transform, so a mismatched pair either encodes twice or not at
/// all.
///
/// An empty `detail` falls through to [`Self::render_masked`], which is
/// the honest thing to do rather than an optimisation: an edit with no
/// active sharpening *is* an ordinary edit, and it should cost exactly
/// what one costs.
///
/// # `colour_key`, and why the caller supplies it
///
/// It is `Invalidation::through(Affects::Colour)` for this edit, mixed
/// with whatever names the photograph — a `VersionId`, typically. When it
/// is unchanged, and the size and the composed shader and its uniforms are
/// unchanged with it, the fused dispatch is **skipped** and the linear
/// intermediate from the previous frame is convolved again. Dragging a
/// sharpening slider then costs the detail passes alone, which is the
/// reuse FR-DEV-3d asks for and the operational meaning of
/// `Affects::Detail`.
///
/// The caller supplies it rather than this pass deriving it because only
/// the caller knows which *image* is on screen. Everything else that goes
/// into the fused dispatch — the shader's structure, its uniform values,
/// the output size — is mixed in here, so a caller cannot make the reuse
/// unsound by supplying a key that is merely coarse. It can only do so by
/// supplying one that fails to distinguish two photographs, which is why
/// the identity of the image is spelled out as its job.
// Eight arguments, and every one of them is a distinct thing the render
// depends on: the image, both halves of the composition, the size, the
// masks and the cache key. Bundling them into a struct would move the
// problem rather than solve it — the caller would fill in the same eight
// fields — and would hide that composing the two halves apart is the one
// mistake this signature exists to make visible.
#[allow(clippy::too_many_arguments)]
pub fn render_detailed(
&mut self,
source: &DemosaicedImage,
shader: &ComposedShader,
width: u32,
height: u32,
masks: Option<&crate::MaskArray>,
detail: &ComposedDetail,
colour_key: u64,
) -> Result<&wgpu::Texture, GpuError> {
if detail.is_empty() {
return self.render_masked(source, shader, width, height, masks);
}
if shader.output_mode != OutputMode::LinearWorking {
return Err(GpuError::ShaderCompilation(
"this detail chain expects a fused pass composed to hand on \
linear working values, but the shader given encodes its own \
output; compose both halves from the same graph"
.into(),
));
}
let (width, height) = (width.max(1), height.max(1));
self.ensure_target(width, height);
let uniforms = Self::fused_uniforms(source, shader);
let key = Self::colour_signature(colour_key, shader, &uniforms, masks);
let reuse = self.colour_key == Some((key, width, height));
// Compile before borrowing anything: `pipeline` and `colour_target`
// both want `&mut self`, and the second holds its borrow across the
// encode below.
self.pipeline(shader)?;
let colour_view = self
.detail
.colour_target(detail.len(), width, height)
.clone();
let mut enc = self
.ctx
.device
.create_command_encoder(&wgpu::CommandEncoderDescriptor {
label: Some("adjust-detail-encoder"),
});
if !reuse {
let params_buf = self
.ctx
.device
.create_buffer_init(&wgpu::util::BufferInitDescriptor {
label: Some("adjust-params"),
contents: bytemuck::cast_slice(&uniforms),
usage: wgpu::BufferUsages::UNIFORM,
});
let bind_group = self
.ctx
.device
.create_bind_group(&wgpu::BindGroupDescriptor {
label: Some("adjust-linear-bg"),
layout: &self.linear_bind_group_layout,
entries: &[
wgpu::BindGroupEntry {
binding: 0,
resource: wgpu::BindingResource::TextureView(source.view()),
},
wgpu::BindGroupEntry {
binding: 1,
resource: params_buf.as_entire_binding(),
},
wgpu::BindGroupEntry {
binding: 2,
resource: wgpu::BindingResource::TextureView(&colour_view),
},
wgpu::BindGroupEntry {
binding: 3,
resource: wgpu::BindingResource::TextureView(
masks.map_or(&self.empty_masks, |m| m.view()),
),
},
],
});
let pipeline = self
.cache
.get(&shader.structure_hash)
.expect("compiled above");
let mut pass = enc.begin_compute_pass(&wgpu::ComputePassDescriptor {
label: Some("adjust-pass"),
timestamp_writes: None,
});
pass.set_pipeline(pipeline);
pass.set_bind_group(0, &bind_group, &[]);
pass.dispatch_workgroups(width.div_ceil(8), height.div_ceil(8), 1);
drop(pass);
self.colour_dispatches += 1;
}
// One encoder for the colour pass and every detail pass, submitted
// once — the shape `MaskPass::render` established. Submission order is
// the whole of the synchronisation: each pass reads what the previous
// one wrote, through the same queue.
let target_view = self.targets[self.current]
.as_ref()
.expect("ensured above")
.view
.clone();
let ran = self
.detail
.encode(&mut enc, detail, &target_view, width, height)?;
self.ctx.queue.submit(Some(enc.finish()));
self.detail_dispatches += ran;
self.colour_key = Some((key, width, height));
Ok(&self.targets[self.current]
.as_ref()
.expect("ensured above")
.texture)
}
/// The fused pass's uniform block, with the source's own values written in.
///
/// Split out because both render paths need exactly this and a second copy
/// would eventually disagree about where the camera matrix goes — which is
/// silent, and corrupts every operation's uniforms downstream of it.
fn fused_uniforms(source: &DemosaicedImage, shader: &ComposedShader) -> Vec<f32> {
// Base uniforms: the camera matrix and as-shot white balance, which
// every generated shader reads regardless of which operations are
// active. Framing's slots follow them and are filled by the composer,
// which is why only the first sixteen are written here.
let mut uniforms = shader.uniforms.clone();
if uniforms.len() < RESERVED_FIELDS {
uniforms.resize(RESERVED_FIELDS, 0.0);
}
let m = source.color_matrix();
let wb = source.as_shot_wb();
// Rows padded to vec4 for std140 alignment.
uniforms[0..4].copy_from_slice(&[m[0], m[1], m[2], 0.0]);
uniforms[4..8].copy_from_slice(&[m[3], m[4], m[5], 0.0]);
uniforms[8..12].copy_from_slice(&[m[6], m[7], m[8], 0.0]);
// The fourth slot is the non-linear flag, not padding: it tells the
// shader whether to linearise the sampled texel before any operation
// runs. See `DemosaicedImage::is_non_linear`.
let non_linear = if source.is_non_linear() { 1.0 } else { 0.0 };
uniforms[12..16].copy_from_slice(&[wb[0], wb[1], wb[2], non_linear]);
uniforms
}
/// TRACES: FR-DEV-3d
/// Everything the fused dispatch depends on, in one integer.
///
/// The caller's edit key, plus the three things the caller does not know
/// about: which pipeline was compiled, what was uploaded to it, and which
/// mask array was bound. Hashing the uniforms rather than trusting the
/// caller's key to cover them is what makes the reuse safe against a
/// caller whose key is coarser than it should be — and the uniforms are
/// parameters and matrix coefficients from the CPU, never rendered floats,
/// so hashing their bit patterns satisfies ARCH §6.13.
fn colour_signature(
caller: u64,
shader: &ComposedShader,
uniforms: &[f32],
masks: Option<&crate::MaskArray>,
) -> u64 {
let mut h: u64 = 0xcbf2_9ce4_8422_2325;
let mut mix = |v: u64| {
for byte in v.to_le_bytes() {
h ^= u64::from(byte);
h = h.wrapping_mul(0x100_0000_01b3);
}
};
mix(caller);
mix(shader.structure_hash);
for v in uniforms {
// Negative zero folded onto zero: the two render identically, and
// a slider that reached zero from below must not miss the cache.
mix(u64::from(if *v == 0.0 { 0 } else { v.to_bits() }));
}
match masks {
None => mix(0),
Some(m) => {
let (w, h) = m.size();
mix(1);
mix(u64::from(w));
mix(u64::from(h));
mix(u64::from(m.layers()));
}
}
h
}
/// How many distinct pipelines are compiled. Exposed for tests asserting
/// that slider movement does not recompile.
pub fn cached_pipelines(&self) -> usize {
self.cache.len()
}
/// How many detail-pass pipelines are compiled. As above, for the stage
/// that runs after this one.
pub fn cached_detail_pipelines(&self) -> usize {
self.detail.cached_pipelines()
}
/// TRACES: FR-DEV-3d
/// Fused colour dispatches encoded since this pass was created.
///
/// Exists to be asserted on. The saving `Affects::Detail` buys — a
/// sharpening slider that does not re-run the colour chain — is invisible
/// in the output by construction, since the picture is meant to be
/// identical either way. A counter is the only thing that can see it.
pub fn colour_dispatches(&self) -> usize {
self.colour_dispatches
}
/// Detail dispatches encoded since this pass was created.
pub fn detail_dispatches(&self) -> usize {
self.detail_dispatches
}
/// How many linear intermediates have been allocated. For tests: see
/// [`crate::MaskPass::allocations`] for the regression this catches.
pub fn detail_allocations(&self) -> usize {
self.detail.allocations()
}
/// The texture the last render wrote, if there has been one.
pub fn output(&self) -> Option<&wgpu::Texture> {
self.targets[self.current].as_ref().map(|t| &t.texture)
@@ -501,7 +851,7 @@ impl AdjustPass {
}
/// Number the lines of generated source, so a compiler error can be located.
fn numbered(src: &str) -> String {
pub(crate) fn numbered(src: &str) -> String {
src.lines()
.enumerate()
.map(|(i, l)| format!("{:>4} | {l}", i + 1))