Files
DarkRoom/core/dr-pipeline/src/detail.rs
T
dtourolleandClaude Opus 5 5323608051 Draw the repairs, before anything sharpens what they removed
A spot set now composes detail passes of its own, one per round, and they
go ahead of every operation's kernel. That placement is the decision worth
recording: a sharpening pass reads a neighbourhood, so sharpening a dust
mark before removing it smears its edge into pixels the repair's disc does
not cover, and what survives is a faint over-sharpened ring around an
otherwise perfect patch. It also disagrees with ARCH §5.2, which draws
spot removal after clarity — docs/spot-removal.md §5.1 is where that is
argued out.

Every length reaching the shader is in render pixels, converted here where
the framing is in scope. Both the centre and the source go through
`Framing::output_at` — the same map the fused pass applies to every pixel
— so a rotated photograph rotates the offset with no trigonometry, and the
radius is found by mapping a point one radius above the centre and
measuring, rather than by multiplying by a ratio this function has no
business knowing about. The tests turn and crop the frame and expect the
mark to stay gone, which is the property that arrangement buys.

compose_full now takes the spot set, because a photograph with a repair
and no sharpening still has a detail stage: a fused pass that encoded its
own output there would quantise twice and bind to a texture of the wrong
format. compose_detail_for takes the source size for the same kind of
reason — a RenderScale describes the region on screen, and a spot is
stored against the photograph.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-26 20:17:51 +02:00

1142 lines
50 KiB
Rust

//! Neighbourhood operations — the ones that must read a pixel they are not
//! writing.
//!
//! # Why this exists at all
//!
//! Every operation in [`crate::operation`] contributes a fragment taking a
//! `vec3<f32>` and returning one. That contract is what makes the fused
//! dispatch possible, and it is also an absolute wall: a fragment is handed a
//! colour, not a coordinate, so it cannot look left. Sharpening, noise
//! reduction, clarity, texture, dehaze and spot removal are all defined by
//! what the *neighbours* are doing, and none of them can be written as a point
//! function of `c` at any price.
//!
//! FR-DEV-3 asks for all six and FR-DEV-8 for spot removal. So the fused pass
//! is not the whole pipeline; it is the *point-operation* stage of it, and
//! this module is the stage that follows.
//!
//! # Where it sits, and why there
//!
//! ```text
//! demosaiced source (camera space, full sensor resolution)
//! |
//! | <- framing prologue: output pixel -> source position
//! v
//! +------------------------------------------+
//! | the fused point-operation pass | one dispatch
//! | white balance, exposure, tone, colour |
//! | the mask layers |
//! | camera RGB -> linear sRGB |
//! +------------------------------------------+
//! | rgba16float, linear, **unclipped**, at render resolution
//! v
//! +------------------------------------------+
//! | the detail stage - this module | one dispatch per pass
//! | sharpen, NR, clarity, texture, spots |
//! +------------------------------------------+
//! | the last pass applies the output transform
//! v
//! rgba8unorm display or export texture
//! ```
//!
//! Four things about that position are decisions rather than convenience, and
//! each of them could defensibly have gone the other way.
//!
//! **After tone, not before.** Sharpening before a tone curve and sharpening
//! after it are different pictures, not the same picture computed two ways: an
//! S-curve steepens the mid-tones, so a halo introduced before it is amplified
//! by whatever slope the curve happens to have at that luminance, and the
//! amount that looked right stops looking right the moment the curve moves.
//! After the curve, the amount the user chose is the amount they see, and it
//! survives every later change to tone. This is also what ARCH §5.2 draws:
//! texture, clarity, spot removal and sharpen/NR sit below the tone curve and
//! the colour mixer.
//!
//! **In linear light, after the camera matrix.** The fused pass works in
//! *camera* space, because white balance and exposure are physically
//! meaningful there and nowhere else. A detail pass is the opposite case: it
//! wants a luminance, and camera RGB has no luminance — the three channels are
//! whatever the CFA's dyes passed, and weighting them 0.2126/0.7152/0.0722
//! would be numerology. So the split is taken *after* the `cam_to_srgb`
//! multiply, where the working space is linear sRGB and a luminance is a
//! luminance.
//!
//! **Before the output transform, and before the clip.** FR-DEV-2 allows
//! exactly one quantisation, at the display or export stage. A detail pass
//! reading an 8-bit display-encoded texture and writing another one would
//! quantise twice and do its arithmetic in a space where a difference of one
//! code value means different things at different brightnesses — which is how
//! sharpening ends up with visible banding in a sky. The intermediate is
//! therefore `rgba16float` and holds linear values that have **not** been
//! clamped to `0..=1`: a recovered highlight is still above one at this point,
//! and clipping it before the sharpener sees it would put a hard edge exactly
//! where the sharpener is most visible. The last detail pass performs the
//! primaries conversion, the clip and the encode, so the single quantisation
//! stays single.
//!
//! **After framing, at render resolution.** The alternative — running detail
//! on the demosaiced source before the framing prologue — is superficially
//! attractive, because a radius in sensor pixels would then mean exactly what
//! it says. It is unaffordable: the source is the full sensor, so a detail
//! pass there costs 24 MP of work for a 2 MP preview and FR-DSP-1 stops being
//! true. Running at render resolution instead makes the cost proportional to
//! what is on screen, and pushes the whole difficulty into one place — the
//! scale — which [`RenderScale`] exists to make explicit rather than implicit.
//!
//! # What this stage deliberately cannot do
//!
//! **There is no per-mask detail.** A mask layer's chain is fused into the
//! point-operation pass; the detail stage runs once, afterwards, over the
//! whole frame. Local sharpening is therefore not expressible here, and
//! [`crate::mask::MaskLayer::active_ops`] filters detail operations out rather
//! than emitting a block that would silently do nothing. Making it possible
//! means giving a detail pass the mask array and a layer index, which is a
//! change to this module's shader preamble and not to its shape — but it is
//! not done, and a caller should not assume it.
//!
//! # Adding a neighbourhood operation
//!
//! Declare it in `ops/<id>.yaml` with `rust:`, exactly as the tone curve does
//! — the schema in `ops/README.md` describes a point function, and stretching
//! it to cover kernels would be a worse language than Rust aimed at one
//! caller. Then implement [`crate::Operation`] as usual for the parameters,
//! descriptor and sidecar, and additionally:
//!
//! ```ignore
//! impl Operation for Sharpen {
//! fn affects(&self) -> Affects { Affects::Detail }
//! fn detail(&self) -> Option<&dyn DetailStage> { Some(self) }
//! fn wgsl_body(&self) -> String { String::new() } // never called
//! // ... descriptor, set_param, param, is_active exactly as usual
//! }
//!
//! impl DetailStage for Sharpen {
//! fn passes(&self, scale: RenderScale) -> Vec<DetailPass> { /* ... */ }
//! }
//! ```
//!
//! Everything else arrives unchanged and for free: the develop panel builds
//! its controls from the descriptor, the sidecar persists the parameters, the
//! history and the presets carry them, and an operation at its defaults
//! contributes no pass at all.
use std::fmt::Write as _;
use dr_types::ColourSpace;
use crate::operation::{Helper, Operation, Uniform};
/// Floats the generated detail uniform block always carries, before an
/// operation's own.
///
/// One `vec4`, which is also the smallest a WGSL uniform struct can be and
/// stay aligned. See [`compose_detail`] for what the lanes hold.
pub const DETAIL_BASE_UNIFORM_FIELDS: usize = 4;
/// TRACES: FR-DSP-1
/// The relationship between the resolution an edit is being **rendered** at
/// and the resolution it will eventually be **exported** at.
///
/// # The problem this type is the answer to
///
/// A point operation is scale-free. Exposure is a multiply, and multiplying by
/// two is multiplying by two whether the frame is 2 000 pixels wide or 24 000.
/// Every operation in the fused pass has this property, which is why nothing
/// in the pipeline has needed to know its own resolution until now.
///
/// A neighbourhood operation has no such luck. "Sharpen with a radius of one
/// pixel" is a statement about a specific grid, and the develop view is not
/// rendering on that grid — FR-DSP-1 has it rendering at whatever the viewport
/// needs, which for a 60 MP frame in a 2 000 px panel is one render pixel per
/// nine source pixels. Tune a radius there, export at full size, and the
/// exported file is sharpened at a ninth of the strength the photographer
/// chose. That is not a rounding difference; it is a different photograph.
///
/// # The rule
///
/// **A length is stored normalised and converted here.** Never store pixels in
/// an edit. This is not a new idea in this codebase — [`crate::mask`] already
/// does it, storing every feather and morphology radius as a fraction of the
/// frame's shorter edge and multiplying up in `dr-gpu` at whatever size the
/// mask is being rasterised at (see `MaskLayer::feather`, and
/// `field_short_edge` in `dr-gpu`'s mask pass). A detail operation follows the
/// same rule through [`Self::frame_fraction`] and gets the same guarantee: the
/// effect covers the same *proportion* of the picture at every size, so what
/// was tuned on screen is what lands in the file.
///
/// # Two units, because there are two kinds of length
///
/// The mask rule is not quite enough on its own, because detail operations
/// split into two families that mean different things by "radius":
///
/// - **Compositional** — clarity, texture, dehaze. The radius is a fraction of
/// the picture, tens of pixels at any size, and [`Self::frame_fraction`] is
/// exactly right. These preview faithfully at any scale.
///
/// - **Acutance** — capture sharpening, luminance noise reduction. The radius
/// is a property of the *sensor*: it is about the lens's circle of confusion
/// and the demosaic's interpolation, both measured in source pixels and
/// neither of which cares how large the viewport is.
/// [`Self::source_pixels`] converts one of those into render pixels.
///
/// # The honest limit
///
/// For the second family the conversion runs out. At a one-ninth proxy a
/// 1.0-source-pixel radius is 0.11 render pixels, and there is no kernel that
/// represents a ninth of a pixel — the information the sharpener would act on
/// was thrown away by the downscale before the pass ever ran. No arrangement
/// of this stage recovers it, which is why every editor that has shipped tells
/// the photographer to judge sharpening at 1:1, and why Lightroom's detail
/// panel contains a 1:1 loupe rather than a scaled preview.
///
/// [`Self::resolves`] reports that condition instead of hiding it, so an
/// operation can fade itself out and an interface can say "zoom to 100% to
/// judge this" — which is the truth, and better than a preview that lies.
/// Zooming is enough: the framing's view rect shrinks while the render target
/// keeps its size, so [`Self::ratio`] climbs back to 1.0 at 1:1 and the
/// preview becomes exact, with no separate full-resolution path to maintain.
#[derive(Debug, Clone, Copy, PartialEq)]
pub struct RenderScale {
render: (u32, u32),
full: (u32, u32),
}
impl RenderScale {
/// `render` is the size being rendered now; `full` is the size the same
/// framed region would have at source resolution.
///
/// Both describe *the region being looked at*, not the whole photograph —
/// so a crop and a zoom are already accounted for by the time they arrive.
/// [`crate::EditGraph::render_scale`] works both out from the framing, and
/// is what a caller should normally use.
pub fn new(render: (u32, u32), full: (u32, u32)) -> Self {
Self {
render: (render.0.max(1), render.1.max(1)),
full: (full.0.max(1), full.1.max(1)),
}
}
/// A scale that is already at source resolution — an export, or a 1:1
/// view. [`Self::ratio`] is 1.0 and nothing is approximated.
pub fn full(render: (u32, u32)) -> Self {
Self::new(render, render)
}
pub fn render_size(&self) -> (u32, u32) {
self.render
}
pub fn full_size(&self) -> (u32, u32) {
self.full
}
/// Render pixels per source pixel. 1.0 at export, below 1.0 on a proxy.
///
/// Averaged over the two axes rather than taken from one. They agree to
/// within a pixel by construction — both sizes describe the same rectangle
/// — but each is separately rounded to an integer, and taking the mean
/// stops a narrow viewport disagreeing with itself.
pub fn ratio(&self) -> f32 {
let x = self.render.0 as f32 / self.full.0 as f32;
let y = self.render.1 as f32 / self.full.1 as f32;
(x + y) * 0.5
}
/// Whether this render is smaller than the file it stands for.
pub fn is_proxy(&self) -> bool {
self.ratio() < 0.999
}
/// A length stated in **source pixels**, in render pixels.
///
/// For the acutance family — sharpening, luminance NR — whose radius is a
/// property of the sensor rather than of the composition.
pub fn source_pixels(&self, radius: f32) -> f32 {
radius * self.ratio()
}
/// A length stated as a **fraction of the frame's shorter edge**, in
/// render pixels.
///
/// For the compositional family — clarity, texture, dehaze — and the same
/// unit `dr-gpu`'s mask rasteriser already converts feathers in. An edit
/// stored this way is resolution-independent by construction.
pub fn frame_fraction(&self, fraction: f32) -> f32 {
fraction * self.render.0.min(self.render.1) as f32
}
/// Whether a radius stated in source pixels survives this render.
///
/// False means the effect is smaller than a pixel here and whatever is
/// drawn is a guess. Report it; do not paper over it — see the type's
/// documentation for why there is nothing better to do.
pub fn resolves(&self, radius_in_source_pixels: f32) -> bool {
self.source_pixels(radius_in_source_pixels) >= 1.0
}
}
/// One dispatch of a neighbourhood operation.
///
/// An operation returns as many of these as it needs. A separable Gaussian is
/// two — horizontal then vertical — and gets the ping-pong between them for
/// free; an unsharp mask wanting its blur held alongside the original would be
/// more, and is the case this shape exists to leave room for.
#[derive(Debug, Clone, PartialEq)]
pub struct DetailPass {
/// A short name, used to label the GPU pass and to make a shader
/// compilation failure say which of an operation's passes broke.
pub label: &'static str,
/// The furthest this pass reads from the pixel it writes, in **render**
/// pixels.
///
/// Declared rather than inferred from the WGSL, because nothing can infer
/// it from the WGSL: the offsets are computed at runtime from uniforms.
/// It is the halo a tile has to be grown by before this pass can be
/// computed tile-wise (ARCH §5.3), and it is the reason a detail operation
/// is not simply "some more shader code" — the scheduler has to know how
/// far the dependency reaches before it can schedule anything at all.
///
/// An understated radius shows as a seam at every tile boundary, which is
/// the kind of artefact that looks like a driver bug. State it honestly.
pub radius: u32,
/// The WGSL body.
///
/// Reads and writes `c`, a `vec3<f32>` of **linear sRGB**, pre-loaded with
/// this pixel's own value. Also in scope:
///
/// - `coord: vec2<i32>` — this pixel.
/// - `tap(coord, offset) -> vec3<f32>` — a neighbour, clamped to the edge
/// of the image, which is what makes a kernel at the border average the
/// pixels that exist rather than fade into black.
/// - `render_dims: vec2<f32>` and `render_scale: f32` — the size being
/// rendered and [`RenderScale::ratio`], for the rare pass that needs
/// them in the shader. Prefer computing lengths on the CPU in
/// [`DetailStage::passes`], where the units are named methods rather
/// than an untyped float.
/// - `aux: f32` and `tap_aux(coord, offset) -> f32` — **one scalar per
/// pixel that survives to the next pass**, pre-loaded with what the
/// previous pass left there and written back out unless the body
/// assigns it.
///
/// # Why `aux` exists
///
/// The ping-pong hands each pass exactly one texture: what the pass before
/// it wrote. That is enough for a chain of filters — a separable blur is
/// two of them — and it is *not* enough for an unsharp mask, which is the
/// shape of sharpening, clarity, texture and dehaze alike. An unsharp mask
/// needs the blur **and** the original in the same place at the same time,
/// and once the first pass has written its blur the original is gone.
///
/// Three channels cannot carry both. Even restricted to the case where the
/// operation only moves luminance — so the colour is a luminance and two
/// chromaticity degrees of freedom — the combining pass needs four
/// numbers: the original luminance, two of chromaticity, and the blurred
/// luminance. Four does not fit in three, and no encoding makes it fit.
///
/// The intermediate is `rgba16float` and its alpha was being written as a
/// constant `1.0` and read by nobody, so the fourth number goes there. A
/// blur pass leaves `c` alone and puts its result in `aux`; the pass after
/// it therefore receives the untouched original *and* the blur, and can
/// subtract one from the other. An operation with no use for the lane says
/// nothing and hands on what it was given.
///
/// The last pass in the chain writes the display texture, whose alpha is
/// opacity rather than scratch space, so `aux` is readable there and not
/// written. That is exactly the right way round: the combining pass is the
/// one that reads it.
///
/// Uniforms are addressed by the bare names declared in [`Self::uniforms`],
/// exactly as a fused fragment addresses its own; the composer rewrites
/// them to their prefixed struct fields.
///
/// **Values are not clipped.** A recovered highlight arrives above 1.0 and
/// an out-of-gamut colour can arrive below 0.0. That is deliberate — see
/// the module documentation — and a kernel that assumes `0..=1` will
/// produce dark rings around specular highlights.
pub wgsl: String,
/// Uniform values this pass's body reads.
pub uniforms: Vec<Uniform>,
/// TRACES: FR-DEV-8
/// Per-instance data, for a pass whose work is a *list* rather than a
/// kernel.
///
/// Reaches the body as `instances: array<vec4<f32>>`, with
/// `instance_count` in scope as a `u32`. Empty for every pass that is a
/// convolution, which is every pass that existed before spot removal.
///
/// # Why not the uniform block
///
/// Because the uniform block is fixed by the pass's *structure*, and this
/// is not: sixty-four repairs and one repair are the same shader with a
/// different buffer behind it. Packing the list into uniforms would need a
/// fixed maximum, paid for on every frame whether the photograph carries
/// one spot or none, and it would need the composer to emit `vec4` fields —
/// a WGSL uniform array has a stride of 16 whatever it holds.
///
/// The property that matters more: with the list in storage the generated
/// source does not mention how many there are, so placing the tenth spot
/// uploads 512 bytes and reuses the compiled pipeline, exactly as moving a
/// slider does for the fused pass.
pub storage: Vec<[f32; 4]>,
}
/// TRACES: FR-DEV-3 | FR-DEV-8
/// An operation that reads pixels other than the one it is writing.
///
/// Implemented *alongside* [`Operation`], never instead of it: the parameters,
/// the descriptor, the panel controls and the sidecar all come from the
/// `Operation` half, and only the execution differs. An operation that
/// implements this must also return [`crate::Affects::Detail`] from
/// `affects()` and `Some(self)` from `Operation::detail()` — the three are
/// checked against each other by a test in [`crate::operation`], because an
/// operation that forgot one of them would be dropped from both stages and
/// simply not happen, with no error anywhere.
pub trait DetailStage: Send + Sync {
/// The passes to run, in order, at this resolution.
///
/// Called per render, so the operation sees the scale it is actually being
/// asked to draw at and converts its own lengths here — in Rust, where
/// [`RenderScale`]'s two conversions are named after the two units, rather
/// than in WGSL where both would be a bare `f32`.
///
/// Returning an empty vector means "nothing to do at this scale", which is
/// the honest answer for an acutance operation on a heavy proxy. It is
/// **not** how an operation says it is neutral: that is `is_active()`, and
/// an inactive operation is never asked.
fn passes(&self, scale: RenderScale) -> Vec<DetailPass>;
}
/// One compile-ready detail pass: complete WGSL and the uniform block for it.
#[derive(Debug, Clone, PartialEq)]
pub struct ComposedDetailPass {
/// `<op id>/<pass label>`, for GPU labels and error messages.
pub label: String,
/// Complete, compilable WGSL.
pub source: String,
/// Uniform values in the order the generated struct declares them.
pub uniforms: Vec<f32>,
/// TRACES: FR-DEV-8
/// The instance list, if this pass declared one. See [`DetailPass::storage`].
pub storage: Vec<[f32; 4]>,
/// See [`DetailPass::radius`].
pub radius: u32,
/// Whether this pass writes the display/export texture rather than another
/// linear intermediate.
///
/// True for exactly the last pass in the chain, which carries the output
/// transform — the primaries conversion, the clip and the encode that the
/// fused pass performs when there is no detail stage at all. Folding them
/// into the last pass rather than adding a resolve dispatch keeps the cost
/// of the stage at one dispatch per pass, not one plus one.
pub writes_output: bool,
/// Identifies this pass's *structure*, for the pipeline cache. Covers the
/// generated source, not the uniform values — so moving a slider uploads a
/// buffer and reuses the compiled pipeline, exactly as the fused pass does.
pub structure_hash: u64,
}
/// The detail stage of one edit, at one resolution.
#[derive(Debug, Clone, Default, PartialEq)]
pub struct ComposedDetail {
pub passes: Vec<ComposedDetailPass>,
}
impl ComposedDetail {
/// Whether the edit has no detail stage — the common case, and the one
/// that must cost nothing.
pub fn is_empty(&self) -> bool {
self.passes.is_empty()
}
pub fn len(&self) -> usize {
self.passes.len()
}
/// The widest halo any pass needs, in render pixels (ARCH §5.3).
pub fn radius(&self) -> u32 {
self.passes.iter().map(|p| p.radius).max().unwrap_or(0)
}
}
/// TRACES: FR-DEV-3 | FR-DSP-1
/// Generate the detail stage for a set of operations at one resolution.
///
/// Operations that declare no [`DetailStage`], or that are at their neutral
/// settings, contribute nothing — the same rule the fused composer follows, so
/// an edit with no sharpening produces an empty chain and `dr-gpu` runs the
/// single dispatch it always did.
///
/// `output` is the space the **last** pass encodes into, and it is a parameter
/// for the same reason it is a parameter to [`crate::compose_with_framing`]: a
/// screen render and a Display P3 export are the same edit and different
/// shaders, and neither is more authoritative than the other.
///
/// # The generated uniform block
///
/// A fixed `vec4` first, then the pass's own scalars, prefixed with the
/// operation id so that a pass never has to know what else is in the block.
/// The lanes of the leading `vec4` are, in order: render width, render height,
/// [`RenderScale::ratio`], and the pass's index within its operation. The
/// first three reach the body as `render_dims` and `render_scale`; the fourth
/// is there because a two-pass operation emitting one body for both directions
/// is a reasonable thing to want, and would otherwise need a uniform of its
/// own purely to say which half it is in.
pub fn compose_detail(
ops: &[Box<dyn Operation>],
scale: RenderScale,
output: ColourSpace,
) -> ComposedDetail {
compose_detail_with(ops, &[], scale, output)
}
/// TRACES: FR-DEV-8
/// The detail stage with a set of repairs ahead of the operations.
///
/// `spots` are already-built passes, from [`crate::SpotSet::passes`], and they
/// go **first** — before sharpening, before noise reduction, before every
/// kernel in `ops`.
///
/// That placement is a decision rather than an ordering convenience. A
/// sharpening kernel reads a neighbourhood, so sharpening a dust mark before
/// removing it smears the mark's edge into pixels the repair's own disc does
/// not cover: what is left afterwards is a faint over-sharpened ring around an
/// otherwise perfect patch, which is exactly the artefact that reads as broken
/// software. Removing the mark first means every later pass sees the
/// photograph the photographer thinks they are sharpening.
///
/// It also means ARCH §5.2's stage list, which draws spot removal after
/// texture and clarity, is not what this does — see `docs/spot-removal.md`
/// §5.1, which is where the disagreement is written down.
pub fn compose_detail_with(
ops: &[Box<dyn Operation>],
spots: &[DetailPass],
scale: RenderScale,
output: ColourSpace,
) -> ComposedDetail {
// Every pass of every active detail operation, flattened, carrying the
// operation it came from for the uniform prefix and the helper set.
let mut planned: Vec<(&'static str, &'static [Helper], DetailPass, usize)> = Vec::new();
for (index, pass) in spots.iter().enumerate() {
planned.push((
crate::spot::SPOT_ID,
crate::spot::SPOT_HELPERS,
pass.clone(),
index,
));
}
for op in ops {
if !op.is_active() {
continue;
}
let Some(stage) = op.detail() else {
continue;
};
let id = op.descriptor().id.0;
for (index, pass) in stage.passes(scale).into_iter().enumerate() {
planned.push((id, op.helpers(), pass, index));
}
}
// An active detail operation that emitted nothing at this scale.
//
// Legal, and the honest answer for an acutance operation on a heavy proxy
// — a one-source-pixel radius is a third of a render pixel there and no
// kernel represents a third of a pixel (see [`RenderScale`]). But it opens
// a hole between the two halves of the composition: [`compose_full`]
// decides to hand on linear working values from the *operations*, which it
// must, having no scale to consult, so the fused pass has already stopped
// short of the output transform. Returning an empty chain here would leave
// that transform undone and bind an `rgba16float` shader to an
// `rgba8unorm` target, which surfaces as a wgpu validation failure a long
// way from the cause.
//
// So the chain is never empty when the fused pass is expecting one: a
// single pass with no body, which reads the intermediate and performs the
// output transform the fused pass skipped. One dispatch, in the uncommon
// case where a photographer has a kernel switched on at a scale that
// cannot draw it — against the alternative of the preview failing outright
// or `compose_full` growing a resolution argument it has no other use for.
if planned.is_empty() && ops.iter().any(|o| o.is_active() && o.detail().is_some()) {
return ComposedDetail {
passes: vec![compose_one(
RESOLVE_ID,
&[],
&DetailPass {
label: "resolve",
radius: 0,
wgsl: String::new(),
uniforms: Vec::new(),
storage: Vec::new(),
},
0,
scale,
output,
true,
)],
};
}
let last = planned.len().saturating_sub(1);
let passes = planned
.into_iter()
.enumerate()
.map(|(position, (id, helpers, pass, index))| {
compose_one(id, helpers, &pass, index, scale, output, position == last)
})
.collect();
ComposedDetail { passes }
}
/// The operation id the resolve pass is labelled with.
///
/// Not an operation: no `ops/*.yaml` declares it and nothing in the chain
/// answers to it. It exists so the generated label reads `detail/resolve`
/// rather than borrowing the id of whichever operation happened to fall
/// through, which would send a reader looking for a bug in that operation.
const RESOLVE_ID: &str = "detail";
#[allow(clippy::too_many_arguments)]
fn compose_one(
id: &str,
helpers: &[Helper],
pass: &DetailPass,
index: usize,
scale: RenderScale,
output: ColourSpace,
writes_output: bool,
) -> ComposedDetailPass {
let prefix = format!("{}_{index}", crate::operation::sanitise(id));
let mut uniform_fields = String::from(
" // x, y: the size being rendered. z: render pixels per source\n\
\x20 // pixel — 1.0 at export, less on a proxy (FR-DSP-1). w: which\n\
\x20 // pass of this operation this is.\n\
\x20 detail_base: vec4<f32>,\n",
);
let (rw, rh) = scale.render_size();
let mut uniform_values = vec![rw as f32, rh as f32, scale.ratio(), index as f32];
debug_assert_eq!(uniform_values.len(), DETAIL_BASE_UNIFORM_FIELDS);
if !pass.uniforms.is_empty() {
let _ = writeln!(uniform_fields, " // {id}/{}", pass.label);
}
for u in &pass.uniforms {
let _ = writeln!(uniform_fields, " {prefix}_{}: f32,", u.name);
uniform_values.push(u.value);
}
// A uniform struct whose size is not a multiple of 16 is rejected by the
// WGSL uniform address space rules — the same padding the fused composer
// applies, for the same reason.
let pad = (4 - (uniform_values.len() % 4)) % 4;
for i in 0..pad {
let _ = writeln!(uniform_fields, " _pad{i}: f32,");
uniform_values.push(0.0);
}
let mut body = pass.wgsl.clone();
for u in &pass.uniforms {
body = crate::operation::rewrite_uniform(&body, u.name, &format!("u.{prefix}_{}", u.name));
}
let mut helper_src = String::new();
let mut seen: Vec<&str> = Vec::new();
for h in helpers {
if seen.contains(&h.name) {
continue;
}
seen.push(h.name);
let _ = writeln!(helper_src, "{}\n", h.source.trim_end());
}
// The storage format and the tail are the *only* difference between an
// intermediate pass and the final one. Everything above — the taps, the
// uniforms, the body — is identical, which is what lets an operation write
// one kernel without knowing whether it happens to be last in the chain.
let (store_format, tail) = if writes_output {
(
"rgba8unorm",
format!(
"{} // Clip to the output gamut and encode. The one quantisation\n\
\x20 // the pipeline performs (FR-DEV-2), and it is here rather than\n\
\x20 // in the fused pass because this is now the last thing to run.\n\
\x20 c = clamp(c, vec3<f32>(0.0), vec3<f32>(1.0));\n\
\x20 textureStore(output, coord, vec4<f32>(encode_output(c), 1.0));",
crate::operation::primaries_conversion(output)
),
)
} else {
(
"rgba16float",
" // Another linear intermediate: no clip and no encode, because\n\
\x20 // the pass after this one still has to read real values.\n\
\x20 //\n\
\x20 // `aux` rides in alpha. A pass that never touches it hands on\n\
\x20 // whatever it was given, so the lane costs an operation that\n\
\x20 // does not want it exactly one copy of a value it already read.\n\
\x20 textureStore(output, coord, vec4<f32>(c, aux));"
.to_string(),
)
};
let encode_fn = if writes_output {
crate::operation::encode_output_fn(output)
} else {
String::new()
};
let label = format!("{id}/{}", pass.label);
let indented = body
.lines()
.map(|l| format!(" {l}"))
.collect::<Vec<_>>()
.join("\n");
let source = format!(
"// GENERATED — do not edit.
//
// Detail pass `{label}` — a neighbourhood operation, which is why it is a
// dispatch of its own rather than a block in the fused shader: it reads pixels
// it is not writing, and the fused contract hands a fragment a colour with no
// way back to a coordinate.
//
// In: linear sRGB, scene-referred, **unclipped**, at render resolution.
// Out: {}
struct Params {{
{uniform_fields}}}
@group(0) @binding(0) var source: texture_2d<f32>;
@group(0) @binding(1) var<uniform> u: Params;
@group(0) @binding(2) var output: texture_storage_2d<{store_format}, write>;
// A pass whose work is a list rather than a kernel reads it here; every other
// pass leaves this bound to a single empty element and never looks at it. See
// `DetailPass::storage` for why the list is not in the uniform block.
@group(0) @binding(3) var<storage, read> instances: array<vec4<f32>>;
// A neighbour, clamped to the edge of the image.
//
// Clamped rather than zero-filled: a kernel straddling the border must average
// the pixels that exist. Returning zero there darkens every edge by a band the
// width of the radius, which reads as a vignette nobody asked for and is the
// classic way a first convolution goes wrong.
fn tap(coord: vec2<i32>, offset: vec2<i32>) -> vec3<f32> {{
let last = vec2<i32>(textureDimensions(source)) - vec2<i32>(1);
return textureLoad(source, clamp(coord + offset, vec2<i32>(0), last), 0).rgb;
}}
// The same neighbour's scratch lane — see `aux` in the body below.
fn tap_aux(coord: vec2<i32>, offset: vec2<i32>) -> f32 {{
let last = vec2<i32>(textureDimensions(source)) - vec2<i32>(1);
return textureLoad(source, clamp(coord + offset, vec2<i32>(0), last), 0).a;
}}
{helper_src}{encode_fn}
@compute @workgroup_size(8, 8, 1)
fn main(@builtin(global_invocation_id) gid: vec3<u32>) {{
let dims = textureDimensions(output);
if (gid.x >= dims.x || gid.y >= dims.y) {{
return;
}}
let coord = vec2<i32>(gid.xy);
// What this render is, relative to the export it has to match.
let render_dims = u.detail_base.xy;
let render_scale = u.detail_base.z;
// How many entries `instances` actually holds, read from the buffer itself
// rather than from a uniform so the two cannot disagree. A pass that
// declared no list is bound to a one-element placeholder and never asks.
let instance_count = arrayLength(&instances);
var c = tap(coord, vec2<i32>(0));
// One scalar per pixel that survives the hand-off from one pass to the
// next, alongside the colour. See `DetailPass::wgsl` for what it is for
// and why three channels were not enough.
var aux = tap_aux(coord, vec2<i32>(0));
{{
{indented}
}}
{tail}
}}
",
if writes_output {
"display-encoded, in the output space."
} else {
"linear sRGB, for the next pass."
},
);
let structure_hash = crate::operation::hash_source(&source);
ComposedDetailPass {
label,
source,
uniforms: uniform_values,
storage: pass.storage.clone(),
radius: pass.radius,
writes_output,
structure_hash,
}
}
// Compiled for this crate's own tests as well as for the feature, so that
// `cargo test -p dr-pipeline` exercises the seam whether or not anybody
// downstream remembered to turn the feature on. A test that quietly does not
// exist is worse than no test, because the absence looks like a pass.
#[cfg(any(test, feature = "detail-probe"))]
pub mod probe;
#[cfg(test)]
mod tests {
use super::*;
#[test]
fn a_full_render_approximates_nothing() {
let s = RenderScale::full((2000, 1300));
assert!(!s.is_proxy());
assert!((s.ratio() - 1.0).abs() < 1e-6);
// A one-pixel sharpening radius is one pixel at export, always.
assert!((s.source_pixels(1.0) - 1.0).abs() < 1e-6);
assert!(s.resolves(1.0));
}
#[test]
fn a_proxy_shrinks_a_source_length_and_says_so() {
// A 6000px frame in a 1500px panel: four source pixels per render
// pixel, so a 1px capture-sharpening radius is a quarter of a render
// pixel and cannot be drawn. This is the case the whole type exists
// for, and the answer has to be "no", not a plausible-looking number.
let s = RenderScale::new((1500, 1000), (6000, 4000));
assert!(s.is_proxy());
assert!((s.ratio() - 0.25).abs() < 1e-6);
assert!(!s.resolves(1.0), "a quarter of a pixel is not a kernel");
assert!(s.resolves(4.0), "four source pixels do survive");
}
#[test]
fn a_frame_fraction_is_the_same_proportion_at_every_size() {
// The mask rule, restated as a test: 1% of the shorter edge is 1% of
// the shorter edge whether the render is a thumbnail or an export.
// This is what makes a clarity radius tuned on screen correct in the
// exported file.
let proxy = RenderScale::new((2000, 1333), (6000, 4000));
let export = RenderScale::full((6000, 4000));
let as_fraction = |s: &RenderScale| {
let (w, h) = s.render_size();
s.frame_fraction(0.01) / w.min(h) as f32
};
assert!((as_fraction(&proxy) - as_fraction(&export)).abs() < 1e-6);
// And in absolute terms it really does scale with the render.
assert!((proxy.frame_fraction(0.01) - 13.33).abs() < 0.5);
assert!((export.frame_fraction(0.01) - 40.0).abs() < 0.5);
}
#[test]
fn zooming_to_one_to_one_makes_the_preview_exact() {
// The reason there is no separate full-resolution preview path: the
// framing's view rect shrinks while the render target keeps its size,
// so the ratio climbs back to 1.0 and a sharpening radius means
// exactly what it will mean in the file.
let fit = RenderScale::new((2000, 1333), (6000, 4000));
let one_to_one = RenderScale::new((2000, 1333), (2000, 1333));
assert!(!fit.resolves(1.0));
assert!(one_to_one.resolves(1.0));
}
use crate::detail::probe::BoxBlur;
use crate::operation::{compose_full, OutputMode};
fn with_blur(radius: f32) -> Vec<Box<dyn Operation>> {
let mut ops = crate::ops::chain();
ops.push(Box::new(BoxBlur::with_radius(radius)));
ops
}
fn fused(ops: &[Box<dyn Operation>]) -> crate::ComposedShader {
compose_full(
ops,
&crate::Framing::new(),
dr_types::ColourSpace::Srgb,
&crate::mask::MaskStack::new(),
&crate::spot::SpotSet::new(),
)
}
#[test]
fn a_detail_operation_contributes_nothing_to_the_fused_shader() {
// The seam itself: a neighbourhood operation is in the graph, is
// active, and yet emits no block in the single dispatch — because it
// physically cannot, and asking it for one would produce an empty
// block that reads as an operation doing nothing.
let shader = fused(&with_blur(0.05));
assert!(
!shader.source.contains("---- detail_probe ----"),
"a detail operation must not appear as a fused fragment"
);
assert!(
!shader.source.contains("detail_probe_radius"),
"nor should it occupy a slot in the fused uniform block"
);
}
#[test]
fn an_active_detail_operation_makes_the_fused_pass_hand_on_linear_values() {
// The other half of the same decision. With no detail stage the fused
// pass encodes and quantises, exactly as it always has; with one, it
// stops at linear working values and the detail chain finishes the
// job. Getting this wrong is not a subtle wrong colour — it is a
// storage format that does not match the texture bound to it.
let neutral = fused(&with_blur(0.0));
assert_eq!(neutral.output_mode, OutputMode::Encoded);
assert!(neutral.source.contains("texture_storage_2d<rgba8unorm"));
assert!(neutral.source.contains("encode_output"));
let blurring = fused(&with_blur(0.05));
assert_eq!(blurring.output_mode, OutputMode::LinearWorking);
assert!(blurring.source.contains("texture_storage_2d<rgba16float"));
assert!(
!blurring.source.contains("fn encode_output"),
"the fused pass must not encode when a detail stage follows: \
FR-DEV-2 allows exactly one quantisation"
);
assert!(
!blurring.source.contains("clamp(c, vec3<f32>(0.0)"),
"nor clip, or the sharpener sees a hard edge at every highlight"
);
}
#[test]
fn a_neutral_detail_operation_costs_the_edit_nothing() {
// The rule the whole pipeline is built on, extended to this stage: an
// operation at its defaults contributes no code, no uniform and no
// dispatch. An unedited photograph must not pay for a sharpener it is
// not using.
let ops = with_blur(0.0);
let composed = compose_detail(
&ops,
RenderScale::full((512, 512)),
dr_types::ColourSpace::Srgb,
);
assert!(composed.is_empty());
assert_eq!(fused(&ops).output_mode, OutputMode::Encoded);
}
#[test]
fn a_separable_blur_becomes_two_passes_and_only_the_last_encodes() {
// The multi-pass case, which is the one the ping-pong exists for. The
// first pass writes a linear intermediate and the second writes the
// display texture — so the output transform happens exactly once, at
// the end, wherever the end happens to be.
let ops = with_blur(0.05);
let composed = compose_detail(
&ops,
RenderScale::full((512, 512)),
dr_types::ColourSpace::Srgb,
);
assert_eq!(composed.len(), 2);
let first = &composed.passes[0];
let last = &composed.passes[1];
assert_eq!(first.label, "detail_probe/horizontal");
assert_eq!(last.label, "detail_probe/vertical");
assert!(!first.writes_output);
assert!(first.source.contains("texture_storage_2d<rgba16float"));
assert!(!first.source.contains("fn encode_output"));
assert!(last.writes_output);
assert!(last.source.contains("texture_storage_2d<rgba8unorm"));
assert!(last.source.contains("fn encode_output"));
// Two passes of one operation are two shaders, so they must not share
// a pipeline-cache entry — the classic way a second pass silently runs
// the first one's code.
assert_ne!(first.structure_hash, last.structure_hash);
}
#[test]
fn a_pass_addresses_its_uniforms_without_knowing_the_block() {
// The same contract the fused composer offers: a body writes `radius`
// and the composer rewrites it to a prefixed struct field, so two
// operations may both call a uniform `radius` and neither has to know.
let ops = with_blur(0.05);
let composed = compose_detail(
&ops,
RenderScale::full((512, 512)),
dr_types::ColourSpace::Srgb,
);
let src = &composed.passes[0].source;
assert!(src.contains("detail_probe_0_radius: f32,"));
assert!(src.contains("let r = i32(u.detail_probe_0_radius);"));
// And the second pass carries its own index, so its uniforms cannot be
// uploaded into the first pass's slots.
assert!(composed.passes[1]
.source
.contains("detail_probe_1_radius: f32,"));
}
#[test]
fn every_pass_declares_a_uniform_block_the_gpu_will_accept() {
// A uniform struct whose size is not a multiple of 16 is rejected
// outright by the WGSL uniform address space rules, and the failure
// arrives as a shader compilation error against generated source.
let ops = with_blur(0.05);
for pass in compose_detail(
&ops,
RenderScale::full((512, 512)),
dr_types::ColourSpace::Srgb,
)
.passes
{
assert_eq!(pass.uniforms.len() % 4, 0, "{}", pass.label);
assert!(pass.uniforms.iter().all(|v| v.is_finite()));
// The base block is first and fixed, so a pass never addresses a
// slot by number and the render size is always in the same place.
assert_eq!(pass.uniforms[0], 512.0);
assert_eq!(pass.uniforms[1], 512.0);
}
}
#[test]
fn the_declared_radius_is_the_halo_a_tile_would_need() {
// ARCH §5.3 schedules tiles, and a tile cannot be computed without
// knowing how far outside itself the pass reads. Nothing can infer it
// from the WGSL — the offsets are computed from uniforms at runtime —
// so the operation states it, and this is the assertion that it states
// the truth rather than zero.
let ops = with_blur(0.05);
let scale = RenderScale::full((400, 400));
let composed = compose_detail(&ops, scale, dr_types::ColourSpace::Srgb);
let expected = BoxBlur::with_radius(0.05).kernel(scale);
assert_eq!(expected, 20, "5% of a 400px edge");
assert_eq!(composed.radius(), expected);
assert!(composed.passes.iter().all(|p| p.radius == expected));
}
#[test]
fn a_normalised_radius_is_the_same_effect_at_every_resolution() {
// FR-DSP-1's hard part, at the level this crate can test it: the same
// edit composed at two sizes produces kernels in the same *proportion*
// to the frame. `dr-gpu`'s `proxy_and_export_agree` checks the pixels
// that fall out of it.
let ops = with_blur(0.04);
let sizes = [(200u32, 200u32), (800, 800), (2400, 2400)];
let fractions: Vec<f32> = sizes
.iter()
.map(|&(w, h)| {
let scale = RenderScale::full((w, h));
let composed = compose_detail(&ops, scale, dr_types::ColourSpace::Srgb);
composed.radius() as f32 / w.min(h) as f32
})
.collect();
for f in &fractions {
assert!(
(f - 0.04).abs() < 0.005,
"the kernel drifted from the declared fraction: {fractions:?}"
);
}
}
#[test]
fn an_active_operation_that_draws_nothing_still_finishes_the_render() {
// The seam between the two composers, and the one case where they
// cannot see each other. `compose_full` decides to hand on linear
// working values from the *operations* — it has no resolution to
// consult — while this composer converts a radius and can legitimately
// decide there is nothing to draw at this size. An empty chain would
// then leave the output transform undone: the fused pass writes
// `rgba16float` and the frontend binds an `rgba8unorm` target to it.
//
// A photographer meets this by turning on capture sharpening or
// luminance noise reduction while the develop view is fitted to a
// large file, which is the normal way to work, so it is not an edge
// case that can be left to fail.
let ops = with_blur(0.001);
let scale = RenderScale::full((400, 400));
assert!(ops.last().expect("the blur").is_active());
assert_eq!(
BoxBlur::with_radius(0.001).passes(scale).len(),
0,
"the premise: a radius too small to draw emits no pass"
);
let composed = compose_detail(&ops, scale, dr_types::ColourSpace::Srgb);
assert_eq!(composed.len(), 1, "the chain must not be empty here");
assert_eq!(composed.radius(), 0, "it reads only the pixel it writes");
let resolve = &composed.passes[0];
assert_eq!(resolve.label, "detail/resolve");
assert!(resolve.writes_output);
assert!(resolve.source.contains("texture_storage_2d<rgba8unorm"));
assert!(resolve.source.contains("fn encode_output"));
// Exactly the fixed base block and no more: a pass with no body has
// nothing of its own to upload, and the block still has to be a
// multiple of sixteen bytes.
assert_eq!(resolve.uniforms.len(), DETAIL_BASE_UNIFORM_FIELDS);
assert_eq!(resolve.uniforms.len() % 4, 0);
// And it really is a copy: the fused pass composed alongside it is the
// one that stopped short, so the two agree about who encodes.
assert_eq!(fused(&ops).output_mode, OutputMode::LinearWorking);
}
#[test]
fn a_pass_can_hand_a_scalar_to_the_next_one_alongside_the_colour() {
// What makes an unsharp mask — sharpening, clarity, texture, dehaze —
// expressible at all in a chain that hands each pass exactly one
// texture. The blur goes in `aux` and the colour rides through
// untouched, so the pass that combines them receives both; without the
// lane, four numbers would have to fit in three channels and the
// operation could only ever be a blur.
//
// A pass that says nothing about `aux` hands on what it was given,
// which is why the box blur below needs no knowledge of it.
let ops = with_blur(0.05);
let composed = compose_detail(
&ops,
RenderScale::full((512, 512)),
dr_types::ColourSpace::Srgb,
);
for pass in &composed.passes {
assert!(
pass.source.contains("fn tap_aux(") && pass.source.contains("var aux = tap_aux("),
"{} cannot read the scratch lane",
pass.label
);
}
assert!(
composed.passes[0]
.source
.contains("textureStore(output, coord, vec4<f32>(c, aux));"),
"an intermediate must carry the lane to the pass after it"
);
// The last pass writes the display texture, whose alpha is opacity and
// not scratch space. Readable there, not written — which is the right
// way round, because the combining pass is the one that reads it.
assert!(composed.passes[1].writes_output);
assert!(!composed.passes[1].source.contains("vec4<f32>(c, aux)"));
}
#[test]
fn an_edit_with_no_detail_operation_composes_no_passes() {
// The property that keeps the cost of this stage at zero for the
// overwhelmingly common edit: no sharpening means no chain, which
// means `dr-gpu` runs the single fused dispatch it always did.
let ops = crate::ops::chain();
let composed = compose_detail(
&ops,
RenderScale::full((64, 64)),
dr_types::ColourSpace::Srgb,
);
assert!(composed.is_empty());
assert_eq!(composed.radius(), 0);
}
}