Clarity's Gaussian sigma is 1.2% of the frame's shorter edge, so its radius is a property of the viewport: 52 render pixels at 4K, two separable passes of 105 taps each over 8.3 M pixels. That measured 33.9 ms — seven times the entire fused point chain, for one slider — and is docs/technical-debt.md TD-4. A detail pass may now declare `output_scale`, and clarity's base is computed on a grid a quarter the size on each axis. The pass that combines needs the blur *and* the full-resolution colour, and a colour that has been through a quarter-scale target is no longer full resolution. So a scaled pass cannot simply join the ping-pong: there are two chains now. The full-resolution one carries the colour and no scaled pass touches it; the reduced one carries the base and reaches the combining pass through a second binding as `reduced_at()`. The reduce is a dispatch of its own rather than something the first blur half does on the way past, and that is the whole difference between this and the strided kernel the module documentation rules out. A stride samples an image that is not band-limited and aliases high-frequency content down into the base, which is then subtracted, and arrives in the output as mottling across smooth gradients. This band-limits first and samples after. What is discarded is content the base could not represent at any resolution, because a Gaussian at sigma = 26 px holds nothing above one cycle per 26 px and the quarter-scale grid carries one per 8 — so the reduced base is not an approximation of the full-resolution one, it is the same function sampled where it is still determined. Which is also why the scale belongs to the band rather than to the stage. Texture's sigma is a decade finer, so the reduce pass's own box would be wider than the Gaussian it was prefiltering; texture never reduces. And clarity steps 4 -> 2 -> 1 as sigma falls, because a quarter of a small sigma is not a Gaussian either — the case that gives up is the one that was already cheap. `radius` stays in each pass's own pixels and `ComposedDetail::radius` multiplies it back up, so 13 reduced pixels at scale 4 still report the 52 render pixels a tile would have to be grown by. The halo a scheduler sees does not move. The halo tests pass unchanged, which was TD-4's stated bar; they render at 1024 px and so exercise the reduced path rather than stepping around it. Added `crossing_the_reduction_threshold_does_not_change_the_picture`, because nothing yet compared the reduced form against a *less* reduced one — every other test measures one form against itself. It renders the same edit either side of the 4 -> 2 step-down and holds the peak excursion to 0.03 stops and the reach to 2% of the frame. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
1266 lines
57 KiB
Rust
1266 lines
57 KiB
Rust
//! Neighbourhood operations — the ones that must read a pixel they are not
|
|
//! writing.
|
|
//!
|
|
//! # Why this exists at all
|
|
//!
|
|
//! Every operation in [`crate::operation`] contributes a fragment taking a
|
|
//! `vec3<f32>` and returning one. That contract is what makes the fused
|
|
//! dispatch possible, and it is also an absolute wall: a fragment is handed a
|
|
//! colour, not a coordinate, so it cannot look left. Sharpening, noise
|
|
//! reduction, clarity, texture, dehaze and spot removal are all defined by
|
|
//! what the *neighbours* are doing, and none of them can be written as a point
|
|
//! function of `c` at any price.
|
|
//!
|
|
//! FR-DEV-3 asks for all six and FR-DEV-8 for spot removal. So the fused pass
|
|
//! is not the whole pipeline; it is the *point-operation* stage of it, and
|
|
//! this module is the stage that follows.
|
|
//!
|
|
//! # Where it sits, and why there
|
|
//!
|
|
//! ```text
|
|
//! demosaiced source (camera space, full sensor resolution)
|
|
//! |
|
|
//! | <- framing prologue: output pixel -> source position
|
|
//! v
|
|
//! +------------------------------------------+
|
|
//! | the fused point-operation pass | one dispatch
|
|
//! | white balance, exposure, tone, colour |
|
|
//! | the mask layers |
|
|
//! | camera RGB -> linear sRGB |
|
|
//! +------------------------------------------+
|
|
//! | rgba16float, linear, **unclipped**, at render resolution
|
|
//! v
|
|
//! +------------------------------------------+
|
|
//! | the detail stage - this module | one dispatch per pass
|
|
//! | sharpen, NR, clarity, texture, spots |
|
|
//! +------------------------------------------+
|
|
//! | the last pass applies the output transform
|
|
//! v
|
|
//! rgba8unorm display or export texture
|
|
//! ```
|
|
//!
|
|
//! Four things about that position are decisions rather than convenience, and
|
|
//! each of them could defensibly have gone the other way.
|
|
//!
|
|
//! **After tone, not before.** Sharpening before a tone curve and sharpening
|
|
//! after it are different pictures, not the same picture computed two ways: an
|
|
//! S-curve steepens the mid-tones, so a halo introduced before it is amplified
|
|
//! by whatever slope the curve happens to have at that luminance, and the
|
|
//! amount that looked right stops looking right the moment the curve moves.
|
|
//! After the curve, the amount the user chose is the amount they see, and it
|
|
//! survives every later change to tone. This is also what ARCH §5.2 draws:
|
|
//! texture, clarity, spot removal and sharpen/NR sit below the tone curve and
|
|
//! the colour mixer.
|
|
//!
|
|
//! **In linear light, after the camera matrix.** The fused pass works in
|
|
//! *camera* space, because white balance and exposure are physically
|
|
//! meaningful there and nowhere else. A detail pass is the opposite case: it
|
|
//! wants a luminance, and camera RGB has no luminance — the three channels are
|
|
//! whatever the CFA's dyes passed, and weighting them 0.2126/0.7152/0.0722
|
|
//! would be numerology. So the split is taken *after* the `cam_to_srgb`
|
|
//! multiply, where the working space is linear sRGB and a luminance is a
|
|
//! luminance.
|
|
//!
|
|
//! **Before the output transform, and before the clip.** FR-DEV-2 allows
|
|
//! exactly one quantisation, at the display or export stage. A detail pass
|
|
//! reading an 8-bit display-encoded texture and writing another one would
|
|
//! quantise twice and do its arithmetic in a space where a difference of one
|
|
//! code value means different things at different brightnesses — which is how
|
|
//! sharpening ends up with visible banding in a sky. The intermediate is
|
|
//! therefore `rgba16float` and holds linear values that have **not** been
|
|
//! clamped to `0..=1`: a recovered highlight is still above one at this point,
|
|
//! and clipping it before the sharpener sees it would put a hard edge exactly
|
|
//! where the sharpener is most visible. The last detail pass performs the
|
|
//! primaries conversion, the clip and the encode, so the single quantisation
|
|
//! stays single.
|
|
//!
|
|
//! **After framing, at render resolution.** The alternative — running detail
|
|
//! on the demosaiced source before the framing prologue — is superficially
|
|
//! attractive, because a radius in sensor pixels would then mean exactly what
|
|
//! it says. It is unaffordable: the source is the full sensor, so a detail
|
|
//! pass there costs 24 MP of work for a 2 MP preview and FR-DSP-1 stops being
|
|
//! true. Running at render resolution instead makes the cost proportional to
|
|
//! what is on screen, and pushes the whole difficulty into one place — the
|
|
//! scale — which [`RenderScale`] exists to make explicit rather than implicit.
|
|
//!
|
|
//! # What this stage deliberately cannot do
|
|
//!
|
|
//! **There is no per-mask detail.** A mask layer's chain is fused into the
|
|
//! point-operation pass; the detail stage runs once, afterwards, over the
|
|
//! whole frame. Local sharpening is therefore not expressible here, and
|
|
//! [`crate::mask::MaskLayer::active_ops`] filters detail operations out rather
|
|
//! than emitting a block that would silently do nothing. Making it possible
|
|
//! means giving a detail pass the mask array and a layer index, which is a
|
|
//! change to this module's shader preamble and not to its shape — but it is
|
|
//! not done, and a caller should not assume it.
|
|
//!
|
|
//! # Adding a neighbourhood operation
|
|
//!
|
|
//! Declare it in `ops/<id>.yaml` with `rust:`, exactly as the tone curve does
|
|
//! — the schema in `ops/README.md` describes a point function, and stretching
|
|
//! it to cover kernels would be a worse language than Rust aimed at one
|
|
//! caller. Then implement [`crate::Operation`] as usual for the parameters,
|
|
//! descriptor and sidecar, and additionally:
|
|
//!
|
|
//! ```ignore
|
|
//! impl Operation for Sharpen {
|
|
//! fn affects(&self) -> Affects { Affects::Detail }
|
|
//! fn detail(&self) -> Option<&dyn DetailStage> { Some(self) }
|
|
//! fn wgsl_body(&self) -> String { String::new() } // never called
|
|
//! // ... descriptor, set_param, param, is_active exactly as usual
|
|
//! }
|
|
//!
|
|
//! impl DetailStage for Sharpen {
|
|
//! fn passes(&self, scale: RenderScale) -> Vec<DetailPass> { /* ... */ }
|
|
//! }
|
|
//! ```
|
|
//!
|
|
//! Everything else arrives unchanged and for free: the develop panel builds
|
|
//! its controls from the descriptor, the sidecar persists the parameters, the
|
|
//! history and the presets carry them, and an operation at its defaults
|
|
//! contributes no pass at all.
|
|
|
|
use std::fmt::Write as _;
|
|
|
|
use dr_types::ColourSpace;
|
|
|
|
use crate::operation::{Helper, Operation, Uniform};
|
|
|
|
/// Floats the generated detail uniform block always carries, before an
|
|
/// operation's own.
|
|
///
|
|
/// One `vec4`, which is also the smallest a WGSL uniform struct can be and
|
|
/// stay aligned. See [`compose_detail`] for what the lanes hold.
|
|
pub const DETAIL_BASE_UNIFORM_FIELDS: usize = 4;
|
|
|
|
/// TRACES: FR-DSP-1
|
|
/// The relationship between the resolution an edit is being **rendered** at
|
|
/// and the resolution it will eventually be **exported** at.
|
|
///
|
|
/// # The problem this type is the answer to
|
|
///
|
|
/// A point operation is scale-free. Exposure is a multiply, and multiplying by
|
|
/// two is multiplying by two whether the frame is 2 000 pixels wide or 24 000.
|
|
/// Every operation in the fused pass has this property, which is why nothing
|
|
/// in the pipeline has needed to know its own resolution until now.
|
|
///
|
|
/// A neighbourhood operation has no such luck. "Sharpen with a radius of one
|
|
/// pixel" is a statement about a specific grid, and the develop view is not
|
|
/// rendering on that grid — FR-DSP-1 has it rendering at whatever the viewport
|
|
/// needs, which for a 60 MP frame in a 2 000 px panel is one render pixel per
|
|
/// nine source pixels. Tune a radius there, export at full size, and the
|
|
/// exported file is sharpened at a ninth of the strength the photographer
|
|
/// chose. That is not a rounding difference; it is a different photograph.
|
|
///
|
|
/// # The rule
|
|
///
|
|
/// **A length is stored normalised and converted here.** Never store pixels in
|
|
/// an edit. This is not a new idea in this codebase — [`crate::mask`] already
|
|
/// does it, storing every feather and morphology radius as a fraction of the
|
|
/// frame's shorter edge and multiplying up in `dr-gpu` at whatever size the
|
|
/// mask is being rasterised at (see `MaskLayer::feather`, and
|
|
/// `field_short_edge` in `dr-gpu`'s mask pass). A detail operation follows the
|
|
/// same rule through [`Self::frame_fraction`] and gets the same guarantee: the
|
|
/// effect covers the same *proportion* of the picture at every size, so what
|
|
/// was tuned on screen is what lands in the file.
|
|
///
|
|
/// # Two units, because there are two kinds of length
|
|
///
|
|
/// The mask rule is not quite enough on its own, because detail operations
|
|
/// split into two families that mean different things by "radius":
|
|
///
|
|
/// - **Compositional** — clarity, texture, dehaze. The radius is a fraction of
|
|
/// the picture, tens of pixels at any size, and [`Self::frame_fraction`] is
|
|
/// exactly right. These preview faithfully at any scale.
|
|
///
|
|
/// - **Acutance** — capture sharpening, luminance noise reduction. The radius
|
|
/// is a property of the *sensor*: it is about the lens's circle of confusion
|
|
/// and the demosaic's interpolation, both measured in source pixels and
|
|
/// neither of which cares how large the viewport is.
|
|
/// [`Self::source_pixels`] converts one of those into render pixels.
|
|
///
|
|
/// # The honest limit
|
|
///
|
|
/// For the second family the conversion runs out. At a one-ninth proxy a
|
|
/// 1.0-source-pixel radius is 0.11 render pixels, and there is no kernel that
|
|
/// represents a ninth of a pixel — the information the sharpener would act on
|
|
/// was thrown away by the downscale before the pass ever ran. No arrangement
|
|
/// of this stage recovers it, which is why every editor that has shipped tells
|
|
/// the photographer to judge sharpening at 1:1, and why Lightroom's detail
|
|
/// panel contains a 1:1 loupe rather than a scaled preview.
|
|
///
|
|
/// [`Self::resolves`] reports that condition instead of hiding it, so an
|
|
/// operation can fade itself out and an interface can say "zoom to 100% to
|
|
/// judge this" — which is the truth, and better than a preview that lies.
|
|
/// Zooming is enough: the framing's view rect shrinks while the render target
|
|
/// keeps its size, so [`Self::ratio`] climbs back to 1.0 at 1:1 and the
|
|
/// preview becomes exact, with no separate full-resolution path to maintain.
|
|
#[derive(Debug, Clone, Copy, PartialEq)]
|
|
pub struct RenderScale {
|
|
render: (u32, u32),
|
|
full: (u32, u32),
|
|
}
|
|
|
|
impl RenderScale {
|
|
/// `render` is the size being rendered now; `full` is the size the same
|
|
/// framed region would have at source resolution.
|
|
///
|
|
/// Both describe *the region being looked at*, not the whole photograph —
|
|
/// so a crop and a zoom are already accounted for by the time they arrive.
|
|
/// [`crate::EditGraph::render_scale`] works both out from the framing, and
|
|
/// is what a caller should normally use.
|
|
pub fn new(render: (u32, u32), full: (u32, u32)) -> Self {
|
|
Self {
|
|
render: (render.0.max(1), render.1.max(1)),
|
|
full: (full.0.max(1), full.1.max(1)),
|
|
}
|
|
}
|
|
|
|
/// A scale that is already at source resolution — an export, or a 1:1
|
|
/// view. [`Self::ratio`] is 1.0 and nothing is approximated.
|
|
pub fn full(render: (u32, u32)) -> Self {
|
|
Self::new(render, render)
|
|
}
|
|
|
|
pub fn render_size(&self) -> (u32, u32) {
|
|
self.render
|
|
}
|
|
|
|
pub fn full_size(&self) -> (u32, u32) {
|
|
self.full
|
|
}
|
|
|
|
/// Render pixels per source pixel. 1.0 at export, below 1.0 on a proxy.
|
|
///
|
|
/// Averaged over the two axes rather than taken from one. They agree to
|
|
/// within a pixel by construction — both sizes describe the same rectangle
|
|
/// — but each is separately rounded to an integer, and taking the mean
|
|
/// stops a narrow viewport disagreeing with itself.
|
|
pub fn ratio(&self) -> f32 {
|
|
let x = self.render.0 as f32 / self.full.0 as f32;
|
|
let y = self.render.1 as f32 / self.full.1 as f32;
|
|
(x + y) * 0.5
|
|
}
|
|
|
|
/// Whether this render is smaller than the file it stands for.
|
|
pub fn is_proxy(&self) -> bool {
|
|
self.ratio() < 0.999
|
|
}
|
|
|
|
/// A length stated in **source pixels**, in render pixels.
|
|
///
|
|
/// For the acutance family — sharpening, luminance NR — whose radius is a
|
|
/// property of the sensor rather than of the composition.
|
|
pub fn source_pixels(&self, radius: f32) -> f32 {
|
|
radius * self.ratio()
|
|
}
|
|
|
|
/// A length stated as a **fraction of the frame's shorter edge**, in
|
|
/// render pixels.
|
|
///
|
|
/// For the compositional family — clarity, texture, dehaze — and the same
|
|
/// unit `dr-gpu`'s mask rasteriser already converts feathers in. An edit
|
|
/// stored this way is resolution-independent by construction.
|
|
pub fn frame_fraction(&self, fraction: f32) -> f32 {
|
|
fraction * self.render.0.min(self.render.1) as f32
|
|
}
|
|
|
|
/// Whether a radius stated in source pixels survives this render.
|
|
///
|
|
/// False means the effect is smaller than a pixel here and whatever is
|
|
/// drawn is a guess. Report it; do not paper over it — see the type's
|
|
/// documentation for why there is nothing better to do.
|
|
pub fn resolves(&self, radius_in_source_pixels: f32) -> bool {
|
|
self.source_pixels(radius_in_source_pixels) >= 1.0
|
|
}
|
|
}
|
|
|
|
/// One dispatch of a neighbourhood operation.
|
|
///
|
|
/// An operation returns as many of these as it needs. A separable Gaussian is
|
|
/// two — horizontal then vertical — and gets the ping-pong between them for
|
|
/// free; an unsharp mask wanting its blur held alongside the original would be
|
|
/// more, and is the case this shape exists to leave room for.
|
|
#[derive(Debug, Clone, PartialEq)]
|
|
pub struct DetailPass {
|
|
/// A short name, used to label the GPU pass and to make a shader
|
|
/// compilation failure say which of an operation's passes broke.
|
|
pub label: &'static str,
|
|
|
|
/// The furthest this pass reads from the pixel it writes, in **render**
|
|
/// pixels.
|
|
///
|
|
/// Declared rather than inferred from the WGSL, because nothing can infer
|
|
/// it from the WGSL: the offsets are computed at runtime from uniforms.
|
|
/// It is the halo a tile has to be grown by before this pass can be
|
|
/// computed tile-wise (ARCH §5.3), and it is the reason a detail operation
|
|
/// is not simply "some more shader code" — the scheduler has to know how
|
|
/// far the dependency reaches before it can schedule anything at all.
|
|
///
|
|
/// An understated radius shows as a seam at every tile boundary, which is
|
|
/// the kind of artefact that looks like a driver bug. State it honestly.
|
|
pub radius: u32,
|
|
|
|
/// How much smaller than the render this pass writes.
|
|
///
|
|
/// `1` is the ordinary case and means "the render size", which is what
|
|
/// every pass did before this field existed. A larger value writes a
|
|
/// target that many times smaller on each axis, into a **second** chain
|
|
/// held alongside the full-resolution one — see [`Self::wgsl`] for how the
|
|
/// two are addressed, and the module documentation for why there are two.
|
|
///
|
|
/// # Why a pass may want this
|
|
///
|
|
/// A blur wide enough to be a *base* — clarity's is 1.2% of the frame,
|
|
/// 52 render pixels at 4K — holds no spatial frequency a quarter-scale
|
|
/// grid cannot represent. Computing it at the render size therefore buys
|
|
/// nothing and costs everything: 105 taps over 8.3 M pixels, twice, which
|
|
/// measured at 34 ms and is where `docs/technical-debt.md` TD-4 came from.
|
|
/// At a quarter it is a sixteenth of the pixels at a quarter of the
|
|
/// radius, and the result is not an approximation of the full-resolution
|
|
/// base — it is the same band-limited function, sampled where it is still
|
|
/// Nyquist-safe.
|
|
///
|
|
/// [`Self::radius`] stays in this pass's **own** pixels, so a pass at
|
|
/// scale 4 with a radius of 13 declares 13, not 52. The halo it implies
|
|
/// for a tile scheduler is `radius * output_scale`, and
|
|
/// [`ComposedDetail::radius`] is what performs that multiplication —
|
|
/// stating the radius in the grid the loop actually runs in is what keeps
|
|
/// the shader and the declaration the same number.
|
|
///
|
|
/// **Never the last pass.** The final pass carries the output transform
|
|
/// and writes the display texture, which is full resolution by
|
|
/// definition; a scaled pass in that position is a codegen bug and
|
|
/// `dr-gpu` refuses it rather than binding a shader to a target of the
|
|
/// wrong size.
|
|
pub output_scale: u32,
|
|
|
|
/// The WGSL body.
|
|
///
|
|
/// Reads and writes `c`, a `vec3<f32>` of **linear sRGB**, pre-loaded with
|
|
/// this pixel's own value. Also in scope:
|
|
///
|
|
/// - `coord: vec2<i32>` — this pixel.
|
|
/// - `tap(coord, offset) -> vec3<f32>` — a neighbour, clamped to the edge
|
|
/// of the image, which is what makes a kernel at the border average the
|
|
/// pixels that exist rather than fade into black.
|
|
/// - `render_dims: vec2<f32>` and `render_scale: f32` — the size being
|
|
/// rendered and [`RenderScale::ratio`], for the rare pass that needs
|
|
/// them in the shader. Prefer computing lengths on the CPU in
|
|
/// [`DetailStage::passes`], where the units are named methods rather
|
|
/// than an untyped float.
|
|
/// - `aux: f32` and `tap_aux(coord, offset) -> f32` — **one scalar per
|
|
/// pixel that survives to the next pass**, pre-loaded with what the
|
|
/// previous pass left there and written back out unless the body
|
|
/// assigns it.
|
|
/// - `reduced_at(coord) -> f32` — the **reduced chain's** scalar at this
|
|
/// pixel,
|
|
/// bilinearly upsampled. Zero unless a scaled pass ran earlier in this
|
|
/// operation; see [`Self::output_scale`].
|
|
///
|
|
/// `coord` is always in *this pass's own* output grid, and `tap` maps it
|
|
/// into the source's grid for you. A pass at [`Self::output_scale`] 4
|
|
/// therefore addresses its own quarter-size target with `coord`, while
|
|
/// `tap(coord, offset)` offsets in **source** pixels — which is what lets
|
|
/// a reduce pass average the 4 x 4 block a single output pixel covers by
|
|
/// looping `offset` over it. Where source and target are the same size the
|
|
/// mapping is the identity, so every pass written before scaling existed
|
|
/// behaves exactly as it did.
|
|
///
|
|
/// # Why `aux` exists
|
|
///
|
|
/// The ping-pong hands each pass exactly one texture: what the pass before
|
|
/// it wrote. That is enough for a chain of filters — a separable blur is
|
|
/// two of them — and it is *not* enough for an unsharp mask, which is the
|
|
/// shape of sharpening, clarity, texture and dehaze alike. An unsharp mask
|
|
/// needs the blur **and** the original in the same place at the same time,
|
|
/// and once the first pass has written its blur the original is gone.
|
|
///
|
|
/// Three channels cannot carry both. Even restricted to the case where the
|
|
/// operation only moves luminance — so the colour is a luminance and two
|
|
/// chromaticity degrees of freedom — the combining pass needs four
|
|
/// numbers: the original luminance, two of chromaticity, and the blurred
|
|
/// luminance. Four does not fit in three, and no encoding makes it fit.
|
|
///
|
|
/// The intermediate is `rgba16float` and its alpha was being written as a
|
|
/// constant `1.0` and read by nobody, so the fourth number goes there. A
|
|
/// blur pass leaves `c` alone and puts its result in `aux`; the pass after
|
|
/// it therefore receives the untouched original *and* the blur, and can
|
|
/// subtract one from the other. An operation with no use for the lane says
|
|
/// nothing and hands on what it was given.
|
|
///
|
|
/// The last pass in the chain writes the display texture, whose alpha is
|
|
/// opacity rather than scratch space, so `aux` is readable there and not
|
|
/// written. That is exactly the right way round: the combining pass is the
|
|
/// one that reads it.
|
|
///
|
|
/// Uniforms are addressed by the bare names declared in [`Self::uniforms`],
|
|
/// exactly as a fused fragment addresses its own; the composer rewrites
|
|
/// them to their prefixed struct fields.
|
|
///
|
|
/// **Values are not clipped.** A recovered highlight arrives above 1.0 and
|
|
/// an out-of-gamut colour can arrive below 0.0. That is deliberate — see
|
|
/// the module documentation — and a kernel that assumes `0..=1` will
|
|
/// produce dark rings around specular highlights.
|
|
pub wgsl: String,
|
|
|
|
/// Uniform values this pass's body reads.
|
|
pub uniforms: Vec<Uniform>,
|
|
|
|
/// TRACES: FR-DEV-8
|
|
/// Per-instance data, for a pass whose work is a *list* rather than a
|
|
/// kernel.
|
|
///
|
|
/// Reaches the body as `instances: array<vec4<f32>>`, with
|
|
/// `instance_count` in scope as a `u32`. Empty for every pass that is a
|
|
/// convolution, which is every pass that existed before spot removal.
|
|
///
|
|
/// # Why not the uniform block
|
|
///
|
|
/// Because the uniform block is fixed by the pass's *structure*, and this
|
|
/// is not: sixty-four repairs and one repair are the same shader with a
|
|
/// different buffer behind it. Packing the list into uniforms would need a
|
|
/// fixed maximum, paid for on every frame whether the photograph carries
|
|
/// one spot or none, and it would need the composer to emit `vec4` fields —
|
|
/// a WGSL uniform array has a stride of 16 whatever it holds.
|
|
///
|
|
/// The property that matters more: with the list in storage the generated
|
|
/// source does not mention how many there are, so placing the tenth spot
|
|
/// uploads 512 bytes and reuses the compiled pipeline, exactly as moving a
|
|
/// slider does for the fused pass.
|
|
pub storage: Vec<[f32; 4]>,
|
|
}
|
|
|
|
/// TRACES: FR-DEV-3 | FR-DEV-8
|
|
/// An operation that reads pixels other than the one it is writing.
|
|
///
|
|
/// Implemented *alongside* [`Operation`], never instead of it: the parameters,
|
|
/// the descriptor, the panel controls and the sidecar all come from the
|
|
/// `Operation` half, and only the execution differs. An operation that
|
|
/// implements this must also return [`crate::Affects::Detail`] from
|
|
/// `affects()` and `Some(self)` from `Operation::detail()` — the three are
|
|
/// checked against each other by a test in [`crate::operation`], because an
|
|
/// operation that forgot one of them would be dropped from both stages and
|
|
/// simply not happen, with no error anywhere.
|
|
pub trait DetailStage: Send + Sync {
|
|
/// The passes to run, in order, at this resolution.
|
|
///
|
|
/// Called per render, so the operation sees the scale it is actually being
|
|
/// asked to draw at and converts its own lengths here — in Rust, where
|
|
/// [`RenderScale`]'s two conversions are named after the two units, rather
|
|
/// than in WGSL where both would be a bare `f32`.
|
|
///
|
|
/// Returning an empty vector means "nothing to do at this scale", which is
|
|
/// the honest answer for an acutance operation on a heavy proxy. It is
|
|
/// **not** how an operation says it is neutral: that is `is_active()`, and
|
|
/// an inactive operation is never asked.
|
|
fn passes(&self, scale: RenderScale) -> Vec<DetailPass>;
|
|
}
|
|
|
|
/// One compile-ready detail pass: complete WGSL and the uniform block for it.
|
|
#[derive(Debug, Clone, PartialEq)]
|
|
pub struct ComposedDetailPass {
|
|
/// `<op id>/<pass label>`, for GPU labels and error messages.
|
|
pub label: String,
|
|
/// Complete, compilable WGSL.
|
|
pub source: String,
|
|
/// Uniform values in the order the generated struct declares them.
|
|
pub uniforms: Vec<f32>,
|
|
/// TRACES: FR-DEV-8
|
|
/// The instance list, if this pass declared one. See [`DetailPass::storage`].
|
|
pub storage: Vec<[f32; 4]>,
|
|
/// See [`DetailPass::radius`].
|
|
pub radius: u32,
|
|
/// See [`DetailPass::output_scale`].
|
|
pub output_scale: u32,
|
|
/// Whether this pass writes the display/export texture rather than another
|
|
/// linear intermediate.
|
|
///
|
|
/// True for exactly the last pass in the chain, which carries the output
|
|
/// transform — the primaries conversion, the clip and the encode that the
|
|
/// fused pass performs when there is no detail stage at all. Folding them
|
|
/// into the last pass rather than adding a resolve dispatch keeps the cost
|
|
/// of the stage at one dispatch per pass, not one plus one.
|
|
pub writes_output: bool,
|
|
/// Identifies this pass's *structure*, for the pipeline cache. Covers the
|
|
/// generated source, not the uniform values — so moving a slider uploads a
|
|
/// buffer and reuses the compiled pipeline, exactly as the fused pass does.
|
|
pub structure_hash: u64,
|
|
}
|
|
|
|
/// The detail stage of one edit, at one resolution.
|
|
#[derive(Debug, Clone, Default, PartialEq)]
|
|
pub struct ComposedDetail {
|
|
pub passes: Vec<ComposedDetailPass>,
|
|
}
|
|
|
|
impl ComposedDetail {
|
|
/// Whether the edit has no detail stage — the common case, and the one
|
|
/// that must cost nothing.
|
|
pub fn is_empty(&self) -> bool {
|
|
self.passes.is_empty()
|
|
}
|
|
|
|
pub fn len(&self) -> usize {
|
|
self.passes.len()
|
|
}
|
|
|
|
/// The widest halo any pass needs, in render pixels (ARCH §5.3).
|
|
///
|
|
/// A pass declares its radius in its own grid, so a scaled pass's has to
|
|
/// be multiplied back up before the maxima are comparable: 13 reduced
|
|
/// pixels at scale 4 reach exactly as far across the photograph as 52
|
|
/// render pixels do, and a scheduler comparing the two unscaled would size
|
|
/// a halo at a quarter of what the pass actually reads.
|
|
pub fn radius(&self) -> u32 {
|
|
self.passes
|
|
.iter()
|
|
.map(|p| p.radius.saturating_mul(p.output_scale))
|
|
.max()
|
|
.unwrap_or(0)
|
|
}
|
|
}
|
|
|
|
/// TRACES: FR-DEV-3 | FR-DSP-1
|
|
/// Generate the detail stage for a set of operations at one resolution.
|
|
///
|
|
/// Operations that declare no [`DetailStage`], or that are at their neutral
|
|
/// settings, contribute nothing — the same rule the fused composer follows, so
|
|
/// an edit with no sharpening produces an empty chain and `dr-gpu` runs the
|
|
/// single dispatch it always did.
|
|
///
|
|
/// `output` is the space the **last** pass encodes into, and it is a parameter
|
|
/// for the same reason it is a parameter to [`crate::compose_with_framing`]: a
|
|
/// screen render and a Display P3 export are the same edit and different
|
|
/// shaders, and neither is more authoritative than the other.
|
|
///
|
|
/// # The generated uniform block
|
|
///
|
|
/// A fixed `vec4` first, then the pass's own scalars, prefixed with the
|
|
/// operation id so that a pass never has to know what else is in the block.
|
|
/// The lanes of the leading `vec4` are, in order: render width, render height,
|
|
/// [`RenderScale::ratio`], and the pass's index within its operation. The
|
|
/// first three reach the body as `render_dims` and `render_scale`; the fourth
|
|
/// is there because a two-pass operation emitting one body for both directions
|
|
/// is a reasonable thing to want, and would otherwise need a uniform of its
|
|
/// own purely to say which half it is in.
|
|
pub fn compose_detail(
|
|
ops: &[Box<dyn Operation>],
|
|
scale: RenderScale,
|
|
output: ColourSpace,
|
|
) -> ComposedDetail {
|
|
compose_detail_with(ops, &[], scale, output)
|
|
}
|
|
|
|
/// TRACES: FR-DEV-8
|
|
/// The detail stage with a set of repairs ahead of the operations.
|
|
///
|
|
/// `spots` are already-built passes, from [`crate::SpotSet::passes`], and they
|
|
/// go **first** — before sharpening, before noise reduction, before every
|
|
/// kernel in `ops`.
|
|
///
|
|
/// That placement is a decision rather than an ordering convenience. A
|
|
/// sharpening kernel reads a neighbourhood, so sharpening a dust mark before
|
|
/// removing it smears the mark's edge into pixels the repair's own disc does
|
|
/// not cover: what is left afterwards is a faint over-sharpened ring around an
|
|
/// otherwise perfect patch, which is exactly the artefact that reads as broken
|
|
/// software. Removing the mark first means every later pass sees the
|
|
/// photograph the photographer thinks they are sharpening.
|
|
///
|
|
/// It also means ARCH §5.2's stage list, which draws spot removal after
|
|
/// texture and clarity, is not what this does — see `docs/spot-removal.md`
|
|
/// §5.1, which is where the disagreement is written down.
|
|
pub fn compose_detail_with(
|
|
ops: &[Box<dyn Operation>],
|
|
spots: &[DetailPass],
|
|
scale: RenderScale,
|
|
output: ColourSpace,
|
|
) -> ComposedDetail {
|
|
// Every pass of every active detail operation, flattened, carrying the
|
|
// operation it came from for the uniform prefix and the helper set.
|
|
//
|
|
// The helper slice borrows from the operation rather than being `'static`:
|
|
// `Operation::helpers` hands out a slice owned by the operation now, so
|
|
// that a node built from a declaration at load time can own its list
|
|
// (FR-PLG-2). The borrow lasts as long as `ops`, which outlives this
|
|
// function's body.
|
|
let mut planned: Vec<(&str, &[Helper], DetailPass, usize)> = Vec::new();
|
|
for (index, pass) in spots.iter().enumerate() {
|
|
planned.push((
|
|
crate::spot::SPOT_ID,
|
|
crate::spot::SPOT_HELPERS,
|
|
pass.clone(),
|
|
index,
|
|
));
|
|
}
|
|
for op in ops {
|
|
if !op.is_active() {
|
|
continue;
|
|
}
|
|
let Some(stage) = op.detail() else {
|
|
continue;
|
|
};
|
|
let id = op.descriptor().id.0;
|
|
for (index, pass) in stage.passes(scale).into_iter().enumerate() {
|
|
planned.push((id, op.helpers(), pass, index));
|
|
}
|
|
}
|
|
|
|
// An active detail operation that emitted nothing at this scale.
|
|
//
|
|
// Legal, and the honest answer for an acutance operation on a heavy proxy
|
|
// — a one-source-pixel radius is a third of a render pixel there and no
|
|
// kernel represents a third of a pixel (see [`RenderScale`]). But it opens
|
|
// a hole between the two halves of the composition: [`compose_full`]
|
|
// decides to hand on linear working values from the *operations*, which it
|
|
// must, having no scale to consult, so the fused pass has already stopped
|
|
// short of the output transform. Returning an empty chain here would leave
|
|
// that transform undone and bind an `rgba16float` shader to an
|
|
// `rgba8unorm` target, which surfaces as a wgpu validation failure a long
|
|
// way from the cause.
|
|
//
|
|
// So the chain is never empty when the fused pass is expecting one: a
|
|
// single pass with no body, which reads the intermediate and performs the
|
|
// output transform the fused pass skipped. One dispatch, in the uncommon
|
|
// case where a photographer has a kernel switched on at a scale that
|
|
// cannot draw it — against the alternative of the preview failing outright
|
|
// or `compose_full` growing a resolution argument it has no other use for.
|
|
if planned.is_empty() && ops.iter().any(|o| o.is_active() && o.detail().is_some()) {
|
|
return ComposedDetail {
|
|
passes: vec![compose_one(
|
|
RESOLVE_ID,
|
|
&[],
|
|
&DetailPass {
|
|
output_scale: 1,
|
|
label: "resolve",
|
|
radius: 0,
|
|
wgsl: String::new(),
|
|
uniforms: Vec::new(),
|
|
storage: Vec::new(),
|
|
},
|
|
0,
|
|
scale,
|
|
output,
|
|
true,
|
|
)],
|
|
};
|
|
}
|
|
|
|
let last = planned.len().saturating_sub(1);
|
|
let passes = planned
|
|
.into_iter()
|
|
.enumerate()
|
|
.map(|(position, (id, helpers, pass, index))| {
|
|
compose_one(id, helpers, &pass, index, scale, output, position == last)
|
|
})
|
|
.collect();
|
|
|
|
ComposedDetail { passes }
|
|
}
|
|
|
|
/// The operation id the resolve pass is labelled with.
|
|
///
|
|
/// Not an operation: no `ops/*.yaml` declares it and nothing in the chain
|
|
/// answers to it. It exists so the generated label reads `detail/resolve`
|
|
/// rather than borrowing the id of whichever operation happened to fall
|
|
/// through, which would send a reader looking for a bug in that operation.
|
|
const RESOLVE_ID: &str = "detail";
|
|
|
|
#[allow(clippy::too_many_arguments)]
|
|
fn compose_one(
|
|
id: &str,
|
|
helpers: &[Helper],
|
|
pass: &DetailPass,
|
|
index: usize,
|
|
scale: RenderScale,
|
|
output: ColourSpace,
|
|
writes_output: bool,
|
|
) -> ComposedDetailPass {
|
|
let prefix = format!("{}_{index}", crate::operation::sanitise(id));
|
|
|
|
let mut uniform_fields = String::from(
|
|
" // x, y: the size being rendered. z: render pixels per source\n\
|
|
\x20 // pixel — 1.0 at export, less on a proxy (FR-DSP-1). w: which\n\
|
|
\x20 // pass of this operation this is.\n\
|
|
\x20 detail_base: vec4<f32>,\n",
|
|
);
|
|
let (rw, rh) = scale.render_size();
|
|
let mut uniform_values = vec![rw as f32, rh as f32, scale.ratio(), index as f32];
|
|
debug_assert_eq!(uniform_values.len(), DETAIL_BASE_UNIFORM_FIELDS);
|
|
|
|
if !pass.uniforms.is_empty() {
|
|
let _ = writeln!(uniform_fields, " // {id}/{}", pass.label);
|
|
}
|
|
for u in &pass.uniforms {
|
|
let _ = writeln!(uniform_fields, " {prefix}_{}: f32,", u.name);
|
|
uniform_values.push(u.value);
|
|
}
|
|
|
|
// A uniform struct whose size is not a multiple of 16 is rejected by the
|
|
// WGSL uniform address space rules — the same padding the fused composer
|
|
// applies, for the same reason.
|
|
let pad = (4 - (uniform_values.len() % 4)) % 4;
|
|
for i in 0..pad {
|
|
let _ = writeln!(uniform_fields, " _pad{i}: f32,");
|
|
uniform_values.push(0.0);
|
|
}
|
|
|
|
let mut body = pass.wgsl.clone();
|
|
for u in &pass.uniforms {
|
|
body = crate::operation::rewrite_uniform(&body, u.name, &format!("u.{prefix}_{}", u.name));
|
|
}
|
|
|
|
let mut helper_src = String::new();
|
|
let mut seen: Vec<&str> = Vec::new();
|
|
for h in helpers {
|
|
if seen.contains(&h.name) {
|
|
continue;
|
|
}
|
|
seen.push(h.name);
|
|
let _ = writeln!(helper_src, "{}\n", h.source.trim_end());
|
|
}
|
|
|
|
// The storage format and the tail are the *only* difference between an
|
|
// intermediate pass and the final one. Everything above — the taps, the
|
|
// uniforms, the body — is identical, which is what lets an operation write
|
|
// one kernel without knowing whether it happens to be last in the chain.
|
|
let (store_format, tail) = if writes_output {
|
|
(
|
|
"rgba8unorm",
|
|
format!(
|
|
"{} // Clip to the output gamut and encode. The one quantisation\n\
|
|
\x20 // the pipeline performs (FR-DEV-2), and it is here rather than\n\
|
|
\x20 // in the fused pass because this is now the last thing to run.\n\
|
|
\x20 c = clamp(c, vec3<f32>(0.0), vec3<f32>(1.0));\n\
|
|
\x20 textureStore(output, coord, vec4<f32>(encode_output(c), 1.0));",
|
|
crate::operation::primaries_conversion(output)
|
|
),
|
|
)
|
|
} else {
|
|
(
|
|
"rgba16float",
|
|
" // Another linear intermediate: no clip and no encode, because\n\
|
|
\x20 // the pass after this one still has to read real values.\n\
|
|
\x20 //\n\
|
|
\x20 // `aux` rides in alpha. A pass that never touches it hands on\n\
|
|
\x20 // whatever it was given, so the lane costs an operation that\n\
|
|
\x20 // does not want it exactly one copy of a value it already read.\n\
|
|
\x20 textureStore(output, coord, vec4<f32>(c, aux));"
|
|
.to_string(),
|
|
)
|
|
};
|
|
|
|
let encode_fn = if writes_output {
|
|
crate::operation::encode_output_fn(output)
|
|
} else {
|
|
String::new()
|
|
};
|
|
|
|
let label = format!("{id}/{}", pass.label);
|
|
let indented = body
|
|
.lines()
|
|
.map(|l| format!(" {l}"))
|
|
.collect::<Vec<_>>()
|
|
.join("\n");
|
|
let source = format!(
|
|
"// GENERATED — do not edit.
|
|
//
|
|
// Detail pass `{label}` — a neighbourhood operation, which is why it is a
|
|
// dispatch of its own rather than a block in the fused shader: it reads pixels
|
|
// it is not writing, and the fused contract hands a fragment a colour with no
|
|
// way back to a coordinate.
|
|
//
|
|
// In: linear sRGB, scene-referred, **unclipped**, at render resolution.
|
|
// Out: {}
|
|
|
|
struct Params {{
|
|
{uniform_fields}}}
|
|
|
|
@group(0) @binding(0) var source: texture_2d<f32>;
|
|
@group(0) @binding(1) var<uniform> u: Params;
|
|
@group(0) @binding(2) var output: texture_storage_2d<{store_format}, write>;
|
|
// A pass whose work is a list rather than a kernel reads it here; every other
|
|
// pass leaves this bound to a single empty element and never looks at it. See
|
|
// `DetailPass::storage` for why the list is not in the uniform block.
|
|
@group(0) @binding(3) var<storage, read> instances: array<vec4<f32>>;
|
|
// The reduced chain — what a scaled pass most recently wrote, at whatever
|
|
// fraction of the render size it declared. Bound to a 1x1 placeholder for
|
|
// every pass that never calls `base`, so that one bind group layout serves a
|
|
// pass which uses it and a pass which has never heard of it.
|
|
@group(0) @binding(4) var reduced: texture_2d<f32>;
|
|
|
|
// Where in `source` this output pixel begins.
|
|
//
|
|
// The ratio is 1 whenever a pass writes what it reads, which is every pass
|
|
// that does not set `output_scale` — the multiply and the divide cancel
|
|
// exactly, so the ordinary case is unchanged and pays two integer operations
|
|
// for the privilege. A scaled pass gets the top-left of the block it covers,
|
|
// which is what makes `tap`'s offsets mean *source* pixels and lets a reduce
|
|
// pass walk its own footprint.
|
|
fn source_origin(coord: vec2<i32>) -> vec2<i32> {{
|
|
let src = vec2<i32>(textureDimensions(source));
|
|
let dst = vec2<i32>(textureDimensions(output));
|
|
return coord * src / max(dst, vec2<i32>(1));
|
|
}}
|
|
|
|
// A neighbour, clamped to the edge of the image.
|
|
//
|
|
// Clamped rather than zero-filled: a kernel straddling the border must average
|
|
// the pixels that exist. Returning zero there darkens every edge by a band the
|
|
// width of the radius, which reads as a vignette nobody asked for and is the
|
|
// classic way a first convolution goes wrong.
|
|
fn tap(coord: vec2<i32>, offset: vec2<i32>) -> vec3<f32> {{
|
|
let last = vec2<i32>(textureDimensions(source)) - vec2<i32>(1);
|
|
let at = source_origin(coord) + offset;
|
|
return textureLoad(source, clamp(at, vec2<i32>(0), last), 0).rgb;
|
|
}}
|
|
|
|
// The same neighbour's scratch lane — see `aux` in the body below.
|
|
fn tap_aux(coord: vec2<i32>, offset: vec2<i32>) -> f32 {{
|
|
let last = vec2<i32>(textureDimensions(source)) - vec2<i32>(1);
|
|
let at = source_origin(coord) + offset;
|
|
return textureLoad(source, clamp(at, vec2<i32>(0), last), 0).a;
|
|
}}
|
|
|
|
// The reduced chain, read at this pass's own resolution.
|
|
//
|
|
// Named `reduced_at` rather than `base` because `base` is a natural local in a
|
|
// body that has just computed one — spot removal already has such a local, and
|
|
// a function shadowed by a variable is a compile error a long way from its
|
|
// cause.
|
|
//
|
|
// Bilinear, and on pixel *centres* rather than corners: the reduce pass took
|
|
// its sample at the centre of the block it averaged, so an upsample that
|
|
// treated the grids as corner-aligned would shift the base by half a reduced
|
|
// pixel — two full pixels at scale 4, which on a wide unsharp mask is a base
|
|
// offset from the image it is subtracted from, and reads as a directional
|
|
// smear along every edge.
|
|
//
|
|
// Nearest would be cheaper and is not enough: the base is subtracted from the
|
|
// full-resolution image, so any blockiness in it appears in the *difference*
|
|
// at full contrast. That is a visible 4-pixel grid over the whole frame.
|
|
fn reduced_at(coord: vec2<i32>) -> f32 {{
|
|
let rd = vec2<f32>(textureDimensions(reduced));
|
|
let dst = vec2<f32>(max(textureDimensions(output), vec2<u32>(1u)));
|
|
let p = (vec2<f32>(coord) + vec2<f32>(0.5)) * rd / dst - vec2<f32>(0.5);
|
|
let last = vec2<i32>(rd) - vec2<i32>(1);
|
|
let base_px = vec2<i32>(floor(p));
|
|
let f = fract(p);
|
|
|
|
let s00 = textureLoad(reduced, clamp(base_px, vec2<i32>(0), last), 0).a;
|
|
let s10 = textureLoad(reduced, clamp(base_px + vec2<i32>(1, 0), vec2<i32>(0), last), 0).a;
|
|
let s01 = textureLoad(reduced, clamp(base_px + vec2<i32>(0, 1), vec2<i32>(0), last), 0).a;
|
|
let s11 = textureLoad(reduced, clamp(base_px + vec2<i32>(1, 1), vec2<i32>(0), last), 0).a;
|
|
|
|
return mix(mix(s00, s10, f.x), mix(s01, s11, f.x), f.y);
|
|
}}
|
|
|
|
{helper_src}{encode_fn}
|
|
@compute @workgroup_size(8, 8, 1)
|
|
fn main(@builtin(global_invocation_id) gid: vec3<u32>) {{
|
|
let dims = textureDimensions(output);
|
|
if (gid.x >= dims.x || gid.y >= dims.y) {{
|
|
return;
|
|
}}
|
|
|
|
let coord = vec2<i32>(gid.xy);
|
|
// What this render is, relative to the export it has to match.
|
|
let render_dims = u.detail_base.xy;
|
|
let render_scale = u.detail_base.z;
|
|
// How many entries `instances` actually holds, read from the buffer itself
|
|
// rather than from a uniform so the two cannot disagree. A pass that
|
|
// declared no list is bound to a one-element placeholder and never asks.
|
|
let instance_count = arrayLength(&instances);
|
|
|
|
var c = tap(coord, vec2<i32>(0));
|
|
// One scalar per pixel that survives the hand-off from one pass to the
|
|
// next, alongside the colour. See `DetailPass::wgsl` for what it is for
|
|
// and why three channels were not enough.
|
|
var aux = tap_aux(coord, vec2<i32>(0));
|
|
|
|
{{
|
|
{indented}
|
|
}}
|
|
|
|
{tail}
|
|
}}
|
|
",
|
|
if writes_output {
|
|
"display-encoded, in the output space."
|
|
} else {
|
|
"linear sRGB, for the next pass."
|
|
},
|
|
);
|
|
|
|
let structure_hash = crate::operation::hash_source(&source);
|
|
|
|
ComposedDetailPass {
|
|
label,
|
|
source,
|
|
uniforms: uniform_values,
|
|
storage: pass.storage.clone(),
|
|
radius: pass.radius,
|
|
// Clamped rather than trusted: a zero would divide by nothing in the
|
|
// dispatch size and a declaration is data, which since FR-PLG-2 can
|
|
// come from a file this build did not write.
|
|
output_scale: pass.output_scale.max(1),
|
|
writes_output,
|
|
structure_hash,
|
|
}
|
|
}
|
|
|
|
// Compiled for this crate's own tests as well as for the feature, so that
|
|
// `cargo test -p dr-pipeline` exercises the seam whether or not anybody
|
|
// downstream remembered to turn the feature on. A test that quietly does not
|
|
// exist is worse than no test, because the absence looks like a pass.
|
|
#[cfg(any(test, feature = "detail-probe"))]
|
|
pub mod probe;
|
|
|
|
#[cfg(test)]
|
|
mod tests {
|
|
use super::*;
|
|
|
|
#[test]
|
|
fn a_full_render_approximates_nothing() {
|
|
let s = RenderScale::full((2000, 1300));
|
|
assert!(!s.is_proxy());
|
|
assert!((s.ratio() - 1.0).abs() < 1e-6);
|
|
// A one-pixel sharpening radius is one pixel at export, always.
|
|
assert!((s.source_pixels(1.0) - 1.0).abs() < 1e-6);
|
|
assert!(s.resolves(1.0));
|
|
}
|
|
|
|
#[test]
|
|
fn a_proxy_shrinks_a_source_length_and_says_so() {
|
|
// A 6000px frame in a 1500px panel: four source pixels per render
|
|
// pixel, so a 1px capture-sharpening radius is a quarter of a render
|
|
// pixel and cannot be drawn. This is the case the whole type exists
|
|
// for, and the answer has to be "no", not a plausible-looking number.
|
|
let s = RenderScale::new((1500, 1000), (6000, 4000));
|
|
assert!(s.is_proxy());
|
|
assert!((s.ratio() - 0.25).abs() < 1e-6);
|
|
assert!(!s.resolves(1.0), "a quarter of a pixel is not a kernel");
|
|
assert!(s.resolves(4.0), "four source pixels do survive");
|
|
}
|
|
|
|
#[test]
|
|
fn a_frame_fraction_is_the_same_proportion_at_every_size() {
|
|
// The mask rule, restated as a test: 1% of the shorter edge is 1% of
|
|
// the shorter edge whether the render is a thumbnail or an export.
|
|
// This is what makes a clarity radius tuned on screen correct in the
|
|
// exported file.
|
|
let proxy = RenderScale::new((2000, 1333), (6000, 4000));
|
|
let export = RenderScale::full((6000, 4000));
|
|
let as_fraction = |s: &RenderScale| {
|
|
let (w, h) = s.render_size();
|
|
s.frame_fraction(0.01) / w.min(h) as f32
|
|
};
|
|
assert!((as_fraction(&proxy) - as_fraction(&export)).abs() < 1e-6);
|
|
// And in absolute terms it really does scale with the render.
|
|
assert!((proxy.frame_fraction(0.01) - 13.33).abs() < 0.5);
|
|
assert!((export.frame_fraction(0.01) - 40.0).abs() < 0.5);
|
|
}
|
|
|
|
#[test]
|
|
fn zooming_to_one_to_one_makes_the_preview_exact() {
|
|
// The reason there is no separate full-resolution preview path: the
|
|
// framing's view rect shrinks while the render target keeps its size,
|
|
// so the ratio climbs back to 1.0 and a sharpening radius means
|
|
// exactly what it will mean in the file.
|
|
let fit = RenderScale::new((2000, 1333), (6000, 4000));
|
|
let one_to_one = RenderScale::new((2000, 1333), (2000, 1333));
|
|
assert!(!fit.resolves(1.0));
|
|
assert!(one_to_one.resolves(1.0));
|
|
}
|
|
|
|
use crate::detail::probe::BoxBlur;
|
|
use crate::operation::{compose_full, OutputMode};
|
|
|
|
fn with_blur(radius: f32) -> Vec<Box<dyn Operation>> {
|
|
let mut ops = crate::ops::chain();
|
|
ops.push(Box::new(BoxBlur::with_radius(radius)));
|
|
ops
|
|
}
|
|
|
|
fn fused(ops: &[Box<dyn Operation>]) -> crate::ComposedShader {
|
|
compose_full(
|
|
ops,
|
|
&crate::Framing::new(),
|
|
dr_types::ColourSpace::Srgb,
|
|
&crate::mask::MaskStack::new(),
|
|
&crate::spot::SpotSet::new(),
|
|
)
|
|
}
|
|
|
|
#[test]
|
|
fn a_detail_operation_contributes_nothing_to_the_fused_shader() {
|
|
// The seam itself: a neighbourhood operation is in the graph, is
|
|
// active, and yet emits no block in the single dispatch — because it
|
|
// physically cannot, and asking it for one would produce an empty
|
|
// block that reads as an operation doing nothing.
|
|
let shader = fused(&with_blur(0.05));
|
|
assert!(
|
|
!shader.source.contains("---- detail_probe ----"),
|
|
"a detail operation must not appear as a fused fragment"
|
|
);
|
|
assert!(
|
|
!shader.source.contains("detail_probe_radius"),
|
|
"nor should it occupy a slot in the fused uniform block"
|
|
);
|
|
}
|
|
|
|
#[test]
|
|
fn an_active_detail_operation_makes_the_fused_pass_hand_on_linear_values() {
|
|
// The other half of the same decision. With no detail stage the fused
|
|
// pass encodes and quantises, exactly as it always has; with one, it
|
|
// stops at linear working values and the detail chain finishes the
|
|
// job. Getting this wrong is not a subtle wrong colour — it is a
|
|
// storage format that does not match the texture bound to it.
|
|
let neutral = fused(&with_blur(0.0));
|
|
assert_eq!(neutral.output_mode, OutputMode::Encoded);
|
|
assert!(neutral.source.contains("texture_storage_2d<rgba8unorm"));
|
|
assert!(neutral.source.contains("encode_output"));
|
|
|
|
let blurring = fused(&with_blur(0.05));
|
|
assert_eq!(blurring.output_mode, OutputMode::LinearWorking);
|
|
assert!(blurring.source.contains("texture_storage_2d<rgba16float"));
|
|
assert!(
|
|
!blurring.source.contains("fn encode_output"),
|
|
"the fused pass must not encode when a detail stage follows: \
|
|
FR-DEV-2 allows exactly one quantisation"
|
|
);
|
|
assert!(
|
|
!blurring.source.contains("clamp(c, vec3<f32>(0.0)"),
|
|
"nor clip, or the sharpener sees a hard edge at every highlight"
|
|
);
|
|
}
|
|
|
|
#[test]
|
|
fn a_neutral_detail_operation_costs_the_edit_nothing() {
|
|
// The rule the whole pipeline is built on, extended to this stage: an
|
|
// operation at its defaults contributes no code, no uniform and no
|
|
// dispatch. An unedited photograph must not pay for a sharpener it is
|
|
// not using.
|
|
let ops = with_blur(0.0);
|
|
let composed = compose_detail(
|
|
&ops,
|
|
RenderScale::full((512, 512)),
|
|
dr_types::ColourSpace::Srgb,
|
|
);
|
|
assert!(composed.is_empty());
|
|
assert_eq!(fused(&ops).output_mode, OutputMode::Encoded);
|
|
}
|
|
|
|
#[test]
|
|
fn a_separable_blur_becomes_two_passes_and_only_the_last_encodes() {
|
|
// The multi-pass case, which is the one the ping-pong exists for. The
|
|
// first pass writes a linear intermediate and the second writes the
|
|
// display texture — so the output transform happens exactly once, at
|
|
// the end, wherever the end happens to be.
|
|
let ops = with_blur(0.05);
|
|
let composed = compose_detail(
|
|
&ops,
|
|
RenderScale::full((512, 512)),
|
|
dr_types::ColourSpace::Srgb,
|
|
);
|
|
assert_eq!(composed.len(), 2);
|
|
|
|
let first = &composed.passes[0];
|
|
let last = &composed.passes[1];
|
|
assert_eq!(first.label, "detail_probe/horizontal");
|
|
assert_eq!(last.label, "detail_probe/vertical");
|
|
|
|
assert!(!first.writes_output);
|
|
assert!(first.source.contains("texture_storage_2d<rgba16float"));
|
|
assert!(!first.source.contains("fn encode_output"));
|
|
|
|
assert!(last.writes_output);
|
|
assert!(last.source.contains("texture_storage_2d<rgba8unorm"));
|
|
assert!(last.source.contains("fn encode_output"));
|
|
|
|
// Two passes of one operation are two shaders, so they must not share
|
|
// a pipeline-cache entry — the classic way a second pass silently runs
|
|
// the first one's code.
|
|
assert_ne!(first.structure_hash, last.structure_hash);
|
|
}
|
|
|
|
#[test]
|
|
fn a_pass_addresses_its_uniforms_without_knowing_the_block() {
|
|
// The same contract the fused composer offers: a body writes `radius`
|
|
// and the composer rewrites it to a prefixed struct field, so two
|
|
// operations may both call a uniform `radius` and neither has to know.
|
|
let ops = with_blur(0.05);
|
|
let composed = compose_detail(
|
|
&ops,
|
|
RenderScale::full((512, 512)),
|
|
dr_types::ColourSpace::Srgb,
|
|
);
|
|
let src = &composed.passes[0].source;
|
|
assert!(src.contains("detail_probe_0_radius: f32,"));
|
|
assert!(src.contains("let r = i32(u.detail_probe_0_radius);"));
|
|
// And the second pass carries its own index, so its uniforms cannot be
|
|
// uploaded into the first pass's slots.
|
|
assert!(composed.passes[1]
|
|
.source
|
|
.contains("detail_probe_1_radius: f32,"));
|
|
}
|
|
|
|
#[test]
|
|
fn every_pass_declares_a_uniform_block_the_gpu_will_accept() {
|
|
// A uniform struct whose size is not a multiple of 16 is rejected
|
|
// outright by the WGSL uniform address space rules, and the failure
|
|
// arrives as a shader compilation error against generated source.
|
|
let ops = with_blur(0.05);
|
|
for pass in compose_detail(
|
|
&ops,
|
|
RenderScale::full((512, 512)),
|
|
dr_types::ColourSpace::Srgb,
|
|
)
|
|
.passes
|
|
{
|
|
assert_eq!(pass.uniforms.len() % 4, 0, "{}", pass.label);
|
|
assert!(pass.uniforms.iter().all(|v| v.is_finite()));
|
|
// The base block is first and fixed, so a pass never addresses a
|
|
// slot by number and the render size is always in the same place.
|
|
assert_eq!(pass.uniforms[0], 512.0);
|
|
assert_eq!(pass.uniforms[1], 512.0);
|
|
}
|
|
}
|
|
|
|
#[test]
|
|
fn the_declared_radius_is_the_halo_a_tile_would_need() {
|
|
// ARCH §5.3 schedules tiles, and a tile cannot be computed without
|
|
// knowing how far outside itself the pass reads. Nothing can infer it
|
|
// from the WGSL — the offsets are computed from uniforms at runtime —
|
|
// so the operation states it, and this is the assertion that it states
|
|
// the truth rather than zero.
|
|
let ops = with_blur(0.05);
|
|
let scale = RenderScale::full((400, 400));
|
|
let composed = compose_detail(&ops, scale, dr_types::ColourSpace::Srgb);
|
|
let expected = BoxBlur::with_radius(0.05).kernel(scale);
|
|
assert_eq!(expected, 20, "5% of a 400px edge");
|
|
assert_eq!(composed.radius(), expected);
|
|
assert!(composed.passes.iter().all(|p| p.radius == expected));
|
|
}
|
|
|
|
#[test]
|
|
fn a_normalised_radius_is_the_same_effect_at_every_resolution() {
|
|
// FR-DSP-1's hard part, at the level this crate can test it: the same
|
|
// edit composed at two sizes produces kernels in the same *proportion*
|
|
// to the frame. `dr-gpu`'s `proxy_and_export_agree` checks the pixels
|
|
// that fall out of it.
|
|
let ops = with_blur(0.04);
|
|
let sizes = [(200u32, 200u32), (800, 800), (2400, 2400)];
|
|
let fractions: Vec<f32> = sizes
|
|
.iter()
|
|
.map(|&(w, h)| {
|
|
let scale = RenderScale::full((w, h));
|
|
let composed = compose_detail(&ops, scale, dr_types::ColourSpace::Srgb);
|
|
composed.radius() as f32 / w.min(h) as f32
|
|
})
|
|
.collect();
|
|
for f in &fractions {
|
|
assert!(
|
|
(f - 0.04).abs() < 0.005,
|
|
"the kernel drifted from the declared fraction: {fractions:?}"
|
|
);
|
|
}
|
|
}
|
|
|
|
#[test]
|
|
fn an_active_operation_that_draws_nothing_still_finishes_the_render() {
|
|
// The seam between the two composers, and the one case where they
|
|
// cannot see each other. `compose_full` decides to hand on linear
|
|
// working values from the *operations* — it has no resolution to
|
|
// consult — while this composer converts a radius and can legitimately
|
|
// decide there is nothing to draw at this size. An empty chain would
|
|
// then leave the output transform undone: the fused pass writes
|
|
// `rgba16float` and the frontend binds an `rgba8unorm` target to it.
|
|
//
|
|
// A photographer meets this by turning on capture sharpening or
|
|
// luminance noise reduction while the develop view is fitted to a
|
|
// large file, which is the normal way to work, so it is not an edge
|
|
// case that can be left to fail.
|
|
let ops = with_blur(0.001);
|
|
let scale = RenderScale::full((400, 400));
|
|
assert!(ops.last().expect("the blur").is_active());
|
|
assert_eq!(
|
|
BoxBlur::with_radius(0.001).passes(scale).len(),
|
|
0,
|
|
"the premise: a radius too small to draw emits no pass"
|
|
);
|
|
|
|
let composed = compose_detail(&ops, scale, dr_types::ColourSpace::Srgb);
|
|
assert_eq!(composed.len(), 1, "the chain must not be empty here");
|
|
assert_eq!(composed.radius(), 0, "it reads only the pixel it writes");
|
|
|
|
let resolve = &composed.passes[0];
|
|
assert_eq!(resolve.label, "detail/resolve");
|
|
assert!(resolve.writes_output);
|
|
assert!(resolve.source.contains("texture_storage_2d<rgba8unorm"));
|
|
assert!(resolve.source.contains("fn encode_output"));
|
|
// Exactly the fixed base block and no more: a pass with no body has
|
|
// nothing of its own to upload, and the block still has to be a
|
|
// multiple of sixteen bytes.
|
|
assert_eq!(resolve.uniforms.len(), DETAIL_BASE_UNIFORM_FIELDS);
|
|
assert_eq!(resolve.uniforms.len() % 4, 0);
|
|
|
|
// And it really is a copy: the fused pass composed alongside it is the
|
|
// one that stopped short, so the two agree about who encodes.
|
|
assert_eq!(fused(&ops).output_mode, OutputMode::LinearWorking);
|
|
}
|
|
|
|
#[test]
|
|
fn a_pass_can_hand_a_scalar_to_the_next_one_alongside_the_colour() {
|
|
// What makes an unsharp mask — sharpening, clarity, texture, dehaze —
|
|
// expressible at all in a chain that hands each pass exactly one
|
|
// texture. The blur goes in `aux` and the colour rides through
|
|
// untouched, so the pass that combines them receives both; without the
|
|
// lane, four numbers would have to fit in three channels and the
|
|
// operation could only ever be a blur.
|
|
//
|
|
// A pass that says nothing about `aux` hands on what it was given,
|
|
// which is why the box blur below needs no knowledge of it.
|
|
let ops = with_blur(0.05);
|
|
let composed = compose_detail(
|
|
&ops,
|
|
RenderScale::full((512, 512)),
|
|
dr_types::ColourSpace::Srgb,
|
|
);
|
|
|
|
for pass in &composed.passes {
|
|
assert!(
|
|
pass.source.contains("fn tap_aux(") && pass.source.contains("var aux = tap_aux("),
|
|
"{} cannot read the scratch lane",
|
|
pass.label
|
|
);
|
|
}
|
|
assert!(
|
|
composed.passes[0]
|
|
.source
|
|
.contains("textureStore(output, coord, vec4<f32>(c, aux));"),
|
|
"an intermediate must carry the lane to the pass after it"
|
|
);
|
|
// The last pass writes the display texture, whose alpha is opacity and
|
|
// not scratch space. Readable there, not written — which is the right
|
|
// way round, because the combining pass is the one that reads it.
|
|
assert!(composed.passes[1].writes_output);
|
|
assert!(!composed.passes[1].source.contains("vec4<f32>(c, aux)"));
|
|
}
|
|
|
|
#[test]
|
|
fn an_edit_with_no_detail_operation_composes_no_passes() {
|
|
// The property that keeps the cost of this stage at zero for the
|
|
// overwhelmingly common edit: no sharpening means no chain, which
|
|
// means `dr-gpu` runs the single fused dispatch it always did.
|
|
let ops = crate::ops::chain();
|
|
let composed = compose_detail(
|
|
&ops,
|
|
RenderScale::full((64, 64)),
|
|
dr_types::ColourSpace::Srgb,
|
|
);
|
|
assert!(composed.is_empty());
|
|
assert_eq!(composed.radius(), 0);
|
|
}
|
|
}
|