Clarity's Gaussian sigma is 1.2% of the frame's shorter edge, so its radius is a property of the viewport: 52 render pixels at 4K, two separable passes of 105 taps each over 8.3 M pixels. That measured 33.9 ms — seven times the entire fused point chain, for one slider — and is docs/technical-debt.md TD-4. A detail pass may now declare `output_scale`, and clarity's base is computed on a grid a quarter the size on each axis. The pass that combines needs the blur *and* the full-resolution colour, and a colour that has been through a quarter-scale target is no longer full resolution. So a scaled pass cannot simply join the ping-pong: there are two chains now. The full-resolution one carries the colour and no scaled pass touches it; the reduced one carries the base and reaches the combining pass through a second binding as `reduced_at()`. The reduce is a dispatch of its own rather than something the first blur half does on the way past, and that is the whole difference between this and the strided kernel the module documentation rules out. A stride samples an image that is not band-limited and aliases high-frequency content down into the base, which is then subtracted, and arrives in the output as mottling across smooth gradients. This band-limits first and samples after. What is discarded is content the base could not represent at any resolution, because a Gaussian at sigma = 26 px holds nothing above one cycle per 26 px and the quarter-scale grid carries one per 8 — so the reduced base is not an approximation of the full-resolution one, it is the same function sampled where it is still determined. Which is also why the scale belongs to the band rather than to the stage. Texture's sigma is a decade finer, so the reduce pass's own box would be wider than the Gaussian it was prefiltering; texture never reduces. And clarity steps 4 -> 2 -> 1 as sigma falls, because a quarter of a small sigma is not a Gaussian either — the case that gives up is the one that was already cheap. `radius` stays in each pass's own pixels and `ComposedDetail::radius` multiplies it back up, so 13 reduced pixels at scale 4 still report the 52 render pixels a tile would have to be grown by. The halo a scheduler sees does not move. The halo tests pass unchanged, which was TD-4's stated bar; they render at 1024 px and so exercise the reduced path rather than stepping around it. Added `crossing_the_reduction_threshold_does_not_change_the_picture`, because nothing yet compared the reduced form against a *less* reduced one — every other test measures one form against itself. It renders the same edit either side of the 4 -> 2 step-down and holds the peak excursion to 0.03 stops and the reach to 2% of the frame. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
928 lines
40 KiB
Rust
928 lines
40 KiB
Rust
//! TRACES: FR-DEV-3
|
|
//! Noise reduction — luminance and chroma, as two independent amounts.
|
|
//!
|
|
//! # Why two controls and not one
|
|
//!
|
|
//! Sensor noise arrives as two quite different faults, and a photographer
|
|
//! treats them differently because they cost different things to remove.
|
|
//!
|
|
//! **Luminance noise** is fine, high-frequency grain in lightness. It sits at
|
|
//! the same spatial frequency as real detail — eyelashes, fabric weave, tree
|
|
//! bark — so the eye cannot be given more smoothing without also being given
|
|
//! less texture. The radius that helps is one or two photosites, and past
|
|
//! about three the picture stops looking like a photograph and starts looking
|
|
//! like a painting. Many photographers deliberately leave some.
|
|
//!
|
|
//! **Chroma noise** is coarse, blotchy and low-frequency: magenta and green
|
|
//! patches tens of pixels across, produced by the demosaic interpolating
|
|
//! between colour-filtered sites that disagree. Nothing in a photograph looks
|
|
//! like it, so it can be smoothed hard — and it has to be, because a radius
|
|
//! of two pixels does not touch a blotch of twenty. Human spatial acuity for
|
|
//! colour is roughly a quarter of that for lightness, which is why a
|
|
//! chroma-only blur that would be obvious in luminance is invisible here, and
|
|
//! is the same fact JPEG chroma subsampling has exploited since 1992.
|
|
//!
|
|
//! So the radius that is *correct* differs between the two by roughly an
|
|
//! order of magnitude. That is the reason these are two amounts rather than
|
|
//! one: a single slider would either under-treat the colour blotches or
|
|
//! destroy the detail, and there is no setting at which it does neither.
|
|
//!
|
|
//! # The split, and why it makes the two amounts genuinely independent
|
|
//!
|
|
//! Each pass decomposes the linear sRGB colour into a luminance and a colour
|
|
//! difference:
|
|
//!
|
|
//! ```text
|
|
//! y = luminance(c) Rec. 709 weights, exact in this space
|
|
//! d = c - vec3(y) luminance(d) == 0, by construction
|
|
//! c = vec3(y) + d exactly, up to floating-point rounding
|
|
//! ```
|
|
//!
|
|
//! `d` carries no lightness at all: the weights sum to one, so subtracting a
|
|
//! grey of the same luminance leaves a vector whose own luminance is zero.
|
|
//! The luminance pass therefore replaces `y` and returns `d` untouched, and
|
|
//! the chroma passes replace `d` and return `y` untouched. Neither can leak
|
|
//! into the other, which is what lets a photographer set the two sliders
|
|
//! independently and get what they say rather than their product.
|
|
//!
|
|
//! This is also why the split is taken *here* rather than in the fused pass:
|
|
//! the detail stage runs after the camera matrix, where the working space is
|
|
//! linear sRGB and a Rec. 709 luminance is a luminance rather than a weighted
|
|
//! sum of whatever the colour filter array's dyes happened to pass.
|
|
//!
|
|
//! # The filter: a bilateral, in two different arrangements
|
|
//!
|
|
//! A plain Gaussian is not an option. Denoising is exactly the problem of
|
|
//! averaging pixels that differ only by noise while refusing to average
|
|
//! pixels that differ because the scene does, and a Gaussian cannot tell the
|
|
//! difference — it removes grain and edges in the same proportion, which is
|
|
//! the smeared look that makes noise reduction recognisable at a glance.
|
|
//!
|
|
//! A **bilateral filter** multiplies the spatial weight by a *range* weight
|
|
//! that falls off with how different the neighbour's value is, so a neighbour
|
|
//! across an edge contributes almost nothing and the edge survives the
|
|
//! average that removes the grain either side of it. It is the conventional
|
|
//! edge-preserving choice; it needs neither a guide image nor the per-window
|
|
//! statistics a guided filter accumulates, and it is one expression per tap —
|
|
//! which matters, because this runs on the frame path (ARCH §6.1).
|
|
//!
|
|
//! It is also, in its exact form, quadratic: a radius *r* costs `(2r+1)²`
|
|
//! taps. That is affordable at the luminance radius and ruinous at the chroma
|
|
//! radius, and the two are arranged differently in consequence:
|
|
//!
|
|
//! | | radius (source px) | arrangement | taps / pixel | dispatches |
|
|
//! |---|---|---|---|---|
|
|
//! | luminance | 1.0 … 2.5 | exact 2-D bilateral | 9 … 49 | 1 |
|
|
//! | chroma | 2.0 … 12.0 | separable bilateral | 10 … 50 | 2 |
|
|
//!
|
|
//! The worst case with both at full strength is **99 taps per pixel across
|
|
//! three dispatches**, against 49 + 625 = 674 for the exact 2-D form of both.
|
|
//! The chroma pass is where all of that saving is.
|
|
//!
|
|
//! **The luminance pass is exact rather than separable** because at these
|
|
//! radii the separable form is not actually cheaper in the way that matters:
|
|
//! r = 2 is 25 taps in one dispatch against 20 taps in two, and the second
|
|
//! dispatch costs a full-frame `rgba16float` write and read that the taps
|
|
//! saved do not pay for. It is also the higher-quality answer, with none of
|
|
//! the axis-aligned streaking the approximation can show.
|
|
//!
|
|
//! **The chroma pass is separable** — a 1-D bilateral along x, then along y,
|
|
//! the approximation Pham and van Vliet published in 2005. It is an
|
|
//! approximation and not an identity: a bilateral's range weights make the
|
|
//! two-dimensional kernel non-separable in principle, and the residual shows
|
|
//! as faint axis-aligned structure along strong diagonal edges. That is
|
|
//! acceptable here for the same reason the large radius is acceptable — it is
|
|
//! in chroma, where the eye's spatial acuity is four times lower — and the
|
|
//! alternative, 625 taps a pixel at 4K, is roughly five gigataps a frame and
|
|
//! not a frame path at all. Where the approximation *would* be visible, in
|
|
//! luminance, it is not used.
|
|
//!
|
|
//! # The threshold, and what it is a fraction of
|
|
//!
|
|
//! A bilateral needs to know how large a difference counts as noise. A single
|
|
//! absolute number in linear light cannot say: linear light puts middle grey
|
|
//! at 0.18, so a threshold tuned for a highlight is roughly a hundred times
|
|
//! too coarse for a shadow and would flatten it completely.
|
|
//!
|
|
//! The threshold is therefore proportional to the square root of the signal:
|
|
//!
|
|
//! ```text
|
|
//! sigma(y) = k * sqrt(max(y, 0) + NOISE_FLOOR)
|
|
//! ```
|
|
//!
|
|
//! which is the photon-noise law — the arrival of light is Poisson, so its
|
|
//! variance equals its mean and its standard deviation goes as the square
|
|
//! root. `NOISE_FLOOR` stands in for the sensor's read noise, which does not
|
|
//! vanish at black, and keeps `sigma` finite there instead of collapsing to
|
|
//! zero and switching the filter off exactly where noise is worst.
|
|
//!
|
|
//! **The honest limitation.** By the time this stage runs, exposure, the tone
|
|
//! curve and the recovery controls have already moved these values, so they
|
|
//! are no longer proportional to photon counts and the law is an
|
|
//! approximation rather than a measurement. It is kept because it is a far
|
|
//! better approximation than a constant — the tone mapping is monotone and
|
|
//! only gently compressive, so the ordering and the rough scaling survive it
|
|
//! — and because the thing that *would* be exact is FR-DEV-3g's learned
|
|
//! denoiser, operating in the raw domain where the noise model still holds.
|
|
//! This is the conventional path that degrades to when no model is present,
|
|
//! and it is not trying to be it.
|
|
//!
|
|
//! # Radius units
|
|
//!
|
|
//! Both radii are stated in **source pixels** and converted through
|
|
//! [`RenderScale::source_pixels`] at every render. Noise is a property of the
|
|
//! sensor and of the demosaic: its grain is about one photosite across
|
|
//! because photosites are what recorded it, and that stays true regardless of
|
|
//! how large the frame is drawn on screen or how many megapixels the body
|
|
//! has. [`RenderScale::frame_fraction`], the other unit, would say the
|
|
//! opposite — that grain covers a fixed proportion of the *picture* — so the
|
|
//! same body's files would need different settings as their pixel count
|
|
//! changed, and a crop would need different settings from the frame it came
|
|
//! out of.
|
|
//!
|
|
//! The consequence is the one [`crate::detail`] documents: on a heavy proxy a
|
|
//! luminance radius of one source pixel is a fraction of a render pixel, the
|
|
//! information it would act on was thrown away by the downscale, and this
|
|
//! operation emits no luminance pass at all rather than drawing a plausible
|
|
//! lie. The chroma radius, ten times larger, still resolves — which is also
|
|
//! true of the fault it treats, since a blotch twenty pixels across survives
|
|
//! being halved.
|
|
use std::sync::{Arc, LazyLock};
|
|
|
|
use crate::descriptor::{
|
|
Attribute, LocalizedKey, OpDescriptor, OpId, ParamDescriptor, ParamId, Scale, Unit,
|
|
};
|
|
use crate::detail::{DetailPass, DetailStage, RenderScale};
|
|
use crate::operation::{Affects, Helper, Operation, Uniform};
|
|
|
|
pub const ID: OpId = OpId("noise_reduction");
|
|
pub const LUMINANCE: ParamId = ParamId("luminance");
|
|
pub const CHROMA: ParamId = ParamId("chroma");
|
|
|
|
static DESCRIPTOR: LazyLock<Arc<OpDescriptor>> = LazyLock::new(|| {
|
|
Arc::new(OpDescriptor {
|
|
id: ID,
|
|
label: LocalizedKey("op.noise_reduction"),
|
|
attributes: vec![Attribute::Detail],
|
|
// Zero to a hundred rather than the symmetric `amount` shape the tonal
|
|
// controls use. There is no meaningful negative: "minus fifty noise
|
|
// reduction" would be adding grain, which is a look rather than a repair
|
|
// and belongs to a different operation carrying `Attribute::Effect`. A
|
|
// control whose left half does nothing is worse than one that stops.
|
|
params: vec![
|
|
ParamDescriptor::scalar(
|
|
"luminance",
|
|
"param.noise_reduction.luminance",
|
|
0.0,
|
|
100.0,
|
|
0.0,
|
|
Unit::None,
|
|
Scale::Linear,
|
|
0,
|
|
),
|
|
ParamDescriptor::scalar(
|
|
"chroma",
|
|
"param.noise_reduction.chroma",
|
|
0.0,
|
|
100.0,
|
|
0.0,
|
|
Unit::None,
|
|
Scale::Linear,
|
|
0,
|
|
),
|
|
],
|
|
})
|
|
});
|
|
|
|
/// The luminance radius at the lowest and the highest amount, in **source**
|
|
/// pixels.
|
|
///
|
|
/// It starts at one rather than at zero because a kernel smaller than a pixel
|
|
/// is not a kernel; the amount fades the *threshold* in from zero instead, so
|
|
/// the control is still continuous at its neutral. It stops at 2.5 because
|
|
/// past roughly three source pixels a luminance average stops removing grain
|
|
/// and starts removing the subject — the point at which every editor's
|
|
/// luminance slider gets described as watercolour.
|
|
const LUMA_RADIUS: (f32, f32) = (1.0, 2.5);
|
|
|
|
/// The chroma radius at the lowest and the highest amount, in **source**
|
|
/// pixels.
|
|
///
|
|
/// An order of magnitude larger, because the fault is an order of magnitude
|
|
/// larger: demosaic-born colour blotches are tens of pixels across and a
|
|
/// two-pixel average does not see them.
|
|
const CHROMA_RADIUS: (f32, f32) = (2.0, 12.0);
|
|
|
|
/// Hard ceilings on the kernel actually dispatched, in **render** pixels.
|
|
///
|
|
/// Necessary because [`RenderScale::ratio`] exceeds one when the view is
|
|
/// zoomed past 1:1 — the render target keeps its size while the region it
|
|
/// covers shrinks — so a radius in source pixels can ask for an arbitrarily
|
|
/// large kernel at high magnification. Without a cap, zooming to 800% would
|
|
/// quietly turn a twelve-pixel chroma radius into a ninety-six-pixel one and
|
|
/// cost sixty-four times the taps, at exactly the moment the user is
|
|
/// inspecting the result closely and most wants the view to stay responsive.
|
|
/// Clamping instead means the effect stops growing past the point where more
|
|
/// of it would be visible anyway.
|
|
const LUMA_KERNEL_CAP: u32 = 3;
|
|
const CHROMA_KERNEL_CAP: u32 = 16;
|
|
|
|
/// The luminance range threshold at full amount, as a coefficient on
|
|
/// `sqrt(signal)`.
|
|
///
|
|
/// At middle grey this is `0.075 * sqrt(0.18) ≈ 0.032`, about three percent
|
|
/// of full scale — roughly eight 8-bit code values, which is the grain of a
|
|
/// high-ISO frame. Much larger and it would start treating real texture as
|
|
/// noise.
|
|
const LUMA_SIGMA: f32 = 0.075;
|
|
|
|
/// The chroma range threshold at full amount, on the length of the colour
|
|
/// difference vector.
|
|
///
|
|
/// Far larger than the luminance threshold because it is allowed to be: a
|
|
/// genuine colour boundary separates colours by much more than this — a
|
|
/// saturated red sits about 0.84 from grey — while chroma noise is a few
|
|
/// hundredths. That gap is the whole reason chroma can be filtered hard
|
|
/// without visible bleeding.
|
|
const CHROMA_SIGMA: f32 = 0.20;
|
|
|
|
/// How tightly the chroma passes are steered by luminance, at the lowest and
|
|
/// the highest amount.
|
|
///
|
|
/// The chroma filter weights a neighbour by *both* how far its colour is and
|
|
/// how far its lightness is. The lightness term is a cross-bilateral guide in
|
|
/// the ordinary sense — the cleaner channel steering the noisier one — and it
|
|
/// is what stops colour crossing a boundary the colour channel itself cannot
|
|
/// see: a dark object against a light background of the same hue.
|
|
///
|
|
/// It does not start at zero. A guide with a zero threshold rejects every
|
|
/// neighbour, and the filter would do nothing however far the colour
|
|
/// threshold was opened. It widens with the amount because a photographer
|
|
/// asking for more chroma denoising has a noisier frame, whose luminance —
|
|
/// the guide itself — is also noisier and would otherwise break the weights.
|
|
const CHROMA_GUIDE_SIGMA: (f32, f32) = (0.06, 0.16);
|
|
|
|
/// The read-noise floor, in linear working units.
|
|
///
|
|
/// Keeps `sigma` finite at black. Roughly 1/400 of full scale, about where a
|
|
/// deep shadow sits after a normal rendering: small enough not to affect a
|
|
/// midtone, large enough that the filter does not switch itself off in the
|
|
/// shadows.
|
|
const NOISE_FLOOR: f32 = 0.0025;
|
|
|
|
/// TRACES: FR-DEV-3
|
|
/// Luminance and chroma noise reduction, as two independent amounts.
|
|
#[derive(Debug, Clone, Copy, Default)]
|
|
pub struct NoiseReduction {
|
|
luminance: f32,
|
|
chroma: f32,
|
|
}
|
|
|
|
impl NoiseReduction {
|
|
pub fn new() -> Self {
|
|
Self::default()
|
|
}
|
|
|
|
/// Both amounts at once, for tests and for a preset applying the pair.
|
|
pub fn with_amounts(luminance: f32, chroma: f32) -> Self {
|
|
Self { luminance, chroma }
|
|
}
|
|
|
|
/// The luminance radius in **source** pixels, or `None` at neutral.
|
|
pub fn luminance_radius(&self) -> Option<f32> {
|
|
(self.luminance > 0.0).then(|| lerp(LUMA_RADIUS, self.luminance / 100.0))
|
|
}
|
|
|
|
/// The chroma radius in **source** pixels, or `None` at neutral.
|
|
pub fn chroma_radius(&self) -> Option<f32> {
|
|
(self.chroma > 0.0).then(|| lerp(CHROMA_RADIUS, self.chroma / 100.0))
|
|
}
|
|
|
|
/// The luminance kernel this render would dispatch, in render pixels.
|
|
///
|
|
/// Zero means "not at this resolution": either the control is neutral, or
|
|
/// the radius is smaller than a render pixel and the detail it would act
|
|
/// on is not present in this render at all (see [`RenderScale::resolves`]).
|
|
///
|
|
/// Exposed so a test — and, in time, an interface offering to zoom to 1:1
|
|
/// — can state the expected kernel without repeating the rounding rule,
|
|
/// which is how a test comes to agree with a bug.
|
|
pub fn luminance_kernel(&self, scale: RenderScale) -> u32 {
|
|
kernel(self.luminance_radius(), scale, LUMA_KERNEL_CAP)
|
|
}
|
|
|
|
/// The chroma kernel this render would dispatch, in render pixels.
|
|
pub fn chroma_kernel(&self, scale: RenderScale) -> u32 {
|
|
kernel(self.chroma_radius(), scale, CHROMA_KERNEL_CAP)
|
|
}
|
|
|
|
/// The luminance range threshold coefficient — the `k` in
|
|
/// `sigma = k * sqrt(signal + floor)`.
|
|
fn luma_sigma(&self) -> f32 {
|
|
LUMA_SIGMA * (self.luminance / 100.0)
|
|
}
|
|
|
|
/// The chroma passes' two thresholds: the luminance guide, then the
|
|
/// colour difference.
|
|
fn chroma_sigmas(&self) -> (f32, f32) {
|
|
let t = self.chroma / 100.0;
|
|
(lerp(CHROMA_GUIDE_SIGMA, t), CHROMA_SIGMA * t)
|
|
}
|
|
}
|
|
|
|
fn lerp((low, high): (f32, f32), t: f32) -> f32 {
|
|
low + (high - low) * t.clamp(0.0, 1.0)
|
|
}
|
|
|
|
/// A radius in source pixels, as the kernel to walk in render pixels.
|
|
///
|
|
/// Zero when [`RenderScale::resolves`] says the radius does not survive this
|
|
/// render. That check rather than the rounding, because the two disagree
|
|
/// exactly where it matters: 0.6 of a render pixel *rounds* to one, and a
|
|
/// one-pixel kernel would then be dispatched to remove grain that the
|
|
/// downscale averaged away before this stage ran. It would cost a dispatch to
|
|
/// draw something that is not in the picture, and — worse — it would look
|
|
/// like an effect, so a photographer would tune against it.
|
|
///
|
|
/// Otherwise rounded rather than truncated, so a 1.4-pixel radius is one pixel
|
|
/// and a 1.6-pixel radius is two; and capped, so that zooming past 1:1 cannot
|
|
/// make the cost of a frame grow without bound.
|
|
fn kernel(radius: Option<f32>, scale: RenderScale, cap: u32) -> u32 {
|
|
let Some(radius) = radius else { return 0 };
|
|
if !scale.resolves(radius) {
|
|
return 0;
|
|
}
|
|
(scale.source_pixels(radius).round().max(1.0) as u32).min(cap)
|
|
}
|
|
|
|
/// The Gaussian spatial falloff for a kernel of this radius, as the
|
|
/// `1 / (2 * sigma^2)` the shader multiplies a squared distance by.
|
|
///
|
|
/// `sigma` is half the radius, which puts the weight at the rim of the kernel
|
|
/// at `exp(-2)`, about 0.135 — small enough that the kernel has no visible
|
|
/// hard edge, large enough that the outermost taps are still doing work
|
|
/// rather than being paid for and discarded.
|
|
fn inv_spatial(kernel: u32) -> f32 {
|
|
let sigma = (kernel as f32 * 0.5).max(0.5);
|
|
1.0 / (2.0 * sigma * sigma)
|
|
}
|
|
|
|
impl Operation for NoiseReduction {
|
|
fn descriptor(&self) -> Arc<OpDescriptor> {
|
|
DESCRIPTOR.clone()
|
|
}
|
|
|
|
fn set_param(&mut self, id: ParamId, value: f32) {
|
|
match id {
|
|
LUMINANCE => self.luminance = value,
|
|
CHROMA => self.chroma = value,
|
|
_ => log::warn!("noise_reduction: unknown parameter {id}"),
|
|
}
|
|
}
|
|
|
|
fn param(&self, id: ParamId) -> f32 {
|
|
match id {
|
|
LUMINANCE => self.luminance,
|
|
CHROMA => self.chroma,
|
|
_ => 0.0,
|
|
}
|
|
}
|
|
|
|
fn is_active(&self) -> bool {
|
|
self.luminance > 0.0 || self.chroma > 0.0
|
|
}
|
|
|
|
/// Never called: a detail operation contributes no fused fragment, and
|
|
/// `compose_full` filters it out before asking.
|
|
fn wgsl_body(&self) -> String {
|
|
String::new()
|
|
}
|
|
|
|
fn uniforms(&self) -> Vec<Uniform> {
|
|
Vec::new()
|
|
}
|
|
|
|
fn affects(&self) -> Affects {
|
|
Affects::Detail
|
|
}
|
|
|
|
fn detail(&self) -> Option<&dyn DetailStage> {
|
|
Some(self)
|
|
}
|
|
|
|
/// The shared `luminance` helper, which both kernels call.
|
|
///
|
|
/// Taken from `ops/_helpers.yaml` rather than defined here, so that
|
|
/// lightness means one thing across the whole pipeline. Its own
|
|
/// documentation calls the Rec. 709 weights an approximation, which they
|
|
/// are in camera space — but the detail stage runs after the camera
|
|
/// matrix, in linear sRGB, where they are exactly the right weights.
|
|
fn helpers(&self) -> &'static [Helper] {
|
|
HELPERS
|
|
}
|
|
}
|
|
|
|
static HELPERS: &[Helper] = &[crate::ops::helpers::LUMINANCE];
|
|
|
|
impl DetailStage for NoiseReduction {
|
|
fn passes(&self, scale: RenderScale) -> Vec<DetailPass> {
|
|
let mut passes = Vec::new();
|
|
|
|
// Luminance first, and the order is not arbitrary: the chroma passes
|
|
// are steered by luminance, and a luminance that has already been
|
|
// denoised is a cleaner guide than a noisy one. Doing it the other way
|
|
// round would make the chroma weights noisier for no gain anywhere.
|
|
// When the luminance control is neutral the guide is simply the
|
|
// luminance as it arrived, which is the honest fallback.
|
|
let luma = self.luminance_kernel(scale);
|
|
if luma > 0 {
|
|
passes.push(DetailPass {
|
|
output_scale: 1,
|
|
label: "luminance",
|
|
radius: luma,
|
|
// A convolution, not a list: nothing to bind at binding 3.
|
|
storage: Vec::new(),
|
|
uniforms: vec![
|
|
Uniform {
|
|
name: "radius",
|
|
value: luma as f32,
|
|
},
|
|
Uniform {
|
|
name: "inv_spatial",
|
|
value: inv_spatial(luma),
|
|
},
|
|
Uniform {
|
|
name: "sigma_k",
|
|
value: self.luma_sigma(),
|
|
},
|
|
Uniform {
|
|
name: "noise_floor",
|
|
value: NOISE_FLOOR,
|
|
},
|
|
],
|
|
wgsl: LUMA_WGSL.to_string(),
|
|
});
|
|
}
|
|
|
|
let chroma = self.chroma_kernel(scale);
|
|
if chroma > 0 {
|
|
let (guide, colour) = self.chroma_sigmas();
|
|
// One body, dispatched twice with the step vector rotated. Writing
|
|
// it as two passes over one kernel rather than as two kernels is
|
|
// what keeps the two halves of a separable filter from drifting
|
|
// apart — the classic way an axis ends up filtered differently
|
|
// from the other and the result acquires a diagonal bias.
|
|
for (index, (sx, sy)) in [(1.0, 0.0), (0.0, 1.0)].into_iter().enumerate() {
|
|
passes.push(DetailPass {
|
|
output_scale: 1,
|
|
label: if index == 0 {
|
|
"chroma-horizontal"
|
|
} else {
|
|
"chroma-vertical"
|
|
},
|
|
radius: chroma,
|
|
// A convolution, not a list: nothing to bind at binding 3.
|
|
storage: Vec::new(),
|
|
uniforms: vec![
|
|
Uniform {
|
|
name: "radius",
|
|
value: chroma as f32,
|
|
},
|
|
Uniform {
|
|
name: "step_x",
|
|
value: sx,
|
|
},
|
|
Uniform {
|
|
name: "step_y",
|
|
value: sy,
|
|
},
|
|
Uniform {
|
|
name: "inv_spatial",
|
|
value: inv_spatial(chroma),
|
|
},
|
|
Uniform {
|
|
name: "guide_k",
|
|
value: guide,
|
|
},
|
|
Uniform {
|
|
name: "chroma_k",
|
|
value: colour,
|
|
},
|
|
Uniform {
|
|
name: "noise_floor",
|
|
value: NOISE_FLOOR,
|
|
},
|
|
],
|
|
wgsl: CHROMA_WGSL.to_string(),
|
|
});
|
|
}
|
|
}
|
|
|
|
passes
|
|
}
|
|
}
|
|
|
|
/// The exact two-dimensional bilateral, acting on luminance alone.
|
|
const LUMA_WGSL: &str = "\
|
|
// A bilateral filter over luminance: the spatial Gaussian every blur has,
|
|
// multiplied by a range term that falls off with how different the
|
|
// neighbour's lightness is. That second factor is the entire difference
|
|
// between denoising and smearing — a neighbour on the far side of an edge
|
|
// contributes essentially nothing, so the edge survives the average that
|
|
// removes the grain either side of it.
|
|
//
|
|
// Exact rather than separable. At the one-to-three-pixel radii a luminance
|
|
// kernel is allowed, the two-pass approximation saves a handful of taps and
|
|
// costs a whole extra full-frame write and read, and it can leave axis-aligned
|
|
// streaking in the one channel the eye reads sharpest.
|
|
let r = i32(radius);
|
|
let y0 = luminance(c);
|
|
|
|
// Photon noise: the standard deviation of a signal goes as its square root, so
|
|
// the threshold has to as well. A constant would be a hundred times too coarse
|
|
// in the shadows relative to the highlights and would flatten them.
|
|
// `noise_floor` stands for read noise and keeps this finite at black.
|
|
let sigma = max(sigma_k * sqrt(max(y0, 0.0) + noise_floor), 1e-5);
|
|
let inv_range = 1.0 / (2.0 * sigma * sigma);
|
|
|
|
var weight_sum = 0.0;
|
|
var luma_sum = 0.0;
|
|
for (var dy = -r; dy <= r; dy = dy + 1) {
|
|
for (var dx = -r; dx <= r; dx = dx + 1) {
|
|
let n = luminance(tap(coord, vec2<i32>(dx, dy)));
|
|
let dl = n - y0;
|
|
let distance2 = f32(dx * dx + dy * dy);
|
|
let w = exp(-(distance2 * inv_spatial + dl * dl * inv_range));
|
|
weight_sum = weight_sum + w;
|
|
luma_sum = luma_sum + n * w;
|
|
}
|
|
}
|
|
|
|
// Substitute the filtered luminance and leave the colour difference exactly as
|
|
// it arrived. `c - vec3(y0)` has zero luminance by construction, so adding the
|
|
// change in lightness back changes lightness and nothing else — which is what
|
|
// keeps this control independent of the chroma one.
|
|
c = c + vec3<f32>(luma_sum / weight_sum - y0);";
|
|
|
|
/// One axis of the separable cross-bilateral, acting on chroma alone.
|
|
const CHROMA_WGSL: &str = "\
|
|
// One axis of a separable bilateral over the colour difference.
|
|
//
|
|
// Separable because the radius is an order of magnitude larger than the
|
|
// luminance one and the exact form is quadratic: at twelve source pixels that
|
|
// is 625 taps a pixel, which is not a frame path. Two one-dimensional passes
|
|
// are fifty, and the axis-aligned residual the approximation leaves is in
|
|
// chroma, where the eye resolves about a quarter of what it resolves in
|
|
// lightness.
|
|
//
|
|
// The weight has two range terms, not one. The colour term is what the filter
|
|
// is for. The luminance term is a *guide*: it stops colour crossing a boundary
|
|
// the colour channel itself cannot see — a dark object against a light
|
|
// background of the same hue — by letting the cleaner channel steer the
|
|
// noisier one.
|
|
let r = i32(radius);
|
|
let step = vec2<i32>(i32(step_x), i32(step_y));
|
|
|
|
let y0 = luminance(c);
|
|
let d0 = c - vec3<f32>(y0);
|
|
|
|
// Both thresholds follow the same square-root-of-signal law, for the reason
|
|
// the luminance pass states.
|
|
let level = sqrt(max(y0, 0.0) + noise_floor);
|
|
let guide = max(guide_k * level, 1e-5);
|
|
let colour = max(chroma_k * level, 1e-5);
|
|
let inv_guide = 1.0 / (2.0 * guide * guide);
|
|
let inv_colour = 1.0 / (2.0 * colour * colour);
|
|
|
|
var weight_sum = 0.0;
|
|
var chroma_sum = vec3<f32>(0.0);
|
|
for (var i = -r; i <= r; i = i + 1) {
|
|
let n = tap(coord, step * i);
|
|
let yn = luminance(n);
|
|
let dn = n - vec3<f32>(yn);
|
|
let dl = yn - y0;
|
|
let dc = dn - d0;
|
|
let w = exp(-(f32(i * i) * inv_spatial + dl * dl * inv_guide + dot(dc, dc) * inv_colour));
|
|
weight_sum = weight_sum + w;
|
|
chroma_sum = chroma_sum + dn * w;
|
|
}
|
|
|
|
// This pixel's own luminance, unchanged, plus the filtered colour difference.
|
|
// Every `dn` has zero luminance, so their weighted mean does too and the
|
|
// reconstructed colour keeps exactly the lightness it arrived with.
|
|
c = vec3<f32>(y0) + chroma_sum / weight_sum;";
|
|
|
|
#[cfg(test)]
|
|
mod tests {
|
|
use super::*;
|
|
use dr_types::ColourSpace;
|
|
|
|
/// A 24 MP frame, and the panel a develop view might show it in.
|
|
const FULL: (u32, u32) = (6000, 4000);
|
|
|
|
fn nr(luminance: f32, chroma: f32) -> NoiseReduction {
|
|
NoiseReduction::with_amounts(luminance, chroma)
|
|
}
|
|
|
|
fn chain_with(op: NoiseReduction) -> Vec<Box<dyn Operation>> {
|
|
vec![Box::new(op)]
|
|
}
|
|
|
|
fn compose(op: NoiseReduction, scale: RenderScale) -> crate::detail::ComposedDetail {
|
|
crate::detail::compose_detail(&chain_with(op), scale, ColourSpace::Srgb)
|
|
}
|
|
|
|
#[test]
|
|
fn neutral_costs_the_edit_nothing() {
|
|
// The rule the whole pipeline rests on. An unedited photograph must
|
|
// not pay for a denoiser it is not using — no pass, no dispatch, and
|
|
// the fused shader ends exactly as it always did.
|
|
let op = nr(0.0, 0.0);
|
|
assert!(!op.is_active());
|
|
assert!(compose(op, RenderScale::full(FULL)).is_empty());
|
|
}
|
|
|
|
#[test]
|
|
fn each_amount_reaches_the_shader_on_its_own() {
|
|
// The point of two controls: either alone must produce its own passes
|
|
// and nothing of the other's. A denoiser that emitted the chroma
|
|
// dispatches whenever luminance was on would cost two thirds of the
|
|
// stage for an effect the user did not ask for — and would be
|
|
// invisible in the picture, because at a zero threshold the chroma
|
|
// filter is very nearly the identity.
|
|
let scale = RenderScale::full(FULL);
|
|
|
|
let labels = |op| {
|
|
compose(op, scale)
|
|
.passes
|
|
.iter()
|
|
.map(|p| p.label.clone())
|
|
.collect::<Vec<String>>()
|
|
};
|
|
|
|
assert_eq!(labels(nr(50.0, 0.0)), ["noise_reduction/luminance"]);
|
|
assert_eq!(
|
|
labels(nr(0.0, 50.0)),
|
|
[
|
|
"noise_reduction/chroma-horizontal",
|
|
"noise_reduction/chroma-vertical"
|
|
]
|
|
);
|
|
|
|
let both = compose(nr(50.0, 50.0), scale);
|
|
assert_eq!(both.len(), 3);
|
|
// The luminance pass runs first, so the chroma guide is the denoised
|
|
// luminance rather than the raw one.
|
|
assert_eq!(both.passes[0].label, "noise_reduction/luminance");
|
|
// And only the last pass in the whole chain performs the output
|
|
// transform, whichever pass that happens to be.
|
|
assert!(!both.passes[0].writes_output);
|
|
assert!(!both.passes[1].writes_output);
|
|
assert!(both.passes[2].writes_output);
|
|
}
|
|
|
|
#[test]
|
|
fn chroma_always_reaches_further_than_luminance() {
|
|
// The reason these are two controls rather than one. Colour blotches
|
|
// are tens of pixels across and grain is one or two, so no single
|
|
// radius treats both — and if this ever inverted, the chroma slider
|
|
// would have become an expensive second luminance slider.
|
|
for amount in [1.0, 25.0, 50.0, 75.0, 100.0] {
|
|
let op = nr(amount, amount);
|
|
let l = op.luminance_radius().expect("active");
|
|
let c = op.chroma_radius().expect("active");
|
|
assert!(c > l * 2.0, "at {amount}: chroma {c} vs luminance {l}");
|
|
}
|
|
}
|
|
|
|
#[test]
|
|
fn a_radius_is_a_count_of_sensor_pixels_not_a_fraction_of_the_frame() {
|
|
// TRACES: FR-DSP-1 — the decision this operation is most likely to
|
|
// get wrong, stated as the property that distinguishes the two units.
|
|
//
|
|
// Noise is made by photosites, so its grain is the same size in
|
|
// *source* pixels however the frame is being rendered. Convert with
|
|
// `source_pixels` and the kernel in render pixels tracks the scale;
|
|
// convert with `frame_fraction` and it would instead be constant for a
|
|
// constant render size, which is a different — and wrong — claim.
|
|
let op = nr(100.0, 100.0);
|
|
let radius = op.chroma_radius().expect("active");
|
|
assert!((radius - 12.0).abs() < 1e-6);
|
|
|
|
// The same photograph at three sizes. Measured back in source pixels,
|
|
// the kernel is the same length every time.
|
|
for render in [(1500u32, 1000u32), (3000, 2000), (6000, 4000)] {
|
|
let scale = RenderScale::new(render, FULL);
|
|
let in_source_pixels = op.chroma_kernel(scale) as f32 / scale.ratio();
|
|
assert!(
|
|
(in_source_pixels - radius).abs() < 0.5,
|
|
"{render:?} denoised {in_source_pixels} source pixels, not {radius}"
|
|
);
|
|
}
|
|
|
|
// And the distinguishing case: one panel, two cameras. A 24 MP file
|
|
// and a 96 MP file shown at the same size have the same *frame
|
|
// fraction* per render pixel but four times the photosites, so the
|
|
// sensor-pixel radius covers a quarter as much of the picture in the
|
|
// second. A frame-fraction radius would have given both the same
|
|
// kernel, which would mean the 96 MP body needed a different setting
|
|
// to remove the same grain.
|
|
let render = (1500, 1000);
|
|
let small = op.chroma_kernel(RenderScale::new(render, (6000, 4000)));
|
|
let large = op.chroma_kernel(RenderScale::new(render, (12000, 8000)));
|
|
assert!(
|
|
small > large,
|
|
"denser sensor, same panel: {small} vs {large} render pixels"
|
|
);
|
|
}
|
|
|
|
#[test]
|
|
fn a_luminance_radius_too_small_to_draw_is_not_drawn() {
|
|
// The honest limit `RenderScale` exists to report. At a quarter-size
|
|
// proxy a 2.5-source-pixel luminance radius is 0.6 render pixels: the
|
|
// grain it would remove was averaged away by the downscale before this
|
|
// stage ran, and there is no kernel that represents a fraction of a
|
|
// pixel. Emitting a pass anyway would burn a dispatch to draw a guess.
|
|
let proxy = RenderScale::new((1500, 1000), FULL);
|
|
let op = nr(100.0, 100.0);
|
|
assert_eq!(op.luminance_kernel(proxy), 0);
|
|
assert!(!proxy.resolves(op.luminance_radius().expect("active")));
|
|
|
|
// Chroma is a different case at the same scale, and the difference is
|
|
// real rather than a rounding accident: a blotch twenty pixels across
|
|
// is still ten pixels across in a half-size proxy, so it both survives
|
|
// the downscale and can still be removed.
|
|
assert_eq!(op.chroma_kernel(proxy), 3);
|
|
let composed = compose(op, proxy);
|
|
assert_eq!(composed.len(), 2, "chroma alone survives a heavy proxy");
|
|
}
|
|
|
|
#[test]
|
|
fn zooming_past_one_to_one_does_not_let_the_kernel_run_away() {
|
|
// A 1:1 view already renders one render pixel per source pixel; at
|
|
// 800% there are eight. Without the cap the chroma kernel would be
|
|
// ninety-six render pixels — sixty-four times the taps — precisely
|
|
// when the user is looking closely and least tolerant of a stall.
|
|
let magnified = RenderScale::new((2000, 2000), (250, 250));
|
|
assert!((magnified.ratio() - 8.0).abs() < 1e-6);
|
|
let op = nr(100.0, 100.0);
|
|
assert_eq!(op.chroma_kernel(magnified), CHROMA_KERNEL_CAP);
|
|
assert_eq!(op.luminance_kernel(magnified), LUMA_KERNEL_CAP);
|
|
}
|
|
|
|
#[test]
|
|
fn the_declared_halo_is_the_kernel_the_shader_walks() {
|
|
// ARCH §5.3 grows a tile by the declared radius before scheduling it.
|
|
// An understated radius shows as a seam at every tile boundary, which
|
|
// looks like a driver bug rather than like an arithmetic error — so
|
|
// the number handed to the scheduler and the number the loop counts to
|
|
// must be the same number, not two that happen to agree today.
|
|
let scale = RenderScale::full(FULL);
|
|
let op = nr(100.0, 100.0);
|
|
for pass in compose(op, scale).passes {
|
|
let declared = pass.radius as f32;
|
|
let walked = pass.uniforms[crate::detail::DETAIL_BASE_UNIFORM_FIELDS];
|
|
assert_eq!(declared, walked, "{}", pass.label);
|
|
}
|
|
assert_eq!(compose(op, scale).radius(), op.chroma_kernel(scale));
|
|
}
|
|
|
|
#[test]
|
|
fn every_pass_declares_a_uniform_block_the_gpu_will_accept() {
|
|
// A uniform struct whose size is not a multiple of sixteen is rejected
|
|
// outright by the WGSL uniform address space rules, and the failure
|
|
// arrives as a compile error against generated source a long way from
|
|
// here.
|
|
for pass in compose(nr(60.0, 60.0), RenderScale::full(FULL)).passes {
|
|
assert_eq!(pass.uniforms.len() % 4, 0, "{}", pass.label);
|
|
assert!(
|
|
pass.uniforms.iter().all(|v| v.is_finite()),
|
|
"{} uploaded a non-finite uniform",
|
|
pass.label
|
|
);
|
|
}
|
|
}
|
|
|
|
#[test]
|
|
fn the_two_chroma_passes_are_one_kernel_along_two_axes() {
|
|
// A separable filter is only separable if both halves are the same
|
|
// filter. The bodies must be identical and the step vectors must be
|
|
// perpendicular unit steps; anything else is two different blurs whose
|
|
// composition is not the two-dimensional one intended.
|
|
let passes = nr(0.0, 80.0).passes(RenderScale::full(FULL));
|
|
assert_eq!(passes.len(), 2);
|
|
assert_eq!(passes[0].wgsl, passes[1].wgsl);
|
|
|
|
let step = |p: &DetailPass| {
|
|
let get = |name| {
|
|
p.uniforms
|
|
.iter()
|
|
.find(|u| u.name == name)
|
|
.expect("declared")
|
|
.value
|
|
};
|
|
(get("step_x"), get("step_y"))
|
|
};
|
|
assert_eq!(step(&passes[0]), (1.0, 0.0));
|
|
assert_eq!(step(&passes[1]), (0.0, 1.0));
|
|
|
|
// Everything else about the two must match, or one axis is filtered
|
|
// harder than the other and a round blotch comes out oval.
|
|
assert_eq!(passes[0].radius, passes[1].radius);
|
|
for name in [
|
|
"radius",
|
|
"inv_spatial",
|
|
"guide_k",
|
|
"chroma_k",
|
|
"noise_floor",
|
|
] {
|
|
let of = |p: &DetailPass| {
|
|
p.uniforms
|
|
.iter()
|
|
.find(|u| u.name == name)
|
|
.expect("declared")
|
|
.value
|
|
};
|
|
assert_eq!(of(&passes[0]), of(&passes[1]), "{name}");
|
|
}
|
|
}
|
|
|
|
#[test]
|
|
fn the_threshold_opens_with_the_amount_and_closes_at_neutral() {
|
|
// The amount is a *threshold* as much as a radius: it decides how
|
|
// large a difference the filter is willing to call noise. If it did
|
|
// not reach zero at the neutral end the control would be
|
|
// discontinuous, and the first pixel of travel on the slider would
|
|
// visibly flatten the image.
|
|
let mut previous = 0.0;
|
|
for amount in [1.0, 10.0, 50.0, 100.0] {
|
|
let sigma = nr(amount, 0.0).luma_sigma();
|
|
assert!(sigma > previous, "at {amount}: {sigma} <= {previous}");
|
|
previous = sigma;
|
|
}
|
|
assert_eq!(nr(0.0, 0.0).luma_sigma(), 0.0);
|
|
|
|
// The chroma guide is the exception, and deliberately so: a guide with
|
|
// a zero threshold rejects every neighbour, so the filter would do
|
|
// nothing at all at low amounts however wide the colour threshold was.
|
|
let (guide, colour) = nr(0.0, 1.0).chroma_sigmas();
|
|
assert!(guide > 0.0, "a zero guide would reject every neighbour");
|
|
assert!(colour > 0.0);
|
|
}
|
|
|
|
#[test]
|
|
fn a_chroma_threshold_is_far_wider_than_a_luminance_one() {
|
|
// Not a tuning detail but the reason the two are separable problems.
|
|
// Chroma noise is a few hundredths from grey and a real colour
|
|
// boundary is most of the way to a primary, so the gap between them is
|
|
// wide enough to filter hard through. Luminance has no such gap, which
|
|
// is why its threshold has to stay tight.
|
|
let (_, colour) = nr(100.0, 100.0).chroma_sigmas();
|
|
assert!(colour > nr(100.0, 100.0).luma_sigma() * 2.0);
|
|
}
|
|
|
|
#[test]
|
|
fn parameters_round_trip_and_an_unknown_one_is_ignored() {
|
|
// What the sidecar, the history and the preset system all rely on.
|
|
let mut op = NoiseReduction::new();
|
|
op.set_param(LUMINANCE, 40.0);
|
|
op.set_param(CHROMA, 70.0);
|
|
assert_eq!(op.param(LUMINANCE), 40.0);
|
|
assert_eq!(op.param(CHROMA), 70.0);
|
|
op.set_param(ParamId("sharpness"), 99.0);
|
|
assert_eq!(op.param(LUMINANCE), 40.0);
|
|
assert_eq!(op.param(ParamId("sharpness")), 0.0);
|
|
}
|
|
|
|
#[test]
|
|
fn the_generated_wgsl_addresses_its_own_uniforms() {
|
|
// The composer prefixes each uniform with the operation id and the
|
|
// pass index, so two operations may both call a uniform `radius` and
|
|
// neither has to know. A body that slipped through unrewritten would
|
|
// fail to compile against the generated struct.
|
|
let composed = compose(nr(50.0, 50.0), RenderScale::full(FULL));
|
|
assert!(composed.passes[0]
|
|
.source
|
|
.contains("noise_reduction_0_sigma_k: f32,"));
|
|
assert!(composed.passes[1]
|
|
.source
|
|
.contains("noise_reduction_1_chroma_k: f32,"));
|
|
assert!(composed.passes[2]
|
|
.source
|
|
.contains("noise_reduction_2_chroma_k: f32,"));
|
|
// The shared luminance helper reaches every pass that calls it.
|
|
for pass in &composed.passes {
|
|
assert!(
|
|
pass.source.contains("fn luminance(c: vec3<f32>)"),
|
|
"{} calls luminance without defining it",
|
|
pass.label
|
|
);
|
|
}
|
|
// Three passes of one operation are three shaders, and must not share
|
|
// a pipeline-cache entry.
|
|
let hashes: std::collections::BTreeSet<u64> =
|
|
composed.passes.iter().map(|p| p.structure_hash).collect();
|
|
assert_eq!(hashes.len(), 3);
|
|
}
|
|
}
|