Let clarity's base be computed where it is still fully determined
Clarity's Gaussian sigma is 1.2% of the frame's shorter edge, so its radius is a property of the viewport: 52 render pixels at 4K, two separable passes of 105 taps each over 8.3 M pixels. That measured 33.9 ms — seven times the entire fused point chain, for one slider — and is docs/technical-debt.md TD-4. A detail pass may now declare `output_scale`, and clarity's base is computed on a grid a quarter the size on each axis. The pass that combines needs the blur *and* the full-resolution colour, and a colour that has been through a quarter-scale target is no longer full resolution. So a scaled pass cannot simply join the ping-pong: there are two chains now. The full-resolution one carries the colour and no scaled pass touches it; the reduced one carries the base and reaches the combining pass through a second binding as `reduced_at()`. The reduce is a dispatch of its own rather than something the first blur half does on the way past, and that is the whole difference between this and the strided kernel the module documentation rules out. A stride samples an image that is not band-limited and aliases high-frequency content down into the base, which is then subtracted, and arrives in the output as mottling across smooth gradients. This band-limits first and samples after. What is discarded is content the base could not represent at any resolution, because a Gaussian at sigma = 26 px holds nothing above one cycle per 26 px and the quarter-scale grid carries one per 8 — so the reduced base is not an approximation of the full-resolution one, it is the same function sampled where it is still determined. Which is also why the scale belongs to the band rather than to the stage. Texture's sigma is a decade finer, so the reduce pass's own box would be wider than the Gaussian it was prefiltering; texture never reduces. And clarity steps 4 -> 2 -> 1 as sigma falls, because a quarter of a small sigma is not a Gaussian either — the case that gives up is the one that was already cheap. `radius` stays in each pass's own pixels and `ComposedDetail::radius` multiplies it back up, so 13 reduced pixels at scale 4 still report the 52 render pixels a tile would have to be grown by. The halo a scheduler sees does not move. The halo tests pass unchanged, which was TD-4's stated bar; they render at 1024 px and so exercise the reduced path rather than stepping around it. Added `crossing_the_reduction_threshold_does_not_change_the_picture`, because nothing yet compared the reduced form against a *less* reduced one — every other test measures one form against itself. It renders the same edit either side of the 4 -> 2 step-down and holds the peak excursion to 0.03 stops and the reach to 2% of the frame. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
@@ -141,16 +141,35 @@
|
||||
//! # What this costs
|
||||
//!
|
||||
//! Clarity's kernel is large — of the order of a hundred taps per pass at
|
||||
//! preview resolution — and the two passes are the honest, exact separable
|
||||
//! Gaussian rather than a sparse approximation of one. A strided kernel would
|
||||
//! be several times cheaper and is deliberately not taken: undersampling an
|
||||
//! image that is not band-limited aliases high-frequency content down into the
|
||||
//! base, the base is then subtracted, and the aliasing arrives in the output as
|
||||
//! preview resolution — and each half is the honest, exact separable Gaussian
|
||||
//! rather than a sparse approximation of one. A strided kernel would be
|
||||
//! several times cheaper and is deliberately not taken: undersampling an image
|
||||
//! that is not band-limited aliases high-frequency content down into the base,
|
||||
//! the base is then subtracted, and the aliasing arrives in the output as
|
||||
//! low-frequency mottling across smooth gradients. Mottled skies are precisely
|
||||
//! the artefact this control must not have. The right optimisation is a base
|
||||
//! computed at reduced resolution, which needs a detail stage that can write a
|
||||
//! smaller target than it reads; that is a change to [`crate::detail`], not to
|
||||
//! this file.
|
||||
//! the artefact this control must not have.
|
||||
//!
|
||||
//! Run at the render size, that measured **34 ms at 4K** — seven times the
|
||||
//! entire fused point chain, for one slider — which is `docs/technical-debt.md`
|
||||
//! TD-4 and is what [`Recipe::base_scale`] now answers. The base is computed on
|
||||
//! a grid a quarter the size on each axis: a sixteenth of the pixels at a
|
||||
//! quarter of the radius.
|
||||
//!
|
||||
//! **This is not the strided kernel wearing a hat**, and the difference is
|
||||
//! exactly the paragraph above. A stride samples an image that is not band-
|
||||
//! limited and aliases; the reduction *band-limits first* — that is what the
|
||||
//! `reduce` pass is for and why it is a separate dispatch — and only then
|
||||
//! samples. What is thrown away is content the base could not represent at any
|
||||
//! resolution, because a Gaussian at σ = 26 px has nothing above one cycle per
|
||||
//! 26 px in it and the quarter-scale grid carries one cycle per 8 px. So the
|
||||
//! reduced base is not an approximation of the full-resolution base; it is the
|
||||
//! same band-limited function, sampled where it is still fully determined.
|
||||
//!
|
||||
//! Which is also why [`LocalContrast::reduction`] steps down and why texture
|
||||
//! never reduces at all. The argument holds only while the reduced grid can
|
||||
//! still carry the Gaussian, and the moment it cannot, the honest answer is
|
||||
//! the full-resolution one — which is the cheap case anyway, because the
|
||||
//! viewport that produced it is small.
|
||||
|
||||
use std::marker::PhantomData;
|
||||
use std::sync::{Arc, LazyLock};
|
||||
@@ -176,6 +195,14 @@ pub const AMOUNT: ParamId = ParamId("amount");
|
||||
/// pipeline.
|
||||
const TRUNCATION: f32 = 2.0;
|
||||
|
||||
/// The smallest σ, in reduced pixels, worth running a Gaussian over.
|
||||
///
|
||||
/// One pixel, which with [`TRUNCATION`] is a five-tap kernel — the narrowest
|
||||
/// that still has a shape. Below it the weights collapse towards a single tap
|
||||
/// and the blur that survives is the reduce pass's box, which is a different
|
||||
/// filter with a different edge response. See [`LocalContrast::reduction`].
|
||||
const MIN_REDUCED_SIGMA: f32 = 1.0;
|
||||
|
||||
/// Everything that makes one of these two controls the control it is.
|
||||
///
|
||||
/// A struct rather than four associated constants so that the differences
|
||||
@@ -200,6 +227,15 @@ pub struct Recipe {
|
||||
/// (4) — true for clarity, false for texture, and that asymmetry is
|
||||
/// deliberate.
|
||||
midtone_taper: bool,
|
||||
/// The most this band's base may be shrunk before it is blurred.
|
||||
///
|
||||
/// A ceiling, not the answer — [`LocalContrast::reduction`] steps it down
|
||||
/// on a viewport too small to carry it. `1` refuses the optimisation
|
||||
/// outright, which is the only correct value for a band whose σ is already
|
||||
/// a few pixels.
|
||||
///
|
||||
/// Must be a power of two: the step-down halves.
|
||||
base_scale: u32,
|
||||
}
|
||||
|
||||
/// The band of spatial frequencies a control acts on.
|
||||
@@ -235,6 +271,22 @@ impl Band for Coarse {
|
||||
threshold: 0.35,
|
||||
gain: 1.0,
|
||||
midtone_taper: true,
|
||||
// A quarter, which is what `docs/technical-debt.md` TD-4 bought back.
|
||||
//
|
||||
// σ is 1.2% of the shorter edge — 26 px at 4K — so the base holds no
|
||||
// spatial frequency anywhere near the quarter-scale Nyquist of one
|
||||
// cycle per 8 px. Computing it there is not an approximation of the
|
||||
// full-resolution base; it is the same band-limited function sampled
|
||||
// where it is still fully determined. What it costs is a sixteenth of
|
||||
// the pixels at a quarter of the radius, about a sixty-fourth of the
|
||||
// work, against the 34 ms this control measured at 4K.
|
||||
//
|
||||
// Not an eighth. σ/8 is 3.2 px at 4K and under two on a 1080p
|
||||
// viewport, which is where the reduce pass's own box filter starts
|
||||
// doing more of the blurring than the Gaussian does — and the halo
|
||||
// behaviour this operation is careful about is a property of the
|
||||
// Gaussian.
|
||||
base_scale: 4,
|
||||
};
|
||||
}
|
||||
|
||||
@@ -256,6 +308,13 @@ impl Band for Fine {
|
||||
// equal numbers on the two sliders should land at comparable strength.
|
||||
gain: 1.25,
|
||||
midtone_taper: false,
|
||||
// Never reduced, and this is the reason the scale belongs to the band
|
||||
// rather than to the stage. Texture's σ is a decade finer — 2.6 px at
|
||||
// 4K — so a quarter-scale grid would not hold its base at all: the
|
||||
// reduce pass's 4x4 box is already wider than the Gaussian it would be
|
||||
// prefiltering, and what came back would be a blur of the wrong width
|
||||
// rather than a cheaper blur of the right one.
|
||||
base_scale: 1,
|
||||
};
|
||||
}
|
||||
|
||||
@@ -385,6 +444,32 @@ impl<B: Band> LocalContrast<B> {
|
||||
(self.sigma(scale) * TRUNCATION).round().max(0.0) as u32
|
||||
}
|
||||
|
||||
/// TRACES: FR-DSP-3
|
||||
/// The factor this render's base is computed at — 1 meaning "the render
|
||||
/// size", as everything did before TD-4.
|
||||
///
|
||||
/// [`Recipe::base_scale`] is a ceiling rather than the answer, because a
|
||||
/// reduced grid still has to hold a Gaussian. At a quarter of a small
|
||||
/// viewport clarity's σ falls under a pixel, and a kernel of one or two
|
||||
/// taps is not a Gaussian — it is the reduce pass's own box filter with a
|
||||
/// rounding error on top, which would make the control change character on
|
||||
/// a window resize rather than merely get cheaper.
|
||||
///
|
||||
/// So the reduction steps down by halves until the reduced σ is worth
|
||||
/// convolving: a quarter on a desktop viewport, a half on a small one,
|
||||
/// none on a thumbnail. Stepping down rather than switching off keeps most
|
||||
/// of the saving in the middle of the range, and the case it gives up on
|
||||
/// is the one that was already cheap — the cost is `radius x pixels` and a
|
||||
/// small viewport is small in both.
|
||||
pub fn reduction(&self, scale: RenderScale) -> u32 {
|
||||
let sigma = self.sigma(scale);
|
||||
let mut reduction = B::RECIPE.base_scale.max(1);
|
||||
while reduction > 1 && sigma / (reduction as f32) < MIN_REDUCED_SIGMA {
|
||||
reduction /= 2;
|
||||
}
|
||||
reduction
|
||||
}
|
||||
|
||||
/// Stops of local contrast at this slider position.
|
||||
fn gain(&self) -> f32 {
|
||||
self.amount / 100.0 * B::RECIPE.gain
|
||||
@@ -481,27 +566,134 @@ impl<B: Band> DetailStage for LocalContrast<B> {
|
||||
value: self.gain(),
|
||||
});
|
||||
|
||||
let reduction = self.reduction(scale);
|
||||
if reduction == 1 {
|
||||
// The full-resolution form, unchanged: blur x into the scratch
|
||||
// lane, then finish along y and apply the mask in one pass.
|
||||
return vec![
|
||||
DetailPass {
|
||||
output_scale: 1,
|
||||
label: "base",
|
||||
radius,
|
||||
// A convolution, not a list: nothing to bind at binding 3.
|
||||
storage: Vec::new(),
|
||||
uniforms: shape,
|
||||
wgsl: BASE_X.to_string(),
|
||||
},
|
||||
DetailPass {
|
||||
output_scale: 1,
|
||||
label: "combine",
|
||||
radius,
|
||||
// A convolution, not a list: nothing to bind at binding 3.
|
||||
storage: Vec::new(),
|
||||
uniforms: combine,
|
||||
wgsl: combine_body(B::RECIPE.midtone_taper, Base::Convolved),
|
||||
},
|
||||
];
|
||||
}
|
||||
|
||||
// The reduced form. Four passes rather than two, and cheaper than the
|
||||
// two by a factor of about `reduction²`: three of them run on a grid
|
||||
// that many times smaller on each axis, and the one that does not is a
|
||||
// single bilinear read.
|
||||
//
|
||||
// The σ and the radius are the band's own, divided — not recomputed
|
||||
// from `RenderScale`, which knows nothing about this grid. Deriving
|
||||
// them from the numbers the full-resolution path uses is what keeps
|
||||
// the two forms the same filter, so that crossing the threshold in
|
||||
// `reduction` does not change the picture.
|
||||
let reduced_sigma = sigma / reduction as f32;
|
||||
// At least one tap either side. `reduction` has already guaranteed
|
||||
// σ >= MIN_REDUCED_SIGMA, so this floor is a belt on top of a brace.
|
||||
let reduced_radius = ((reduced_sigma * TRUNCATION).round() as u32).max(1);
|
||||
let reduced_shape = vec![
|
||||
Uniform {
|
||||
name: "radius",
|
||||
value: reduced_radius as f32,
|
||||
},
|
||||
Uniform {
|
||||
name: "inv_variance",
|
||||
value: 1.0 / (reduced_sigma * reduced_sigma),
|
||||
},
|
||||
];
|
||||
|
||||
// The combining pass no longer convolves anything, so it needs neither
|
||||
// the radius nor the variance — only the two numbers the unsharp mask
|
||||
// itself is made of.
|
||||
let mask = vec![
|
||||
Uniform {
|
||||
name: "threshold",
|
||||
value: B::RECIPE.threshold,
|
||||
},
|
||||
Uniform {
|
||||
name: "gain",
|
||||
value: self.gain(),
|
||||
},
|
||||
];
|
||||
|
||||
vec![
|
||||
DetailPass {
|
||||
label: "base",
|
||||
radius,
|
||||
// A convolution, not a list: nothing to bind at binding 3.
|
||||
output_scale: reduction,
|
||||
label: "reduce",
|
||||
// Reads only the block it writes, so it reaches no further
|
||||
// than the pixel it is producing and a tile needs no halo for
|
||||
// it. The halo the *chain* needs comes from the blurs below.
|
||||
radius: 0,
|
||||
storage: Vec::new(),
|
||||
uniforms: shape,
|
||||
wgsl: BASE_X.to_string(),
|
||||
uniforms: vec![Uniform {
|
||||
name: "reduction",
|
||||
value: reduction as f32,
|
||||
}],
|
||||
wgsl: REDUCE.to_string(),
|
||||
},
|
||||
DetailPass {
|
||||
label: "combine",
|
||||
radius,
|
||||
// A convolution, not a list: nothing to bind at binding 3.
|
||||
output_scale: reduction,
|
||||
label: "base-x",
|
||||
radius: reduced_radius,
|
||||
storage: Vec::new(),
|
||||
uniforms: combine,
|
||||
wgsl: combine_body(B::RECIPE.midtone_taper),
|
||||
uniforms: reduced_shape.clone(),
|
||||
wgsl: reduced_blur(Axis::X),
|
||||
},
|
||||
DetailPass {
|
||||
output_scale: reduction,
|
||||
label: "base-y",
|
||||
radius: reduced_radius,
|
||||
storage: Vec::new(),
|
||||
uniforms: reduced_shape,
|
||||
wgsl: reduced_blur(Axis::Y),
|
||||
},
|
||||
DetailPass {
|
||||
output_scale: 1,
|
||||
label: "combine",
|
||||
// One bilinear read of the reduced chain, which reaches one
|
||||
// reduced pixel — `reduction` render pixels — around itself.
|
||||
// Stated rather than left at zero because an understated
|
||||
// radius is a tile seam, and a seam is worth more than the
|
||||
// three lines it costs to be accurate here.
|
||||
radius: reduction,
|
||||
storage: Vec::new(),
|
||||
uniforms: mask,
|
||||
wgsl: combine_body(B::RECIPE.midtone_taper, Base::Reduced),
|
||||
},
|
||||
]
|
||||
}
|
||||
}
|
||||
|
||||
/// Which way a separable half runs.
|
||||
enum Axis {
|
||||
X,
|
||||
Y,
|
||||
}
|
||||
|
||||
/// Where the combining pass finds the base it subtracts.
|
||||
enum Base {
|
||||
/// Convolved along y by the combining pass itself, out of the scratch
|
||||
/// lane the previous pass wrote. The full-resolution form.
|
||||
Convolved,
|
||||
/// Already finished, on the reduced chain, and read back up.
|
||||
Reduced,
|
||||
}
|
||||
|
||||
/// Half of the base, along x.
|
||||
///
|
||||
/// Deliberately does not touch `c`: the pass after this one needs the
|
||||
@@ -533,13 +725,114 @@ for (var i = -r; i <= r; i = i + 1) {
|
||||
// exist.
|
||||
aux = sum / weight;";
|
||||
|
||||
/// Band-limit the image onto the reduced grid, in log luminance.
|
||||
///
|
||||
/// A pass of its own rather than something the first blur half does on the
|
||||
/// way past, because it is a different filter doing a different job: this one
|
||||
/// exists so that the Gaussian's *input* is representable on the coarse grid.
|
||||
/// Sampling every fourth pixel instead would alias — a shimmer that changes
|
||||
/// when the viewport is resized, which is the classic way a mip-based blur
|
||||
/// goes wrong and is very hard to attribute to a clarity slider.
|
||||
///
|
||||
/// A box over exactly the block the output pixel covers. Not a wider or
|
||||
/// prettier prefilter: the Gaussian that follows is 8σ wide on this grid, so
|
||||
/// what a better prefilter would buy is a correction of a fraction of a
|
||||
/// reduced pixel to a curve four pixels across, and it would cost taps on the
|
||||
/// only pass here that reads the full-resolution image.
|
||||
///
|
||||
/// **In log luminance, not linear.** The base is a mean of logarithms — that
|
||||
/// is what makes `detail` a ratio and the whole operation exposure-invariant
|
||||
/// (see the module documentation, halo control 1). Averaging linear values
|
||||
/// here and taking the logarithm later is a different number, and the
|
||||
/// difference is precisely the local contrast this operation exists to
|
||||
/// measure: it would be quietly subtracted out of every block.
|
||||
const REDUCE: &str = "\
|
||||
// The colour rides through untouched, as it does in every pass of this
|
||||
// operation — though here it is untouched and also unused: the reduced chain
|
||||
// carries a scalar, and the colour the combining pass subtracts from is the
|
||||
// full-resolution one it reads from the other chain.
|
||||
let s = i32(reduction);
|
||||
var sum = 0.0;
|
||||
for (var y = 0; y < s; y = y + 1) {
|
||||
for (var x = 0; x < s; x = x + 1) {
|
||||
sum = sum + log_luma(tap(coord, vec2<i32>(x, y)));
|
||||
}
|
||||
}
|
||||
aux = sum / f32(s * s);";
|
||||
|
||||
/// One half of the reduced separable Gaussian.
|
||||
///
|
||||
/// Reads the scratch lane rather than the colour, which is the one line that
|
||||
/// differs from [`BASE_X`]: by this point the log-luminance conversion has
|
||||
/// already been done, once, by the reduce pass. Doing it again per tap would
|
||||
/// be a logarithm inside the inner loop for a value that cannot have changed.
|
||||
fn reduced_blur(axis: Axis) -> String {
|
||||
let offset = match axis {
|
||||
Axis::X => "vec2<i32>(i, 0)",
|
||||
Axis::Y => "vec2<i32>(0, i)",
|
||||
};
|
||||
format!(
|
||||
"\
|
||||
// Half of a separable Gaussian over the reduced base, in log luminance.
|
||||
//
|
||||
// Weights are evaluated rather than tabulated, as in `BASE_X` and for the same
|
||||
// reason — and here the loop is a quarter as long, which is the whole point.
|
||||
let r = i32(radius);
|
||||
var sum = 0.0;
|
||||
var weight = 0.0;
|
||||
for (var i = -r; i <= r; i = i + 1) {{
|
||||
let f = f32(i);
|
||||
let w = exp(-0.5 * f * f * inv_variance);
|
||||
sum = sum + w * tap_aux(coord, {offset});
|
||||
weight = weight + w;
|
||||
}}
|
||||
// Normalised by the weights actually summed, so a kernel clamped at the
|
||||
// border averages the pixels that exist rather than fading towards zero.
|
||||
aux = sum / weight;"
|
||||
)
|
||||
}
|
||||
|
||||
/// The second pass: finish the base along y, then apply the mask.
|
||||
///
|
||||
/// Generated rather than constant because the midtone taper is present for
|
||||
/// clarity and absent for texture. Emitting the line only where it applies
|
||||
/// keeps texture's shader honest about not having one, and saves it a uniform
|
||||
/// and two helper functions it would never call.
|
||||
fn combine_body(midtone_taper: bool) -> String {
|
||||
fn combine_body(midtone_taper: bool, base: Base) -> String {
|
||||
// Where the base comes from — the one thing that differs between the two
|
||||
// forms. Everything below this line is the unsharp mask itself, written
|
||||
// once, so the reduced form cannot drift into being a different operation
|
||||
// from the full-resolution one it replaces.
|
||||
let base = match base {
|
||||
Base::Convolved => "\
|
||||
// The other half of the base, then the unsharp mask itself.
|
||||
//
|
||||
// `tap_aux` reads the previous pass's log-luminance blur, while `c` is still
|
||||
// the colour the colour pass produced — which is the arrangement that makes an
|
||||
// unsharp mask expressible in a chain that hands on one texture per pass.
|
||||
let r = i32(radius);
|
||||
var sum = 0.0;
|
||||
var weight = 0.0;
|
||||
for (var i = -r; i <= r; i = i + 1) {
|
||||
let f = f32(i);
|
||||
let w = exp(-0.5 * f * f * inv_variance);
|
||||
sum = sum + w * tap_aux(coord, vec2<i32>(0, i));
|
||||
weight = weight + w;
|
||||
}
|
||||
let base = sum / weight;"
|
||||
.to_string(),
|
||||
Base::Reduced => "\
|
||||
// The base, finished on the reduced chain and read back up bilinearly. `c` is
|
||||
// the full-resolution colour, straight off the other chain — which is why the
|
||||
// reduced passes had to leave that chain alone, and why this pass reads two
|
||||
// textures rather than one.
|
||||
//
|
||||
// One read where the full-resolution form runs a 105-tap convolution. That
|
||||
// difference *is* TD-4.
|
||||
let base = reduced_at(coord);"
|
||||
.to_string(),
|
||||
};
|
||||
|
||||
let weight = if midtone_taper {
|
||||
"\n\
|
||||
// Clarity only: tapered to nothing at both ends of the range. See\n\
|
||||
@@ -556,21 +849,7 @@ fn combine_body(midtone_taper: bool) -> String {
|
||||
|
||||
format!(
|
||||
"\
|
||||
// The other half of the base, then the unsharp mask itself.
|
||||
//
|
||||
// `tap_aux` reads the previous pass's log-luminance blur, while `c` is still
|
||||
// the colour the colour pass produced — which is the arrangement that makes an
|
||||
// unsharp mask expressible in a chain that hands on one texture per pass.
|
||||
let r = i32(radius);
|
||||
var sum = 0.0;
|
||||
var weight = 0.0;
|
||||
for (var i = -r; i <= r; i = i + 1) {{
|
||||
let f = f32(i);
|
||||
let w = exp(-0.5 * f * f * inv_variance);
|
||||
sum = sum + w * tap_aux(coord, vec2<i32>(0, i));
|
||||
weight = weight + w;
|
||||
}}
|
||||
let base = sum / weight;
|
||||
{base}
|
||||
|
||||
// Local contrast, in **stops**. Both terms are logarithms, so this is a ratio:
|
||||
// `detail` says how much brighter this pixel is than its surroundings, and
|
||||
@@ -632,14 +911,34 @@ mod tests {
|
||||
// forty-pixel blur for a control sitting at zero.
|
||||
let scale = RenderScale::full((2000, 1500));
|
||||
let only_clarity = composed(50.0, 0.0, scale);
|
||||
assert_eq!(only_clarity.len(), 2, "one operation, two passes");
|
||||
// Four at this viewport, because clarity's base is computed reduced
|
||||
// here — reduce, two blur halves, combine. The number is the band's
|
||||
// and the viewport's, not a constant; what this test is about is that
|
||||
// *all* of them belong to clarity.
|
||||
assert_eq!(only_clarity.len(), 4, "one operation, its own passes");
|
||||
assert!(only_clarity
|
||||
.passes
|
||||
.iter()
|
||||
.all(|p| p.label.starts_with("clarity/")));
|
||||
|
||||
// Both on: clarity's four plus texture's two. Texture stays at the
|
||||
// render size whatever the viewport — its band is a decade finer, so
|
||||
// a reduced grid could not hold its base — and that asymmetry is the
|
||||
// reason the scale belongs to the band rather than to the stage.
|
||||
let both = composed(50.0, 50.0, scale);
|
||||
assert_eq!(both.len(), 4);
|
||||
assert_eq!(both.len(), 6);
|
||||
assert_eq!(
|
||||
both.passes
|
||||
.iter()
|
||||
.filter(|p| p.label.starts_with("texture/"))
|
||||
.count(),
|
||||
2
|
||||
);
|
||||
assert!(both
|
||||
.passes
|
||||
.iter()
|
||||
.filter(|p| p.label.starts_with("texture/"))
|
||||
.all(|p| p.output_scale == 1));
|
||||
}
|
||||
|
||||
#[test]
|
||||
@@ -759,21 +1058,46 @@ mod tests {
|
||||
// passes on one texture per pass. If the blur pass ever writes `c`,
|
||||
// the combining pass has nothing to subtract the base *from* and the
|
||||
// operation silently becomes a blur.
|
||||
let passes = Clarity::with_amount(50.0).passes(RenderScale::full((2000, 1500)));
|
||||
let base = &passes[0];
|
||||
let combine = &passes[1];
|
||||
// Both forms, because they are two chains and the property has to
|
||||
// hold in each. A thumbnail is too small to carry a reduced base and
|
||||
// takes the two-pass path; a desktop viewport takes the four-pass one.
|
||||
for scale in [
|
||||
RenderScale::full((160, 120)),
|
||||
RenderScale::full((2000, 1500)),
|
||||
] {
|
||||
let passes = Clarity::with_amount(50.0).passes(scale);
|
||||
let (combine, blurs) = passes.split_last().expect("clarity is active");
|
||||
|
||||
assert!(base.wgsl.contains("aux = sum / weight;"));
|
||||
for blur in blurs {
|
||||
assert!(
|
||||
!blur.wgsl.contains("c = "),
|
||||
"{} must leave the colour alone: {}",
|
||||
blur.label,
|
||||
blur.wgsl
|
||||
);
|
||||
assert!(
|
||||
blur.wgsl.contains("aux ="),
|
||||
"{} must put its result in the scratch lane",
|
||||
blur.label
|
||||
);
|
||||
}
|
||||
|
||||
// However the base was arrived at, the mask subtracts it from the
|
||||
// colour the *colour* pass produced. That is the whole property:
|
||||
// an unsharp mask needs the blur and the original together, and
|
||||
// neither chain may have overwritten the original on the way.
|
||||
assert!(combine.wgsl.contains("log_luma(c) - base"));
|
||||
}
|
||||
|
||||
// And the two forms differ in exactly one place — where `base` came
|
||||
// from. A reduced combine does no convolution at all.
|
||||
let reduced = Clarity::with_amount(50.0).passes(RenderScale::full((2000, 1500)));
|
||||
let combine = reduced.last().expect("clarity is active");
|
||||
assert!(combine.wgsl.contains("let base = reduced_at(coord);"));
|
||||
assert!(
|
||||
!base.wgsl.contains("c = "),
|
||||
"the blur pass must leave the colour alone: {}",
|
||||
base.wgsl
|
||||
!combine.wgsl.contains("tap_aux("),
|
||||
"a reduced combine has nothing left to convolve"
|
||||
);
|
||||
assert!(
|
||||
combine.wgsl.contains("tap_aux("),
|
||||
"the combining pass must read the blur out of the scratch lane"
|
||||
);
|
||||
assert!(combine.wgsl.contains("log_luma(c) - base"));
|
||||
}
|
||||
|
||||
#[test]
|
||||
@@ -795,8 +1119,29 @@ mod tests {
|
||||
let clarity = Clarity::with_amount(50.0).kernel(scale);
|
||||
assert_eq!(composed.radius(), clarity, "the widest pass sets the halo");
|
||||
for pass in &composed.passes {
|
||||
// The reduce pass is the one honest exception: it reads exactly
|
||||
// the block it writes and no further, so a tile computing it needs
|
||||
// no halo at all. Every other pass reaches somewhere and must say
|
||||
// so.
|
||||
if pass.label.ends_with("/reduce") {
|
||||
assert_eq!(pass.radius, 0, "the reduce pass reads only its own block");
|
||||
continue;
|
||||
}
|
||||
assert!(pass.radius > 0, "{} declared no reach", pass.label);
|
||||
}
|
||||
|
||||
// The equality above is the claim worth restating: a reduced base
|
||||
// reaches exactly as far across the photograph as the full-resolution
|
||||
// one it replaces. `radius x output_scale`, 9 x 4 against 36, which is
|
||||
// what makes the reduction invisible to a tile scheduler.
|
||||
let reduced = Clarity::with_amount(50.0);
|
||||
assert_eq!(reduced.reduction(scale), 4);
|
||||
let base_x = composed
|
||||
.passes
|
||||
.iter()
|
||||
.find(|p| p.label.ends_with("/base-x"))
|
||||
.expect("a reduced chain has an x half");
|
||||
assert_eq!(base_x.radius * base_x.output_scale, clarity);
|
||||
}
|
||||
|
||||
#[test]
|
||||
@@ -805,7 +1150,8 @@ mod tests {
|
||||
// saturates at the threshold, so the largest overshoot a full-travel
|
||||
// slider can produce is `gain * threshold` stops however violent the
|
||||
// edge — a bound that holds by construction rather than by tuning.
|
||||
let combine = &Clarity::with_amount(100.0).passes(RenderScale::full((2000, 1500)))[1];
|
||||
let passes = Clarity::with_amount(100.0).passes(RenderScale::full((2000, 1500)));
|
||||
let combine = passes.last().expect("clarity is active");
|
||||
assert!(combine
|
||||
.wgsl
|
||||
.contains("threshold * tanh(detail / threshold)"));
|
||||
@@ -824,7 +1170,8 @@ mod tests {
|
||||
// Both at full travel, which is what makes them a bound and not a
|
||||
// measurement: no picture, and no edge in any picture, can produce more.
|
||||
assert!((bound(&combine.uniforms) - 0.35).abs() < 1e-6);
|
||||
let texture = &Texture::with_amount(100.0).passes(RenderScale::full((2000, 1500)))[1];
|
||||
let texture_passes = Texture::with_amount(100.0).passes(RenderScale::full((2000, 1500)));
|
||||
let texture = texture_passes.last().expect("texture is active");
|
||||
assert!((bound(&texture.uniforms) - 1.25).abs() < 1e-6);
|
||||
}
|
||||
|
||||
@@ -834,8 +1181,10 @@ mod tests {
|
||||
// highlight and fabric in a shadow, which is exactly where clarity
|
||||
// must not.
|
||||
let scale = RenderScale::full((2000, 1500));
|
||||
let clarity = &Clarity::with_amount(50.0).passes(scale)[1];
|
||||
let texture = &Texture::with_amount(50.0).passes(scale)[1];
|
||||
let clarity_passes = Clarity::with_amount(50.0).passes(scale);
|
||||
let texture_passes = Texture::with_amount(50.0).passes(scale);
|
||||
let clarity = clarity_passes.last().expect("clarity is active");
|
||||
let texture = texture_passes.last().expect("texture is active");
|
||||
assert!(clarity.wgsl.contains("midtone_weight(luminance(c))"));
|
||||
assert!(!texture.wgsl.contains("midtone_weight"));
|
||||
|
||||
@@ -865,15 +1214,24 @@ mod tests {
|
||||
let up = Clarity::with_amount(50.0).passes(scale);
|
||||
let down = Clarity::with_amount(-50.0).passes(scale);
|
||||
let gain = |p: &[DetailPass]| {
|
||||
p[1].uniforms
|
||||
p.last()
|
||||
.expect("clarity is active")
|
||||
.uniforms
|
||||
.iter()
|
||||
.find(|u| u.name == "gain")
|
||||
.unwrap()
|
||||
.value
|
||||
};
|
||||
assert!((gain(&up) + gain(&down)).abs() < 1e-6);
|
||||
// Same kernel either way — the direction is a sign, not a scale.
|
||||
assert_eq!(up[0].radius, down[0].radius);
|
||||
// Same kernel either way — the direction is a sign, not a scale. Read
|
||||
// off the blur halves rather than the first pass, which since the
|
||||
// reduced base exists is a downscale carrying no radius at all.
|
||||
let radii = |p: &[DetailPass]| {
|
||||
p.iter()
|
||||
.map(|d| (d.radius, d.output_scale))
|
||||
.collect::<Vec<_>>()
|
||||
};
|
||||
assert_eq!(radii(&up), radii(&down));
|
||||
}
|
||||
|
||||
#[test]
|
||||
|
||||
Reference in New Issue
Block a user