Let clarity's base be computed where it is still fully determined
Clarity's Gaussian sigma is 1.2% of the frame's shorter edge, so its radius is a property of the viewport: 52 render pixels at 4K, two separable passes of 105 taps each over 8.3 M pixels. That measured 33.9 ms — seven times the entire fused point chain, for one slider — and is docs/technical-debt.md TD-4. A detail pass may now declare `output_scale`, and clarity's base is computed on a grid a quarter the size on each axis. The pass that combines needs the blur *and* the full-resolution colour, and a colour that has been through a quarter-scale target is no longer full resolution. So a scaled pass cannot simply join the ping-pong: there are two chains now. The full-resolution one carries the colour and no scaled pass touches it; the reduced one carries the base and reaches the combining pass through a second binding as `reduced_at()`. The reduce is a dispatch of its own rather than something the first blur half does on the way past, and that is the whole difference between this and the strided kernel the module documentation rules out. A stride samples an image that is not band-limited and aliases high-frequency content down into the base, which is then subtracted, and arrives in the output as mottling across smooth gradients. This band-limits first and samples after. What is discarded is content the base could not represent at any resolution, because a Gaussian at sigma = 26 px holds nothing above one cycle per 26 px and the quarter-scale grid carries one per 8 — so the reduced base is not an approximation of the full-resolution one, it is the same function sampled where it is still determined. Which is also why the scale belongs to the band rather than to the stage. Texture's sigma is a decade finer, so the reduce pass's own box would be wider than the Gaussian it was prefiltering; texture never reduces. And clarity steps 4 -> 2 -> 1 as sigma falls, because a quarter of a small sigma is not a Gaussian either — the case that gives up is the one that was already cheap. `radius` stays in each pass's own pixels and `ComposedDetail::radius` multiplies it back up, so 13 reduced pixels at scale 4 still report the 52 render pixels a tile would have to be grown by. The halo a scheduler sees does not move. The halo tests pass unchanged, which was TD-4's stated bar; they render at 1024 px and so exercise the reduced path rather than stepping around it. Added `crossing_the_reduction_threshold_does_not_change_the_picture`, because nothing yet compared the reduced form against a *less* reduced one — every other test measures one form against itself. It renders the same edit either side of the 4 -> 2 step-down and holds the peak excursion to 0.03 stops and the reach to 2% of the frame. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
@@ -301,6 +301,40 @@ pub struct DetailPass {
|
||||
/// the kind of artefact that looks like a driver bug. State it honestly.
|
||||
pub radius: u32,
|
||||
|
||||
/// How much smaller than the render this pass writes.
|
||||
///
|
||||
/// `1` is the ordinary case and means "the render size", which is what
|
||||
/// every pass did before this field existed. A larger value writes a
|
||||
/// target that many times smaller on each axis, into a **second** chain
|
||||
/// held alongside the full-resolution one — see [`Self::wgsl`] for how the
|
||||
/// two are addressed, and the module documentation for why there are two.
|
||||
///
|
||||
/// # Why a pass may want this
|
||||
///
|
||||
/// A blur wide enough to be a *base* — clarity's is 1.2% of the frame,
|
||||
/// 52 render pixels at 4K — holds no spatial frequency a quarter-scale
|
||||
/// grid cannot represent. Computing it at the render size therefore buys
|
||||
/// nothing and costs everything: 105 taps over 8.3 M pixels, twice, which
|
||||
/// measured at 34 ms and is where `docs/technical-debt.md` TD-4 came from.
|
||||
/// At a quarter it is a sixteenth of the pixels at a quarter of the
|
||||
/// radius, and the result is not an approximation of the full-resolution
|
||||
/// base — it is the same band-limited function, sampled where it is still
|
||||
/// Nyquist-safe.
|
||||
///
|
||||
/// [`Self::radius`] stays in this pass's **own** pixels, so a pass at
|
||||
/// scale 4 with a radius of 13 declares 13, not 52. The halo it implies
|
||||
/// for a tile scheduler is `radius * output_scale`, and
|
||||
/// [`ComposedDetail::radius`] is what performs that multiplication —
|
||||
/// stating the radius in the grid the loop actually runs in is what keeps
|
||||
/// the shader and the declaration the same number.
|
||||
///
|
||||
/// **Never the last pass.** The final pass carries the output transform
|
||||
/// and writes the display texture, which is full resolution by
|
||||
/// definition; a scaled pass in that position is a codegen bug and
|
||||
/// `dr-gpu` refuses it rather than binding a shader to a target of the
|
||||
/// wrong size.
|
||||
pub output_scale: u32,
|
||||
|
||||
/// The WGSL body.
|
||||
///
|
||||
/// Reads and writes `c`, a `vec3<f32>` of **linear sRGB**, pre-loaded with
|
||||
@@ -319,6 +353,19 @@ pub struct DetailPass {
|
||||
/// pixel that survives to the next pass**, pre-loaded with what the
|
||||
/// previous pass left there and written back out unless the body
|
||||
/// assigns it.
|
||||
/// - `reduced_at(coord) -> f32` — the **reduced chain's** scalar at this
|
||||
/// pixel,
|
||||
/// bilinearly upsampled. Zero unless a scaled pass ran earlier in this
|
||||
/// operation; see [`Self::output_scale`].
|
||||
///
|
||||
/// `coord` is always in *this pass's own* output grid, and `tap` maps it
|
||||
/// into the source's grid for you. A pass at [`Self::output_scale`] 4
|
||||
/// therefore addresses its own quarter-size target with `coord`, while
|
||||
/// `tap(coord, offset)` offsets in **source** pixels — which is what lets
|
||||
/// a reduce pass average the 4 x 4 block a single output pixel covers by
|
||||
/// looping `offset` over it. Where source and target are the same size the
|
||||
/// mapping is the identity, so every pass written before scaling existed
|
||||
/// behaves exactly as it did.
|
||||
///
|
||||
/// # Why `aux` exists
|
||||
///
|
||||
@@ -424,6 +471,8 @@ pub struct ComposedDetailPass {
|
||||
pub storage: Vec<[f32; 4]>,
|
||||
/// See [`DetailPass::radius`].
|
||||
pub radius: u32,
|
||||
/// See [`DetailPass::output_scale`].
|
||||
pub output_scale: u32,
|
||||
/// Whether this pass writes the display/export texture rather than another
|
||||
/// linear intermediate.
|
||||
///
|
||||
@@ -457,8 +506,18 @@ impl ComposedDetail {
|
||||
}
|
||||
|
||||
/// The widest halo any pass needs, in render pixels (ARCH §5.3).
|
||||
///
|
||||
/// A pass declares its radius in its own grid, so a scaled pass's has to
|
||||
/// be multiplied back up before the maxima are comparable: 13 reduced
|
||||
/// pixels at scale 4 reach exactly as far across the photograph as 52
|
||||
/// render pixels do, and a scheduler comparing the two unscaled would size
|
||||
/// a halo at a quarter of what the pass actually reads.
|
||||
pub fn radius(&self) -> u32 {
|
||||
self.passes.iter().map(|p| p.radius).max().unwrap_or(0)
|
||||
self.passes
|
||||
.iter()
|
||||
.map(|p| p.radius.saturating_mul(p.output_scale))
|
||||
.max()
|
||||
.unwrap_or(0)
|
||||
}
|
||||
}
|
||||
|
||||
@@ -572,6 +631,7 @@ pub fn compose_detail_with(
|
||||
RESOLVE_ID,
|
||||
&[],
|
||||
&DetailPass {
|
||||
output_scale: 1,
|
||||
label: "resolve",
|
||||
radius: 0,
|
||||
wgsl: String::new(),
|
||||
@@ -723,6 +783,25 @@ struct Params {{
|
||||
// pass leaves this bound to a single empty element and never looks at it. See
|
||||
// `DetailPass::storage` for why the list is not in the uniform block.
|
||||
@group(0) @binding(3) var<storage, read> instances: array<vec4<f32>>;
|
||||
// The reduced chain — what a scaled pass most recently wrote, at whatever
|
||||
// fraction of the render size it declared. Bound to a 1x1 placeholder for
|
||||
// every pass that never calls `base`, so that one bind group layout serves a
|
||||
// pass which uses it and a pass which has never heard of it.
|
||||
@group(0) @binding(4) var reduced: texture_2d<f32>;
|
||||
|
||||
// Where in `source` this output pixel begins.
|
||||
//
|
||||
// The ratio is 1 whenever a pass writes what it reads, which is every pass
|
||||
// that does not set `output_scale` — the multiply and the divide cancel
|
||||
// exactly, so the ordinary case is unchanged and pays two integer operations
|
||||
// for the privilege. A scaled pass gets the top-left of the block it covers,
|
||||
// which is what makes `tap`'s offsets mean *source* pixels and lets a reduce
|
||||
// pass walk its own footprint.
|
||||
fn source_origin(coord: vec2<i32>) -> vec2<i32> {{
|
||||
let src = vec2<i32>(textureDimensions(source));
|
||||
let dst = vec2<i32>(textureDimensions(output));
|
||||
return coord * src / max(dst, vec2<i32>(1));
|
||||
}}
|
||||
|
||||
// A neighbour, clamped to the edge of the image.
|
||||
//
|
||||
@@ -732,13 +811,48 @@ struct Params {{
|
||||
// classic way a first convolution goes wrong.
|
||||
fn tap(coord: vec2<i32>, offset: vec2<i32>) -> vec3<f32> {{
|
||||
let last = vec2<i32>(textureDimensions(source)) - vec2<i32>(1);
|
||||
return textureLoad(source, clamp(coord + offset, vec2<i32>(0), last), 0).rgb;
|
||||
let at = source_origin(coord) + offset;
|
||||
return textureLoad(source, clamp(at, vec2<i32>(0), last), 0).rgb;
|
||||
}}
|
||||
|
||||
// The same neighbour's scratch lane — see `aux` in the body below.
|
||||
fn tap_aux(coord: vec2<i32>, offset: vec2<i32>) -> f32 {{
|
||||
let last = vec2<i32>(textureDimensions(source)) - vec2<i32>(1);
|
||||
return textureLoad(source, clamp(coord + offset, vec2<i32>(0), last), 0).a;
|
||||
let at = source_origin(coord) + offset;
|
||||
return textureLoad(source, clamp(at, vec2<i32>(0), last), 0).a;
|
||||
}}
|
||||
|
||||
// The reduced chain, read at this pass's own resolution.
|
||||
//
|
||||
// Named `reduced_at` rather than `base` because `base` is a natural local in a
|
||||
// body that has just computed one — spot removal already has such a local, and
|
||||
// a function shadowed by a variable is a compile error a long way from its
|
||||
// cause.
|
||||
//
|
||||
// Bilinear, and on pixel *centres* rather than corners: the reduce pass took
|
||||
// its sample at the centre of the block it averaged, so an upsample that
|
||||
// treated the grids as corner-aligned would shift the base by half a reduced
|
||||
// pixel — two full pixels at scale 4, which on a wide unsharp mask is a base
|
||||
// offset from the image it is subtracted from, and reads as a directional
|
||||
// smear along every edge.
|
||||
//
|
||||
// Nearest would be cheaper and is not enough: the base is subtracted from the
|
||||
// full-resolution image, so any blockiness in it appears in the *difference*
|
||||
// at full contrast. That is a visible 4-pixel grid over the whole frame.
|
||||
fn reduced_at(coord: vec2<i32>) -> f32 {{
|
||||
let rd = vec2<f32>(textureDimensions(reduced));
|
||||
let dst = vec2<f32>(max(textureDimensions(output), vec2<u32>(1u)));
|
||||
let p = (vec2<f32>(coord) + vec2<f32>(0.5)) * rd / dst - vec2<f32>(0.5);
|
||||
let last = vec2<i32>(rd) - vec2<i32>(1);
|
||||
let base_px = vec2<i32>(floor(p));
|
||||
let f = fract(p);
|
||||
|
||||
let s00 = textureLoad(reduced, clamp(base_px, vec2<i32>(0), last), 0).a;
|
||||
let s10 = textureLoad(reduced, clamp(base_px + vec2<i32>(1, 0), vec2<i32>(0), last), 0).a;
|
||||
let s01 = textureLoad(reduced, clamp(base_px + vec2<i32>(0, 1), vec2<i32>(0), last), 0).a;
|
||||
let s11 = textureLoad(reduced, clamp(base_px + vec2<i32>(1, 1), vec2<i32>(0), last), 0).a;
|
||||
|
||||
return mix(mix(s00, s10, f.x), mix(s01, s11, f.x), f.y);
|
||||
}}
|
||||
|
||||
{helper_src}{encode_fn}
|
||||
@@ -786,6 +900,10 @@ fn main(@builtin(global_invocation_id) gid: vec3<u32>) {{
|
||||
uniforms: uniform_values,
|
||||
storage: pass.storage.clone(),
|
||||
radius: pass.radius,
|
||||
// Clamped rather than trusted: a zero would divide by nothing in the
|
||||
// dispatch size and a declaration is data, which since FR-PLG-2 can
|
||||
// come from a file this build did not write.
|
||||
output_scale: pass.output_scale.max(1),
|
||||
writes_output,
|
||||
structure_hash,
|
||||
}
|
||||
|
||||
Reference in New Issue
Block a user