Erode dehaze's window in one pass per axis, and recover in the second

Dehaze cost 22.9 ms of a 2560x1600 frame on the reference laptop RTX 3050,
and 54.1 ms at 3840x2160, with the memory clock held at 810 MHz by the power
cap (graphics 1762 MHz). It ran five passes: a run and a span erosion along
x, the same along y, and the recovery. At those clocks a detail pass costs
what it reads and writes, not what it taps: a pass with an empty body -
one render-sized rgba16float read and write - measured 4.0 ms, and each
dehaze pass 4.4-4.6 ms, so the taps were about 2 ms of the 22 and the four
hand-offs between passes were the rest.

Each axis is now one pass that takes the minimum over the whole window
directly, and the recovery rides in the y pass, which already holds the
veil and the pixel's own colour. That is 36 texture reads per pixel at
2560x1600 in place of 12, nearly all of them cache hits, and two passes in
place of five.

The picture is the same bits. A minimum is exact in any order, and the
window is the one Split always covered, the surplus pixel on the far side
included (Split::first and Split::width). The veil crossing the removed
hand-offs was already exactly representable in rgba16float - a minimum of
channels read from rgba16float, floored at zero - so storing it between
passes never rounded anything that the fused form now keeps unrounded.

Measured with a scratch probe that renders the synthetic 60 MP frame from
examples/frame_budget.rs, only a detail parameter moving so the fused pass
is reused, 30 frames per scene after six of warm-up, five runs of each
binary alternated, median of the per-run p50:

  scene                   before     after
  dehaze      2560 fit    22.88 ms    9.06 ms
  dehaze      2560 1:1    23.41 ms    9.52 ms
  dehaze      3840 fit    54.09 ms   28.12 ms
  all detail  2560 fit    53.11 ms   39.97 ms  (NR, sharpen, clarity,
  all detail  2560 1:1    67.48 ms   56.42 ms   texture, dehaze)
  every op    2560 fit    57.59 ms   44.19 ms  (with film)
  every op    2560 1:1    71.83 ms   57.93 ms
  controls without dehaze (NR, sharpen, clarity, texture): within +-2%

The rgba8 output hashed identically before and after for every scene -
dehaze alone, all five detail operations, every operation with film, and
each other detail operation alone - at fit and 1:1, at 2560x1600,
3840x2160, 1917x1203 and 333x211: 64 of 64.
This commit is contained in:
2026-09-26 14:18:42 -04:00
parent 9b580c3720
commit bee5c5866f
2 changed files with 120 additions and 146 deletions
+118 -144
View File
@@ -116,12 +116,34 @@
//! mottling across smooth gradients. This decomposition samples nothing — it
//! evaluates the exact minimum over every pixel of the window, in two steps.
//!
//! # Why it is now two passes and not five
//!
//! The decomposition was run as four erosion passes — run and span along x,
//! then along y — and a fifth for the recovery. On the reference laptop the
//! taps turned out not to be what a pass costs: with the memory clock held at
//! 810 MHz by the power cap, a pass that only reads the render-sized
//! `rgba16float` intermediate and writes the other one costs about 4 ms at
//! 2560 x 1600, and dehaze's five came to 22 ms, of which the taps were about
//! 2. So each axis is now one pass that takes the minimum over the whole
//! window directly — 36 texture reads per pixel at that size instead of 12,
//! nearly all of them served by the cache — and the recovery rides in the y
//! pass, which already has the veil and this pixel's colour in hand. Two
//! passes: 9.2 ms.
//!
//! It is the same picture, bit for bit. A minimum is exact in any order, so
//! the minimum over the window's pixels is one value however it is grouped,
//! and the window is the one [`Split`] always covered, surplus pixel
//! included. The veil reaching the recovery was always exactly representable
//! in the `rgba16float` lane it crossed — a minimum of channel values that
//! were themselves read from `rgba16float` — so no rounding was lost by not
//! storing it between passes.
//!
//! It is also why this operation does not use the reduced chain that TD-4 gave
//! clarity. The runner holds one reduced buffer, so every scaled pass in a
//! chain must declare the same `output_scale`; clarity's steps down with the
//! viewport, so a second operation choosing its own would disagree with it at
//! some window sizes and not others. Four cheap full-resolution passes cost
//! less than that coupling, and the decomposition is what makes them cheap.
//! some window sizes and not others. Two full-resolution passes cost less than
//! that coupling.
//!
//! # The artefact this does not fix
//!
@@ -275,22 +297,23 @@ impl Dehaze {
///
/// See the module documentation: eroding by a contiguous run and then by a set
/// of points spaced one run apart erodes by the sum of the two, which is the
/// whole window. This is the arithmetic of that split, in one place, because
/// both axes need it and a second copy is a second chance to get the centring
/// wrong.
/// whole window. The passes no longer run the two stages apart (see "Why it is
/// now two passes"), but the window they read is still the one this split
/// covers — [`Self::first`] and [`Self::width`] — surplus pixel included,
/// because that is the window every edit made so far was tuned against.
#[derive(Debug, Clone, Copy, PartialEq)]
pub struct Split {
/// Length of the contiguous run the first pass takes the minimum over.
pub run: u32,
/// How many runs the second pass chains together, spaced `run` apart.
pub span: u32,
/// What the second pass subtracts from its offsets to centre the window.
/// How far before the pixel being written the window starts.
///
/// The composite covers `run * span` pixels, which is at least the window
/// asked for and can be one or two more; the surplus falls on the far side
/// rather than being trimmed, because trimming it would need a third pass
/// and a patch a pixel wider on one side is not a visible difference in a
/// field this smooth.
/// rather than being trimmed, because trimming it would have needed a
/// third pass and a patch a pixel wider on one side is not a visible
/// difference in a field this smooth.
pub shift: i32,
}
@@ -311,14 +334,25 @@ impl Split {
}
}
/// The furthest the second pass reads, in pixels.
/// The window's first offset from the pixel being written: `-shift`.
pub fn first(&self) -> i32 {
-self.shift
}
/// How many pixels the window covers, `run * span` — the patch asked for
/// and the one or two surplus pixels on the far side the split leaves.
pub fn width(&self) -> u32 {
self.run * self.span
}
/// The furthest the window reads from the pixel being written, in pixels.
///
/// Stated rather than assumed symmetric: the composite window is centred
/// to within a pixel and not exactly, so the two directions can differ by
/// one. An understated radius is a seam at every tile boundary (ARCH
/// §5.3), which is the kind of artefact that looks like a driver bug.
pub fn reach(&self) -> u32 {
let far = (self.span.saturating_sub(1) * self.run) as i32 - self.shift;
/// Stated rather than assumed symmetric: the window is centred to within a
/// pixel and not exactly, so the two directions can differ by one. An
/// understated radius is a seam at every tile boundary (ARCH §5.3), which
/// is the kind of artefact that looks like a driver bug.
pub fn extent(&self) -> u32 {
let far = self.width() as i32 - 1 - self.shift;
self.shift.max(far).max(0) as u32
}
}
@@ -367,83 +401,50 @@ impl DetailStage for Dehaze {
fn passes(&self, scale: RenderScale) -> Vec<DetailPass> {
let split = Split::of(self.patch(scale));
// The run stage's uniforms are the same on both axes, and so are the
// span stage's. Only the offset expression differs, which is what
// `erode_run` and `erode_span` take as an argument — each filter
// Both axes erode over the same window; only the offset expression
// differs, which is what `erode` takes as an argument — one filter
// written once, so the two axes cannot drift into being different
// filters.
let run = vec![Uniform {
name: "run",
value: split.run as f32,
}];
let span = vec![
let window = vec![
Uniform {
name: "span",
value: split.span as f32,
name: "first",
value: split.first() as f32,
},
Uniform {
name: "stride",
value: split.run as f32,
},
Uniform {
name: "shift",
value: split.shift as f32,
name: "width",
value: split.width() as f32,
},
];
let mut recovery = window.clone();
recovery.extend([
Uniform {
name: "omega",
value: self.omega(),
},
Uniform {
name: "min_transmission",
value: MIN_TRANSMISSION,
},
]);
vec![
DetailPass {
output_scale: 1,
label: "veil-run-x",
// The run starts at this pixel and walks forward, so it reads
// `run - 1` beyond itself and nothing behind.
radius: split.run.saturating_sub(1),
label: "veil-x",
radius: split.extent(),
storage: Vec::new(),
uniforms: run.clone(),
wgsl: erode_run(Axis::X),
uniforms: window,
wgsl: erode(Axis::X),
},
DetailPass {
output_scale: 1,
label: "veil-span-x",
radius: split.reach(),
label: "veil-y-clear",
radius: split.extent(),
storage: Vec::new(),
uniforms: span.clone(),
wgsl: erode_span(Axis::X),
},
DetailPass {
output_scale: 1,
label: "veil-run-y",
radius: split.run.saturating_sub(1),
storage: Vec::new(),
uniforms: run,
wgsl: erode_run(Axis::Y),
},
DetailPass {
output_scale: 1,
label: "veil-span-y",
radius: split.reach(),
storage: Vec::new(),
uniforms: span,
wgsl: erode_span(Axis::Y),
},
DetailPass {
output_scale: 1,
label: "clear",
// Reads only the pixel it writes: the veil arrived in the
// scratch lane four passes ago.
radius: 0,
storage: Vec::new(),
uniforms: vec![
Uniform {
name: "omega",
value: self.omega(),
},
Uniform {
name: "min_transmission",
value: MIN_TRANSMISSION,
},
],
wgsl: CLEAR.to_string(),
uniforms: recovery,
// The erosion in a block of its own, so its locals do not
// collide with the recovery's.
wgsl: format!("{{\n{}\n}}\n\n{CLEAR}", erode(Axis::Y)),
},
]
}
@@ -466,33 +467,33 @@ impl Axis {
}
}
/// The first stage: the minimum over a contiguous run.
/// The minimum over the window along one axis.
///
/// Along x it reads the colour and reduces it to the dark channel; along y the
/// dark channel is already in the scratch lane, so it reads that instead.
/// Doing the channel minimum again on the second axis would be reducing a
/// scalar and would quietly discard the x erosion.
/// dark channel's x minimum is already in the scratch lane, so it reads that
/// instead. Doing the channel minimum again on the second axis would be
/// reducing a scalar and would quietly discard the x erosion.
///
/// Neither stage touches `c`. The recovery needs the original colour *and* the
/// veil in the same place at the same time, and the ping-pong hands each pass
/// only what the pass before it wrote — so the veil travels in `aux` and the
/// colour rides through untouched. See [`DetailPass::wgsl`].
fn erode_run(axis: Axis) -> String {
/// The x pass does not touch `c`. The recovery needs the original colour *and*
/// the veil in the same place at the same time, and the ping-pong hands each
/// pass only what the pass before it wrote — so the veil travels in `aux` and
/// the colour rides through untouched. See [`DetailPass::wgsl`].
fn erode(axis: Axis) -> String {
let source = match axis {
Axis::X => "dark_channel(tap(coord, OFFSET))",
Axis::Y => "tap_aux(coord, OFFSET)",
};
let first = source.replace("OFFSET", &axis.offset("0"));
let rest = source.replace("OFFSET", &axis.offset("i"));
let head = source.replace("OFFSET", &axis.offset("o"));
let rest = source.replace("OFFSET", &axis.offset("o + i"));
format!(
"\
// Half of the erosion's first stage: the minimum over `run` contiguous pixels,
// walking forward from this one. The second stage chains these together, and
// the two structuring elements add up to the whole patch — which is why this
// one is not centred and does not need to be.
let n = i32(run);
var veil = {first};
// The minimum over every pixel of the window along this axis, from `first`
// for `width` pixels. A minimum is exact in any order, so this is the same
// value, bit for bit, as chaining a run and a span over the same pixels.
let o = i32(first);
let n = i32(width);
var veil = {head};
for (var i = 1; i < n; i = i + 1) {{
veil = min(veil, {rest});
}}
@@ -500,38 +501,10 @@ aux = veil;"
)
}
/// The second stage: the minimum over `span` points spaced `stride` apart.
///
/// Each of those points already holds the minimum over the run that starts
/// there, so this reads the whole window while touching `span` pixels of it.
/// `shift` is what centres the composite on the pixel being written; without
/// it the veil would be measured from a patch lying entirely to one side, and
/// the correction would appear to lag the picture by half a patch.
fn erode_span(axis: Axis) -> String {
let offset = axis.offset("j * s - o");
let first = axis.offset("-o");
format!(
"\
// The erosion's second stage. `span` taps, spaced a whole run apart, each
// standing for the run that begins at it — so the minimum over the patch costs
// `run + span` taps rather than the `run * span` pixels it covers, and it is
// the exact minimum over all of them rather than a sample of them.
let n = i32(span);
let s = i32(stride);
let o = i32(shift);
var veil = tap_aux(coord, {first});
for (var j = 1; j < n; j = j + 1) {{
veil = min(veil, tap_aux(coord, {offset}));
}}
aux = veil;"
)
}
/// The recovery: invert the scattering model with the transmission the erosion
/// implies.
const CLEAR: &str = "\
// The veil the four erosion passes measured: the smallest channel anywhere in
// The veil the two erosions measured: the smallest channel anywhere in
// the patch around this pixel, which the dark-channel prior reads as the
// airlight that has been composited over the scene here.
//
@@ -581,7 +554,7 @@ mod tests {
#[test]
fn dehaze_starts_neutral_and_costs_nothing() {
// The rule the whole pipeline rests on. An unedited photograph must not
// pay for a slider nobody has touched — and this one is five dispatches
// pay for a slider nobody has touched — and this one is two dispatches
// when it is on, so "nothing" here is a worthwhile amount of nothing.
assert!(!Dehaze::new().is_active());
assert!(composed(0.0, RenderScale::full((2000, 1500))).is_empty());
@@ -623,7 +596,7 @@ mod tests {
// develop view about it would read as a bug.
let tiny = RenderScale::full((48, 32));
assert_eq!(Dehaze::with_amount(50.0).patch(tiny), 1);
assert_eq!(composed(50.0, tiny).len(), 5);
assert_eq!(composed(50.0, tiny).len(), 2);
}
#[test]
@@ -648,7 +621,8 @@ mod tests {
// Centred to within the pixel the odd surplus leaves over: the window
// covers [-31, 32] around the pixel being written.
assert_eq!(split.shift, 31);
assert_eq!(split.reach(), 31);
assert_eq!((split.first(), split.width()), (-31, 64));
assert_eq!(split.extent(), 32);
}
#[test]
@@ -662,33 +636,32 @@ mod tests {
assert_eq!(split.span, 2);
assert!(split.run * split.span >= 3, "the window is not covered");
assert_eq!(split.shift, 1);
assert_eq!(split.reach(), 1);
// [-1, 2]: the surplus pixel is on the far side.
assert_eq!((split.first(), split.width()), (-1, 4));
assert_eq!(split.extent(), 2);
}
#[test]
fn the_chain_is_four_erosions_and_a_recovery() {
fn the_chain_is_an_erosion_per_axis_the_second_carrying_the_recovery() {
// The shape of the operation, asserted where it is cheap to assert.
// The erosions leave the colour alone and hand the veil forward in the
// scratch lane; only the last pass touches `c`, which is what makes an
// The x erosion leaves the colour alone and hands the veil forward in
// the scratch lane; only the y pass touches `c`, which is what makes an
// unsharp-mask-shaped operation expressible in a chain that hands each
// pass exactly one texture.
// pass exactly one texture. Two passes and not five: each extra pass
// is a render-sized read and write, which is what a pass costs (see
// "Why it is now two passes").
let composed = composed(60.0, RenderScale::full((2000, 1500)));
let labels: Vec<&str> = composed.passes.iter().map(|p| p.label.as_str()).collect();
assert_eq!(
labels,
[
"dehaze/veil-run-x",
"dehaze/veil-span-x",
"dehaze/veil-run-y",
"dehaze/veil-span-y",
"dehaze/clear",
]
);
assert_eq!(labels, ["dehaze/veil-x", "dehaze/veil-y-clear"]);
// Both declare the whole window they read.
let split = Split::of(Dehaze::with_amount(60.0).patch(RenderScale::full((2000, 1500))));
assert!(composed.passes.iter().all(|p| p.radius == split.extent()));
// Only the last writes the display texture, so the output transform
// happens exactly once (FR-DEV-2).
assert!(composed.passes[..4].iter().all(|p| !p.writes_output));
assert!(composed.passes[4].writes_output);
assert!(!composed.passes[0].writes_output);
assert!(composed.passes[1].writes_output);
// Nothing here uses the reduced chain — see the module documentation
// for why a second operation cannot pick its own `output_scale` while
@@ -730,7 +703,8 @@ mod tests {
// nothing to say so.
let composed = composed(60.0, RenderScale::full((2000, 1500)));
assert!(composed.passes[0].source.contains("dark_channel(tap(coord"));
assert!(!composed.passes[2].source.contains("dark_channel(tap(coord"));
assert!(!composed.passes[1].source.contains("dark_channel(tap(coord"));
assert!(composed.passes[1].source.contains("tap_aux(coord"));
// The helper is still emitted for every pass of the operation, and it
// must define the function it is named for or the shader fails to
// compile a long way from here.