Clarity's Gaussian sigma is 1.2% of the frame's shorter edge, so its radius is a property of the viewport: 52 render pixels at 4K, two separable passes of 105 taps each over 8.3 M pixels. That measured 33.9 ms — seven times the entire fused point chain, for one slider — and is docs/technical-debt.md TD-4. A detail pass may now declare `output_scale`, and clarity's base is computed on a grid a quarter the size on each axis. The pass that combines needs the blur *and* the full-resolution colour, and a colour that has been through a quarter-scale target is no longer full resolution. So a scaled pass cannot simply join the ping-pong: there are two chains now. The full-resolution one carries the colour and no scaled pass touches it; the reduced one carries the base and reaches the combining pass through a second binding as `reduced_at()`. The reduce is a dispatch of its own rather than something the first blur half does on the way past, and that is the whole difference between this and the strided kernel the module documentation rules out. A stride samples an image that is not band-limited and aliases high-frequency content down into the base, which is then subtracted, and arrives in the output as mottling across smooth gradients. This band-limits first and samples after. What is discarded is content the base could not represent at any resolution, because a Gaussian at sigma = 26 px holds nothing above one cycle per 26 px and the quarter-scale grid carries one per 8 — so the reduced base is not an approximation of the full-resolution one, it is the same function sampled where it is still determined. Which is also why the scale belongs to the band rather than to the stage. Texture's sigma is a decade finer, so the reduce pass's own box would be wider than the Gaussian it was prefiltering; texture never reduces. And clarity steps 4 -> 2 -> 1 as sigma falls, because a quarter of a small sigma is not a Gaussian either — the case that gives up is the one that was already cheap. `radius` stays in each pass's own pixels and `ComposedDetail::radius` multiplies it back up, so 13 reduced pixels at scale 4 still report the 52 render pixels a tile would have to be grown by. The halo a scheduler sees does not move. The halo tests pass unchanged, which was TD-4's stated bar; they render at 1024 px and so exercise the reduced path rather than stepping around it. Added `crossing_the_reduction_threshold_does_not_change_the_picture`, because nothing yet compared the reduced form against a *less* reduced one — every other test measures one form against itself. It renders the same edit either side of the 4 -> 2 step-down and holds the peak excursion to 0.03 stops and the reach to 2% of the frame. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
171 lines
6.3 KiB
Rust
171 lines
6.3 KiB
Rust
//! TRACES: FR-DEV-8
|
|
//! The instance binding: a detail pass whose work is a list, not a kernel.
|
|
//!
|
|
//! Spot removal needs a detail pass to read a variable number of records —
|
|
//! sixty-four repairs and one repair are the same shader with a different
|
|
//! buffer behind it. That is binding 3, and this file proves the three things
|
|
//! about it that a picture would not tell you clearly:
|
|
//!
|
|
//! - the data uploaded is the data the shader reads, in order;
|
|
//! - a pass that declares no list still runs, bound to the placeholder;
|
|
//! - the same shader with a *different* list does not recompile, which is what
|
|
//! keeps placing a spot as cheap as moving a slider.
|
|
//!
|
|
//! The passes here are synthetic on purpose. `spot_removal.rs` asserts the
|
|
//! repair; this asserts the plumbing, so a failure in one does not have to be
|
|
//! read to work out which of the two broke.
|
|
|
|
use dr_gpu::{AdjustPass, DemosaicedImage, GpuContext};
|
|
use dr_pipeline::detail::{ComposedDetail, ComposedDetailPass};
|
|
use dr_pipeline::{Affects, EditGraph};
|
|
use dr_types::ColourSpace;
|
|
|
|
const SIZE: u32 = 8;
|
|
|
|
fn ctx() -> Option<GpuContext> {
|
|
match pollster::block_on(GpuContext::new_headless()) {
|
|
Ok(c) => Some(c),
|
|
Err(e) => {
|
|
eprintln!("skipping: no GPU adapter ({e})");
|
|
None
|
|
}
|
|
}
|
|
}
|
|
|
|
/// A flat mid-grey frame, so anything the pass adds is the whole answer.
|
|
fn grey(ctx: &GpuContext) -> DemosaicedImage {
|
|
let data: Vec<u8> = (0..SIZE * SIZE).flat_map(|_| [0u8, 0, 0, 255]).collect();
|
|
DemosaicedImage::from_rgba8(ctx, &data, SIZE, SIZE).expect("upload")
|
|
}
|
|
|
|
/// A pass that sums the instance list into the red channel and writes the
|
|
/// output. Deliberately trivial: the value on screen is then a direct readout
|
|
/// of what arrived in the buffer.
|
|
fn summing_pass(storage: Vec<[f32; 4]>, structure: u64) -> ComposedDetailPass {
|
|
let source = "
|
|
@group(0) @binding(0) var source: texture_2d<f32>;
|
|
struct Params { detail_base: vec4<f32> }
|
|
@group(0) @binding(1) var<uniform> u: Params;
|
|
@group(0) @binding(2) var output: texture_storage_2d<rgba8unorm, write>;
|
|
@group(0) @binding(3) var<storage, read> instances: array<vec4<f32>>;
|
|
|
|
@compute @workgroup_size(8, 8, 1)
|
|
fn main(@builtin(global_invocation_id) gid: vec3<u32>) {
|
|
let dims = textureDimensions(output);
|
|
if (gid.x >= dims.x || gid.y >= dims.y) { return; }
|
|
|
|
// Weighted by index, so a buffer read back to front fails this rather
|
|
// than passing by symmetry.
|
|
var total = 0.0;
|
|
let n = arrayLength(&instances);
|
|
for (var i = 0u; i < n; i = i + 1u) {
|
|
total = total + instances[i].x * f32(i + 1u);
|
|
}
|
|
textureStore(output, vec2<i32>(gid.xy), vec4<f32>(total, f32(n) / 255.0, 0.0, 1.0));
|
|
}
|
|
"
|
|
.to_string();
|
|
|
|
ComposedDetailPass {
|
|
output_scale: 1,
|
|
label: "test/instances".to_string(),
|
|
source,
|
|
uniforms: vec![SIZE as f32, SIZE as f32, 1.0, 0.0],
|
|
storage,
|
|
radius: 0,
|
|
writes_output: true,
|
|
// Any distinct number: the hash is a cache key, and these tests are
|
|
// what decide whether two chains share a pipeline.
|
|
structure_hash: structure,
|
|
}
|
|
}
|
|
|
|
fn render(pass: &mut AdjustPass, source: &DemosaicedImage, chain: &ComposedDetail) -> Vec<u8> {
|
|
// The fused half has to be composed knowing a detail stage follows it, or
|
|
// it encodes its own output and the chain would quantise twice — a mismatch
|
|
// `render_detailed` refuses outright. The probe is the graph that says so;
|
|
// its own passes are not used, since the chain here is hand-built.
|
|
let mut graph = EditGraph::with_detail_probe();
|
|
graph.set_param(
|
|
dr_pipeline::descriptor::OpId("detail_probe"),
|
|
dr_pipeline::descriptor::ParamId("radius"),
|
|
0.05,
|
|
);
|
|
let shader = graph.compose_for(ColourSpace::Srgb);
|
|
let key = graph.invalidation().through(Affects::Colour);
|
|
pass.render_detailed(source, &shader, SIZE, SIZE, None, chain, key)
|
|
.expect("render");
|
|
pass.export_pixels().expect("readback").0
|
|
}
|
|
|
|
/// The list arrives whole, in order, and the shader can tell how long it is.
|
|
#[test]
|
|
fn a_pass_reads_the_list_it_was_given() {
|
|
let Some(ctx) = ctx() else { return };
|
|
let source = grey(&ctx);
|
|
let mut pass = AdjustPass::new(&ctx);
|
|
|
|
// 0.1·1 + 0.2·2 + 0.3·3 = 1.4, which clips to 1.0 — so instead: values
|
|
// chosen to land at a quarter, unambiguously distinguishable from both the
|
|
// "read nothing" answer of 0 and the "read them unweighted" answer of 0.15.
|
|
let chain = ComposedDetail {
|
|
passes: vec![summing_pass(
|
|
vec![[0.05, 0.0, 0.0, 0.0], [0.1, 0.0, 0.0, 0.0]],
|
|
1,
|
|
)],
|
|
};
|
|
|
|
let pixels = render(&mut pass, &source, &chain);
|
|
let (red, green) = (pixels[0], pixels[1]);
|
|
|
|
// 0.05·1 + 0.1·2 = 0.25, written straight to an rgba8 target.
|
|
assert!(
|
|
red.abs_diff((0.25 * 255.0) as u8) <= 1,
|
|
"the shader summed {red}, not the list it was handed"
|
|
);
|
|
assert_eq!(green, 2, "arrayLength saw both entries");
|
|
}
|
|
|
|
/// A convolution declares no list and must still run: it is bound to the
|
|
/// placeholder rather than to nothing, because a zero-length storage buffer
|
|
/// cannot be bound at all and a second bind group layout for the difference
|
|
/// would be two layouts to keep in step.
|
|
#[test]
|
|
fn a_pass_with_no_list_still_runs() {
|
|
let Some(ctx) = ctx() else { return };
|
|
let source = grey(&ctx);
|
|
let mut pass = AdjustPass::new(&ctx);
|
|
|
|
let chain = ComposedDetail {
|
|
passes: vec![summing_pass(Vec::new(), 2)],
|
|
};
|
|
|
|
let pixels = render(&mut pass, &source, &chain);
|
|
assert_eq!(pixels[0], 0, "the placeholder is zeroed");
|
|
assert_eq!(pixels[1], 1, "and is exactly one element long");
|
|
}
|
|
|
|
/// The property that makes placing the tenth spot as cheap as moving a slider:
|
|
/// the list is in the buffer, not in the source, so the pipeline is compiled
|
|
/// once however many entries arrive.
|
|
#[test]
|
|
fn changing_the_list_does_not_recompile() {
|
|
let Some(ctx) = ctx() else { return };
|
|
let source = grey(&ctx);
|
|
let mut pass = AdjustPass::new(&ctx);
|
|
|
|
for count in 1..=6 {
|
|
let list = (0..count).map(|_| [0.01, 0.0, 0.0, 0.0]).collect();
|
|
let chain = ComposedDetail {
|
|
passes: vec![summing_pass(list, 3)],
|
|
};
|
|
render(&mut pass, &source, &chain);
|
|
}
|
|
|
|
assert_eq!(
|
|
pass.cached_detail_pipelines(),
|
|
1,
|
|
"six different lists, one compiled pipeline"
|
|
);
|
|
}
|