Files
DarkRoom/core/dr-gpu/tests/detail_instances.rs
T
dtourolleandClaude Opus 5 bff95e25ad Let clarity's base be computed where it is still fully determined
Clarity's Gaussian sigma is 1.2% of the frame's shorter edge, so its radius
is a property of the viewport: 52 render pixels at 4K, two separable passes
of 105 taps each over 8.3 M pixels. That measured 33.9 ms — seven times the
entire fused point chain, for one slider — and is docs/technical-debt.md TD-4.

A detail pass may now declare `output_scale`, and clarity's base is computed
on a grid a quarter the size on each axis.

The pass that combines needs the blur *and* the full-resolution colour, and a
colour that has been through a quarter-scale target is no longer full
resolution. So a scaled pass cannot simply join the ping-pong: there are two
chains now. The full-resolution one carries the colour and no scaled pass
touches it; the reduced one carries the base and reaches the combining pass
through a second binding as `reduced_at()`.

The reduce is a dispatch of its own rather than something the first blur half
does on the way past, and that is the whole difference between this and the
strided kernel the module documentation rules out. A stride samples an image
that is not band-limited and aliases high-frequency content down into the
base, which is then subtracted, and arrives in the output as mottling across
smooth gradients. This band-limits first and samples after. What is discarded
is content the base could not represent at any resolution, because a Gaussian
at sigma = 26 px holds nothing above one cycle per 26 px and the quarter-scale
grid carries one per 8 — so the reduced base is not an approximation of the
full-resolution one, it is the same function sampled where it is still
determined.

Which is also why the scale belongs to the band rather than to the stage.
Texture's sigma is a decade finer, so the reduce pass's own box would be wider
than the Gaussian it was prefiltering; texture never reduces. And clarity
steps 4 -> 2 -> 1 as sigma falls, because a quarter of a small sigma is not a
Gaussian either — the case that gives up is the one that was already cheap.

`radius` stays in each pass's own pixels and `ComposedDetail::radius` multiplies
it back up, so 13 reduced pixels at scale 4 still report the 52 render pixels a
tile would have to be grown by. The halo a scheduler sees does not move.

The halo tests pass unchanged, which was TD-4's stated bar; they render at
1024 px and so exercise the reduced path rather than stepping around it. Added
`crossing_the_reduction_threshold_does_not_change_the_picture`, because
nothing yet compared the reduced form against a *less* reduced one — every
other test measures one form against itself. It renders the same edit either
side of the 4 -> 2 step-down and holds the peak excursion to 0.03 stops and
the reach to 2% of the frame.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-29 13:12:29 +02:00

171 lines
6.3 KiB
Rust

//! TRACES: FR-DEV-8
//! The instance binding: a detail pass whose work is a list, not a kernel.
//!
//! Spot removal needs a detail pass to read a variable number of records —
//! sixty-four repairs and one repair are the same shader with a different
//! buffer behind it. That is binding 3, and this file proves the three things
//! about it that a picture would not tell you clearly:
//!
//! - the data uploaded is the data the shader reads, in order;
//! - a pass that declares no list still runs, bound to the placeholder;
//! - the same shader with a *different* list does not recompile, which is what
//! keeps placing a spot as cheap as moving a slider.
//!
//! The passes here are synthetic on purpose. `spot_removal.rs` asserts the
//! repair; this asserts the plumbing, so a failure in one does not have to be
//! read to work out which of the two broke.
use dr_gpu::{AdjustPass, DemosaicedImage, GpuContext};
use dr_pipeline::detail::{ComposedDetail, ComposedDetailPass};
use dr_pipeline::{Affects, EditGraph};
use dr_types::ColourSpace;
const SIZE: u32 = 8;
fn ctx() -> Option<GpuContext> {
match pollster::block_on(GpuContext::new_headless()) {
Ok(c) => Some(c),
Err(e) => {
eprintln!("skipping: no GPU adapter ({e})");
None
}
}
}
/// A flat mid-grey frame, so anything the pass adds is the whole answer.
fn grey(ctx: &GpuContext) -> DemosaicedImage {
let data: Vec<u8> = (0..SIZE * SIZE).flat_map(|_| [0u8, 0, 0, 255]).collect();
DemosaicedImage::from_rgba8(ctx, &data, SIZE, SIZE).expect("upload")
}
/// A pass that sums the instance list into the red channel and writes the
/// output. Deliberately trivial: the value on screen is then a direct readout
/// of what arrived in the buffer.
fn summing_pass(storage: Vec<[f32; 4]>, structure: u64) -> ComposedDetailPass {
let source = "
@group(0) @binding(0) var source: texture_2d<f32>;
struct Params { detail_base: vec4<f32> }
@group(0) @binding(1) var<uniform> u: Params;
@group(0) @binding(2) var output: texture_storage_2d<rgba8unorm, write>;
@group(0) @binding(3) var<storage, read> instances: array<vec4<f32>>;
@compute @workgroup_size(8, 8, 1)
fn main(@builtin(global_invocation_id) gid: vec3<u32>) {
let dims = textureDimensions(output);
if (gid.x >= dims.x || gid.y >= dims.y) { return; }
// Weighted by index, so a buffer read back to front fails this rather
// than passing by symmetry.
var total = 0.0;
let n = arrayLength(&instances);
for (var i = 0u; i < n; i = i + 1u) {
total = total + instances[i].x * f32(i + 1u);
}
textureStore(output, vec2<i32>(gid.xy), vec4<f32>(total, f32(n) / 255.0, 0.0, 1.0));
}
"
.to_string();
ComposedDetailPass {
output_scale: 1,
label: "test/instances".to_string(),
source,
uniforms: vec![SIZE as f32, SIZE as f32, 1.0, 0.0],
storage,
radius: 0,
writes_output: true,
// Any distinct number: the hash is a cache key, and these tests are
// what decide whether two chains share a pipeline.
structure_hash: structure,
}
}
fn render(pass: &mut AdjustPass, source: &DemosaicedImage, chain: &ComposedDetail) -> Vec<u8> {
// The fused half has to be composed knowing a detail stage follows it, or
// it encodes its own output and the chain would quantise twice — a mismatch
// `render_detailed` refuses outright. The probe is the graph that says so;
// its own passes are not used, since the chain here is hand-built.
let mut graph = EditGraph::with_detail_probe();
graph.set_param(
dr_pipeline::descriptor::OpId("detail_probe"),
dr_pipeline::descriptor::ParamId("radius"),
0.05,
);
let shader = graph.compose_for(ColourSpace::Srgb);
let key = graph.invalidation().through(Affects::Colour);
pass.render_detailed(source, &shader, SIZE, SIZE, None, chain, key)
.expect("render");
pass.export_pixels().expect("readback").0
}
/// The list arrives whole, in order, and the shader can tell how long it is.
#[test]
fn a_pass_reads_the_list_it_was_given() {
let Some(ctx) = ctx() else { return };
let source = grey(&ctx);
let mut pass = AdjustPass::new(&ctx);
// 0.1·1 + 0.2·2 + 0.3·3 = 1.4, which clips to 1.0 — so instead: values
// chosen to land at a quarter, unambiguously distinguishable from both the
// "read nothing" answer of 0 and the "read them unweighted" answer of 0.15.
let chain = ComposedDetail {
passes: vec![summing_pass(
vec![[0.05, 0.0, 0.0, 0.0], [0.1, 0.0, 0.0, 0.0]],
1,
)],
};
let pixels = render(&mut pass, &source, &chain);
let (red, green) = (pixels[0], pixels[1]);
// 0.05·1 + 0.1·2 = 0.25, written straight to an rgba8 target.
assert!(
red.abs_diff((0.25 * 255.0) as u8) <= 1,
"the shader summed {red}, not the list it was handed"
);
assert_eq!(green, 2, "arrayLength saw both entries");
}
/// A convolution declares no list and must still run: it is bound to the
/// placeholder rather than to nothing, because a zero-length storage buffer
/// cannot be bound at all and a second bind group layout for the difference
/// would be two layouts to keep in step.
#[test]
fn a_pass_with_no_list_still_runs() {
let Some(ctx) = ctx() else { return };
let source = grey(&ctx);
let mut pass = AdjustPass::new(&ctx);
let chain = ComposedDetail {
passes: vec![summing_pass(Vec::new(), 2)],
};
let pixels = render(&mut pass, &source, &chain);
assert_eq!(pixels[0], 0, "the placeholder is zeroed");
assert_eq!(pixels[1], 1, "and is exactly one element long");
}
/// The property that makes placing the tenth spot as cheap as moving a slider:
/// the list is in the buffer, not in the source, so the pipeline is compiled
/// once however many entries arrive.
#[test]
fn changing_the_list_does_not_recompile() {
let Some(ctx) = ctx() else { return };
let source = grey(&ctx);
let mut pass = AdjustPass::new(&ctx);
for count in 1..=6 {
let list = (0..count).map(|_| [0.01, 0.0, 0.0, 0.0]).collect();
let chain = ComposedDetail {
passes: vec![summing_pass(list, 3)],
};
render(&mut pass, &source, &chain);
}
assert_eq!(
pass.cached_detail_pipelines(),
1,
"six different lists, one compiled pipeline"
);
}