Run the view transform after the detail stage, in a pass of its own

The fused pass stops at "linear working values" when a sharpener, a
blur or a repair follows, and the detail passes convolve what it hands
on. Until now it handed on the rendering: the base curve, and since the
last commit the view transform, ran before the store. So every kernel
worked on display-referred values while its comments promised the
opposite — D19's second finding.

A fused pass composed for a detail stage now stops before the view
transform, and carries a second shader, `ComposedShader::view`, composed
from the same inputs. It runs the same prologue, for the positions a
fragment reads (a film's grain seeds from `source_px`) and the corners
it blacks out, takes its colour from the detail stage's result bound
where the sample cache would be, and runs the view transform, the
output transform and the mask reveal. `render_detailed` dispatches it
after the last detail pass, in the same encoder.

So no detail pass encodes any more. Every pass writes an intermediate,
the last one included, which retires three things that existed only to
make the last pass encode: `writes_output` and the runner's second
layout, the body-less resolve pass for an active kernel with nothing to
draw at this scale, and capture sharpening's pass-through, which now
emits no pass at all. An empty chain is a whole render: the view pass
reads the fused result directly. The detail stage no longer takes an
output space either, so `compose_detail_for` folds into
`compose_detail` and the space is named once, on the fused half.

The cost is one full-render read and write per frame when a detail
stage exists, and a third intermediate for a one-pass chain.
This commit is contained in:
2026-09-27 16:52:54 -04:00
parent 92afaebd34
commit c07f81edcb
19 changed files with 618 additions and 566 deletions
+22 -10
View File
@@ -38,15 +38,17 @@ fn grey(ctx: &GpuContext) -> DemosaicedImage {
DemosaicedImage::from_rgba8(ctx, &data, SIZE, SIZE).expect("upload")
}
/// A pass that sums the instance list into the red channel and writes the
/// output. Deliberately trivial: the value on screen is then a direct readout
/// of what arrived in the buffer.
/// A pass that sums the instance list into the red channel and writes a
/// linear intermediate, which the view pass then encodes (D19). Deliberately
/// trivial: the value on screen is then a direct readout of what arrived in the
/// buffer, through the sRGB encode — the source is an 8-bit upload, so the
/// view transform is skipped for it and the encode is the only thing between.
fn summing_pass(storage: Vec<[f32; 4]>, structure: u64) -> ComposedDetailPass {
let source = "
@group(0) @binding(0) var source: texture_2d<f32>;
struct Params { detail_base: vec4<f32> }
@group(0) @binding(1) var<uniform> u: Params;
@group(0) @binding(2) var output: texture_storage_2d<rgba8unorm, write>;
@group(0) @binding(2) var output: texture_storage_2d<rgba16float, write>;
@group(0) @binding(3) var<storage, read> instances: array<vec4<f32>>;
@compute @workgroup_size(8, 8, 1)
@@ -61,7 +63,7 @@ fn main(@builtin(global_invocation_id) gid: vec3<u32>) {
for (var i = 0u; i < n; i = i + 1u) {
total = total + instances[i].x * f32(i + 1u);
}
textureStore(output, vec2<i32>(gid.xy), vec4<f32>(total, f32(n) / 255.0, 0.0, 1.0));
textureStore(output, vec2<i32>(gid.xy), vec4<f32>(total, f32(n) * 0.1, 0.0, 1.0));
}
"
.to_string();
@@ -73,13 +75,17 @@ fn main(@builtin(global_invocation_id) gid: vec3<u32>) {
uniforms: vec![SIZE as f32, SIZE as f32, 1.0, 0.0],
storage,
radius: 0,
writes_output: true,
// Any distinct number: the hash is a cache key, and these tests are
// what decide whether two chains share a pipeline.
structure_hash: structure,
}
}
/// A linear value as the view pass leaves it in the 8-bit output.
fn encoded(linear: f32) -> u8 {
(dr_types::Transfer::Srgb.encode(linear) * 255.0).round() as u8
}
fn render(pass: &mut AdjustPass, source: &DemosaicedImage, chain: &ComposedDetail) -> Vec<u8> {
// The fused half has to be composed knowing a detail stage follows it, or
// it encodes its own output and the chain would quantise twice — a mismatch
@@ -118,12 +124,15 @@ fn a_pass_reads_the_list_it_was_given() {
let pixels = render(&mut pass, &source, &chain);
let (red, green) = (pixels[0], pixels[1]);
// 0.05·1 + 0.1·2 = 0.25, written straight to an rgba8 target.
// 0.05·1 + 0.1·2 = 0.25.
assert!(
red.abs_diff((0.25 * 255.0) as u8) <= 1,
red.abs_diff(encoded(0.25)) <= 1,
"the shader summed {red}, not the list it was handed"
);
assert_eq!(green, 2, "arrayLength saw both entries");
assert!(
green.abs_diff(encoded(0.2)) <= 1,
"arrayLength saw both entries"
);
}
/// A convolution declares no list and must still run: it is bound to the
@@ -142,7 +151,10 @@ fn a_pass_with_no_list_still_runs() {
let pixels = render(&mut pass, &source, &chain);
assert_eq!(pixels[0], 0, "the placeholder is zeroed");
assert_eq!(pixels[1], 1, "and is exactly one element long");
assert!(
pixels[1].abs_diff(encoded(0.1)) <= 1,
"and is exactly one element long"
);
}
/// The property that makes placing the tenth spot as cheap as moving a slider: