Run the view transform after the detail stage, in a pass of its own
The fused pass stops at "linear working values" when a sharpener, a blur or a repair follows, and the detail passes convolve what it hands on. Until now it handed on the rendering: the base curve, and since the last commit the view transform, ran before the store. So every kernel worked on display-referred values while its comments promised the opposite — D19's second finding. A fused pass composed for a detail stage now stops before the view transform, and carries a second shader, `ComposedShader::view`, composed from the same inputs. It runs the same prologue, for the positions a fragment reads (a film's grain seeds from `source_px`) and the corners it blacks out, takes its colour from the detail stage's result bound where the sample cache would be, and runs the view transform, the output transform and the mask reveal. `render_detailed` dispatches it after the last detail pass, in the same encoder. So no detail pass encodes any more. Every pass writes an intermediate, the last one included, which retires three things that existed only to make the last pass encode: `writes_output` and the runner's second layout, the body-less resolve pass for an active kernel with nothing to draw at this scale, and capture sharpening's pass-through, which now emits no pass at all. An empty chain is a whole render: the view pass reads the fused result directly. The detail stage no longer takes an output space either, so `compose_detail_for` folds into `compose_detail` and the space is named once, on the fused half. The cost is one full-render read and write per frame when a detail stage exists, and a third intermediate for a one-pass chain.
This commit is contained in:
@@ -38,15 +38,17 @@ fn grey(ctx: &GpuContext) -> DemosaicedImage {
|
||||
DemosaicedImage::from_rgba8(ctx, &data, SIZE, SIZE).expect("upload")
|
||||
}
|
||||
|
||||
/// A pass that sums the instance list into the red channel and writes the
|
||||
/// output. Deliberately trivial: the value on screen is then a direct readout
|
||||
/// of what arrived in the buffer.
|
||||
/// A pass that sums the instance list into the red channel and writes a
|
||||
/// linear intermediate, which the view pass then encodes (D19). Deliberately
|
||||
/// trivial: the value on screen is then a direct readout of what arrived in the
|
||||
/// buffer, through the sRGB encode — the source is an 8-bit upload, so the
|
||||
/// view transform is skipped for it and the encode is the only thing between.
|
||||
fn summing_pass(storage: Vec<[f32; 4]>, structure: u64) -> ComposedDetailPass {
|
||||
let source = "
|
||||
@group(0) @binding(0) var source: texture_2d<f32>;
|
||||
struct Params { detail_base: vec4<f32> }
|
||||
@group(0) @binding(1) var<uniform> u: Params;
|
||||
@group(0) @binding(2) var output: texture_storage_2d<rgba8unorm, write>;
|
||||
@group(0) @binding(2) var output: texture_storage_2d<rgba16float, write>;
|
||||
@group(0) @binding(3) var<storage, read> instances: array<vec4<f32>>;
|
||||
|
||||
@compute @workgroup_size(8, 8, 1)
|
||||
@@ -61,7 +63,7 @@ fn main(@builtin(global_invocation_id) gid: vec3<u32>) {
|
||||
for (var i = 0u; i < n; i = i + 1u) {
|
||||
total = total + instances[i].x * f32(i + 1u);
|
||||
}
|
||||
textureStore(output, vec2<i32>(gid.xy), vec4<f32>(total, f32(n) / 255.0, 0.0, 1.0));
|
||||
textureStore(output, vec2<i32>(gid.xy), vec4<f32>(total, f32(n) * 0.1, 0.0, 1.0));
|
||||
}
|
||||
"
|
||||
.to_string();
|
||||
@@ -73,13 +75,17 @@ fn main(@builtin(global_invocation_id) gid: vec3<u32>) {
|
||||
uniforms: vec![SIZE as f32, SIZE as f32, 1.0, 0.0],
|
||||
storage,
|
||||
radius: 0,
|
||||
writes_output: true,
|
||||
// Any distinct number: the hash is a cache key, and these tests are
|
||||
// what decide whether two chains share a pipeline.
|
||||
structure_hash: structure,
|
||||
}
|
||||
}
|
||||
|
||||
/// A linear value as the view pass leaves it in the 8-bit output.
|
||||
fn encoded(linear: f32) -> u8 {
|
||||
(dr_types::Transfer::Srgb.encode(linear) * 255.0).round() as u8
|
||||
}
|
||||
|
||||
fn render(pass: &mut AdjustPass, source: &DemosaicedImage, chain: &ComposedDetail) -> Vec<u8> {
|
||||
// The fused half has to be composed knowing a detail stage follows it, or
|
||||
// it encodes its own output and the chain would quantise twice — a mismatch
|
||||
@@ -118,12 +124,15 @@ fn a_pass_reads_the_list_it_was_given() {
|
||||
let pixels = render(&mut pass, &source, &chain);
|
||||
let (red, green) = (pixels[0], pixels[1]);
|
||||
|
||||
// 0.05·1 + 0.1·2 = 0.25, written straight to an rgba8 target.
|
||||
// 0.05·1 + 0.1·2 = 0.25.
|
||||
assert!(
|
||||
red.abs_diff((0.25 * 255.0) as u8) <= 1,
|
||||
red.abs_diff(encoded(0.25)) <= 1,
|
||||
"the shader summed {red}, not the list it was handed"
|
||||
);
|
||||
assert_eq!(green, 2, "arrayLength saw both entries");
|
||||
assert!(
|
||||
green.abs_diff(encoded(0.2)) <= 1,
|
||||
"arrayLength saw both entries"
|
||||
);
|
||||
}
|
||||
|
||||
/// A convolution declares no list and must still run: it is bound to the
|
||||
@@ -142,7 +151,10 @@ fn a_pass_with_no_list_still_runs() {
|
||||
|
||||
let pixels = render(&mut pass, &source, &chain);
|
||||
assert_eq!(pixels[0], 0, "the placeholder is zeroed");
|
||||
assert_eq!(pixels[1], 1, "and is exactly one element long");
|
||||
assert!(
|
||||
pixels[1].abs_diff(encoded(0.1)) <= 1,
|
||||
"and is exactly one element long"
|
||||
);
|
||||
}
|
||||
|
||||
/// The property that makes placing the tenth spot as cheap as moving a slider:
|
||||
|
||||
Reference in New Issue
Block a user