Run the view transform after the detail stage, in a pass of its own

The fused pass stops at "linear working values" when a sharpener, a
blur or a repair follows, and the detail passes convolve what it hands
on. Until now it handed on the rendering: the base curve, and since the
last commit the view transform, ran before the store. So every kernel
worked on display-referred values while its comments promised the
opposite — D19's second finding.

A fused pass composed for a detail stage now stops before the view
transform, and carries a second shader, `ComposedShader::view`, composed
from the same inputs. It runs the same prologue, for the positions a
fragment reads (a film's grain seeds from `source_px`) and the corners
it blacks out, takes its colour from the detail stage's result bound
where the sample cache would be, and runs the view transform, the
output transform and the mask reveal. `render_detailed` dispatches it
after the last detail pass, in the same encoder.

So no detail pass encodes any more. Every pass writes an intermediate,
the last one included, which retires three things that existed only to
make the last pass encode: `writes_output` and the runner's second
layout, the body-less resolve pass for an active kernel with nothing to
draw at this scale, and capture sharpening's pass-through, which now
emits no pass at all. An empty chain is a whole render: the view pass
reads the fused result directly. The detail stage no longer takes an
output space either, so `compose_detail_for` folds into
`compose_detail` and the space is named once, on the fused half.

The cost is one full-render read and write per frame when a detail
stage exists, and a third intermediate for a one-pass chain.
This commit is contained in:
2026-09-27 16:52:54 -04:00
parent 92afaebd34
commit c07f81edcb
19 changed files with 618 additions and 566 deletions
+191 -64
View File
@@ -562,6 +562,17 @@ pub struct ComposedShader {
/// the interpolating paths, whose sample is a blend of four texels and
/// not representable exactly in the source's own format.
pub sample_key: Option<u64>,
/// TRACES: FR-DEV-3j
/// The view pass, for a shader that stops at [`OutputMode::LinearWorking`].
///
/// The view transform runs after the detail stage (D19): a sharpener or
/// a blur must convolve scene-linear values, not a display rendering. So
/// when a detail stage follows, this shader stops short of the view
/// transform, and this second shader — the same prologue for the
/// positions a fragment reads, the colour taken from the detail stage's
/// result bound as `sampled`, then the view transform and the output
/// transform — writes the display texture. `None` for every other mode.
pub view: Option<Box<ComposedShader>>,
}
/// Fields the generated uniform struct always carries, before op uniforms.
@@ -700,7 +711,7 @@ pub fn compose_full_revealing(
reveal: Option<&crate::mask::Reveal>,
) -> ComposedShader {
compose_inner(
ops, framing, output, masks, spots, warps, reveal, None, false,
ops, framing, output, masks, spots, warps, reveal, None, false, false,
)
}
@@ -770,6 +781,7 @@ fn compose_camera_tap(
None,
Some(OutputMode::CameraLinear),
smooth,
false,
)
}
@@ -784,6 +796,7 @@ fn compose_inner(
reveal: Option<&crate::mask::Reveal>,
forced: Option<OutputMode>,
smooth: bool,
view_pass: bool,
) -> ComposedShader {
// The lens corrections, composed into one coordinate transform. Beside
// `framing` because they are the other half of the same stage: framing
@@ -819,13 +832,22 @@ fn compose_inner(
// sense above: `compose_camera_probe` is the only function that passes
// it, with an empty operation list, and the mode it forces has its own
// storage format and its own render entry on the GPU side.
let output_mode = forced.unwrap_or(
if ops.iter().any(|o| o.is_active() && o.detail().is_some()) || !spots.is_neutral() {
OutputMode::LinearWorking
} else {
OutputMode::Encoded
},
);
//
// TRACES: FR-DEV-3j
// The view pass is the other exception. It follows the detail stage and
// writes the display texture, so it is `Encoded` whatever the chain holds:
// it exists only because a detail stage does.
let output_mode = if view_pass {
OutputMode::Encoded
} else {
forced.unwrap_or(
if ops.iter().any(|o| o.is_active() && o.detail().is_some()) || !spots.is_neutral() {
OutputMode::LinearWorking
} else {
OutputMode::Encoded
},
)
};
// Whether an operation has taken over the rendering. Decided from the
// operations for the same reason `output_mode` is: a caller that got it
@@ -862,6 +884,11 @@ fn compose_inner(
// Whether the composer emits a view transform — which is every render but
// the camera-space tap and one a rendering operation has taken over.
let views = !op_renders && output_mode != OutputMode::CameraLinear;
// And whether *this* shader is where it goes. Not the fused pass when a
// detail stage follows: the view transform maps into a display range, and
// a sharpener handed a display range is what D19 set out to stop. The
// view pass after the detail stage carries it instead.
let views_here = views && output_mode == OutputMode::Encoded;
// Framing's block follows the base one at a fixed offset, for the same
// reason: the prologue is emitted whether or not any operation is active,
@@ -917,7 +944,7 @@ fn compose_inner(
// none — `compose(&[])` in a test, or a probe. Absent altogether where
// nothing is to be rendered: see `views`.
let default_view = crate::ops::ViewTransform::new();
let view: Option<&dyn Operation> = views.then(|| {
let view: Option<&dyn Operation> = views_here.then(|| {
point(Stage::View)
.next()
.unwrap_or(&default_view as &dyn Operation)
@@ -930,12 +957,17 @@ fn compose_inner(
Matrix,
Orphans,
}
let steps = point(Stage::Camera)
.map(Step::Op)
.chain(std::iter::once(Step::Matrix))
.chain(point(Stage::Scene).map(Step::Op))
.chain(std::iter::once(Step::Orphans))
.chain(view.map(Step::Op));
//
// The view pass holds the view transform and nothing else: everything
// before it has already run, in the fused pass and the detail stage.
let mut steps: Vec<Step> = Vec::new();
if !view_pass {
steps.extend(point(Stage::Camera).map(Step::Op));
steps.push(Step::Matrix);
steps.extend(point(Stage::Scene).map(Step::Op));
steps.push(Step::Orphans);
}
steps.extend(view.map(Step::Op));
for step in steps {
let op = match step {
@@ -1021,7 +1053,15 @@ fn compose_inner(
// rather than among the operations — see `mask::LayerShader::reveal`.
// Empty for every composition nobody is looking at a mask through, which
// is all of them but the screen's.
let reveal_block = layers.reveal.clone();
//
// And after the detail stage, when there is one: on the linear
// intermediate a flat tint would be sharpened and then rendered as some
// other colour. So it rides on whichever shader writes the display.
let reveal_block = if output_mode == OutputMode::Encoded {
layers.reveal.clone()
} else {
String::new()
};
uniform_fields.push_str(&layers.uniform_fields);
uniform_values.extend_from_slice(&layers.uniform_values);
for h in &layers.helpers {
@@ -1157,6 +1197,70 @@ fn compose_inner(
// Formatted with Rust's `Display` so the shader reads the same threshold
// the probe checks against; see `CLIP_ONSET`.
let clip_onset = CLIP_ONSET;
// What the operations are handed. For the fused pass, the source texel
// made linear and balanced as shot; for the view pass, the scene as the
// detail stage left it, which is already all of that.
let head = if view_pass {
" // The view pass (FR-DEV-3j): everything up to the view transform ran
// in the fused pass and the detail stage, and what they left is in the
// texture bound as `sampled`, at this pixel. The prologue above ran only
// for the positions it publishes — `source_px` for a film's grain, and
// the corners it blacks out — and its colour is discarded.
let non_linear = u.as_shot_wb.w > 0.5;
c = textureLoad(sampled, vec2<i32>(gid.xy), 0).rgb;
"
.to_string()
} else {
format!(
" // A non-linear source is already display-encoded; undo that so the
// operations below see linear colour whatever the source was.
let non_linear = u.as_shot_wb.w > 0.5;
if (non_linear) {{
c = decode_srgb(c);
}}
// As-shot white balance. Applied unconditionally, before any operation,
// because it is part of *interpreting* the sensor rather than an edit: a
// Bayer sensor's green photosites collect far more signal than its red
// and blue, so raw camera-space values are strongly green and no amount
// of later correction recovers a neutral image from them. The white
// balance operation, when active, applies its own offset on top of this.
//
// A non-linear source has already had this applied in-camera; the uniform
// is neutral there, so this is a multiply by one rather than a branch.
// How close this pixel was to saturation before any balance was applied.
// A photosite at its white level carries no colour information — every
// channel simply stopped counting — so the balance below must not be
// allowed to tint it.
let clipped = smoothstep({clip_onset}, 1.0, max(c.r, max(c.g, c.b)));
c = c * u.as_shot_wb.rgb;
// **Highlight desaturation, and without it every blown sky is magenta.**
//
// A fully clipped pixel arrives as (1, 1, 1). The as-shot multipliers are
// not neutral — on a Canon 6D they are (1.93, 1.00, 1.68) — so balancing
// sends it to exactly that, and the camera matrix then produces R 2.88,
// G 0.51, B 2.03. Red and blue clip at one and green does not, which is
// magenta. The balance is correct; the input was not a colour.
//
// So a saturated pixel is pulled back toward the neutral its raw values
// actually represent, fading in over the last 1.5% of range. Smoothly,
// because a hard switch puts a visible edge around every highlight where
// the two treatments meet — a rim light on skin is the worst case, and it
// is the one people notice.
//
// The neutral chosen is the balanced grey of the same brightness, so the
// highlight keeps its luminance and loses only the cast.
if (clipped > 0.0) {{
let neutral = vec3<f32>(max(c.r, max(c.g, c.b)));
c = mix(c, neutral, clipped);
}}
"
)
};
let source = format!(
"// GENERATED — do not edit.
//
@@ -1213,51 +1317,7 @@ fn main(@builtin(global_invocation_id) gid: vec3<u32>) {{
}}
{prologue}
// A non-linear source is already display-encoded; undo that so the
// operations below see linear colour whatever the source was.
let non_linear = u.as_shot_wb.w > 0.5;
if (non_linear) {{
c = decode_srgb(c);
}}
// As-shot white balance. Applied unconditionally, before any operation,
// because it is part of *interpreting* the sensor rather than an edit: a
// Bayer sensor's green photosites collect far more signal than its red
// and blue, so raw camera-space values are strongly green and no amount
// of later correction recovers a neutral image from them. The white
// balance operation, when active, applies its own offset on top of this.
//
// A non-linear source has already had this applied in-camera; the uniform
// is neutral there, so this is a multiply by one rather than a branch.
// How close this pixel was to saturation before any balance was applied.
// A photosite at its white level carries no colour information — every
// channel simply stopped counting — so the balance below must not be
// allowed to tint it.
let clipped = smoothstep({clip_onset}, 1.0, max(c.r, max(c.g, c.b)));
c = c * u.as_shot_wb.rgb;
// **Highlight desaturation, and without it every blown sky is magenta.**
//
// A fully clipped pixel arrives as (1, 1, 1). The as-shot multipliers are
// not neutral — on a Canon 6D they are (1.93, 1.00, 1.68) — so balancing
// sends it to exactly that, and the camera matrix then produces R 2.88,
// G 0.51, B 2.03. Red and blue clip at one and green does not, which is
// magenta. The balance is correct; the input was not a colour.
//
// So a saturated pixel is pulled back toward the neutral its raw values
// actually represent, fading in over the last 1.5% of range. Smoothly,
// because a hard switch puts a visible edge around every highlight where
// the two treatments meet — a rim light on skin is the worst case, and it
// is the one people notice.
//
// The neutral chosen is the balanced grey of the same brightness, so the
// highlight keeps its luminance and loses only the cast.
if (clipped > 0.0) {{
let neutral = vec3<f32>(max(c.r, max(c.g, c.b)));
c = mix(c, neutral, clipped);
}}
{body}
{head}{body}
{rendering_tail}{to_output}{reveal_block}
{store}
}}
@@ -1294,12 +1354,25 @@ fn main(@builtin(global_invocation_id) gid: vec3<u32>) {{
output as u64,
);
// TRACES: FR-DEV-3j
// The fused pass that stops for a detail stage hands its view transform
// to a pass of its own, composed here from the same inputs so the two
// halves cannot come from different edits.
let view = (output_mode == OutputMode::LinearWorking).then(|| {
Box::new(compose_inner(
ops, framing, output, masks, spots, warps, reveal, None, smooth, true,
))
});
ComposedShader {
source,
uniforms: uniform_values,
structure_hash,
output_mode,
sample_key,
// The view pass reads the intermediate, not the source, so the
// sample cache has nothing to say about it.
sample_key: if view_pass { None } else { sample_key },
view,
}
}
@@ -2512,6 +2585,54 @@ mod tests {
}
}
#[test]
fn a_detail_stage_puts_the_view_transform_after_it() {
// TRACES: FR-DEV-3j | FR-DEV-2
// D19's second finding: the fused pass stopped at "linear working
// values" for a detail stage, but after the base curve, so every
// sharpener convolved a display rendering. The fused pass must now
// stop *before* the view transform, and the view pass — which reads
// the detail stage's result — carries it, the output transform and the
// mask reveal, and nothing else.
let mut ops = crate::ops::chain();
ops.push(Box::new(crate::detail::probe::BoxBlur::with_radius(0.05)));
let mut exposure = crate::ops::Exposure::new();
exposure.set_param(crate::ops::exposure::EXPOSURE, 1.0);
ops.push(Box::new(exposure));
let fused = compose(&ops);
assert_eq!(fused.output_mode, OutputMode::LinearWorking);
assert!(!fused.source.contains("view_sigmoid"));
assert!(fused.source.contains("---- exposure ----"));
let view = fused.view.as_deref().expect("a view pass");
assert_eq!(view.output_mode, OutputMode::Encoded);
assert!(view.view.is_none());
assert!(view
.source
.contains("c = textureLoad(sampled, vec2<i32>(gid.xy), 0).rgb;"));
assert!(view.source.contains("c = view_sigmoid("));
assert!(view.source.contains("fn encode_output"));
assert_eq!(
view.source.matches("---- ").count(),
1,
"the view pass must not run an operation the fused pass already ran"
);
assert!(!view.source.contains("u.as_shot_wb.rgb"));
assert!(!view.source.contains("camera profile: the matrix"));
assert!(view.sample_key.is_none());
}
#[test]
fn an_edit_with_no_detail_stage_has_no_view_pass() {
// TRACES: FR-DEV-3j
// The common case costs what it always did: one dispatch.
let fused = compose(&crate::ops::chain());
assert_eq!(fused.output_mode, OutputMode::Encoded);
assert!(fused.view.is_none());
assert!(fused.source.contains("c = view_sigmoid("));
}
#[test]
fn a_detail_operation_never_contributes_a_fused_uniform() {
// Slot order in the generated block is emission order, and nothing
@@ -2520,9 +2641,15 @@ mod tests {
// every later operation's uniforms out from under its shader. The
// filter in `compose_full` prevents it; this is the assertion that the
// filter is on the right side of the loop.
//
// Asserted on the field names rather than on a count: since D19 the
// edit with a detail operation also moves the view transform's
// uniforms out of this shader and into its view pass, so the two
// blocks differ in size for a reason that is not this one.
let mut ops = crate::ops::chain();
ops.push(Box::new(crate::detail::probe::BoxBlur::with_radius(0.05)));
let before = compose(&crate::ops::chain()).uniforms.len();
assert_eq!(compose(&ops).uniforms.len(), before);
let fused = compose(&ops);
assert!(!fused.source.contains("detail_probe_"), "{}", fused.source);
assert!(fused.uniforms.len() <= compose(&crate::ops::chain()).uniforms.len());
}
}