Run the view transform after the detail stage, in a pass of its own

The fused pass stops at "linear working values" when a sharpener, a
blur or a repair follows, and the detail passes convolve what it hands
on. Until now it handed on the rendering: the base curve, and since the
last commit the view transform, ran before the store. So every kernel
worked on display-referred values while its comments promised the
opposite — D19's second finding.

A fused pass composed for a detail stage now stops before the view
transform, and carries a second shader, `ComposedShader::view`, composed
from the same inputs. It runs the same prologue, for the positions a
fragment reads (a film's grain seeds from `source_px`) and the corners
it blacks out, takes its colour from the detail stage's result bound
where the sample cache would be, and runs the view transform, the
output transform and the mask reveal. `render_detailed` dispatches it
after the last detail pass, in the same encoder.

So no detail pass encodes any more. Every pass writes an intermediate,
the last one included, which retires three things that existed only to
make the last pass encode: `writes_output` and the runner's second
layout, the body-less resolve pass for an active kernel with nothing to
draw at this scale, and capture sharpening's pass-through, which now
emits no pass at all. An empty chain is a whole render: the view pass
reads the fused result directly. The detail stage no longer takes an
output space either, so `compose_detail_for` folds into
`compose_detail` and the space is named once, on the fused half.

The cost is one full-render read and write per frame when a detail
stage exists, and a third intermediate for a one-pass chain.
This commit is contained in:
2026-09-27 16:52:54 -04:00
parent 92afaebd34
commit c07f81edcb
19 changed files with 618 additions and 566 deletions
+122 -30
View File
@@ -122,6 +122,8 @@ pub struct AdjustPass {
colour_dispatches: usize,
/// Detail dispatches encoded.
detail_dispatches: usize,
/// View passes encoded — one per render with a detail stage (D19).
view_dispatches: usize,
}
struct Target {
@@ -653,6 +655,7 @@ impl AdjustPass {
sample: SampleCache::new(ctx),
colour_dispatches: 0,
detail_dispatches: 0,
view_dispatches: 0,
}
}
@@ -1048,13 +1051,14 @@ impl AdjustPass {
/// Render one frame with a neighbourhood stage.
///
/// `shader` and `detail` must be the two halves of **one** composition —
/// `EditGraph::compose_for` and `EditGraph::compose_detail_for` on the same
/// graph, at the same output space. The fused pass stops at linear working
/// values when a detail stage exists and the last detail pass performs the
/// output transform, so a mismatched pair either encodes twice or not at
/// all.
/// `EditGraph::compose_for` and `EditGraph::compose_detail` on the same
/// graph. The fused pass stops at linear working values when a detail
/// stage exists, and its view pass ([`ComposedShader::view`]) performs the
/// view transform and the output transform after the last detail pass
/// (D19), so a mismatched pair either encodes twice or not at all.
///
/// An empty `detail` falls through to [`Self::render_masked`], which is
/// An encoded shader with an empty `detail` falls through to
/// [`Self::render_masked`], which is
/// the honest thing to do rather than an optimisation: an edit with no
/// active sharpening *is* an ordinary edit, and it should cost exactly
/// what one costs.
@@ -1094,17 +1098,25 @@ impl AdjustPass {
detail: &ComposedDetail,
colour_key: u64,
) -> Result<&wgpu::Texture, GpuError> {
if detail.is_empty() {
// An edit with no detail stage encodes in the fused pass, and is an
// ordinary render. Decided from the shader rather than from the chain:
// an active kernel too fine for this render emits no pass, and its
// fused pass has still stopped at linear values for the view pass.
if shader.output_mode == OutputMode::Encoded && detail.is_empty() {
return self.render_masked(source, shader, width, height, masks);
}
if shader.output_mode != OutputMode::LinearWorking {
return Err(GpuError::ShaderCompilation(
"this detail chain expects a fused pass composed to hand on \
linear working values, but the shader given encodes its own \
output; compose both halves from the same graph"
.into(),
));
}
let view = match (shader.output_mode, shader.view.as_deref()) {
(OutputMode::LinearWorking, Some(view)) => view,
_ => {
return Err(GpuError::ShaderCompilation(
"this detail chain expects a fused pass composed to hand on \
linear working values, with its view pass, but the shader \
given encodes its own output; compose both halves from the \
same graph"
.into(),
));
}
};
let (width, height) = (width.max(1), height.max(1));
self.ensure_target(width, height);
@@ -1117,6 +1129,7 @@ impl AdjustPass {
// both want `&mut self`, and the second holds its borrow across the
// encode below.
self.pipeline(shader)?;
self.pipeline(view)?;
let colour_view = self
.detail
.colour_target(detail.len(), width, height)
@@ -1207,20 +1220,12 @@ impl AdjustPass {
self.colour_dispatches += 1;
}
// One encoder for the colour pass and every detail pass, submitted
// once — the shape `MaskPass::render` established. Submission order is
// the whole of the synchronisation: each pass reads what the previous
// one wrote, through the same queue.
let target_view = self.targets[self.current]
.as_ref()
.expect("ensured above")
.view
.clone();
let ran = match self
.detail
.encode(&mut enc, detail, &target_view, width, height)
{
Ok(ran) => ran,
// One encoder for the colour pass, every detail pass and the view pass,
// submitted once — the shape `MaskPass::render` established.
// Submission order is the whole of the synchronisation: each pass
// reads what the previous one wrote, through the same queue.
let (ran, result) = match self.detail.encode(&mut enc, detail, width, height) {
Ok(done) => done,
Err(e) => {
// Nothing is submitted, so a cache this frame was to write
// holds nothing, and must not be read as though it did.
@@ -1230,6 +1235,86 @@ impl AdjustPass {
return Err(e);
}
};
// TRACES: FR-DEV-3j
// The view pass: the view transform and the output transform, after
// every kernel (D19). Its own uniform block, filled from the source
// like the fused pass's — it reads the non-linear flag and the film
// settings there — with the sample cache off, because the colour it
// reads is the detail stage's result, bound where the cache would be.
let view_uniforms = Self::fused_uniforms(source, view);
let view_params = self
.ctx
.device
.create_buffer_init(&wgpu::util::BufferInitDescriptor {
label: Some("adjust-view-params"),
contents: bytemuck::cast_slice(&view_uniforms),
usage: wgpu::BufferUsages::UNIFORM,
});
let (_, no_sample_out) = self.sample.views(SampleUse::Direct);
let target_view = self.targets[self.current]
.as_ref()
.expect("ensured above")
.view
.clone();
let view_bind_group = self
.ctx
.device
.create_bind_group(&wgpu::BindGroupDescriptor {
label: Some("adjust-view-bg"),
layout: &self.bind_group_layout,
entries: &[
wgpu::BindGroupEntry {
binding: 0,
resource: wgpu::BindingResource::TextureView(source.view()),
},
wgpu::BindGroupEntry {
binding: 1,
resource: view_params.as_entire_binding(),
},
wgpu::BindGroupEntry {
binding: 2,
resource: wgpu::BindingResource::TextureView(&target_view),
},
wgpu::BindGroupEntry {
binding: 3,
resource: wgpu::BindingResource::TextureView(
masks.map_or(&self.empty_masks, |m| m.view()),
),
},
wgpu::BindGroupEntry {
binding: 4,
resource: wgpu::BindingResource::TextureView(&film_curves),
},
wgpu::BindGroupEntry {
binding: 5,
resource: wgpu::BindingResource::TextureView(&film_lut),
},
wgpu::BindGroupEntry {
binding: 6,
resource: wgpu::BindingResource::TextureView(&result),
},
wgpu::BindGroupEntry {
binding: 7,
resource: wgpu::BindingResource::TextureView(&no_sample_out),
},
],
});
{
let pipeline = self
.cache
.get(&view.structure_hash)
.expect("compiled above");
let mut pass = enc.begin_compute_pass(&wgpu::ComputePassDescriptor {
label: Some("adjust-view-pass"),
timestamp_writes: None,
});
pass.set_pipeline(pipeline);
pass.set_bind_group(0, &view_bind_group, &[]);
pass.dispatch_workgroups(width.div_ceil(8), height.div_ceil(8), 1);
}
self.view_dispatches += 1;
self.ctx.queue.submit(Some(enc.finish()));
self.detail_dispatches += ran;
self.colour_key = Some((key, width, height));
@@ -1378,6 +1463,13 @@ impl AdjustPass {
self.detail_dispatches
}
/// TRACES: FR-DEV-3j
/// View passes encoded since this pass was created: one for every render
/// that had a detail stage, since the view transform follows it.
pub fn view_dispatches(&self) -> usize {
self.view_dispatches
}
/// How many linear intermediates have been allocated. For tests: see
/// [`crate::MaskPass::allocations`] for the regression this catches.
pub fn detail_allocations(&self) -> usize {
@@ -2351,7 +2443,7 @@ mod tests {
// find those operations in neither stage and fail for a reason that is
// not a defect. Shadows the smaller size deliberately.
let (w, h) = g.output_size(512, 512);
let detail = g.compose_detail_for((512, 512), (w, h), dr_types::ColourSpace::Srgb);
let detail = g.compose_detail((512, 512), (w, h));
assert!(
!detail.is_empty(),
"the detail half composed nothing, so nothing of it was compiled"
+32 -46
View File
@@ -43,12 +43,12 @@
//! dispatch is skipped, and dragging a sharpening slider costs the detail
//! passes alone (FR-DEV-3d).
//!
//! The remaining passes alternate between slots 1 and 2, and the last one
//! writes the display texture directly rather than an intermediate — so a
//! chain of *N* passes costs *N* dispatches and not *N* + 1, and there is no
//! resolve pass to pay for. That leaves the allocation at `1 + min(N-1, 2)`
//! textures: one for a single-pass operation, two for a separable blur, three
//! however long the chain gets after that.
//! The passes alternate between slots 1 and 2, the last one included: since
//! D19 it hands its result to the adjust pass's **view pass**, which performs
//! the view transform and the output transform after every kernel, so no
//! detail pass writes the display texture. A chain of *N* passes costs *N*
//! dispatches plus that one, and the allocation is `1 + min(N, 2)` textures.
//! An empty chain costs the view pass alone, reading slot 0.
//!
//! # The reduced chain, and why a second one was needed
//!
@@ -186,10 +186,9 @@ impl Intermediates {
/// intermediate against a fresh colour result and never be told.
pub(crate) struct DetailRunner {
ctx: GpuContext,
/// Layout for a pass writing another linear intermediate.
/// Layout for every pass: each writes a linear intermediate, the last
/// one included, and the adjust pass's view pass reads the last (D19).
to_linear: Layout,
/// Layout for the last pass, which writes the display texture.
to_output: Layout,
/// Compiled pipelines by pass structure hash.
cache: HashMap<u64, wgpu::ComputePipeline>,
pool: Intermediates,
@@ -255,7 +254,6 @@ impl DetailRunner {
Self {
ctx: ctx.clone(),
to_linear: Layout::new(ctx, INTERMEDIATE_FORMAT, "detail-linear"),
to_output: Layout::new(ctx, crate::AdjustPass::FORMAT, "detail-output"),
cache: HashMap::new(),
pool: Intermediates::new(),
reduced: Intermediates::new(),
@@ -274,27 +272,30 @@ impl DetailRunner {
width: u32,
height: u32,
) -> &wgpu::TextureView {
// One for the colour pass's result, then one per hand-off between
// detail passes, capped at two because a ping-pong needs no more: the
// last pass writes the display texture rather than an intermediate.
let needed = 1 + passes.saturating_sub(1).min(2);
// One for the colour pass's result, then one per pass, capped at two
// because a ping-pong needs no more. The last pass writes an
// intermediate like the others since D19 — the view pass reads it —
// so a one-pass chain needs two slots where it used to need one.
let needed = 1 + passes.min(2);
self.pool.ensure(&self.ctx, needed, width, height);
&self.pool.slots[0].view
}
/// Encode every pass of `chain`, the last one writing `output`.
/// Encode every pass of `chain`, and return how many ran and the view
/// the last one wrote — slot 0, the colour pass's own result, for an
/// empty chain.
///
/// The caller must already have run the fused colour pass into
/// [`Self::colour_target`] — or established that a previous frame's is
/// still valid, which is the whole point of keeping slot 0.
/// still valid, which is the whole point of keeping slot 0 — and reads the
/// returned view in the view pass that finishes the render (D19).
pub(crate) fn encode(
&mut self,
encoder: &mut wgpu::CommandEncoder,
chain: &ComposedDetail,
output: &wgpu::TextureView,
width: u32,
height: u32,
) -> Result<usize, GpuError> {
) -> Result<(usize, wgpu::TextureView), GpuError> {
for pass in &chain.passes {
self.compile(pass)?;
}
@@ -340,17 +341,15 @@ impl DetailRunner {
for pass in chain.passes.iter() {
let scaled = pass.output_scale > 1;
// The last pass carries the output transform into the display
// texture, which is the render size by definition. A scaled pass
// there would bind a shader dispatching over a quarter-size grid
// to a full-size target and write a quarter of the picture — a
// wrong image rather than a validation failure, so it is caught
// here and named.
if scaled && pass.writes_output {
// The last pass hands the view pass its input, which is read at
// the render size by definition. A scaled pass there would leave
// the result in the reduced chain and the view pass would read the
// full-size slot before it — a wrong image rather than a
// validation failure, so it is caught here and named.
if scaled && std::ptr::eq(pass, chain.passes.last().expect("iterating")) {
return Err(GpuError::ShaderCompilation(format!(
"detail pass {} declares output_scale {} and is last in \
the chain; the output transform is written at the render \
size",
the chain; the view pass reads the render size",
pass.label, pass.output_scale
)));
}
@@ -365,17 +364,14 @@ impl DetailRunner {
};
// Read what the previous pass in *this pass's own chain* wrote;
// write the next slot of it, or the display texture if this is the
// last pass. Alternating slots is what stops a pass reading the
// write the next slot of it. Alternating slots is what stops a pass reading the
// texture it is writing — on a compute pass that is not an error
// the driver reports, merely a picture that depends on scheduling.
let source = match (scaled, carried) {
(true, Some(slot)) => &self.reduced.slots[slot].view,
_ => &self.pool.slots[full].view,
};
let destination = if pass.writes_output {
output
} else if scaled {
let destination = if scaled {
&self.reduced.slots[reduced_writes % 2].view
} else {
&self.pool.slots[1 + (full_writes % 2)].view
@@ -387,11 +383,7 @@ impl DetailRunner {
Some(slot) => &self.reduced.slots[slot].view,
None => &self.no_reduced,
};
let layout = if pass.writes_output {
&self.to_output
} else {
&self.to_linear
};
let layout = &self.to_linear;
let params = self
.ctx
@@ -464,9 +456,7 @@ impl DetailRunner {
compute.dispatch_workgroups(dispatch_w.div_ceil(8), dispatch_h.div_ceil(8), 1);
drop(compute);
if pass.writes_output {
// Nothing downstream to hand anything to.
} else if scaled {
if scaled {
carried = Some(reduced_writes % 2);
reduced_writes += 1;
} else {
@@ -480,7 +470,7 @@ impl DetailRunner {
}
}
Ok(chain.passes.len())
Ok((chain.passes.len(), self.pool.slots[full].view.clone()))
}
/// Compile one pass, or leave the cached pipeline in place.
@@ -508,11 +498,7 @@ impl DetailRunner {
source: wgpu::ShaderSource::Wgsl(pass.source.as_str().into()),
});
let layout = if pass.writes_output {
&self.to_output
} else {
&self.to_linear
};
let layout = &self.to_linear;
let pipeline = self
.ctx