Run the view transform after the detail stage, in a pass of its own
The fused pass stops at "linear working values" when a sharpener, a blur or a repair follows, and the detail passes convolve what it hands on. Until now it handed on the rendering: the base curve, and since the last commit the view transform, ran before the store. So every kernel worked on display-referred values while its comments promised the opposite — D19's second finding. A fused pass composed for a detail stage now stops before the view transform, and carries a second shader, `ComposedShader::view`, composed from the same inputs. It runs the same prologue, for the positions a fragment reads (a film's grain seeds from `source_px`) and the corners it blacks out, takes its colour from the detail stage's result bound where the sample cache would be, and runs the view transform, the output transform and the mask reveal. `render_detailed` dispatches it after the last detail pass, in the same encoder. So no detail pass encodes any more. Every pass writes an intermediate, the last one included, which retires three things that existed only to make the last pass encode: `writes_output` and the runner's second layout, the body-less resolve pass for an active kernel with nothing to draw at this scale, and capture sharpening's pass-through, which now emits no pass at all. An empty chain is a whole render: the view pass reads the fused result directly. The detail stage no longer takes an output space either, so `compose_detail_for` folds into `compose_detail` and the space is named once, on the fused half. The cost is one full-render read and write per frame when a detail stage exists, and a third intermediate for a one-pass chain.
This commit is contained in:
+32
-46
@@ -43,12 +43,12 @@
|
||||
//! dispatch is skipped, and dragging a sharpening slider costs the detail
|
||||
//! passes alone (FR-DEV-3d).
|
||||
//!
|
||||
//! The remaining passes alternate between slots 1 and 2, and the last one
|
||||
//! writes the display texture directly rather than an intermediate — so a
|
||||
//! chain of *N* passes costs *N* dispatches and not *N* + 1, and there is no
|
||||
//! resolve pass to pay for. That leaves the allocation at `1 + min(N-1, 2)`
|
||||
//! textures: one for a single-pass operation, two for a separable blur, three
|
||||
//! however long the chain gets after that.
|
||||
//! The passes alternate between slots 1 and 2, the last one included: since
|
||||
//! D19 it hands its result to the adjust pass's **view pass**, which performs
|
||||
//! the view transform and the output transform after every kernel, so no
|
||||
//! detail pass writes the display texture. A chain of *N* passes costs *N*
|
||||
//! dispatches plus that one, and the allocation is `1 + min(N, 2)` textures.
|
||||
//! An empty chain costs the view pass alone, reading slot 0.
|
||||
//!
|
||||
//! # The reduced chain, and why a second one was needed
|
||||
//!
|
||||
@@ -186,10 +186,9 @@ impl Intermediates {
|
||||
/// intermediate against a fresh colour result and never be told.
|
||||
pub(crate) struct DetailRunner {
|
||||
ctx: GpuContext,
|
||||
/// Layout for a pass writing another linear intermediate.
|
||||
/// Layout for every pass: each writes a linear intermediate, the last
|
||||
/// one included, and the adjust pass's view pass reads the last (D19).
|
||||
to_linear: Layout,
|
||||
/// Layout for the last pass, which writes the display texture.
|
||||
to_output: Layout,
|
||||
/// Compiled pipelines by pass structure hash.
|
||||
cache: HashMap<u64, wgpu::ComputePipeline>,
|
||||
pool: Intermediates,
|
||||
@@ -255,7 +254,6 @@ impl DetailRunner {
|
||||
Self {
|
||||
ctx: ctx.clone(),
|
||||
to_linear: Layout::new(ctx, INTERMEDIATE_FORMAT, "detail-linear"),
|
||||
to_output: Layout::new(ctx, crate::AdjustPass::FORMAT, "detail-output"),
|
||||
cache: HashMap::new(),
|
||||
pool: Intermediates::new(),
|
||||
reduced: Intermediates::new(),
|
||||
@@ -274,27 +272,30 @@ impl DetailRunner {
|
||||
width: u32,
|
||||
height: u32,
|
||||
) -> &wgpu::TextureView {
|
||||
// One for the colour pass's result, then one per hand-off between
|
||||
// detail passes, capped at two because a ping-pong needs no more: the
|
||||
// last pass writes the display texture rather than an intermediate.
|
||||
let needed = 1 + passes.saturating_sub(1).min(2);
|
||||
// One for the colour pass's result, then one per pass, capped at two
|
||||
// because a ping-pong needs no more. The last pass writes an
|
||||
// intermediate like the others since D19 — the view pass reads it —
|
||||
// so a one-pass chain needs two slots where it used to need one.
|
||||
let needed = 1 + passes.min(2);
|
||||
self.pool.ensure(&self.ctx, needed, width, height);
|
||||
&self.pool.slots[0].view
|
||||
}
|
||||
|
||||
/// Encode every pass of `chain`, the last one writing `output`.
|
||||
/// Encode every pass of `chain`, and return how many ran and the view
|
||||
/// the last one wrote — slot 0, the colour pass's own result, for an
|
||||
/// empty chain.
|
||||
///
|
||||
/// The caller must already have run the fused colour pass into
|
||||
/// [`Self::colour_target`] — or established that a previous frame's is
|
||||
/// still valid, which is the whole point of keeping slot 0.
|
||||
/// still valid, which is the whole point of keeping slot 0 — and reads the
|
||||
/// returned view in the view pass that finishes the render (D19).
|
||||
pub(crate) fn encode(
|
||||
&mut self,
|
||||
encoder: &mut wgpu::CommandEncoder,
|
||||
chain: &ComposedDetail,
|
||||
output: &wgpu::TextureView,
|
||||
width: u32,
|
||||
height: u32,
|
||||
) -> Result<usize, GpuError> {
|
||||
) -> Result<(usize, wgpu::TextureView), GpuError> {
|
||||
for pass in &chain.passes {
|
||||
self.compile(pass)?;
|
||||
}
|
||||
@@ -340,17 +341,15 @@ impl DetailRunner {
|
||||
for pass in chain.passes.iter() {
|
||||
let scaled = pass.output_scale > 1;
|
||||
|
||||
// The last pass carries the output transform into the display
|
||||
// texture, which is the render size by definition. A scaled pass
|
||||
// there would bind a shader dispatching over a quarter-size grid
|
||||
// to a full-size target and write a quarter of the picture — a
|
||||
// wrong image rather than a validation failure, so it is caught
|
||||
// here and named.
|
||||
if scaled && pass.writes_output {
|
||||
// The last pass hands the view pass its input, which is read at
|
||||
// the render size by definition. A scaled pass there would leave
|
||||
// the result in the reduced chain and the view pass would read the
|
||||
// full-size slot before it — a wrong image rather than a
|
||||
// validation failure, so it is caught here and named.
|
||||
if scaled && std::ptr::eq(pass, chain.passes.last().expect("iterating")) {
|
||||
return Err(GpuError::ShaderCompilation(format!(
|
||||
"detail pass {} declares output_scale {} and is last in \
|
||||
the chain; the output transform is written at the render \
|
||||
size",
|
||||
the chain; the view pass reads the render size",
|
||||
pass.label, pass.output_scale
|
||||
)));
|
||||
}
|
||||
@@ -365,17 +364,14 @@ impl DetailRunner {
|
||||
};
|
||||
|
||||
// Read what the previous pass in *this pass's own chain* wrote;
|
||||
// write the next slot of it, or the display texture if this is the
|
||||
// last pass. Alternating slots is what stops a pass reading the
|
||||
// write the next slot of it. Alternating slots is what stops a pass reading the
|
||||
// texture it is writing — on a compute pass that is not an error
|
||||
// the driver reports, merely a picture that depends on scheduling.
|
||||
let source = match (scaled, carried) {
|
||||
(true, Some(slot)) => &self.reduced.slots[slot].view,
|
||||
_ => &self.pool.slots[full].view,
|
||||
};
|
||||
let destination = if pass.writes_output {
|
||||
output
|
||||
} else if scaled {
|
||||
let destination = if scaled {
|
||||
&self.reduced.slots[reduced_writes % 2].view
|
||||
} else {
|
||||
&self.pool.slots[1 + (full_writes % 2)].view
|
||||
@@ -387,11 +383,7 @@ impl DetailRunner {
|
||||
Some(slot) => &self.reduced.slots[slot].view,
|
||||
None => &self.no_reduced,
|
||||
};
|
||||
let layout = if pass.writes_output {
|
||||
&self.to_output
|
||||
} else {
|
||||
&self.to_linear
|
||||
};
|
||||
let layout = &self.to_linear;
|
||||
|
||||
let params = self
|
||||
.ctx
|
||||
@@ -464,9 +456,7 @@ impl DetailRunner {
|
||||
compute.dispatch_workgroups(dispatch_w.div_ceil(8), dispatch_h.div_ceil(8), 1);
|
||||
drop(compute);
|
||||
|
||||
if pass.writes_output {
|
||||
// Nothing downstream to hand anything to.
|
||||
} else if scaled {
|
||||
if scaled {
|
||||
carried = Some(reduced_writes % 2);
|
||||
reduced_writes += 1;
|
||||
} else {
|
||||
@@ -480,7 +470,7 @@ impl DetailRunner {
|
||||
}
|
||||
}
|
||||
|
||||
Ok(chain.passes.len())
|
||||
Ok((chain.passes.len(), self.pool.slots[full].view.clone()))
|
||||
}
|
||||
|
||||
/// Compile one pass, or leave the cached pipeline in place.
|
||||
@@ -508,11 +498,7 @@ impl DetailRunner {
|
||||
source: wgpu::ShaderSource::Wgsl(pass.source.as_str().into()),
|
||||
});
|
||||
|
||||
let layout = if pass.writes_output {
|
||||
&self.to_output
|
||||
} else {
|
||||
&self.to_linear
|
||||
};
|
||||
let layout = &self.to_linear;
|
||||
|
||||
let pipeline = self
|
||||
.ctx
|
||||
|
||||
Reference in New Issue
Block a user