Commit Graph
3 Commits
Author SHA1 Message Date
dtourolle c07f81edcb Run the view transform after the detail stage, in a pass of its own
The fused pass stops at "linear working values" when a sharpener, a
blur or a repair follows, and the detail passes convolve what it hands
on. Until now it handed on the rendering: the base curve, and since the
last commit the view transform, ran before the store. So every kernel
worked on display-referred values while its comments promised the
opposite — D19's second finding.

A fused pass composed for a detail stage now stops before the view
transform, and carries a second shader, `ComposedShader::view`, composed
from the same inputs. It runs the same prologue, for the positions a
fragment reads (a film's grain seeds from `source_px`) and the corners
it blacks out, takes its colour from the detail stage's result bound
where the sample cache would be, and runs the view transform, the
output transform and the mask reveal. `render_detailed` dispatches it
after the last detail pass, in the same encoder.

So no detail pass encodes any more. Every pass writes an intermediate,
the last one included, which retires three things that existed only to
make the last pass encode: `writes_output` and the runner's second
layout, the body-less resolve pass for an active kernel with nothing to
draw at this scale, and capture sharpening's pass-through, which now
emits no pass at all. An empty chain is a whole render: the view pass
reads the fused result directly. The detail stage no longer takes an
output space either, so `compose_detail_for` folds into
`compose_detail` and the space is named once, on the fused half.

The cost is one full-render read and write per frame when a detail
stage exists, and a third intermediate for a one-pass chain.
2026-09-27 16:52:54 -04:00
dtourolleandClaude Opus 5 bff95e25ad Let clarity's base be computed where it is still fully determined
Clarity's Gaussian sigma is 1.2% of the frame's shorter edge, so its radius
is a property of the viewport: 52 render pixels at 4K, two separable passes
of 105 taps each over 8.3 M pixels. That measured 33.9 ms — seven times the
entire fused point chain, for one slider — and is docs/technical-debt.md TD-4.

A detail pass may now declare `output_scale`, and clarity's base is computed
on a grid a quarter the size on each axis.

The pass that combines needs the blur *and* the full-resolution colour, and a
colour that has been through a quarter-scale target is no longer full
resolution. So a scaled pass cannot simply join the ping-pong: there are two
chains now. The full-resolution one carries the colour and no scaled pass
touches it; the reduced one carries the base and reaches the combining pass
through a second binding as `reduced_at()`.

The reduce is a dispatch of its own rather than something the first blur half
does on the way past, and that is the whole difference between this and the
strided kernel the module documentation rules out. A stride samples an image
that is not band-limited and aliases high-frequency content down into the
base, which is then subtracted, and arrives in the output as mottling across
smooth gradients. This band-limits first and samples after. What is discarded
is content the base could not represent at any resolution, because a Gaussian
at sigma = 26 px holds nothing above one cycle per 26 px and the quarter-scale
grid carries one per 8 — so the reduced base is not an approximation of the
full-resolution one, it is the same function sampled where it is still
determined.

Which is also why the scale belongs to the band rather than to the stage.
Texture's sigma is a decade finer, so the reduce pass's own box would be wider
than the Gaussian it was prefiltering; texture never reduces. And clarity
steps 4 -> 2 -> 1 as sigma falls, because a quarter of a small sigma is not a
Gaussian either — the case that gives up is the one that was already cheap.

`radius` stays in each pass's own pixels and `ComposedDetail::radius` multiplies
it back up, so 13 reduced pixels at scale 4 still report the 52 render pixels a
tile would have to be grown by. The halo a scheduler sees does not move.

The halo tests pass unchanged, which was TD-4's stated bar; they render at
1024 px and so exercise the reduced path rather than stepping around it. Added
`crossing_the_reduction_threshold_does_not_change_the_picture`, because
nothing yet compared the reduced form against a *less* reduced one — every
other test measures one form against itself. It renders the same edit either
side of the 4 -> 2 step-down and holds the peak excursion to 0.03 stops and
the reach to 2% of the frame.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-29 13:12:29 +02:00
dtourolleandClaude Opus 5 6997c0f7ac Let a detail pass carry a list, not only a kernel
Every neighbourhood pass so far has been a convolution, whose whole
description fits in the uniform block because its structure fixes how many
numbers it needs. Spot removal is not that shape: sixty-four repairs and
one repair are the same shader with a different buffer behind it.

So a pass may declare `storage`, which arrives at binding 3 as
`array<vec4<f32>>` with `arrayLength` in scope. The alternative — packing
the list into uniforms — needs a fixed maximum paid for on every frame, a
composer that can emit vec4 fields because a uniform array's stride is 16
whatever it holds, and it gives the next operation that wants a table
nothing to build on.

The property worth having is what stays out of the generated source: the
count is in the buffer, so placing the tenth spot uploads 512 bytes and
reuses the compiled pipeline, exactly as moving a slider does for the
fused pass. `changing_the_list_does_not_recompile` is that, asserted.

One bind group entry rather than two more layouts, and one placeholder
buffer allocated in `new` rather than sixteen bytes per pass per frame —
a zero-length storage buffer cannot be bound, and per-frame allocation is
what this module's documentation exists to refuse.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-26 20:00:24 +02:00