Commit Graph
4 Commits
Author SHA1 Message Date
dtourolle c07f81edcb Run the view transform after the detail stage, in a pass of its own
The fused pass stops at "linear working values" when a sharpener, a
blur or a repair follows, and the detail passes convolve what it hands
on. Until now it handed on the rendering: the base curve, and since the
last commit the view transform, ran before the store. So every kernel
worked on display-referred values while its comments promised the
opposite — D19's second finding.

A fused pass composed for a detail stage now stops before the view
transform, and carries a second shader, `ComposedShader::view`, composed
from the same inputs. It runs the same prologue, for the positions a
fragment reads (a film's grain seeds from `source_px`) and the corners
it blacks out, takes its colour from the detail stage's result bound
where the sample cache would be, and runs the view transform, the
output transform and the mask reveal. `render_detailed` dispatches it
after the last detail pass, in the same encoder.

So no detail pass encodes any more. Every pass writes an intermediate,
the last one included, which retires three things that existed only to
make the last pass encode: `writes_output` and the runner's second
layout, the body-less resolve pass for an active kernel with nothing to
draw at this scale, and capture sharpening's pass-through, which now
emits no pass at all. An empty chain is a whole render: the view pass
reads the fused result directly. The detail stage no longer takes an
output space either, so `compose_detail_for` folds into
`compose_detail` and the space is named once, on the fused half.

The cost is one full-render read and write per frame when a detail
stage exists, and a third intermediate for a one-pass chain.
2026-09-27 16:52:54 -04:00
dtourolle bee5c5866f Erode dehaze's window in one pass per axis, and recover in the second
Dehaze cost 22.9 ms of a 2560x1600 frame on the reference laptop RTX 3050,
and 54.1 ms at 3840x2160, with the memory clock held at 810 MHz by the power
cap (graphics 1762 MHz). It ran five passes: a run and a span erosion along
x, the same along y, and the recovery. At those clocks a detail pass costs
what it reads and writes, not what it taps: a pass with an empty body -
one render-sized rgba16float read and write - measured 4.0 ms, and each
dehaze pass 4.4-4.6 ms, so the taps were about 2 ms of the 22 and the four
hand-offs between passes were the rest.

Each axis is now one pass that takes the minimum over the whole window
directly, and the recovery rides in the y pass, which already holds the
veil and the pixel's own colour. That is 36 texture reads per pixel at
2560x1600 in place of 12, nearly all of them cache hits, and two passes in
place of five.

The picture is the same bits. A minimum is exact in any order, and the
window is the one Split always covered, the surplus pixel on the far side
included (Split::first and Split::width). The veil crossing the removed
hand-offs was already exactly representable in rgba16float - a minimum of
channels read from rgba16float, floored at zero - so storing it between
passes never rounded anything that the fused form now keeps unrounded.

Measured with a scratch probe that renders the synthetic 60 MP frame from
examples/frame_budget.rs, only a detail parameter moving so the fused pass
is reused, 30 frames per scene after six of warm-up, five runs of each
binary alternated, median of the per-run p50:

  scene                   before     after
  dehaze      2560 fit    22.88 ms    9.06 ms
  dehaze      2560 1:1    23.41 ms    9.52 ms
  dehaze      3840 fit    54.09 ms   28.12 ms
  all detail  2560 fit    53.11 ms   39.97 ms  (NR, sharpen, clarity,
  all detail  2560 1:1    67.48 ms   56.42 ms   texture, dehaze)
  every op    2560 fit    57.59 ms   44.19 ms  (with film)
  every op    2560 1:1    71.83 ms   57.93 ms
  controls without dehaze (NR, sharpen, clarity, texture): within +-2%

The rgba8 output hashed identically before and after for every scene -
dehaze alone, all five detail operations, every operation with film, and
each other detail operation alone - at fit and 1:1, at 2560x1600,
3840x2160, 1917x1203 and 333x211: 64 of 64.
2026-09-26 14:18:42 -04:00
dtourolle 84fade99ec Put the developer docs under docs/dev and index the folder for users first
docs/ had 26 developer documents flat beside the manual, and the two
audiences are very differently sized: most readers want the manual and
the gesture reference, a few want the register, the designs and the
measurements. The manual and gestures.md stay at the top; everything for
someone changing the code moves to docs/dev/, and the two documents that
name their own successors — the v0.1 milestone and the UI-refinement plan
— go to docs/dev/archive/ rather than being deleted, since both are still
cited. docs/README.md is the index, users first.

Every reference follows: code comments, Cargo manifests, the workflows,
the pre-commit hook, the bench and traceability tools (which locate the
repo root by docs/dev/requirements.md now), packaging, the Docker READMEs,
CLAUDE.md, CONTRIBUTING.md and the README. The matrix links one level
deeper and is regenerated. Links out of the moved documents into the tree
gain a level; a link checker over every Markdown file finds none broken.
2026-09-20 21:16:03 +02:00
dtourolle 81b1ae8c42 Measure the haze from the picture, and divide it back out
Four files named dehaze as a member of the compositional detail family —
`detail.rs` twice, `dr-gpu`'s detail module, `ops/README.md` and
`capture_sharpen.rs` — and no such node existed. Every one of them was
describing the family by listing clarity, texture and a control the
photographer could not reach.

Haze is the one degradation the controls already in the chain cannot
remove, and the reason is spatial rather than tonal. Scattering
composites an airlight over the scene in proportion to distance, so the
lift is per-pixel: a black point that clears the mountains crushes the
foreground, and a contrast curve that clears the mountains does the
same. So the node has to estimate the transmission at every pixel, which
is the dark-channel prior — the local minimum over the channels and over
a patch is the airlight that has been added there — and then invert the
scattering model with it.

The airlight is taken as neutral and as unit, which removes the one part
of the published method this stage cannot perform. Estimating it
properly is a whole-frame reduction, and the detail chain has none: it
hands each pass the pass before it. It is also unnecessary, because
white balance is the first node in the chain and has already driven the
illuminant to grey, so only the magnitude is unknown — and an unknown
magnitude on the veil is a scale factor on the amount slider, which the
photographer is setting by eye regardless.

The patch is a fraction of the frame's shorter edge, through
`RenderScale::frame_fraction`, and never a count of pixels. It has to be
wide enough to contain something dark and narrow enough that what it
measures is still local, and both of those are statements about how much
of the composition it covers — so it must cover the same proportion of
the picture on a proxy as in the export, or the file is sharpened for a
patch three times narrower than the one that was tuned on screen.

Affording it needs an identity a Gaussian does not have. Erosions
compose by adding their structuring elements, so the minimum over a run
of d followed by the minimum over k points spaced d apart is the exact
minimum over the whole kd window. At the square root that is 16 taps
rather than 61 at 4K, and it is the same filter rather than an
approximation of one — which is the difference from the strided kernel
`local_contrast` refuses, where sampling an image that is not
band-limited aliases into the base and comes back as mottling.

It runs first among the compositional detail nodes, at order 125: after
noise reduction, because dividing by a transmission below one amplifies
the noise in the veiled distance by exactly the factor it recovers the
contrast by, and before clarity and texture, coarse before fine, so that
their base is computed on the picture the veil has left rather than on a
modelling about to be divided out.

What it cannot honour is the placement dehaze most wants. It shifts
colour — it subtracts a grey term and rescales, so saturation changes
wherever the veil is thick — and the colour work would ideally be
correcting the picture that leaves here. The detail stage runs as a
group after every point operation, because a neighbourhood pass is a
separate dispatch over a texture the fused pass has finished writing, so
an order placing this node ahead of `vibrance` would be a lie the chain
cannot tell. Interleaving would mean splitting the fused pass in half
around it, at the cost of a second full-frame dispatch and intermediate
for every edit in the catalogue whether it dehazes or not. The
declaration records that rather than leaving it to be rediscovered.

FR-DEV-18 is added to the requirements register alongside it. The tag
had nowhere to point, and an orphan tag fails the traceability gate
rather than quietly counting for nothing.
2026-09-06 19:01:47 +02:00