Record what the learned denoise shipped as, and what was measured

The grain blend replaces the Amount of §7.2, and why its objection to a
blend does not hold for brightness alone; the Hexagon is out (int8 -6 to
-9 dB); §11 holds the data, the noise model taken from the library, the
model, the validation table, the blind estimate's reach and the speed.
This commit is contained in:
2026-10-03 11:51:02 -04:00
parent 960014803a
commit 4eb7cf77f5
+74 -9
View File
@@ -1,8 +1,9 @@
# Learned denoise — joint demosaic and denoise on the mosaic
Design for **FR-DEV-3g** ([requirements.md](requirements.md)), the learned stage
[outstanding.md §3](outstanding.md) says is missing. Draft of 2026-09-27: nothing here is built,
and every figure marked *estimate* is waiting for the measurement that replaces it.
[outstanding.md §3](outstanding.md) says is missing. Drafted 2026-09-27; a first version shipped
in 0.21.0, and §11 records what was built and measured. Figures still marked *estimate* are
waiting for the measurement that replaces them.
---
@@ -288,17 +289,27 @@ is ~120 MB, and a derived file inside a synced tree is exactly what
with progress over the canvas — the same pattern as a photograph that is only on the server.
- Export needs the result and computes it if the cache has lost it.
### 7.2 The Amount control
### 7.2 The grain control
A Denoise toggle and one Amount slider in develop. Moving the slider runs inference on the
**visible viewport only** (~1 MP, a fraction of a second — *estimate*) so the photographer judges
on the real result; releasing it queues the whole frame. There is no per-frame blend between the
two paths: blending the classical output back in re-adds the noise the network removed.
What shipped is a switch and a **Keep grain** slider, not the Amount described first. The slider
blends the two demosaics per pixel — but only the *brightness* of their difference: `out =
denoised + grain · ΔY / wb`, with `ΔY` the luminance of `wb · (classical − denoised)`. Taken after
the as-shot balance and handed back divided by it, the grain is neutral in the finished picture.
The objection that stood here — that blending the classical output back in re-adds the noise —
holds for a plain mix, which also brings back the classical path's colour speckle and false
colour. A luminance-only blend returns film-like grain and nothing else, and it needs no
inference: one elementwise GPU pass (`dr_gpu::GrainBlend`) per slider value, producing a new
source the adjust pass draws. Comparing the two on real 6D frames, the user chose this one.
The σ-map Amount (§3.3) still works — `NoiseModel::scaled` — and stays available for a later
"strength" control; its cost is a re-run of the network.
### 7.3 Runtime
Through `dr-inference-engine`, as the other models run ([inference.md](inference.md)): TensorRT or
CUDA fp16 on the laptop, MIGraphX on the desktop, ORT CPU everywhere, QNN on the tablet. Work is
Through `dr-inference-engine`, as the other models run ([inference.md](inference.md)), as
`Role::Denoiser`: TensorRT or CUDA fp16 on the laptop, MIGraphX on the desktop, ORT CPU everywhere.
**Not the Hexagon** — see §11 — so the tablet runs it on its CPU. Work is
scheduled in the `Background` class so a slider never waits on it (architecture §5.3).
## 8. Speed and the tablet
@@ -356,3 +367,57 @@ MIT architecture, so this model adds no third-party licence to D13.
4. Whether a Lightroom or DxO comparison is available for §6.2.
5. A borrowed X-Trans body, or X-Trans experimental in v1.
## 11. What shipped in 0.21.0, and what was measured
**Data.** 500 distinct ISO 50–100 6D frames from the library, over 121 shooting days (bursts and
near-duplicate perceptual hashes dropped; 55 frames from held-out days for validation). Read through
`dr-gpu`'s `mosaic_dump` example — `dr-decode` and the app's own hot-pixel pass — so the network's
input is the mosaic the classical demosaic reads. Truth by 2×2 binning with a Catmull-Rom quarter-pixel
shift of red and blue (§4.2). Training lives in `darkroom-denoise`, beside `darkroom-infill`.
**Noise model (§5), from the library instead of a capture.** Shot gain and read variance per ISO from
Adobe's `NoiseProfile` in the converted DNGs; read noise checked against each frame's masked border
(agreement within 2–3 % from ISO 125 to 25600); read-noise *shape* taken from the border as quantiles
on a tail-dense grid (excess kurtosis up to ~10 at high ISO), with only the photosites the app's
hot-pixel rule would remove left out; row noise from the border's row means; **column noise** from
the masked rows above the image — about a third of its variance is this sensor's fixed pattern.
Third stops are their own rows: ISO 160, 320 and 640 are quieter than their neighbours, as §5.1
expected. Training without the column noise left the 6D's vertical stripes in (0.90 DN of 1.01);
with it, 0.17 DN.
**Model.** Not NAFNet: its channel attention averages over the whole input, which breaks exact
tiling. A U-Net of 3×3 convolutions, ReLU, strided and transposed convolutions and additive skips —
3.2 M parameters, 48 GMAC per raw megapixel, receptive field 185 photosites (counted from the
layers; a perturbation probe under-read it as 157 because a switched-off ReLU hides a path).
Tiles of 1408 keep their central 1024 behind a 192 halo, exactly.
**Results.** PSNR after the display transform, held-out days, step 60 000:
| ISO | Network | Bilinear | Bilinear on a clean mosaic |
|---|---|---|---|
| 400 | 41.6 | 36.8 | 40.0 |
| 1600 | 40.8 | 33.8 | 40.0 |
| 6400 | 39.5 | 29.3 | 40.0 |
| 25600 | 37.8 | 24.6 | 40.0 |
Unbiased in linear light on real frames (shadow level within 1 % of a heavily averaged bilinear).
Checked against the app's own render for channel and axis order (`tools/check_against_app.py`).
**Precision (§8).** fp16: 0.00 dB at every ISO. int8 QDQ, calibrated on training tiles: −6 to −9 dB
— the shadow steps §8 feared losing are lost. So the Hexagon refuses the role and the tablet runs f32
on its CPU; the residual head of §8 is the route back.
**Noise for any Bayer body (§3.3).** Table, then `NoiseProfile`, then the frame itself: read, row and
column noise from its masked border, the shot gain alone estimated from the quietest flat patches.
On 130 6D frames the estimate is within ±10 % of the table from ISO 1000 up and scattered below. The
network loses under 0.3 dB for σ off by 15–20 % and twice as much for under- as for over-estimating;
the estimate leans high. Every Bayer body is offered the switch; develop says which source was used.
**Speed, a whole 6D frame (20 MP).** TensorRT fp16 3.1 s, ONNX Runtime CPU 14.4 s, on the laptop —
measured while the GPU sat power-capped at an 810 MHz memory clock; uncapped is expected to be
about four times faster. The Rust path reproduces the training repository's output to 2.5e-4 at
worst; TensorRT fp16 is 75 dB from f32.
**Not yet:** the result is not cached across sessions (§7.1) — reopening recomputes; the tripod real
pairs of §6.1; X-Trans (§9); the hand-written WGSL path, for which `export.py` already writes the
weights blob and a manifest a shader can follow.