diff --git a/docs/dev/denoise.md b/docs/dev/denoise.md index cddde95..480eb07 100644 --- a/docs/dev/denoise.md +++ b/docs/dev/denoise.md @@ -1,8 +1,9 @@ # Learned denoise — joint demosaic and denoise on the mosaic Design for **FR-DEV-3g** ([requirements.md](requirements.md)), the learned stage -[outstanding.md §3](outstanding.md) says is missing. Draft of 2026-09-27: nothing here is built, -and every figure marked *estimate* is waiting for the measurement that replaces it. +[outstanding.md §3](outstanding.md) says is missing. Drafted 2026-09-27; a first version shipped +in 0.21.0, and §11 records what was built and measured. Figures still marked *estimate* are +waiting for the measurement that replaces them. --- @@ -288,17 +289,27 @@ is ~120 MB, and a derived file inside a synced tree is exactly what with progress over the canvas — the same pattern as a photograph that is only on the server. - Export needs the result and computes it if the cache has lost it. -### 7.2 The Amount control +### 7.2 The grain control -A Denoise toggle and one Amount slider in develop. Moving the slider runs inference on the -**visible viewport only** (~1 MP, a fraction of a second — *estimate*) so the photographer judges -on the real result; releasing it queues the whole frame. There is no per-frame blend between the -two paths: blending the classical output back in re-adds the noise the network removed. +What shipped is a switch and a **Keep grain** slider, not the Amount described first. The slider +blends the two demosaics per pixel — but only the *brightness* of their difference: `out = +denoised + grain · ΔY / wb`, with `ΔY` the luminance of `wb · (classical − denoised)`. Taken after +the as-shot balance and handed back divided by it, the grain is neutral in the finished picture. + +The objection that stood here — that blending the classical output back in re-adds the noise — +holds for a plain mix, which also brings back the classical path's colour speckle and false +colour. A luminance-only blend returns film-like grain and nothing else, and it needs no +inference: one elementwise GPU pass (`dr_gpu::GrainBlend`) per slider value, producing a new +source the adjust pass draws. Comparing the two on real 6D frames, the user chose this one. + +The σ-map Amount (§3.3) still works — `NoiseModel::scaled` — and stays available for a later +"strength" control; its cost is a re-run of the network. ### 7.3 Runtime -Through `dr-inference-engine`, as the other models run ([inference.md](inference.md)): TensorRT or -CUDA fp16 on the laptop, MIGraphX on the desktop, ORT CPU everywhere, QNN on the tablet. Work is +Through `dr-inference-engine`, as the other models run ([inference.md](inference.md)), as +`Role::Denoiser`: TensorRT or CUDA fp16 on the laptop, MIGraphX on the desktop, ORT CPU everywhere. +**Not the Hexagon** — see §11 — so the tablet runs it on its CPU. Work is scheduled in the `Background` class so a slider never waits on it (architecture §5.3). ## 8. Speed and the tablet @@ -356,3 +367,57 @@ MIT architecture, so this model adds no third-party licence to D13. 4. Whether a Lightroom or DxO comparison is available for §6.2. 5. A borrowed X-Trans body, or X-Trans experimental in v1. +## 11. What shipped in 0.21.0, and what was measured + +**Data.** 500 distinct ISO 50–100 6D frames from the library, over 121 shooting days (bursts and +near-duplicate perceptual hashes dropped; 55 frames from held-out days for validation). Read through +`dr-gpu`'s `mosaic_dump` example — `dr-decode` and the app's own hot-pixel pass — so the network's +input is the mosaic the classical demosaic reads. Truth by 2×2 binning with a Catmull-Rom quarter-pixel +shift of red and blue (§4.2). Training lives in `darkroom-denoise`, beside `darkroom-infill`. + +**Noise model (§5), from the library instead of a capture.** Shot gain and read variance per ISO from +Adobe's `NoiseProfile` in the converted DNGs; read noise checked against each frame's masked border +(agreement within 2–3 % from ISO 125 to 25600); read-noise *shape* taken from the border as quantiles +on a tail-dense grid (excess kurtosis up to ~10 at high ISO), with only the photosites the app's +hot-pixel rule would remove left out; row noise from the border's row means; **column noise** from +the masked rows above the image — about a third of its variance is this sensor's fixed pattern. +Third stops are their own rows: ISO 160, 320 and 640 are quieter than their neighbours, as §5.1 +expected. Training without the column noise left the 6D's vertical stripes in (0.90 DN of 1.01); +with it, 0.17 DN. + +**Model.** Not NAFNet: its channel attention averages over the whole input, which breaks exact +tiling. A U-Net of 3×3 convolutions, ReLU, strided and transposed convolutions and additive skips — +3.2 M parameters, 48 GMAC per raw megapixel, receptive field 185 photosites (counted from the +layers; a perturbation probe under-read it as 157 because a switched-off ReLU hides a path). +Tiles of 1408 keep their central 1024 behind a 192 halo, exactly. + +**Results.** PSNR after the display transform, held-out days, step 60 000: + +| ISO | Network | Bilinear | Bilinear on a clean mosaic | +|---|---|---|---| +| 400 | 41.6 | 36.8 | 40.0 | +| 1600 | 40.8 | 33.8 | 40.0 | +| 6400 | 39.5 | 29.3 | 40.0 | +| 25600 | 37.8 | 24.6 | 40.0 | + +Unbiased in linear light on real frames (shadow level within 1 % of a heavily averaged bilinear). +Checked against the app's own render for channel and axis order (`tools/check_against_app.py`). + +**Precision (§8).** fp16: 0.00 dB at every ISO. int8 QDQ, calibrated on training tiles: −6 to −9 dB +— the shadow steps §8 feared losing are lost. So the Hexagon refuses the role and the tablet runs f32 +on its CPU; the residual head of §8 is the route back. + +**Noise for any Bayer body (§3.3).** Table, then `NoiseProfile`, then the frame itself: read, row and +column noise from its masked border, the shot gain alone estimated from the quietest flat patches. +On 130 6D frames the estimate is within ±10 % of the table from ISO 1000 up and scattered below. The +network loses under 0.3 dB for σ off by 15–20 % and twice as much for under- as for over-estimating; +the estimate leans high. Every Bayer body is offered the switch; develop says which source was used. + +**Speed, a whole 6D frame (20 MP).** TensorRT fp16 3.1 s, ONNX Runtime CPU 14.4 s, on the laptop — +measured while the GPU sat power-capped at an 810 MHz memory clock; uncapped is expected to be +about four times faster. The Rust path reproduces the training repository's output to 2.5e-4 at +worst; TensorRT fp16 is 75 dB from f32. + +**Not yet:** the result is not cached across sessions (§7.1) — reopening recomputes; the tripod real +pairs of §6.1; X-Trans (§9); the hand-written WGSL path, for which `export.py` already writes the +weights blob and a manifest a shader can follow.