diff --git a/docs/dev/denoise.md b/docs/dev/denoise.md new file mode 100644 index 0000000..cddde95 --- /dev/null +++ b/docs/dev/denoise.md @@ -0,0 +1,358 @@ +# Learned denoise — joint demosaic and denoise on the mosaic + +Design for **FR-DEV-3g** ([requirements.md](requirements.md)), the learned stage +[outstanding.md §3](outstanding.md) says is missing. Draft of 2026-09-27: nothing here is built, +and every figure marked *estimate* is waiting for the measurement that replaces it. + +--- + +## 1. What we are matching + +Lightroom's Denoise (April 2023, Eric Chan's "Denoise demystified") is the reference, and three +facts about it set the shape of this design: + +- **It runs on the mosaic.** The network takes Bayer or X-Trans photosites before any demosaic + and emits full RGB: denoise and demosaic are one learned step. It descends from Adobe's 2019 + learned demosaic (Raw Details). A photograph that is already demosaiced is not eligible. +- **It is run once, not per frame.** The result is written as a new linear DNG beside the + original, and every later edit reads that file. The amount is chosen once, from a preview crop. +- **It is trained on synthetic pairs.** Clean raws with sensor-modelled noise added, not + photographed pairs. + +The reason the mosaic is the right place is physical: before the demosaic, noise is independent +per photosite with a known distribution (shot plus read). After it, the interpolation has +correlated that noise into colour blotches many photosites across, which classical noise reduction +cannot separate from texture. The same step removes demosaic artefacts +— maze, zipper, false colour, X-Trans worms (FR-RAW-5). + +We match the first and third facts and not the second: our result is a cache, not a file in the +library (§7). + +## 2. Where it sits + +[architecture.md §5.2](architecture.md) already reserves the slot. The learned stage **replaces +the demosaic box** when it is on; nothing else in the chain moves. + +``` +RawImage ─► hot/dead photosites ─► black/white levels ─► ┬─ demosaic (classical) ─┬─► camera profile ─► … + └─ learned demosaic+NR ──┘ + (cached, §7) +``` + +- **In:** the repaired, normalised mosaic, from the same buffer `Demosaicer::run` reads. The hot + pixel pass stays in front: an outlier of 50σ is outside anything the noise model generates, and + a network shown one invents a structure around it. +- **Out:** linear camera RGB, f16, full resolution — exactly the texture the classical demosaic + produces, so the camera profile, the raw histogram and every operation below it are unchanged. +- **Off by default, per photograph.** The classical path stays the default and the fallback; the + stage's absence degrades gracefully, as FR-DEV-3g requires. + +## 3. The model + +### 3.1 The 12×12 → 4×4 question + +The proposal: a network that reads a 12×12 window of photosites and predicts the RGB of the +central 4×4, slid across the frame in steps of four. + +**The output half is right. The input half is too small by a factor of five or more.** + +*What is right about it.* Predicting a block aligned to the colour-filter period keeps the phase +fixed: every prediction sees the same arrangement of red, green and blue around it, so the network +never has to work out where it stands in the pattern. It also makes tiling trivial and exact. +Both properties are kept below — as the head of the network and as the tiling contract (§3.4). + +*What is wrong with it.* A denoiser can only average away noise it can see around the pixel, and +at high ISO it needs to see a long way: + +- The Canon 6D at ISO 6400 (clip ≈ 1,200 e⁻, read noise ≈ 2 e⁻ — *estimate*, §5 measures it) has a + mid-tone of ~150 e⁻, shot SNR ≈ 12, and a shadow three stops down of ~19 e⁻, SNR ≈ 4. +- A shadow that looks clean wants SNR ≈ 40: a factor of 10, which is ~100 independent same-colour + samples in a flat area. Red and blue are a quarter of the photosites, so that is ~400 + photosites: a **20×20 window just for a flat shadow**, 40×40 two stops further down. +- A 12×12 window holds 36 red photosites. Averaged perfectly, that is a factor of 6 on red and + blue in a flat area, and less everywhere there is structure. +- Chroma blotches are low-frequency noise — 16 to 64 photosites across. A window smaller than the + blotch cannot tell it from a colour change. + +Demosaic alone is content with 12×12: good classical demosaics read 5×5 to 9×9. So the proposal is +a good demosaic network and a weak denoiser — which is a useful ablation (experiment E1, §6.3). + +*What it costs.* Adjacent 12×12 windows with a 4×4 output overlap nine-fold, so a network +evaluated per window recomputes each photosite's features nine times. A convolutional network is +the same computation with that work shared: it is "predict the central block from its +neighbourhood" evaluated everywhere at once. + +### 3.2 The shape + +``` +mosaic (H×W) ──space-to-depth 2×2──► 4 ch @ H/2 × W/2 ┐ +noise map σ(x) ─space-to-depth 2×2──► 4 ch @ H/2 × W/2 ┴► U-Net ─► 12 ch @ H/2 × W/2 ─depth-to-space─► RGB @ H×W + (2×2 block × RGB per position) +``` + +- **Packing.** Bayer is packed 2×2 into four channels at half resolution, so every input position + is one whole quad and every output position is the 2×2 block of RGB it covers — the proposal's + head, at the Bayer period. (A 4×4 packing with a 48-channel head is the same thing at a coarser + stride and is a free parameter.) +- **Phase unification.** Every body's pattern is cropped by a row or a column to RGGB before + packing, and the output is un-cropped. Flips are only used for augmentation in the CFA-preserving + form (Liu et al., "Bayer pattern unification and augmentation", 2019). +- **Body.** A U-Net with four downsamplings and NAFNet blocks (Chen et al., 2022; MIT). The + receptive field at the raw scale is several hundred photosites, which covers §3.1's worst case + with room. +- **Two sizes.** **M** (widths 32-64-128-256, ~6 M parameters, ~60 GMAC per raw megapixel — + *estimate*) is the desktop model and the one trained first. **S** (widths 16-32-64-128, fewer + bottleneck blocks, ~1 M parameters, ~12 GMAC/MP) is distilled from M for the tablet (§8). + +### 3.3 Conditioning on the noise + +The network is told how noisy each photosite is, rather than learning one model per ISO: + +- A per-photosite standard-deviation map, `σ(x) = √(K·x + σ_r²)` from the body's gain `K` and read + noise `σ_r` at that ISO, packed alongside the mosaic (FFDNet's arrangement, Zhang et al., 2018). +- **This is what makes it camera-general.** A body it was never trained on only has to supply + `K` and `σ_r`. Three sources, in order of preference: a calibration table for the body (§5); the + DNG `NoiseProfile` tag, which Adobe's converter writes; a blind estimate from the photograph's + own flat regions (Foi et al., 2008), which always exists. +- **It is also the Amount control.** Scaling the map up tells the network there is more noise than + there is and it smooths harder; scaling it down preserves more grain. Changing the amount re-runs + inference (§7.2), which is why it is set on a preview crop, as Lightroom does. + +The alternative — PMRID's k-sigma transform, which maps every ISO onto one noise level — is +simpler and gives no Amount control. It is the fallback if conditioning underperforms. + +### 3.4 Tiling + +A 20 MP frame does not go through a network in one piece on either device. Inference tiles the +mosaic into 512×512 input tiles with a 64-photosite halo on every side and keeps the central +384×384 of each output: the proposal's "12 in, 4 out", scaled up. Halo and tile sizes must be +multiples of 2 (the CFA phase) and of 16 (four downsamplings at half resolution), so the seams +land at identical positions in every tile's own coordinates. + +This is inference-local tiling and does not depend on FR-DSP-2's render-path tiling, which stays +under the challenge [outstanding.md §4](outstanding.md) records. + +## 4. Training data + +### 4.1 What the library holds + +From the reference catalog, 2026-09-27: 17,255 catalogued RAWs (9,345 DNG, 7,910 CR2), **all but +seven from one body, the Canon EOS 6D** (RGGB Bayer, 5472×3648, AA filter), 166 shooting days from +2015 to 2026. + +| ISO | Frames | Use | +|---|---|---| +| ≤ 200 | 5,065 | Clean sources for synthetic pairs | +| 201–1600 | 7,807 | Low-noise end of the eval set | +| 1601–6400 | 3,379 | Real-noise eval set; noise-model check (§5.3) | +| > 6400 | 562 | The hard cases, by eye | + +There are **no X-Trans raws**, which matters for §9. The catalog does not hold shutter speed, so +selection needs the files' EXIF. Whether the DNGs are mosaic (converted CR2) or linear must be +checked before they are counted as sources: a linear DNG has no photosites to learn from. + +### 4.2 How a training pair is made + +1. **Clean source.** A base-ISO 6D frame, black-subtracted and normalised. +2. **Full-colour truth by binning.** Each plane is resampled by half a photosite so the four + planes share a centre, then every 2×2 quad becomes one RGB pixel (R, mean of the two G, B): + a true full-colour image at 2736×1824 with no interpolation in it. This is the only way to have + ground truth for the demosaic half. +3. **Re-mosaic.** That RGB image is sampled back into an RGGB mosaic. (It can equally be sampled + into X-Trans, §9.) +4. **Darken and add noise.** Scale the signal by `1/g` for a target ISO `100·g`, then add noise + from the calibrated model at that ISO (§5): Poisson shot, Tukey-lambda read noise, row noise + and quantisation — the ELD model (Wei et al., CVPR 2020). The input is this mosaic; the target + is the clean RGB at the same scale. +5. **Augment.** Random blur (Gaussian, σ 0–0.7 px) before re-mosaicking, because a binned image is + sharper per pixel than the AA-filtered sensor the model will see; exposure jitter; white-balance + gains within the body's range; CFA-preserving flips. + +**Why the target's own noise is tolerable.** A base-ISO frame is not noise-free, and binning only +halves the green noise; red and blue keep theirs. But darkening by `g` scales signal and target +noise together, while the added shot noise grows as `√g`. At ISO 3200 the input is ≈ 5.7× noisier +than its target, at ISO 800 only ≈ 2.8×. L1 against a noisy target converges on the median, which +is unbiased for symmetric noise. The low-ISO end is the one at risk of learning to keep grain: if +it does, bin 4×4 instead (red and blue noise halved, 1368×912 per source) for those samples. + +**Why not the native mosaic as the target.** That trains denoise alone, with base-ISO noise baked +into the answer ("noisier2noise") and no demosaic truth at all. + +### 4.3 How much + +The limit is scene diversity, not pixel count; every source yields an effectively unlimited +number of pairs through random crops, ISO and noise draws. + +| Figure | Value | Reasoning | +|---|---|---| +| Sources, train | **3,000** | 5,065 base-ISO frames, less bursts (perceptual-hash dedup), heavy clipping, motion blur and linear DNGs. For scale: ELD reaches state of the art trained on ~230 scenes; SID has ~5,000 pairs of ~400 scenes | +| Sources, validation | 200 | Split by shooting day, not by frame, so no scene is on both sides | +| Pixels | ~15 Gpx of RGB truth | 3,000 × 5 MP after binning | +| Crops per step | 8–16 × 256×256 photosites | Fits a 6 GB RTX 3050 at fp16 with M | +| Stored | ~20 GB | 24 random 512×512 crops per source, uint16, zstd. Keeping whole CR2s would be ~75 GB | +| Training | 200–400 k steps, one to two nights per run on the 3050 — *estimate*; expect three to five runs | | + +Stratify the selection: across all 166 days, and deliberately include faces and hair (the library +has 19k detected faces, and skin is where over-smoothing shows first), foliage, fabric, text, and +any base-ISO tripod night work. + +### 4.4 Reading raws the same way in training and in the app + +The training data must be decoded by **the same decoder the app uses**. rawpy (LibRaw) and +`dr_decode::Rawler` can disagree on black level, white level, active area and therefore CFA phase, +and a network trained on one pattern phase and run on another produces colour moiré everywhere. +A `dr-decode` example that dumps the mosaic and its metadata as `.npy` is the only source the +training repo reads — not rawpy, as `darkroom-infill`'s `develop-raws.py` does. + +## 5. The noise model and its calibration + +### 5.1 What is measured + +Per ISO: gain `K` (DN per electron), read-noise distribution (Gaussian σ and Tukey-λ shape), +row-noise σ, black-level offset and any fixed pattern. Canon's third-stop ISOs on bodies of the +6D's generation are digital gains of the full stops, so noise does not scale smoothly between +them; **every third stop is calibrated**, not interpolated. + +### 5.2 The capture (one hour, once per body) + +- **Darks.** Lens cap on, viewfinder covered, manual. Five frames at 1/4000 s and five at 1/30 s + at every third stop from ISO 100 to 25600. They give read noise, row noise and the black-level + pattern; the two shutter speeds confirm dark current is negligible. +- **Flats.** An evenly lit white wall, defocused, at every full stop: pairs at six exposure levels + from 1/64 of clip to 3/4 of it. The variance of each pair's difference against their mean is + the photon transfer curve, whose slope is `K`. + +### 5.3 The check + +Fit the same `(K, σ_r)` blindly from flat regions of the library's 3,379 ISO 1601–6400 frames +(§3.3's third source). If it disagrees with the calibration by more than ~10%, one of them is +wrong — and it tells us how far the blind estimate can be trusted for bodies with no calibration. + +## 6. Evaluation + +### 6.1 Real pairs (the test set) + +Synthetic validation says whether the model learned the synthetic problem; only photographed pairs +say whether it learned the real one. On a tripod, with remote release and mirror lock-up, manual +focus and white balance: **12 scenes** — low-light interior, a night street, fabric, foliage, fine +text, a colour chart if one is to hand, and a still subject with skin and hair. At each, four +ISO 100 frames at a long exposure (averaged: the reference), then ISO 1600, 3200, 6400, 12800 and +25600 at the same aperture with the shutter shortened by the ISO ratio. A per-channel linear fit +against the reference absorbs residual exposure mismatch (ELD's protocol). + +Plus 100 real library frames above ISO 3200 with no reference, judged by eye side by side. + +### 6.2 Baseline and metrics + +The baseline is today's path: the classical demosaic plus `ops/noise_reduction.rs` tuned by hand +per ISO on the validation set. If a Lightroom or DxO trial is to hand, their output on the same +twelve scenes is the ceiling, for our comparison only. + +Metrics, measured after a fixed tone curve (the camera profile and an sRGB curve) and not in linear +light, where the highlights would dominate: PSNR and SSIM per ISO; chroma bias on flat patches, +because denoisers desaturate; a slanted-edge MTF for detail; and maze or zipper artefacts on the +resolution target at ISO 100. + +### 6.3 Experiments that answer design questions + +| | Question | Runs | +|---|---|---| +| E1 | How much context does denoise need? (§3.1) | Same data, receptive field 12, 36, 100, 300+ photosites; PSNR per ISO against it | +| E2 | Noise-map conditioning or k-sigma? (§3.3) | M both ways | +| E3 | Bin 2×2 or 4×4 for truth? (§4.2) | Compare at ISO 400–800, where it matters | +| E4 | Is the blind noise estimate good enough? (§5.3) | Inference with calibrated vs blind maps on the real pairs | + +### 6.4 Acceptance + +- On the real pairs, ≥ 3 dB over the baseline at ISO 6400, and **no ISO at which it is worse**, + ISO 100 included — at base ISO it has to be at least as good a demosaic as the classical one. +- Mean chroma error on flat patches under ΔE 1. +- No maze, zipper or false colour on the resolution target that the classical demosaic does not + also show. +- A 20 MP frame in ≤ 3 s on the laptop's GPU and ≤ 30 s on its CPU (§8). + +## 7. In the application + +### 7.1 A cache, not a new file + +Lightroom writes a DNG into the library. We do not: the library is synced, a 20 MP linear RGB file +is ~120 MB, and a derived file inside a synced tree is exactly what +[storage.md](storage.md) refuses. Instead: + +- The sidecar records the intent — denoise on, amount, model id — as the rest of the edit is + recorded, so it syncs and another device reproduces it. +- The result is a local cache entry: f16 linear camera RGB, zstd, keyed on + `(file identity, decoder version, model id, amount, noise source)`. ~60–80 MB per frame + (*estimate*), LRU under a budget (default 5 GB, §10). +- On open, the classical demosaic shows at once and the learned result swaps in when it is ready, + with progress over the canvas — the same pattern as a photograph that is only on the server. +- Export needs the result and computes it if the cache has lost it. + +### 7.2 The Amount control + +A Denoise toggle and one Amount slider in develop. Moving the slider runs inference on the +**visible viewport only** (~1 MP, a fraction of a second — *estimate*) so the photographer judges +on the real result; releasing it queues the whole frame. There is no per-frame blend between the +two paths: blending the classical output back in re-adds the noise the network removed. + +### 7.3 Runtime + +Through `dr-inference-engine`, as the other models run ([inference.md](inference.md)): TensorRT or +CUDA fp16 on the laptop, MIGraphX on the desktop, ORT CPU everywhere, QNN on the tablet. Work is +scheduled in the `Background` class so a slider never waits on it (architecture §5.3). + +## 8. Speed and the tablet + +M at ~60 GMAC/MP is ~1.2 TMAC for a 20 MP frame (*estimate*). On the RTX 3050 at fp16 that is +about a second; on 20 CPU threads, tens of seconds. + +The tablet's Hexagon is fast — scrfd_10g's ~10 GFLOP in 3.2 ms, [inference.md §1.1](inference.md) — +but **accepts int8 only**, and int8 is hostile to this task: a 14-bit signal quantised to 256 +levels loses the shadow steps the model exists to recover. Two ways round it, to be measured in +this order: + +1. **Predict the residual, not the image.** S emits the correction to a cheap bilinear demosaic + computed in float outside the graph. The residual spans a few σ, which 256 levels resolve; the + addition happens in float. With a variance-stabilising transform (Anscombe) on the input. +2. **16-bit activations** (QNN's A16W8), if the partition log shows the HTP running them. + +If neither holds S's quality within 0.5 dB of fp32 on the real pairs, **v1 is desktop-only** and the +tablet shows the classical path. The sidecar still records the intent, so a desktop can render the +learned result for a photograph edited on the tablet. + +## 9. X-Trans + +The requirements tie this stage to FR-RAW-5, and the library has no Fuji raws. What we can do +without a Fuji body: + +- **Training does not need one.** §4.2 step 3 samples the binned RGB truth into any pattern. + X-Trans packs 6×6 into 36 channels at a sixth of the resolution, with a 108-channel head: the + same design at the X-Trans period. It is a separate model. +- **Noise does.** A calibration capture (§5.2) or, failing that, the blind estimate — plus the + DNG `NoiseProfile` of converted Fuji files. +- **The test set does.** raw.pixls.us has CC0 samples per body but no tripod ISO ladders. A few + hours with a borrowed X-Trans body and the §6.1 protocol is the honest version; without it, + X-Trans ships marked experimental. + +## 10. Plan and open decisions + +| Phase | Work | Output | +|---|---|---| +| P0 | Calibration capture; the `dr-decode` dump example; source selection and crop store | Noise tables, ~20 GB of crops, the 12-scene test set | +| P1 | M on Bayer; eval harness; E1–E4 | A model that passes §6.4 on the laptop | +| P2 | The stage in `dr-gpu`, cache, sidecar field, develop controls, export | A photograph denoised in the app | +| P3 | S distilled; int8 and the residual head on the tablet | Tablet in or out of v1 (§8) | +| P4 | X-Trans model | Experimental unless a body is borrowed | + +Training lives in a sibling repo, `darkroom-denoise`, next to `darkroom-infill` and reusing its +hydration tools. The weights are trained from scratch on the author's own photographs with an +MIT architecture, so this model adds no third-party licence to D13. + +**Decisions wanted before P1:** + +1. Bin 2×2 or 4×4 for the truth, or both (E3 answers it, but the crop store is built once). +2. Cache budget and location. +3. Whether the tablet is in v1's scope or explicitly deferred behind §8's measurement. +4. Whether a Lightroom or DxO comparison is available for §6.2. +5. A borrowed X-Trans body, or X-Trans experimental in v1. +