# Learned denoise — joint demosaic and denoise on the mosaic Design for **FR-DEV-3g** ([requirements.md](requirements.md)), the learned stage [outstanding.md §3](outstanding.md) says is missing. Drafted 2026-09-27; a first version shipped in 0.21.0, and §11 records what was built and measured. Figures still marked *estimate* are waiting for the measurement that replaces them. --- ## 1. What we are matching Lightroom's Denoise (April 2023, Eric Chan's "Denoise demystified") is the reference, and three facts about it set the shape of this design: - **It runs on the mosaic.** The network takes Bayer or X-Trans photosites before any demosaic and emits full RGB: denoise and demosaic are one learned step. It descends from Adobe's 2019 learned demosaic (Raw Details). A photograph that is already demosaiced is not eligible. - **It is run once, not per frame.** The result is written as a new linear DNG beside the original, and every later edit reads that file. The amount is chosen once, from a preview crop. - **It is trained on synthetic pairs.** Clean raws with sensor-modelled noise added, not photographed pairs. The reason the mosaic is the right place is physical: before the demosaic, noise is independent per photosite with a known distribution (shot plus read). After it, the interpolation has correlated that noise into colour blotches many photosites across, which classical noise reduction cannot separate from texture. The same step removes demosaic artefacts — maze, zipper, false colour, X-Trans worms (FR-RAW-5). We match the first and third facts and not the second: our result is a cache, not a file in the library (§7). ## 2. Where it sits [architecture.md §5.2](architecture.md) already reserves the slot. The learned stage **replaces the demosaic box** when it is on; nothing else in the chain moves. ``` RawImage ─► hot/dead photosites ─► black/white levels ─► ┬─ demosaic (classical) ─┬─► camera profile ─► … └─ learned demosaic+NR ──┘ (cached, §7) ``` - **In:** the repaired, normalised mosaic, from the same buffer `Demosaicer::run` reads. The hot pixel pass stays in front: an outlier of 50σ is outside anything the noise model generates, and a network shown one invents a structure around it. - **Out:** linear camera RGB, f16, full resolution — exactly the texture the classical demosaic produces, so the camera profile, the raw histogram and every operation below it are unchanged. - **Off by default, per photograph.** The classical path stays the default and the fallback; the stage's absence degrades gracefully, as FR-DEV-3g requires. ## 3. The model ### 3.1 The 12×12 → 4×4 question The proposal: a network that reads a 12×12 window of photosites and predicts the RGB of the central 4×4, slid across the frame in steps of four. **The output half is right. The input half is too small by a factor of five or more.** *What is right about it.* Predicting a block aligned to the colour-filter period keeps the phase fixed: every prediction sees the same arrangement of red, green and blue around it, so the network never has to work out where it stands in the pattern. It also makes tiling trivial and exact. Both properties are kept below — as the head of the network and as the tiling contract (§3.4). *What is wrong with it.* A denoiser can only average away noise it can see around the pixel, and at high ISO it needs to see a long way: - The Canon 6D at ISO 6400 (clip ≈ 1,200 e⁻, read noise ≈ 2 e⁻ — *estimate*, §5 measures it) has a mid-tone of ~150 e⁻, shot SNR ≈ 12, and a shadow three stops down of ~19 e⁻, SNR ≈ 4. - A shadow that looks clean wants SNR ≈ 40: a factor of 10, which is ~100 independent same-colour samples in a flat area. Red and blue are a quarter of the photosites, so that is ~400 photosites: a **20×20 window just for a flat shadow**, 40×40 two stops further down. - A 12×12 window holds 36 red photosites. Averaged perfectly, that is a factor of 6 on red and blue in a flat area, and less everywhere there is structure. - Chroma blotches are low-frequency noise — 16 to 64 photosites across. A window smaller than the blotch cannot tell it from a colour change. Demosaic alone is content with 12×12: good classical demosaics read 5×5 to 9×9. So the proposal is a good demosaic network and a weak denoiser — which is a useful ablation (experiment E1, §6.3). *What it costs.* Adjacent 12×12 windows with a 4×4 output overlap nine-fold, so a network evaluated per window recomputes each photosite's features nine times. A convolutional network is the same computation with that work shared: it is "predict the central block from its neighbourhood" evaluated everywhere at once. ### 3.2 The shape ``` mosaic (H×W) ──space-to-depth 2×2──► 4 ch @ H/2 × W/2 ┐ noise map σ(x) ─space-to-depth 2×2──► 4 ch @ H/2 × W/2 ┴► U-Net ─► 12 ch @ H/2 × W/2 ─depth-to-space─► RGB @ H×W (2×2 block × RGB per position) ``` - **Packing.** Bayer is packed 2×2 into four channels at half resolution, so every input position is one whole quad and every output position is the 2×2 block of RGB it covers — the proposal's head, at the Bayer period. (A 4×4 packing with a 48-channel head is the same thing at a coarser stride and is a free parameter.) - **Phase unification.** Every body's pattern is cropped by a row or a column to RGGB before packing, and the output is un-cropped. Flips are only used for augmentation in the CFA-preserving form (Liu et al., "Bayer pattern unification and augmentation", 2019). - **Body.** A U-Net with four downsamplings and NAFNet blocks (Chen et al., 2022; MIT). The receptive field at the raw scale is several hundred photosites, which covers §3.1's worst case with room. - **Two sizes.** **M** (widths 32-64-128-256, ~6 M parameters, ~60 GMAC per raw megapixel — *estimate*) is the desktop model and the one trained first. **S** (widths 16-32-64-128, fewer bottleneck blocks, ~1 M parameters, ~12 GMAC/MP) is distilled from M for the tablet (§8). ### 3.3 Conditioning on the noise The network is told how noisy each photosite is, rather than learning one model per ISO: - A per-photosite standard-deviation map, `σ(x) = √(K·x + σ_r²)` from the body's gain `K` and read noise `σ_r` at that ISO, packed alongside the mosaic (FFDNet's arrangement, Zhang et al., 2018). - **This is what makes it camera-general.** A body it was never trained on only has to supply `K` and `σ_r`. Three sources, in order of preference: a calibration table for the body (§5); the DNG `NoiseProfile` tag, which Adobe's converter writes; a blind estimate from the photograph's own flat regions (Foi et al., 2008), which always exists. - **It is also the Amount control.** Scaling the map up tells the network there is more noise than there is and it smooths harder; scaling it down preserves more grain. Changing the amount re-runs inference (§7.2), which is why it is set on a preview crop, as Lightroom does. The alternative — PMRID's k-sigma transform, which maps every ISO onto one noise level — is simpler and gives no Amount control. It is the fallback if conditioning underperforms. ### 3.4 Tiling A 20 MP frame does not go through a network in one piece on either device. Inference tiles the mosaic into 512×512 input tiles with a 64-photosite halo on every side and keeps the central 384×384 of each output: the proposal's "12 in, 4 out", scaled up. Halo and tile sizes must be multiples of 2 (the CFA phase) and of 16 (four downsamplings at half resolution), so the seams land at identical positions in every tile's own coordinates. This is inference-local tiling and does not depend on FR-DSP-2's render-path tiling, which stays under the challenge [outstanding.md §4](outstanding.md) records. ## 4. Training data ### 4.1 What the library holds From the reference catalog, 2026-09-27: 17,255 catalogued RAWs (9,345 DNG, 7,910 CR2), **all but seven from one body, the Canon EOS 6D** (RGGB Bayer, 5472×3648, AA filter), 166 shooting days from 2015 to 2026. | ISO | Frames | Use | |---|---|---| | ≤ 200 | 5,065 | Clean sources for synthetic pairs | | 201–1600 | 7,807 | Low-noise end of the eval set | | 1601–6400 | 3,379 | Real-noise eval set; noise-model check (§5.3) | | > 6400 | 562 | The hard cases, by eye | There are **no X-Trans raws**, which matters for §9. The catalog does not hold shutter speed, so selection needs the files' EXIF. Whether the DNGs are mosaic (converted CR2) or linear must be checked before they are counted as sources: a linear DNG has no photosites to learn from. ### 4.2 How a training pair is made 1. **Clean source.** A base-ISO 6D frame, black-subtracted and normalised. 2. **Full-colour truth by binning.** Each plane is resampled by half a photosite so the four planes share a centre, then every 2×2 quad becomes one RGB pixel (R, mean of the two G, B): a true full-colour image at 2736×1824 with no interpolation in it. This is the only way to have ground truth for the demosaic half. 3. **Re-mosaic.** That RGB image is sampled back into an RGGB mosaic. (It can equally be sampled into X-Trans, §9.) 4. **Darken and add noise.** Scale the signal by `1/g` for a target ISO `100·g`, then add noise from the calibrated model at that ISO (§5): Poisson shot, Tukey-lambda read noise, row noise and quantisation — the ELD model (Wei et al., CVPR 2020). The input is this mosaic; the target is the clean RGB at the same scale. 5. **Augment.** Random blur (Gaussian, σ 0–0.7 px) before re-mosaicking, because a binned image is sharper per pixel than the AA-filtered sensor the model will see; exposure jitter; white-balance gains within the body's range; CFA-preserving flips. **Why the target's own noise is tolerable.** A base-ISO frame is not noise-free, and binning only halves the green noise; red and blue keep theirs. But darkening by `g` scales signal and target noise together, while the added shot noise grows as `√g`. At ISO 3200 the input is ≈ 5.7× noisier than its target, at ISO 800 only ≈ 2.8×. L1 against a noisy target converges on the median, which is unbiased for symmetric noise. The low-ISO end is the one at risk of learning to keep grain: if it does, bin 4×4 instead (red and blue noise halved, 1368×912 per source) for those samples. **Why not the native mosaic as the target.** That trains denoise alone, with base-ISO noise baked into the answer ("noisier2noise") and no demosaic truth at all. ### 4.3 How much The limit is scene diversity, not pixel count; every source yields an effectively unlimited number of pairs through random crops, ISO and noise draws. | Figure | Value | Reasoning | |---|---|---| | Sources, train | **3,000** | 5,065 base-ISO frames, less bursts (perceptual-hash dedup), heavy clipping, motion blur and linear DNGs. For scale: ELD reaches state of the art trained on ~230 scenes; SID has ~5,000 pairs of ~400 scenes | | Sources, validation | 200 | Split by shooting day, not by frame, so no scene is on both sides | | Pixels | ~15 Gpx of RGB truth | 3,000 × 5 MP after binning | | Crops per step | 8–16 × 256×256 photosites | Fits a 6 GB RTX 3050 at fp16 with M | | Stored | ~20 GB | 24 random 512×512 crops per source, uint16, zstd. Keeping whole CR2s would be ~75 GB | | Training | 200–400 k steps, one to two nights per run on the 3050 — *estimate*; expect three to five runs | | Stratify the selection: across all 166 days, and deliberately include faces and hair (the library has 19k detected faces, and skin is where over-smoothing shows first), foliage, fabric, text, and any base-ISO tripod night work. ### 4.4 Reading raws the same way in training and in the app The training data must be decoded by **the same decoder the app uses**. rawpy (LibRaw) and `dr_decode::Rawler` can disagree on black level, white level, active area and therefore CFA phase, and a network trained on one pattern phase and run on another produces colour moiré everywhere. A `dr-decode` example that dumps the mosaic and its metadata as `.npy` is the only source the training repo reads — not rawpy, as `darkroom-infill`'s `develop-raws.py` does. ## 5. The noise model and its calibration ### 5.1 What is measured Per ISO: gain `K` (DN per electron), read-noise distribution (Gaussian σ and Tukey-λ shape), row-noise σ, black-level offset and any fixed pattern. Canon's third-stop ISOs on bodies of the 6D's generation are digital gains of the full stops, so noise does not scale smoothly between them; **every third stop is calibrated**, not interpolated. ### 5.2 The capture (one hour, once per body) - **Darks.** Lens cap on, viewfinder covered, manual. Five frames at 1/4000 s and five at 1/30 s at every third stop from ISO 100 to 25600. They give read noise, row noise and the black-level pattern; the two shutter speeds confirm dark current is negligible. - **Flats.** An evenly lit white wall, defocused, at every full stop: pairs at six exposure levels from 1/64 of clip to 3/4 of it. The variance of each pair's difference against their mean is the photon transfer curve, whose slope is `K`. ### 5.3 The check Fit the same `(K, σ_r)` blindly from flat regions of the library's 3,379 ISO 1601–6400 frames (§3.3's third source). If it disagrees with the calibration by more than ~10%, one of them is wrong — and it tells us how far the blind estimate can be trusted for bodies with no calibration. ## 6. Evaluation ### 6.1 Real pairs (the test set) Synthetic validation says whether the model learned the synthetic problem; only photographed pairs say whether it learned the real one. On a tripod, with remote release and mirror lock-up, manual focus and white balance: **12 scenes** — low-light interior, a night street, fabric, foliage, fine text, a colour chart if one is to hand, and a still subject with skin and hair. At each, four ISO 100 frames at a long exposure (averaged: the reference), then ISO 1600, 3200, 6400, 12800 and 25600 at the same aperture with the shutter shortened by the ISO ratio. A per-channel linear fit against the reference absorbs residual exposure mismatch (ELD's protocol). Plus 100 real library frames above ISO 3200 with no reference, judged by eye side by side. ### 6.2 Baseline and metrics The baseline is today's path: the classical demosaic plus `ops/noise_reduction.rs` tuned by hand per ISO on the validation set. If a Lightroom or DxO trial is to hand, their output on the same twelve scenes is the ceiling, for our comparison only. Metrics, measured after a fixed tone curve (the camera profile and an sRGB curve) and not in linear light, where the highlights would dominate: PSNR and SSIM per ISO; chroma bias on flat patches, because denoisers desaturate; a slanted-edge MTF for detail; and maze or zipper artefacts on the resolution target at ISO 100. ### 6.3 Experiments that answer design questions | | Question | Runs | |---|---|---| | E1 | How much context does denoise need? (§3.1) | Same data, receptive field 12, 36, 100, 300+ photosites; PSNR per ISO against it | | E2 | Noise-map conditioning or k-sigma? (§3.3) | M both ways | | E3 | Bin 2×2 or 4×4 for truth? (§4.2) | Compare at ISO 400–800, where it matters | | E4 | Is the blind noise estimate good enough? (§5.3) | Inference with calibrated vs blind maps on the real pairs | ### 6.4 Acceptance - On the real pairs, ≥ 3 dB over the baseline at ISO 6400, and **no ISO at which it is worse**, ISO 100 included — at base ISO it has to be at least as good a demosaic as the classical one. - Mean chroma error on flat patches under ΔE 1. - No maze, zipper or false colour on the resolution target that the classical demosaic does not also show. - A 20 MP frame in ≤ 3 s on the laptop's GPU and ≤ 30 s on its CPU (§8). ## 7. In the application ### 7.1 A cache, not a new file Lightroom writes a DNG into the library. We do not: the library is synced, a 20 MP linear RGB file is ~120 MB, and a derived file inside a synced tree is exactly what [storage.md](storage.md) refuses. Instead: - The sidecar records the intent — denoise on, amount, model id — as the rest of the edit is recorded, so it syncs and another device reproduces it. - The result is a local cache entry: f16 linear camera RGB, zstd, keyed on `(file identity, decoder version, model id, amount, noise source)`. ~60–80 MB per frame (*estimate*), LRU under a budget (default 5 GB, §10). - On open, the classical demosaic shows at once and the learned result swaps in when it is ready, with progress over the canvas — the same pattern as a photograph that is only on the server. - Export needs the result and computes it if the cache has lost it. ### 7.2 The grain control What shipped is a switch and a **Keep grain** slider, not the Amount described first. The slider blends the two demosaics per pixel — but only the *brightness* of their difference: `out = denoised + grain · ΔY / wb`, with `ΔY` the luminance of `wb · (classical − denoised)`. Taken after the as-shot balance and handed back divided by it, the grain is neutral in the finished picture. The objection that stood here — that blending the classical output back in re-adds the noise — holds for a plain mix, which also brings back the classical path's colour speckle and false colour. A luminance-only blend returns film-like grain and nothing else, and it needs no inference: one elementwise GPU pass (`dr_gpu::GrainBlend`) per slider value, producing a new source the adjust pass draws. Comparing the two on real 6D frames, the user chose this one. The σ-map Amount (§3.3) still works — `NoiseModel::scaled` — and stays available for a later "strength" control; its cost is a re-run of the network. ### 7.3 Runtime Through `dr-inference-engine`, as the other models run ([inference.md](inference.md)), as `Role::Denoiser`: TensorRT or CUDA fp16 on the laptop, MIGraphX on the desktop, ORT CPU everywhere. **Not the Hexagon** — see §11 — so the tablet runs it on its CPU. Work is scheduled in the `Background` class so a slider never waits on it (architecture §5.3). ## 8. Speed and the tablet M at ~60 GMAC/MP is ~1.2 TMAC for a 20 MP frame (*estimate*). On the RTX 3050 at fp16 that is about a second; on 20 CPU threads, tens of seconds. The tablet's Hexagon is fast — scrfd_10g's ~10 GFLOP in 3.2 ms, [inference.md §1.1](inference.md) — but **accepts int8 only**, and int8 is hostile to this task: a 14-bit signal quantised to 256 levels loses the shadow steps the model exists to recover. Two ways round it, to be measured in this order: 1. **Predict the residual, not the image.** S emits the correction to a cheap bilinear demosaic computed in float outside the graph. The residual spans a few σ, which 256 levels resolve; the addition happens in float. With a variance-stabilising transform (Anscombe) on the input. 2. **16-bit activations** (QNN's A16W8), if the partition log shows the HTP running them. If neither holds S's quality within 0.5 dB of fp32 on the real pairs, **v1 is desktop-only** and the tablet shows the classical path. The sidecar still records the intent, so a desktop can render the learned result for a photograph edited on the tablet. ## 9. X-Trans The requirements tie this stage to FR-RAW-5, and the library has no Fuji raws. What we can do without a Fuji body: - **Training does not need one.** §4.2 step 3 samples the binned RGB truth into any pattern. X-Trans packs 6×6 into 36 channels at a sixth of the resolution, with a 108-channel head: the same design at the X-Trans period. It is a separate model. - **Noise does.** A calibration capture (§5.2) or, failing that, the blind estimate — plus the DNG `NoiseProfile` of converted Fuji files. - **The test set does.** raw.pixls.us has CC0 samples per body but no tripod ISO ladders. A few hours with a borrowed X-Trans body and the §6.1 protocol is the honest version; without it, X-Trans ships marked experimental. ## 10. Plan and open decisions | Phase | Work | Output | |---|---|---| | P0 | Calibration capture; the `dr-decode` dump example; source selection and crop store | Noise tables, ~20 GB of crops, the 12-scene test set | | P1 | M on Bayer; eval harness; E1–E4 | A model that passes §6.4 on the laptop | | P2 | The stage in `dr-gpu`, cache, sidecar field, develop controls, export | A photograph denoised in the app | | P3 | S distilled; int8 and the residual head on the tablet | Tablet in or out of v1 (§8) | | P4 | X-Trans model | Experimental unless a body is borrowed | Training lives in a sibling repo, `darkroom-denoise`, next to `darkroom-infill` and reusing its hydration tools. The weights are trained from scratch on the author's own photographs with an MIT architecture, so this model adds no third-party licence to D13. **Decisions wanted before P1:** 1. Bin 2×2 or 4×4 for the truth, or both (E3 answers it, but the crop store is built once). 2. Cache budget and location. 3. Whether the tablet is in v1's scope or explicitly deferred behind §8's measurement. 4. Whether a Lightroom or DxO comparison is available for §6.2. 5. A borrowed X-Trans body, or X-Trans experimental in v1. ## 11. What shipped in 0.21.0, and what was measured **Data.** 500 distinct ISO 50–100 6D frames from the library, over 121 shooting days (bursts and near-duplicate perceptual hashes dropped; 55 frames from held-out days for validation). Read through `dr-gpu`'s `mosaic_dump` example — `dr-decode` and the app's own hot-pixel pass — so the network's input is the mosaic the classical demosaic reads. Truth by 2×2 binning with a Catmull-Rom quarter-pixel shift of red and blue (§4.2). Training lives in `darkroom-denoise`, beside `darkroom-infill`. **Noise model (§5), from the library instead of a capture.** Shot gain and read variance per ISO from Adobe's `NoiseProfile` in the converted DNGs; read noise checked against each frame's masked border (agreement within 2–3 % from ISO 125 to 25600); read-noise *shape* taken from the border as quantiles on a tail-dense grid (excess kurtosis up to ~10 at high ISO), with only the photosites the app's hot-pixel rule would remove left out; row noise from the border's row means; **column noise** from the masked rows above the image — about a third of its variance is this sensor's fixed pattern. Third stops are their own rows: ISO 160, 320 and 640 are quieter than their neighbours, as §5.1 expected. Training without the column noise left the 6D's vertical stripes in (0.90 DN of 1.01); with it, 0.17 DN. **Model.** Not NAFNet: its channel attention averages over the whole input, which breaks exact tiling. A U-Net of 3×3 convolutions, ReLU, strided and transposed convolutions and additive skips — 3.2 M parameters, 48 GMAC per raw megapixel, receptive field 185 photosites (counted from the layers; a perturbation probe under-read it as 157 because a switched-off ReLU hides a path). Tiles of 1408 keep their central 1024 behind a 192 halo, exactly. **Results.** PSNR after the display transform, held-out days, step 60 000: | ISO | Network | Bilinear | Bilinear on a clean mosaic | |---|---|---|---| | 400 | 41.6 | 36.8 | 40.0 | | 1600 | 40.8 | 33.8 | 40.0 | | 6400 | 39.5 | 29.3 | 40.0 | | 25600 | 37.8 | 24.6 | 40.0 | Unbiased in linear light on real frames (shadow level within 1 % of a heavily averaged bilinear). Checked against the app's own render for channel and axis order (`tools/check_against_app.py`). **Precision (§8).** fp16: 0.00 dB at every ISO. int8 QDQ, calibrated on training tiles: −6 to −9 dB — the shadow steps §8 feared losing are lost. So the Hexagon refuses the role and the tablet runs f32 on its CPU; the residual head of §8 is the route back. **Noise for any Bayer body (§3.3).** Table, then `NoiseProfile`, then the frame itself: read, row and column noise from its masked border, the shot gain alone estimated from the quietest flat patches. On 130 6D frames the estimate is within ±10 % of the table from ISO 1000 up and scattered below. The network loses under 0.3 dB for σ off by 15–20 % and twice as much for under- as for over-estimating; the estimate leans high. Every Bayer body is offered the switch; develop says which source was used. **Speed, a whole 6D frame (20 MP).** TensorRT fp16 3.1 s, ONNX Runtime CPU 14.4 s, on the laptop — measured while the GPU sat power-capped at an 810 MHz memory clock; uncapped is expected to be about four times faster. The Rust path reproduces the training repository's output to 2.5e-4 at worst; TensorRT fp16 is 75 dB from f32. **Not yet:** ~~the result is not cached across sessions (§7.1) — reopening recomputes;~~ done after 0.21.0, §12; the tripod real pairs of §6.1; X-Trans (§9); the hand-written WGSL path, for which `export.py` already writes the weights blob and a manifest a shader can follow. ## 12. On by default, with a strength, and cached (after 0.21.0) The photographer asked for the learned demosaic to be how a raw is developed, not an option found under Detail. So: - **On by default, at full strength, on every device.** The switch is `switch_on`, so an untouched photograph writes nothing and is developed from the network everywhere; turning it off is the edit. Which hardware runs it is the inference engine's choice (inference.md), not this setting's: the default does not depend on what a device is believed to manage. - **Strength replaces Keep grain.** 0–100, default 100, and grain = 100 − strength, so it is the same luminance-only blend of §7.2 and moving it is one GPU pass, never a re-run. An edit saved by 0.21.0 stored `grain`; it is still read, as its inverse, and never written. - **First in the panel**, above the lens corrections: it decides what every control below is applied to. Its attribute is still Detail, so it also stays where the Detail tab shows it. - **Cached on disk** (§7.1): the network's output for a file, as half floats (about 120 MB for 20 MP — no compressor to link on Android), keyed on a SHA-256 of the file's bytes and the model file's name and size, oldest first past a 5 GB budget, beside the inference engine's cache under the data root. The strength is applied afterwards and is not in the key. A reopened photograph, and an export of one already developed, read it back instead of recomputing. What it costs: every raw opened runs the network once, with the classical demosaic shown until the result lands, and a first export of an unopened raw runs it too. Every raw renders differently from 0.21.0 unless switched off.