A design draft for the learned stage outstanding.md §3 lists as missing: a joint demosaic-and-denoise network that runs on the mosaic, in the slot architecture.md §5.2 reserves for it, with its result kept as a cache rather than a new file in the library. It covers the model and tiling, synthetic training pairs from the library and a per-body noise calibration, evaluation on real pairs, the Amount control, speed on the tablet, X-Trans, and the decisions still open. Nothing in it is built; figures marked "estimate" wait for the measurements that replace them.
359 lines
20 KiB
Markdown
359 lines
20 KiB
Markdown
# Learned denoise — joint demosaic and denoise on the mosaic
|
||
|
||
Design for **FR-DEV-3g** ([requirements.md](requirements.md)), the learned stage
|
||
[outstanding.md §3](outstanding.md) says is missing. Draft of 2026-09-27: nothing here is built,
|
||
and every figure marked *estimate* is waiting for the measurement that replaces it.
|
||
|
||
---
|
||
|
||
## 1. What we are matching
|
||
|
||
Lightroom's Denoise (April 2023, Eric Chan's "Denoise demystified") is the reference, and three
|
||
facts about it set the shape of this design:
|
||
|
||
- **It runs on the mosaic.** The network takes Bayer or X-Trans photosites before any demosaic
|
||
and emits full RGB: denoise and demosaic are one learned step. It descends from Adobe's 2019
|
||
learned demosaic (Raw Details). A photograph that is already demosaiced is not eligible.
|
||
- **It is run once, not per frame.** The result is written as a new linear DNG beside the
|
||
original, and every later edit reads that file. The amount is chosen once, from a preview crop.
|
||
- **It is trained on synthetic pairs.** Clean raws with sensor-modelled noise added, not
|
||
photographed pairs.
|
||
|
||
The reason the mosaic is the right place is physical: before the demosaic, noise is independent
|
||
per photosite with a known distribution (shot plus read). After it, the interpolation has
|
||
correlated that noise into colour blotches many photosites across, which classical noise reduction
|
||
cannot separate from texture. The same step removes demosaic artefacts
|
||
— maze, zipper, false colour, X-Trans worms (FR-RAW-5).
|
||
|
||
We match the first and third facts and not the second: our result is a cache, not a file in the
|
||
library (§7).
|
||
|
||
## 2. Where it sits
|
||
|
||
[architecture.md §5.2](architecture.md) already reserves the slot. The learned stage **replaces
|
||
the demosaic box** when it is on; nothing else in the chain moves.
|
||
|
||
```
|
||
RawImage ─► hot/dead photosites ─► black/white levels ─► ┬─ demosaic (classical) ─┬─► camera profile ─► …
|
||
└─ learned demosaic+NR ──┘
|
||
(cached, §7)
|
||
```
|
||
|
||
- **In:** the repaired, normalised mosaic, from the same buffer `Demosaicer::run` reads. The hot
|
||
pixel pass stays in front: an outlier of 50σ is outside anything the noise model generates, and
|
||
a network shown one invents a structure around it.
|
||
- **Out:** linear camera RGB, f16, full resolution — exactly the texture the classical demosaic
|
||
produces, so the camera profile, the raw histogram and every operation below it are unchanged.
|
||
- **Off by default, per photograph.** The classical path stays the default and the fallback; the
|
||
stage's absence degrades gracefully, as FR-DEV-3g requires.
|
||
|
||
## 3. The model
|
||
|
||
### 3.1 The 12×12 → 4×4 question
|
||
|
||
The proposal: a network that reads a 12×12 window of photosites and predicts the RGB of the
|
||
central 4×4, slid across the frame in steps of four.
|
||
|
||
**The output half is right. The input half is too small by a factor of five or more.**
|
||
|
||
*What is right about it.* Predicting a block aligned to the colour-filter period keeps the phase
|
||
fixed: every prediction sees the same arrangement of red, green and blue around it, so the network
|
||
never has to work out where it stands in the pattern. It also makes tiling trivial and exact.
|
||
Both properties are kept below — as the head of the network and as the tiling contract (§3.4).
|
||
|
||
*What is wrong with it.* A denoiser can only average away noise it can see around the pixel, and
|
||
at high ISO it needs to see a long way:
|
||
|
||
- The Canon 6D at ISO 6400 (clip ≈ 1,200 e⁻, read noise ≈ 2 e⁻ — *estimate*, §5 measures it) has a
|
||
mid-tone of ~150 e⁻, shot SNR ≈ 12, and a shadow three stops down of ~19 e⁻, SNR ≈ 4.
|
||
- A shadow that looks clean wants SNR ≈ 40: a factor of 10, which is ~100 independent same-colour
|
||
samples in a flat area. Red and blue are a quarter of the photosites, so that is ~400
|
||
photosites: a **20×20 window just for a flat shadow**, 40×40 two stops further down.
|
||
- A 12×12 window holds 36 red photosites. Averaged perfectly, that is a factor of 6 on red and
|
||
blue in a flat area, and less everywhere there is structure.
|
||
- Chroma blotches are low-frequency noise — 16 to 64 photosites across. A window smaller than the
|
||
blotch cannot tell it from a colour change.
|
||
|
||
Demosaic alone is content with 12×12: good classical demosaics read 5×5 to 9×9. So the proposal is
|
||
a good demosaic network and a weak denoiser — which is a useful ablation (experiment E1, §6.3).
|
||
|
||
*What it costs.* Adjacent 12×12 windows with a 4×4 output overlap nine-fold, so a network
|
||
evaluated per window recomputes each photosite's features nine times. A convolutional network is
|
||
the same computation with that work shared: it is "predict the central block from its
|
||
neighbourhood" evaluated everywhere at once.
|
||
|
||
### 3.2 The shape
|
||
|
||
```
|
||
mosaic (H×W) ──space-to-depth 2×2──► 4 ch @ H/2 × W/2 ┐
|
||
noise map σ(x) ─space-to-depth 2×2──► 4 ch @ H/2 × W/2 ┴► U-Net ─► 12 ch @ H/2 × W/2 ─depth-to-space─► RGB @ H×W
|
||
(2×2 block × RGB per position)
|
||
```
|
||
|
||
- **Packing.** Bayer is packed 2×2 into four channels at half resolution, so every input position
|
||
is one whole quad and every output position is the 2×2 block of RGB it covers — the proposal's
|
||
head, at the Bayer period. (A 4×4 packing with a 48-channel head is the same thing at a coarser
|
||
stride and is a free parameter.)
|
||
- **Phase unification.** Every body's pattern is cropped by a row or a column to RGGB before
|
||
packing, and the output is un-cropped. Flips are only used for augmentation in the CFA-preserving
|
||
form (Liu et al., "Bayer pattern unification and augmentation", 2019).
|
||
- **Body.** A U-Net with four downsamplings and NAFNet blocks (Chen et al., 2022; MIT). The
|
||
receptive field at the raw scale is several hundred photosites, which covers §3.1's worst case
|
||
with room.
|
||
- **Two sizes.** **M** (widths 32-64-128-256, ~6 M parameters, ~60 GMAC per raw megapixel —
|
||
*estimate*) is the desktop model and the one trained first. **S** (widths 16-32-64-128, fewer
|
||
bottleneck blocks, ~1 M parameters, ~12 GMAC/MP) is distilled from M for the tablet (§8).
|
||
|
||
### 3.3 Conditioning on the noise
|
||
|
||
The network is told how noisy each photosite is, rather than learning one model per ISO:
|
||
|
||
- A per-photosite standard-deviation map, `σ(x) = √(K·x + σ_r²)` from the body's gain `K` and read
|
||
noise `σ_r` at that ISO, packed alongside the mosaic (FFDNet's arrangement, Zhang et al., 2018).
|
||
- **This is what makes it camera-general.** A body it was never trained on only has to supply
|
||
`K` and `σ_r`. Three sources, in order of preference: a calibration table for the body (§5); the
|
||
DNG `NoiseProfile` tag, which Adobe's converter writes; a blind estimate from the photograph's
|
||
own flat regions (Foi et al., 2008), which always exists.
|
||
- **It is also the Amount control.** Scaling the map up tells the network there is more noise than
|
||
there is and it smooths harder; scaling it down preserves more grain. Changing the amount re-runs
|
||
inference (§7.2), which is why it is set on a preview crop, as Lightroom does.
|
||
|
||
The alternative — PMRID's k-sigma transform, which maps every ISO onto one noise level — is
|
||
simpler and gives no Amount control. It is the fallback if conditioning underperforms.
|
||
|
||
### 3.4 Tiling
|
||
|
||
A 20 MP frame does not go through a network in one piece on either device. Inference tiles the
|
||
mosaic into 512×512 input tiles with a 64-photosite halo on every side and keeps the central
|
||
384×384 of each output: the proposal's "12 in, 4 out", scaled up. Halo and tile sizes must be
|
||
multiples of 2 (the CFA phase) and of 16 (four downsamplings at half resolution), so the seams
|
||
land at identical positions in every tile's own coordinates.
|
||
|
||
This is inference-local tiling and does not depend on FR-DSP-2's render-path tiling, which stays
|
||
under the challenge [outstanding.md §4](outstanding.md) records.
|
||
|
||
## 4. Training data
|
||
|
||
### 4.1 What the library holds
|
||
|
||
From the reference catalog, 2026-09-27: 17,255 catalogued RAWs (9,345 DNG, 7,910 CR2), **all but
|
||
seven from one body, the Canon EOS 6D** (RGGB Bayer, 5472×3648, AA filter), 166 shooting days from
|
||
2015 to 2026.
|
||
|
||
| ISO | Frames | Use |
|
||
|---|---|---|
|
||
| ≤ 200 | 5,065 | Clean sources for synthetic pairs |
|
||
| 201–1600 | 7,807 | Low-noise end of the eval set |
|
||
| 1601–6400 | 3,379 | Real-noise eval set; noise-model check (§5.3) |
|
||
| > 6400 | 562 | The hard cases, by eye |
|
||
|
||
There are **no X-Trans raws**, which matters for §9. The catalog does not hold shutter speed, so
|
||
selection needs the files' EXIF. Whether the DNGs are mosaic (converted CR2) or linear must be
|
||
checked before they are counted as sources: a linear DNG has no photosites to learn from.
|
||
|
||
### 4.2 How a training pair is made
|
||
|
||
1. **Clean source.** A base-ISO 6D frame, black-subtracted and normalised.
|
||
2. **Full-colour truth by binning.** Each plane is resampled by half a photosite so the four
|
||
planes share a centre, then every 2×2 quad becomes one RGB pixel (R, mean of the two G, B):
|
||
a true full-colour image at 2736×1824 with no interpolation in it. This is the only way to have
|
||
ground truth for the demosaic half.
|
||
3. **Re-mosaic.** That RGB image is sampled back into an RGGB mosaic. (It can equally be sampled
|
||
into X-Trans, §9.)
|
||
4. **Darken and add noise.** Scale the signal by `1/g` for a target ISO `100·g`, then add noise
|
||
from the calibrated model at that ISO (§5): Poisson shot, Tukey-lambda read noise, row noise
|
||
and quantisation — the ELD model (Wei et al., CVPR 2020). The input is this mosaic; the target
|
||
is the clean RGB at the same scale.
|
||
5. **Augment.** Random blur (Gaussian, σ 0–0.7 px) before re-mosaicking, because a binned image is
|
||
sharper per pixel than the AA-filtered sensor the model will see; exposure jitter; white-balance
|
||
gains within the body's range; CFA-preserving flips.
|
||
|
||
**Why the target's own noise is tolerable.** A base-ISO frame is not noise-free, and binning only
|
||
halves the green noise; red and blue keep theirs. But darkening by `g` scales signal and target
|
||
noise together, while the added shot noise grows as `√g`. At ISO 3200 the input is ≈ 5.7× noisier
|
||
than its target, at ISO 800 only ≈ 2.8×. L1 against a noisy target converges on the median, which
|
||
is unbiased for symmetric noise. The low-ISO end is the one at risk of learning to keep grain: if
|
||
it does, bin 4×4 instead (red and blue noise halved, 1368×912 per source) for those samples.
|
||
|
||
**Why not the native mosaic as the target.** That trains denoise alone, with base-ISO noise baked
|
||
into the answer ("noisier2noise") and no demosaic truth at all.
|
||
|
||
### 4.3 How much
|
||
|
||
The limit is scene diversity, not pixel count; every source yields an effectively unlimited
|
||
number of pairs through random crops, ISO and noise draws.
|
||
|
||
| Figure | Value | Reasoning |
|
||
|---|---|---|
|
||
| Sources, train | **3,000** | 5,065 base-ISO frames, less bursts (perceptual-hash dedup), heavy clipping, motion blur and linear DNGs. For scale: ELD reaches state of the art trained on ~230 scenes; SID has ~5,000 pairs of ~400 scenes |
|
||
| Sources, validation | 200 | Split by shooting day, not by frame, so no scene is on both sides |
|
||
| Pixels | ~15 Gpx of RGB truth | 3,000 × 5 MP after binning |
|
||
| Crops per step | 8–16 × 256×256 photosites | Fits a 6 GB RTX 3050 at fp16 with M |
|
||
| Stored | ~20 GB | 24 random 512×512 crops per source, uint16, zstd. Keeping whole CR2s would be ~75 GB |
|
||
| Training | 200–400 k steps, one to two nights per run on the 3050 — *estimate*; expect three to five runs | |
|
||
|
||
Stratify the selection: across all 166 days, and deliberately include faces and hair (the library
|
||
has 19k detected faces, and skin is where over-smoothing shows first), foliage, fabric, text, and
|
||
any base-ISO tripod night work.
|
||
|
||
### 4.4 Reading raws the same way in training and in the app
|
||
|
||
The training data must be decoded by **the same decoder the app uses**. rawpy (LibRaw) and
|
||
`dr_decode::Rawler` can disagree on black level, white level, active area and therefore CFA phase,
|
||
and a network trained on one pattern phase and run on another produces colour moiré everywhere.
|
||
A `dr-decode` example that dumps the mosaic and its metadata as `.npy` is the only source the
|
||
training repo reads — not rawpy, as `darkroom-infill`'s `develop-raws.py` does.
|
||
|
||
## 5. The noise model and its calibration
|
||
|
||
### 5.1 What is measured
|
||
|
||
Per ISO: gain `K` (DN per electron), read-noise distribution (Gaussian σ and Tukey-λ shape),
|
||
row-noise σ, black-level offset and any fixed pattern. Canon's third-stop ISOs on bodies of the
|
||
6D's generation are digital gains of the full stops, so noise does not scale smoothly between
|
||
them; **every third stop is calibrated**, not interpolated.
|
||
|
||
### 5.2 The capture (one hour, once per body)
|
||
|
||
- **Darks.** Lens cap on, viewfinder covered, manual. Five frames at 1/4000 s and five at 1/30 s
|
||
at every third stop from ISO 100 to 25600. They give read noise, row noise and the black-level
|
||
pattern; the two shutter speeds confirm dark current is negligible.
|
||
- **Flats.** An evenly lit white wall, defocused, at every full stop: pairs at six exposure levels
|
||
from 1/64 of clip to 3/4 of it. The variance of each pair's difference against their mean is
|
||
the photon transfer curve, whose slope is `K`.
|
||
|
||
### 5.3 The check
|
||
|
||
Fit the same `(K, σ_r)` blindly from flat regions of the library's 3,379 ISO 1601–6400 frames
|
||
(§3.3's third source). If it disagrees with the calibration by more than ~10%, one of them is
|
||
wrong — and it tells us how far the blind estimate can be trusted for bodies with no calibration.
|
||
|
||
## 6. Evaluation
|
||
|
||
### 6.1 Real pairs (the test set)
|
||
|
||
Synthetic validation says whether the model learned the synthetic problem; only photographed pairs
|
||
say whether it learned the real one. On a tripod, with remote release and mirror lock-up, manual
|
||
focus and white balance: **12 scenes** — low-light interior, a night street, fabric, foliage, fine
|
||
text, a colour chart if one is to hand, and a still subject with skin and hair. At each, four
|
||
ISO 100 frames at a long exposure (averaged: the reference), then ISO 1600, 3200, 6400, 12800 and
|
||
25600 at the same aperture with the shutter shortened by the ISO ratio. A per-channel linear fit
|
||
against the reference absorbs residual exposure mismatch (ELD's protocol).
|
||
|
||
Plus 100 real library frames above ISO 3200 with no reference, judged by eye side by side.
|
||
|
||
### 6.2 Baseline and metrics
|
||
|
||
The baseline is today's path: the classical demosaic plus `ops/noise_reduction.rs` tuned by hand
|
||
per ISO on the validation set. If a Lightroom or DxO trial is to hand, their output on the same
|
||
twelve scenes is the ceiling, for our comparison only.
|
||
|
||
Metrics, measured after a fixed tone curve (the camera profile and an sRGB curve) and not in linear
|
||
light, where the highlights would dominate: PSNR and SSIM per ISO; chroma bias on flat patches,
|
||
because denoisers desaturate; a slanted-edge MTF for detail; and maze or zipper artefacts on the
|
||
resolution target at ISO 100.
|
||
|
||
### 6.3 Experiments that answer design questions
|
||
|
||
| | Question | Runs |
|
||
|---|---|---|
|
||
| E1 | How much context does denoise need? (§3.1) | Same data, receptive field 12, 36, 100, 300+ photosites; PSNR per ISO against it |
|
||
| E2 | Noise-map conditioning or k-sigma? (§3.3) | M both ways |
|
||
| E3 | Bin 2×2 or 4×4 for truth? (§4.2) | Compare at ISO 400–800, where it matters |
|
||
| E4 | Is the blind noise estimate good enough? (§5.3) | Inference with calibrated vs blind maps on the real pairs |
|
||
|
||
### 6.4 Acceptance
|
||
|
||
- On the real pairs, ≥ 3 dB over the baseline at ISO 6400, and **no ISO at which it is worse**,
|
||
ISO 100 included — at base ISO it has to be at least as good a demosaic as the classical one.
|
||
- Mean chroma error on flat patches under ΔE 1.
|
||
- No maze, zipper or false colour on the resolution target that the classical demosaic does not
|
||
also show.
|
||
- A 20 MP frame in ≤ 3 s on the laptop's GPU and ≤ 30 s on its CPU (§8).
|
||
|
||
## 7. In the application
|
||
|
||
### 7.1 A cache, not a new file
|
||
|
||
Lightroom writes a DNG into the library. We do not: the library is synced, a 20 MP linear RGB file
|
||
is ~120 MB, and a derived file inside a synced tree is exactly what
|
||
[storage.md](storage.md) refuses. Instead:
|
||
|
||
- The sidecar records the intent — denoise on, amount, model id — as the rest of the edit is
|
||
recorded, so it syncs and another device reproduces it.
|
||
- The result is a local cache entry: f16 linear camera RGB, zstd, keyed on
|
||
`(file identity, decoder version, model id, amount, noise source)`. ~60–80 MB per frame
|
||
(*estimate*), LRU under a budget (default 5 GB, §10).
|
||
- On open, the classical demosaic shows at once and the learned result swaps in when it is ready,
|
||
with progress over the canvas — the same pattern as a photograph that is only on the server.
|
||
- Export needs the result and computes it if the cache has lost it.
|
||
|
||
### 7.2 The Amount control
|
||
|
||
A Denoise toggle and one Amount slider in develop. Moving the slider runs inference on the
|
||
**visible viewport only** (~1 MP, a fraction of a second — *estimate*) so the photographer judges
|
||
on the real result; releasing it queues the whole frame. There is no per-frame blend between the
|
||
two paths: blending the classical output back in re-adds the noise the network removed.
|
||
|
||
### 7.3 Runtime
|
||
|
||
Through `dr-inference-engine`, as the other models run ([inference.md](inference.md)): TensorRT or
|
||
CUDA fp16 on the laptop, MIGraphX on the desktop, ORT CPU everywhere, QNN on the tablet. Work is
|
||
scheduled in the `Background` class so a slider never waits on it (architecture §5.3).
|
||
|
||
## 8. Speed and the tablet
|
||
|
||
M at ~60 GMAC/MP is ~1.2 TMAC for a 20 MP frame (*estimate*). On the RTX 3050 at fp16 that is
|
||
about a second; on 20 CPU threads, tens of seconds.
|
||
|
||
The tablet's Hexagon is fast — scrfd_10g's ~10 GFLOP in 3.2 ms, [inference.md §1.1](inference.md) —
|
||
but **accepts int8 only**, and int8 is hostile to this task: a 14-bit signal quantised to 256
|
||
levels loses the shadow steps the model exists to recover. Two ways round it, to be measured in
|
||
this order:
|
||
|
||
1. **Predict the residual, not the image.** S emits the correction to a cheap bilinear demosaic
|
||
computed in float outside the graph. The residual spans a few σ, which 256 levels resolve; the
|
||
addition happens in float. With a variance-stabilising transform (Anscombe) on the input.
|
||
2. **16-bit activations** (QNN's A16W8), if the partition log shows the HTP running them.
|
||
|
||
If neither holds S's quality within 0.5 dB of fp32 on the real pairs, **v1 is desktop-only** and the
|
||
tablet shows the classical path. The sidecar still records the intent, so a desktop can render the
|
||
learned result for a photograph edited on the tablet.
|
||
|
||
## 9. X-Trans
|
||
|
||
The requirements tie this stage to FR-RAW-5, and the library has no Fuji raws. What we can do
|
||
without a Fuji body:
|
||
|
||
- **Training does not need one.** §4.2 step 3 samples the binned RGB truth into any pattern.
|
||
X-Trans packs 6×6 into 36 channels at a sixth of the resolution, with a 108-channel head: the
|
||
same design at the X-Trans period. It is a separate model.
|
||
- **Noise does.** A calibration capture (§5.2) or, failing that, the blind estimate — plus the
|
||
DNG `NoiseProfile` of converted Fuji files.
|
||
- **The test set does.** raw.pixls.us has CC0 samples per body but no tripod ISO ladders. A few
|
||
hours with a borrowed X-Trans body and the §6.1 protocol is the honest version; without it,
|
||
X-Trans ships marked experimental.
|
||
|
||
## 10. Plan and open decisions
|
||
|
||
| Phase | Work | Output |
|
||
|---|---|---|
|
||
| P0 | Calibration capture; the `dr-decode` dump example; source selection and crop store | Noise tables, ~20 GB of crops, the 12-scene test set |
|
||
| P1 | M on Bayer; eval harness; E1–E4 | A model that passes §6.4 on the laptop |
|
||
| P2 | The stage in `dr-gpu`, cache, sidecar field, develop controls, export | A photograph denoised in the app |
|
||
| P3 | S distilled; int8 and the residual head on the tablet | Tablet in or out of v1 (§8) |
|
||
| P4 | X-Trans model | Experimental unless a body is borrowed |
|
||
|
||
Training lives in a sibling repo, `darkroom-denoise`, next to `darkroom-infill` and reusing its
|
||
hydration tools. The weights are trained from scratch on the author's own photographs with an
|
||
MIT architecture, so this model adds no third-party licence to D13.
|
||
|
||
**Decisions wanted before P1:**
|
||
|
||
1. Bin 2×2 or 4×4 for the truth, or both (E3 answers it, but the crop store is built once).
|
||
2. Cache budget and location.
|
||
3. Whether the tablet is in v1's scope or explicitly deferred behind §8's measurement.
|
||
4. Whether a Lightroom or DxO comparison is available for §6.2.
|
||
5. A borrowed X-Trans body, or X-Trans experimental in v1.
|
||
|