Files
DarkRoom/docs/dev/denoise.md
T
dtourolle c2cfacd7d3 Record the new Best in the spec and the manual
denoise.md §15: why Best became one network, how it compares with the
mixture and Medium on real photographs and the chart, the candidates that
fell short, the file names, the saved-edit numbering, and the timings on
the 3050 -- 0.51-0.54 s whole-frame, 0.95 s in tiles, against 2.60 s for
the mixture in tiles. The manual lists Bilinear, Fast and Best, says an
edit made with Medium opens with Best, and gives the new time; Medium's
close-up goes.
2026-10-07 07:14:52 -04:00

38 KiB
Raw Blame History

Learned denoise — joint demosaic and denoise on the mosaic

Design for FR-DEV-3g (requirements.md), the learned stage outstanding.md §3 says is missing. Drafted 2026-09-27; a first version shipped in 0.21.0, and §11 records what was built and measured. Figures still marked estimate are waiting for the measurement that replaces them.


1. What we are matching

Lightroom's Denoise (April 2023, Eric Chan's "Denoise demystified") is the reference, and three facts about it set the shape of this design:

  • It runs on the mosaic. The network takes Bayer or X-Trans photosites before any demosaic and emits full RGB: denoise and demosaic are one learned step. It descends from Adobe's 2019 learned demosaic (Raw Details). A photograph that is already demosaiced is not eligible.
  • It is run once, not per frame. The result is written as a new linear DNG beside the original, and every later edit reads that file. The amount is chosen once, from a preview crop.
  • It is trained on synthetic pairs. Clean raws with sensor-modelled noise added, not photographed pairs.

The reason the mosaic is the right place is physical: before the demosaic, noise is independent per photosite with a known distribution (shot plus read). After it, the interpolation has correlated that noise into colour blotches many photosites across, which classical noise reduction cannot separate from texture. The same step removes demosaic artefacts — maze, zipper, false colour, X-Trans worms (FR-RAW-5).

We match the first and third facts and not the second: our result is a cache, not a file in the library (§7).

2. Where it sits

architecture.md §5.2 already reserves the slot. The learned stage replaces the demosaic box when it is on; nothing else in the chain moves.

RawImage ─► hot/dead photosites ─► black/white levels ─► ┬─ demosaic (classical) ─┬─► camera profile ─► …
                                                         └─ learned demosaic+NR ──┘
                                                              (cached, §7)
  • In: the repaired, normalised mosaic, from the same buffer Demosaicer::run reads. The hot pixel pass stays in front: an outlier of 50σ is outside anything the noise model generates, and a network shown one invents a structure around it.
  • Out: linear camera RGB, f16, full resolution — exactly the texture the classical demosaic produces, so the camera profile, the raw histogram and every operation below it are unchanged.
  • Off by default, per photograph. The classical path stays the default and the fallback; the stage's absence degrades gracefully, as FR-DEV-3g requires.

3. The model

3.1 The 12×12 → 4×4 question

The proposal: a network that reads a 12×12 window of photosites and predicts the RGB of the central 4×4, slid across the frame in steps of four.

The output half is right. The input half is too small by a factor of five or more.

What is right about it. Predicting a block aligned to the colour-filter period keeps the phase fixed: every prediction sees the same arrangement of red, green and blue around it, so the network never has to work out where it stands in the pattern. It also makes tiling trivial and exact. Both properties are kept below — as the head of the network and as the tiling contract (§3.4).

What is wrong with it. A denoiser can only average away noise it can see around the pixel, and at high ISO it needs to see a long way:

  • The Canon 6D at ISO 6400 (clip ≈ 1,200 e⁻, read noise ≈ 2 e⁻ — estimate, §5 measures it) has a mid-tone of ~150 e⁻, shot SNR ≈ 12, and a shadow three stops down of ~19 e⁻, SNR ≈ 4.
  • A shadow that looks clean wants SNR ≈ 40: a factor of 10, which is ~100 independent same-colour samples in a flat area. Red and blue are a quarter of the photosites, so that is ~400 photosites: a 20×20 window just for a flat shadow, 40×40 two stops further down.
  • A 12×12 window holds 36 red photosites. Averaged perfectly, that is a factor of 6 on red and blue in a flat area, and less everywhere there is structure.
  • Chroma blotches are low-frequency noise — 16 to 64 photosites across. A window smaller than the blotch cannot tell it from a colour change.

Demosaic alone is content with 12×12: good classical demosaics read 5×5 to 9×9. So the proposal is a good demosaic network and a weak denoiser — which is a useful ablation (experiment E1, §6.3).

What it costs. Adjacent 12×12 windows with a 4×4 output overlap nine-fold, so a network evaluated per window recomputes each photosite's features nine times. A convolutional network is the same computation with that work shared: it is "predict the central block from its neighbourhood" evaluated everywhere at once.

3.2 The shape

mosaic (H×W) ──space-to-depth 2×2──► 4 ch @ H/2 × W/2  ┐
noise map σ(x) ─space-to-depth 2×2──► 4 ch @ H/2 × W/2  ┴► U-Net ─► 12 ch @ H/2 × W/2 ─depth-to-space─► RGB @ H×W
                                                                  (2×2 block × RGB per position)
  • Packing. Bayer is packed 2×2 into four channels at half resolution, so every input position is one whole quad and every output position is the 2×2 block of RGB it covers — the proposal's head, at the Bayer period. (A 4×4 packing with a 48-channel head is the same thing at a coarser stride and is a free parameter.)
  • Phase unification. Every body's pattern is cropped by a row or a column to RGGB before packing, and the output is un-cropped. Flips are only used for augmentation in the CFA-preserving form (Liu et al., "Bayer pattern unification and augmentation", 2019).
  • Body. A U-Net with four downsamplings and NAFNet blocks (Chen et al., 2022; MIT). The receptive field at the raw scale is several hundred photosites, which covers §3.1's worst case with room.
  • Two sizes. M (widths 32-64-128-256, ~6 M parameters, ~60 GMAC per raw megapixel — estimate) is the desktop model and the one trained first. S (widths 16-32-64-128, fewer bottleneck blocks, ~1 M parameters, ~12 GMAC/MP) is distilled from M for the tablet (§8).

3.3 Conditioning on the noise

The network is told how noisy each photosite is, rather than learning one model per ISO:

  • A per-photosite standard-deviation map, σ(x) = √(K·x + σ_r²) from the body's gain K and read noise σ_r at that ISO, packed alongside the mosaic (FFDNet's arrangement, Zhang et al., 2018).
  • This is what makes it camera-general. A body it was never trained on only has to supply K and σ_r. Three sources, in order of preference: a calibration table for the body (§5); the DNG NoiseProfile tag, which Adobe's converter writes; a blind estimate from the photograph's own flat regions (Foi et al., 2008), which always exists.
  • It is also the Amount control. Scaling the map up tells the network there is more noise than there is and it smooths harder; scaling it down preserves more grain. Changing the amount re-runs inference (§7.2), which is why it is set on a preview crop, as Lightroom does.

The alternative — PMRID's k-sigma transform, which maps every ISO onto one noise level — is simpler and gives no Amount control. It is the fallback if conditioning underperforms.

3.4 Tiling

A 20 MP frame does not go through a network in one piece on either device. Inference tiles the mosaic into 512×512 input tiles with a 64-photosite halo on every side and keeps the central 384×384 of each output: the proposal's "12 in, 4 out", scaled up. Halo and tile sizes must be multiples of 2 (the CFA phase) and of 16 (four downsamplings at half resolution), so the seams land at identical positions in every tile's own coordinates.

This is inference-local tiling and does not depend on FR-DSP-2's render-path tiling, which stays under the challenge outstanding.md §4 records.

4. Training data

4.1 What the library holds

From the reference catalog, 2026-09-27: 17,255 catalogued RAWs (9,345 DNG, 7,910 CR2), all but seven from one body, the Canon EOS 6D (RGGB Bayer, 5472×3648, AA filter), 166 shooting days from 2015 to 2026.

ISO Frames Use
≤ 200 5,065 Clean sources for synthetic pairs
201–1600 7,807 Low-noise end of the eval set
1601–6400 3,379 Real-noise eval set; noise-model check (§5.3)
> 6400 562 The hard cases, by eye

There are no X-Trans raws, which matters for §9. The catalog does not hold shutter speed, so selection needs the files' EXIF. Whether the DNGs are mosaic (converted CR2) or linear must be checked before they are counted as sources: a linear DNG has no photosites to learn from.

4.2 How a training pair is made

  1. Clean source. A base-ISO 6D frame, black-subtracted and normalised.
  2. Full-colour truth by binning. Each plane is resampled by half a photosite so the four planes share a centre, then every 2×2 quad becomes one RGB pixel (R, mean of the two G, B): a true full-colour image at 2736×1824 with no interpolation in it. This is the only way to have ground truth for the demosaic half.
  3. Re-mosaic. That RGB image is sampled back into an RGGB mosaic. (It can equally be sampled into X-Trans, §9.)
  4. Darken and add noise. Scale the signal by 1/g for a target ISO 100·g, then add noise from the calibrated model at that ISO (§5): Poisson shot, Tukey-lambda read noise, row noise and quantisation — the ELD model (Wei et al., CVPR 2020). The input is this mosaic; the target is the clean RGB at the same scale.
  5. Augment. Random blur (Gaussian, σ 0–0.7 px) before re-mosaicking, because a binned image is sharper per pixel than the AA-filtered sensor the model will see; exposure jitter; white-balance gains within the body's range; CFA-preserving flips.

Why the target's own noise is tolerable. A base-ISO frame is not noise-free, and binning only halves the green noise; red and blue keep theirs. But darkening by g scales signal and target noise together, while the added shot noise grows as √g. At ISO 3200 the input is ≈ 5.7× noisier than its target, at ISO 800 only ≈ 2.8×. L1 against a noisy target converges on the median, which is unbiased for symmetric noise. The low-ISO end is the one at risk of learning to keep grain: if it does, bin 4×4 instead (red and blue noise halved, 1368×912 per source) for those samples.

Why not the native mosaic as the target. That trains denoise alone, with base-ISO noise baked into the answer ("noisier2noise") and no demosaic truth at all.

4.3 How much

The limit is scene diversity, not pixel count; every source yields an effectively unlimited number of pairs through random crops, ISO and noise draws.

Figure Value Reasoning
Sources, train 3,000 5,065 base-ISO frames, less bursts (perceptual-hash dedup), heavy clipping, motion blur and linear DNGs. For scale: ELD reaches state of the art trained on ~230 scenes; SID has ~5,000 pairs of ~400 scenes
Sources, validation 200 Split by shooting day, not by frame, so no scene is on both sides
Pixels ~15 Gpx of RGB truth 3,000 × 5 MP after binning
Crops per step 8–16 × 256×256 photosites Fits a 6 GB RTX 3050 at fp16 with M
Stored ~20 GB 24 random 512×512 crops per source, uint16, zstd. Keeping whole CR2s would be ~75 GB
Training 200–400 k steps, one to two nights per run on the 3050 — estimate; expect three to five runs

Stratify the selection: across all 166 days, and deliberately include faces and hair (the library has 19k detected faces, and skin is where over-smoothing shows first), foliage, fabric, text, and any base-ISO tripod night work.

4.4 Reading raws the same way in training and in the app

The training data must be decoded by the same decoder the app uses. rawpy (LibRaw) and dr_decode::Rawler can disagree on black level, white level, active area and therefore CFA phase, and a network trained on one pattern phase and run on another produces colour moiré everywhere. A dr-decode example that dumps the mosaic and its metadata as .npy is the only source the training repo reads — not rawpy, as darkroom-infill's develop-raws.py does.

5. The noise model and its calibration

5.1 What is measured

Per ISO: gain K (DN per electron), read-noise distribution (Gaussian σ and Tukey-λ shape), row-noise σ, black-level offset and any fixed pattern. Canon's third-stop ISOs on bodies of the 6D's generation are digital gains of the full stops, so noise does not scale smoothly between them; every third stop is calibrated, not interpolated.

5.2 The capture (one hour, once per body)

  • Darks. Lens cap on, viewfinder covered, manual. Five frames at 1/4000 s and five at 1/30 s at every third stop from ISO 100 to 25600. They give read noise, row noise and the black-level pattern; the two shutter speeds confirm dark current is negligible.
  • Flats. An evenly lit white wall, defocused, at every full stop: pairs at six exposure levels from 1/64 of clip to 3/4 of it. The variance of each pair's difference against their mean is the photon transfer curve, whose slope is K.

5.3 The check

Fit the same (K, σ_r) blindly from flat regions of the library's 3,379 ISO 1601–6400 frames (§3.3's third source). If it disagrees with the calibration by more than ~10%, one of them is wrong — and it tells us how far the blind estimate can be trusted for bodies with no calibration.

6. Evaluation

6.1 Real pairs (the test set)

Synthetic validation says whether the model learned the synthetic problem; only photographed pairs say whether it learned the real one. On a tripod, with remote release and mirror lock-up, manual focus and white balance: 12 scenes — low-light interior, a night street, fabric, foliage, fine text, a colour chart if one is to hand, and a still subject with skin and hair. At each, four ISO 100 frames at a long exposure (averaged: the reference), then ISO 1600, 3200, 6400, 12800 and 25600 at the same aperture with the shutter shortened by the ISO ratio. A per-channel linear fit against the reference absorbs residual exposure mismatch (ELD's protocol).

Plus 100 real library frames above ISO 3200 with no reference, judged by eye side by side.

6.2 Baseline and metrics

The baseline is today's path: the classical demosaic plus ops/noise_reduction.rs tuned by hand per ISO on the validation set. If a Lightroom or DxO trial is to hand, their output on the same twelve scenes is the ceiling, for our comparison only.

Metrics, measured after a fixed tone curve (the camera profile and an sRGB curve) and not in linear light, where the highlights would dominate: PSNR and SSIM per ISO; chroma bias on flat patches, because denoisers desaturate; a slanted-edge MTF for detail; and maze or zipper artefacts on the resolution target at ISO 100.

6.3 Experiments that answer design questions

Question Runs
E1 How much context does denoise need? (§3.1) Same data, receptive field 12, 36, 100, 300+ photosites; PSNR per ISO against it
E2 Noise-map conditioning or k-sigma? (§3.3) M both ways
E3 Bin 2×2 or 4×4 for truth? (§4.2) Compare at ISO 400–800, where it matters
E4 Is the blind noise estimate good enough? (§5.3) Inference with calibrated vs blind maps on the real pairs

6.4 Acceptance

  • On the real pairs, ≥ 3 dB over the baseline at ISO 6400, and no ISO at which it is worse, ISO 100 included — at base ISO it has to be at least as good a demosaic as the classical one.
  • Mean chroma error on flat patches under ΔE 1.
  • No maze, zipper or false colour on the resolution target that the classical demosaic does not also show.
  • A 20 MP frame in ≤ 3 s on the laptop's GPU and ≤ 30 s on its CPU (§8).

7. In the application

7.1 A cache, not a new file

Lightroom writes a DNG into the library. We do not: the library is synced, a 20 MP linear RGB file is ~120 MB, and a derived file inside a synced tree is exactly what storage.md refuses. Instead:

  • The sidecar records the intent — denoise on, amount, model id — as the rest of the edit is recorded, so it syncs and another device reproduces it.
  • The result is a local cache entry: f16 linear camera RGB, zstd, keyed on (file identity, decoder version, model id, amount, noise source). ~60–80 MB per frame (estimate), LRU under a budget (default 5 GB, §10).
  • On open, the classical demosaic shows at once and the learned result swaps in when it is ready, with progress over the canvas — the same pattern as a photograph that is only on the server.
  • Export needs the result and computes it if the cache has lost it.

7.2 The grain control

What shipped is a switch and a Keep grain slider, not the Amount described first. The slider blends the two demosaics per pixel — but only the brightness of their difference: out = denoised + grain · ΔY / wb, with ΔY the luminance of wb · (classical − denoised). Taken after the as-shot balance and handed back divided by it, the grain is neutral in the finished picture.

The objection that stood here — that blending the classical output back in re-adds the noise — holds for a plain mix, which also brings back the classical path's colour speckle and false colour. A luminance-only blend returns film-like grain and nothing else, and it needs no inference: one elementwise GPU pass (dr_gpu::GrainBlend) per slider value, producing a new source the adjust pass draws. Comparing the two on real 6D frames, the user chose this one.

The σ-map Amount (§3.3) still works — NoiseModel::scaled — and stays available for a later "strength" control; its cost is a re-run of the network.

7.3 Runtime

Through dr-inference-engine, as the other models run (inference.md), as Role::Denoiser: TensorRT or CUDA fp16 on the laptop, MIGraphX on the desktop, ORT CPU everywhere. Not the Hexagon — see §11 — so the tablet runs it on its CPU. Work is scheduled in the Background class so a slider never waits on it (architecture §5.3).

8. Speed and the tablet

M at ~60 GMAC/MP is ~1.2 TMAC for a 20 MP frame (estimate). On the RTX 3050 at fp16 that is about a second; on 20 CPU threads, tens of seconds.

The tablet's Hexagon is fast — scrfd_10g's ~10 GFLOP in 3.2 ms, inference.md §1.1 — but accepts int8 only, and int8 is hostile to this task: a 14-bit signal quantised to 256 levels loses the shadow steps the model exists to recover. Two ways round it, to be measured in this order:

  1. Predict the residual, not the image. S emits the correction to a cheap bilinear demosaic computed in float outside the graph. The residual spans a few σ, which 256 levels resolve; the addition happens in float. With a variance-stabilising transform (Anscombe) on the input.
  2. 16-bit activations (QNN's A16W8), if the partition log shows the HTP running them.

If neither holds S's quality within 0.5 dB of fp32 on the real pairs, v1 is desktop-only and the tablet shows the classical path. The sidecar still records the intent, so a desktop can render the learned result for a photograph edited on the tablet.

Measured 2026-10-04 (inference.md §1.5): the second way holds, without the first. The shipped network, with its Bayer packing re-spelled as SpaceToDepth so QNN can hold it (the 6-D reshape it replaces is exact but past the HTP's rank limit), at A16W16 — 16-bit activations and weights — scores within 0.00 dB of f32 at ISO 400–25600 on the tablet's own HTP, and within 0.09 dB with the 6D's noise model scaled ×0.5, ×2 and ×4 to stand in for other sensors. A16W8 holds the 6D (worst −0.19 dB at ISO 25600) but not ×4 noise at 25600 (−0.52 dB), so A16W16 is what ships. int8 loses 4.7–9.2 dB and fp16 is refused outright. A 1408 tile takes 95 ms on the Hexagon against 1510 ms on the tablet's CPU: about 2.3 s for a 20 MP frame. Calibration ranges come from 96 training-day tiles across every ISO, a third of them with that scaled noise; coverage of other bodies is that synthetic bracket, not their raws.

9. X-Trans

The requirements tie this stage to FR-RAW-5, and the library has no Fuji raws. What we can do without a Fuji body:

  • Training does not need one. §4.2 step 3 samples the binned RGB truth into any pattern. X-Trans packs 6×6 into 36 channels at a sixth of the resolution, with a 108-channel head: the same design at the X-Trans period. It is a separate model.
  • Noise does. A calibration capture (§5.2) or, failing that, the blind estimate — plus the DNG NoiseProfile of converted Fuji files.
  • The test set does. raw.pixls.us has CC0 samples per body but no tripod ISO ladders. A few hours with a borrowed X-Trans body and the §6.1 protocol is the honest version; without it, X-Trans ships marked experimental.

10. Plan and open decisions

Phase Work Output
P0 Calibration capture; the dr-decode dump example; source selection and crop store Noise tables, ~20 GB of crops, the 12-scene test set
P1 M on Bayer; eval harness; E1–E4 A model that passes §6.4 on the laptop
P2 The stage in dr-gpu, cache, sidecar field, develop controls, export A photograph denoised in the app
P3 S distilled; int8 and the residual head on the tablet Tablet in or out of v1 (§8)
P4 X-Trans model Experimental unless a body is borrowed

Training lives in a sibling repo, darkroom-denoise, next to darkroom-infill and reusing its hydration tools. The weights are trained from scratch on the author's own photographs with an MIT architecture, so this model adds no third-party licence to D13.

Decisions wanted before P1:

  1. Bin 2×2 or 4×4 for the truth, or both (E3 answers it, but the crop store is built once).
  2. Cache budget and location.
  3. Whether the tablet is in v1's scope or explicitly deferred behind §8's measurement.
  4. Whether a Lightroom or DxO comparison is available for §6.2.
  5. A borrowed X-Trans body, or X-Trans experimental in v1.

11. What shipped in 0.21.0, and what was measured

Data. 500 distinct ISO 50–100 6D frames from the library, over 121 shooting days (bursts and near-duplicate perceptual hashes dropped; 55 frames from held-out days for validation). Read through dr-gpu's mosaic_dump example — dr-decode and the app's own hot-pixel pass — so the network's input is the mosaic the classical demosaic reads. Truth by 2×2 binning with a Catmull-Rom quarter-pixel shift of red and blue (§4.2). Training lives in darkroom-denoise, beside darkroom-infill.

Noise model (§5), from the library instead of a capture. Shot gain and read variance per ISO from Adobe's NoiseProfile in the converted DNGs; read noise checked against each frame's masked border (agreement within 2–3 % from ISO 125 to 25600); read-noise shape taken from the border as quantiles on a tail-dense grid (excess kurtosis up to ~10 at high ISO), with only the photosites the app's hot-pixel rule would remove left out; row noise from the border's row means; column noise from the masked rows above the image — about a third of its variance is this sensor's fixed pattern. Third stops are their own rows: ISO 160, 320 and 640 are quieter than their neighbours, as §5.1 expected. Training without the column noise left the 6D's vertical stripes in (0.90 DN of 1.01); with it, 0.17 DN.

Model. Not NAFNet: its channel attention averages over the whole input, which breaks exact tiling. A U-Net of 3×3 convolutions, ReLU, strided and transposed convolutions and additive skips — 3.2 M parameters, 48 GMAC per raw megapixel, receptive field 185 photosites (counted from the layers; a perturbation probe under-read it as 157 because a switched-off ReLU hides a path). Tiles of 1408 keep their central 1024 behind a 192 halo, exactly.

Results. PSNR after the display transform, held-out days, step 60 000:

ISO Network Bilinear Bilinear on a clean mosaic
400 41.6 36.8 40.0
1600 40.8 33.8 40.0
6400 39.5 29.3 40.0
25600 37.8 24.6 40.0

Unbiased in linear light on real frames (shadow level within 1 % of a heavily averaged bilinear). Checked against the app's own render for channel and axis order (tools/check_against_app.py).

Precision (§8). fp16: 0.00 dB at every ISO. int8 QDQ, calibrated on training tiles: −6 to −9 dB — the shadow steps §8 feared losing are lost. So the Hexagon refuses the role and the tablet runs f32 on its CPU; the residual head of §8 is the route back.

Noise for any Bayer body (§3.3). Table, then NoiseProfile, then the frame itself: read, row and column noise from its masked border, the shot gain alone estimated from the quietest flat patches. On 130 6D frames the estimate is within ±10 % of the table from ISO 1000 up and scattered below. The network loses under 0.3 dB for σ off by 15–20 % and twice as much for under- as for over-estimating; the estimate leans high. Every Bayer body is offered the switch; develop says which source was used.

Speed, a whole 6D frame (20 MP). TensorRT fp16 3.1 s, ONNX Runtime CPU 14.4 s, on the laptop — measured while the GPU sat power-capped at an 810 MHz memory clock; uncapped is expected to be about four times faster. The Rust path reproduces the training repository's output to 2.5e-4 at worst; TensorRT fp16 is 75 dB from f32.

Not yet: the result is not cached across sessions (§7.1) — reopening recomputes; done after 0.21.0, §12; the tripod real pairs of §6.1; X-Trans (§9); the hand-written WGSL path, for which export.py already writes the weights blob and a manifest a shader can follow.

12. On by default, with a strength, and cached (after 0.21.0)

The photographer asked for the learned demosaic to be how a raw is developed, not an option found under Detail. So:

  • On by default, at full strength, on every device. The switch is switch_on, so an untouched photograph writes nothing and is developed from the network everywhere; turning it off is the edit. Which hardware runs it is the inference engine's choice (inference.md), not this setting's: the default does not depend on what a device is believed to manage.
  • Strength replaces Keep grain. 0–100, default 100, and grain = 100 − strength, so it is the same luminance-only blend of §7.2 and moving it is one GPU pass, never a re-run. An edit saved by 0.21.0 stored grain; it is still read, as its inverse, and never written.
  • First in the panel, above the lens corrections: it decides what every control below is applied to. Its attribute is still Detail, so it also stays where the Detail tab shows it.
  • Cached on disk (§7.1): the network's output for a file, as half floats (about 120 MB for 20 MP — no compressor to link on Android), keyed on a SHA-256 of the file's bytes and the model file's name and size, oldest first past a 5 GB budget, beside the inference engine's cache under the data root. The strength is applied afterwards and is not in the key. A reopened photograph, and an export of one already developed, read it back instead of recomputing.

What it costs: every raw opened runs the network once, with the classical demosaic shown until the result lands, and a first export of an unopened raw runs it too. Every raw renders differently from 0.21.0 unless switched off.

13. Three networks and a method (after 0.22.0)

The photographer asked for a choice between quality and time. Method replaces the Apply switch: Bilinear, Fast, Medium, Best, by index in that order, default Best. An untouched raw writes nothing and develops through Best. apply is still read and never written: 0 is Bilinear, 1 keeps a network already chosen or is the default. A number past the list, from a newer build, reads as the default. A build before this one ignores method and develops through its own network, which is the most an older peer can do.

The networks (darkroom-denoise, every one trained on the same data and noise as §11, plus 1,201 further frames cropped from the library and 6,000 drawn scenes — polygons, lines of one to four photosites, text, gratings — rendered at 4× through a random affine and smooth displacement, so edges fall off the photosite grid):

Method File Network Parameters GMAC / MP Halo
Best mosaic-best-1408.onnx two U-Nets of §11's shape (a flat expert from m2, an edge expert from the ×100 edge-weighted run) and a 128 k-parameter gate that blends them per photosite 6.4 M 110 256
Medium mosaic-medium-1408.onnx §11's U-Net, distilled from Best (75 % its output, 25 % the truth) 3.2 M 48 192
Fast mosaic-fast-1408.onnx widths 16-32-64-128, blocks 1-1-1-2, distilled the same way 0.93 M 11 192

The gate learned on its own to trust the edge expert at 0.77–0.88 on edges and not at all on flat areas. The mixture's receptive field is the experts' plus the gate's, so it keeps the centre of a 1408 tile past a 256 halo, where the single networks keep 1024 past 192. dr_denoise::Shipped carries each file's halo, and TileNet::halo hands it to the tiler.

Quality. PSNR after the display transform on 1,842 held-out crops, and the width of a hard edge on the drawn chart at ISO 6400 (truth 0.80 photosites; lower is sharper):

ISO 400 1600 6400 25600 Edge width
§11's network 40.61 39.77 38.43 36.67 1.77
Best 40.69 39.84 38.47 36.71 0.82
Medium 40.59 39.75 38.40 36.65 1.30
Fast 39.90 39.06 37.54 35.33 1.84
Bilinear 36.15 33.26 28.72 23.83 2.15

On photographs the three are close; on hard edges Best is half as wide as §11's network and Medium most of the way there. Fast costs a dB at high ISO and edges as soft as §11's.

Speed, a whole 20 MP 6D frame, the network alone, TensorRT fp16 on the laptop's RTX 3050 (uncapped: memory at 5 GHz), engine already built: Best 2.48 s, Medium 0.79 s, Fast 0.57 s. Decode and the hot-pixel pass add about 0.5 s. The first build of each TensorRT engine takes 80 s (Fast) to 190 s (Best), in the background at first launch, cached after.

Before the network, two passes changed since §11.

  • A noise-aware repair (dr_denoise::repair) after the app's hot-pixel pass: a photosite more than 8σ beyond every same-colour neighbour and every adjacent photosite, and more than twice each adjacent one, is clamped to the brightest of its same-colour neighbours; a dead one, to the darkest. The ratio test is what spares a point of light, whose neighbours are lit too. The networks were trained behind the same pass (the Python and Rust agree: 935 repairs on an ISO 25600 frame).
  • The tiler feeds the network without waiting: tiles are gathered on every core by a producer thread one tile ahead, and the output is written back in parallel from the runtime's own buffer. 0.14 s of tiler for a frame, which is what keeps Fast under a second.

The Hexagon. Each network has an .a16w16.onnx sibling made by tools/quantise-models.sh --ranges, the ranges from darkroom-3e's gate (96 training tiles, a third at noise ×2 and ×4). On the 6D gate A16W16 loses 0.00 dB for all three; with the noise scaled ×0.5–×4 at most 0.11 dB. A16W8 holds the gate (≤ 0.27 dB) but loses 0.63 dB on Medium at ×4, so A16W16 stays the form.

Cache. Each network keys its own results (§7.1 keys on the model's file name), and the file is hashed once at open, so changing the method never re-reads it. Choosing Bilinear keeps the network's result in memory for the way back; changing to another network drops it, and coming back reads the cache.

Packaging. All six files in the APK (BUNDLED, 23 entries, +44.6 MB, ~41 MB compressed); the three f32 networks in the Arch package and the Windows installer, which stage models/denoise by directory.

14. A whole frame, not 1408² tiles (after 0.23.0)

A fixed 1408² tile is exact only past its halo, and Best's halo is 256: of every 1408² it computes it keeps 896², 2.47 photosites of work for each one kept (Medium and Fast keep 1024², 1.89×). On a GPU the network can instead run over the whole frame and its reflected border in one call, which is exact by the same argument (§3.4) and wastes only the border.

The networks are re-exported with any height and width (mosaic-{best,medium,fast}.onnx beside the 1408 files; darkroom-denoise tools/export_whole.py), from the checkpoints the shipped files came from. The tool refuses unless each matches its 1408 file at 1408² (max |Δ| = 0 for all three), matches torch at 592 × 848, and equals tiled inference over the reflected frame in f64 (≤ 7e-16). The APK leaves them out: the Hexagon takes fixed shapes.

Where they run. Role::WholeDenoiser is served by TensorRT and the CUDA provider only, the rungs where a new input size costs nothing at run time; MIGraphX, OpenVINO and CoreML compile per shape, the Hexagon takes fixed shapes, and the CPU would hold gigabytes of f32 activations. Everywhere else whole_frame_limit() is None and the 1408² tiles run as before. TensorRT gets an optimisation profile up to WHOLE_FRAME_MAX (4608 × 3328) — without one a dynamic input compiles a new engine per size at run time — through the runtime's V2 options, since ort's builder has none, and keeps the engine in a directory per model and profile (ONNX Runtime's cache key leaves the shape out).

The limit is the card's memory. TensorRT plans its memory for the profile's largest shape. A profile up to a whole 6D frame with Best's border (4608 × 6656) asked for 4.9–5.9 GB and would not build on the 6 GB RTX 3050. At 15 MP the tiler (tile::plan) cuts the frame into the fewest equal tiles under the limit: a 6D frame is two of 4160 × 3248, 27 MP of work for 20 MP kept, against 49 MP in 1408² tiles. If a plan's first call fails, as a GPU out of memory does, its kept centre is halved and the frame planned again.

Measured 2026-10-06 on _MG_8862 (6D, ISO 8000, 20 MP), RTX 3050 Laptop, TensorRT fp16, P3 / 5001 MHz, another session's paused training holding 1.3 GB:

Best Network time Against the tiles
1408² tiles 2.60 s —
Whole frame, two 4160 × 3248 tiles 1.37 s max |Δ| 0.0029, mean 1.1e-5 — fp16's own spread (GPU tiles against CPU tiles: 0.0025)

The first build of the whole-frame engine took 28 minutes, in the background at first launch, with the 1408² tiles serving meanwhile — against about 3 minutes for the fixed one; the profile's range is what it tunes across. A cached engine loads in about a second.

In PyTorch fp16 the same network over the whole 20 MP frame in one call took 3.5× less than in tiles, so a card that holds a whole frame gains more than the 6 GB one does; WHOLE_FRAME_MAX is a constant sized for 6 GB until the limit follows the card's memory.

15. Best becomes one network (0.24)

The photographer's goal for 0.24 was Best's quality in under a second on the laptop. Whole frames (§14) took the mixture from 2.60 s to 1.37 s and no further on a 6 GB card, so the other half was a single network that holds the mixture's quality at a third of its work. Methods are now Bilinear, Fast and Best; Medium and the mixture are retired.

The network is fb-combo (darkroom-denoise, 2026-10-07): §11's shape (32-64-128-192, blocks 1-1-2-2, 3.2 M parameters, 48 GMAC/MP, halo 192), 20 000 steps from fb-edges2 ← student-m, taught by the mixture at a half share, with 10 % drawn scenes and 25 % crops from the edge-rich cells of the training frames (branch edge-sampling). Scored on real photographs — the chart overstated the mixture's lead (a chart-sharp network was softer than Medium on real edges) — on the validation crops in the top quarter for sharp detail:

Edge PSNR, ISO 1600 / 6400 / 25600 Sharpness kept Smooth areas Held-out PSNR, ISO 400 / 1600 / 6400 / 25600 Chart edge
Mixture (Best to 0.23) 30.71 / 29.93 / 28.55 0.899 / 0.868 / 0.782 42.61 / 41.71 / 40.12 40.69 / 39.84 / 38.47 / 36.71 0.82
fb-combo (Best from 0.24) 30.67 / 29.87 / 28.49 0.902 / 0.874 / 0.792 42.54 / 41.57 / 39.85 40.64 / 39.78 / 38.38 / 36.54 0.89
Medium (to 0.23) 30.43 / 29.68 / 28.38 0.896 / 0.862 / 0.776 42.59 / 41.69 / 40.08 40.59 / 39.75 / 38.40 / 36.65 1.30

Edges within 0.04–0.06 dB and more sharpness kept at every ISO; the known shortfall is smooth areas at ISO 25600, 0.27 dB. The photographer took it as it stood at 20 000 of a planned 30 000 steps. Others tried on the way, each short of the mixture on real photographs: fb-sharp (drawn scenes, chart-sharp but Medium's real edges), fb-edges (half edge-rich crops: edges close, ISO 25600 flats −0.24 dB), fb-edges2 (a quarter: 0.03–0.14 dB short everywhere, chart 1.33–1.47), and a from-scratch 24-48-96-128 between Fast and Medium.

Files. mosaic-hq-1408.onnx, mosaic-hq.onnx (any size) and mosaic-hq-1408.a16w16.onnx for the Hexagon. A new name, not Medium's or Best's: the result cache keys a model by name and size, and this one is byte for byte Medium's size. The tablet form lost 0.00 dB in simulated QDQ at every ISO and at most 0.09 dB across the ×0.5–×4 noise bracket (A16W8 0.08 / 0.26 dB; int8 −10.6 dB); not yet confirmed on the tablet itself.

Saved edits keep their numbers: 2, which was Medium, is now Best; 3, which was Best, is past the end and reads as the default, Best. Both land on the new network with no migration.

Measured 2026-10-07, _MG_8862, RTX 3050 Laptop, TensorRT fp16, P3 / 5001 MHz, nothing else on the card:

Best Network time Peak GPU memory
mixture, 1408² tiles (0.23) 2.60 s —
mixture, whole frame (§14) 1.37 s —
fb-combo, 1408² tiles 0.95 s 0.55 GB
fb-combo, whole frame (two 4160 × 3248) 0.51–0.54 s 1.75 GB

Decode and the hot-pixel pass add 0.4–0.5 s, so a photograph is about a second end to end. Whole frame against tiles: max |Δ| 0.0029, 90 dB apart — fp16's spread. The whole-frame engine's first build took 12 minutes (the mixture's 28); from the cache it loads in about a second, so the session keeps the engine's ordinary 30 s idle decay rather than unloading after each photograph: at 1.75 GB it fits beside the develop view on a 6 GB card, and an unload would cost the next photograph a second.

The manual's close-up for Best is still the mixture's render, which this network matches to within the table above; it is re-recorded with the next pass of tools/manual/record.sh.