Record what the learned denoise shipped as, and what was measured
The grain blend replaces the Amount of §7.2, and why its objection to a blend does not hold for brightness alone; the Hexagon is out (int8 -6 to -9 dB); §11 holds the data, the noise model taken from the library, the model, the validation table, the blind estimate's reach and the speed.
This commit is contained in:
+74
-9
@@ -1,8 +1,9 @@
|
||||
# Learned denoise — joint demosaic and denoise on the mosaic
|
||||
|
||||
Design for **FR-DEV-3g** ([requirements.md](requirements.md)), the learned stage
|
||||
[outstanding.md §3](outstanding.md) says is missing. Draft of 2026-09-27: nothing here is built,
|
||||
and every figure marked *estimate* is waiting for the measurement that replaces it.
|
||||
[outstanding.md §3](outstanding.md) says is missing. Drafted 2026-09-27; a first version shipped
|
||||
in 0.21.0, and §11 records what was built and measured. Figures still marked *estimate* are
|
||||
waiting for the measurement that replaces them.
|
||||
|
||||
---
|
||||
|
||||
@@ -288,17 +289,27 @@ is ~120 MB, and a derived file inside a synced tree is exactly what
|
||||
with progress over the canvas — the same pattern as a photograph that is only on the server.
|
||||
- Export needs the result and computes it if the cache has lost it.
|
||||
|
||||
### 7.2 The Amount control
|
||||
### 7.2 The grain control
|
||||
|
||||
A Denoise toggle and one Amount slider in develop. Moving the slider runs inference on the
|
||||
**visible viewport only** (~1 MP, a fraction of a second — *estimate*) so the photographer judges
|
||||
on the real result; releasing it queues the whole frame. There is no per-frame blend between the
|
||||
two paths: blending the classical output back in re-adds the noise the network removed.
|
||||
What shipped is a switch and a **Keep grain** slider, not the Amount described first. The slider
|
||||
blends the two demosaics per pixel — but only the *brightness* of their difference: `out =
|
||||
denoised + grain · ΔY / wb`, with `ΔY` the luminance of `wb · (classical − denoised)`. Taken after
|
||||
the as-shot balance and handed back divided by it, the grain is neutral in the finished picture.
|
||||
|
||||
The objection that stood here — that blending the classical output back in re-adds the noise —
|
||||
holds for a plain mix, which also brings back the classical path's colour speckle and false
|
||||
colour. A luminance-only blend returns film-like grain and nothing else, and it needs no
|
||||
inference: one elementwise GPU pass (`dr_gpu::GrainBlend`) per slider value, producing a new
|
||||
source the adjust pass draws. Comparing the two on real 6D frames, the user chose this one.
|
||||
|
||||
The σ-map Amount (§3.3) still works — `NoiseModel::scaled` — and stays available for a later
|
||||
"strength" control; its cost is a re-run of the network.
|
||||
|
||||
### 7.3 Runtime
|
||||
|
||||
Through `dr-inference-engine`, as the other models run ([inference.md](inference.md)): TensorRT or
|
||||
CUDA fp16 on the laptop, MIGraphX on the desktop, ORT CPU everywhere, QNN on the tablet. Work is
|
||||
Through `dr-inference-engine`, as the other models run ([inference.md](inference.md)), as
|
||||
`Role::Denoiser`: TensorRT or CUDA fp16 on the laptop, MIGraphX on the desktop, ORT CPU everywhere.
|
||||
**Not the Hexagon** — see §11 — so the tablet runs it on its CPU. Work is
|
||||
scheduled in the `Background` class so a slider never waits on it (architecture §5.3).
|
||||
|
||||
## 8. Speed and the tablet
|
||||
@@ -356,3 +367,57 @@ MIT architecture, so this model adds no third-party licence to D13.
|
||||
4. Whether a Lightroom or DxO comparison is available for §6.2.
|
||||
5. A borrowed X-Trans body, or X-Trans experimental in v1.
|
||||
|
||||
## 11. What shipped in 0.21.0, and what was measured
|
||||
|
||||
**Data.** 500 distinct ISO 50–100 6D frames from the library, over 121 shooting days (bursts and
|
||||
near-duplicate perceptual hashes dropped; 55 frames from held-out days for validation). Read through
|
||||
`dr-gpu`'s `mosaic_dump` example — `dr-decode` and the app's own hot-pixel pass — so the network's
|
||||
input is the mosaic the classical demosaic reads. Truth by 2×2 binning with a Catmull-Rom quarter-pixel
|
||||
shift of red and blue (§4.2). Training lives in `darkroom-denoise`, beside `darkroom-infill`.
|
||||
|
||||
**Noise model (§5), from the library instead of a capture.** Shot gain and read variance per ISO from
|
||||
Adobe's `NoiseProfile` in the converted DNGs; read noise checked against each frame's masked border
|
||||
(agreement within 2–3 % from ISO 125 to 25600); read-noise *shape* taken from the border as quantiles
|
||||
on a tail-dense grid (excess kurtosis up to ~10 at high ISO), with only the photosites the app's
|
||||
hot-pixel rule would remove left out; row noise from the border's row means; **column noise** from
|
||||
the masked rows above the image — about a third of its variance is this sensor's fixed pattern.
|
||||
Third stops are their own rows: ISO 160, 320 and 640 are quieter than their neighbours, as §5.1
|
||||
expected. Training without the column noise left the 6D's vertical stripes in (0.90 DN of 1.01);
|
||||
with it, 0.17 DN.
|
||||
|
||||
**Model.** Not NAFNet: its channel attention averages over the whole input, which breaks exact
|
||||
tiling. A U-Net of 3×3 convolutions, ReLU, strided and transposed convolutions and additive skips —
|
||||
3.2 M parameters, 48 GMAC per raw megapixel, receptive field 185 photosites (counted from the
|
||||
layers; a perturbation probe under-read it as 157 because a switched-off ReLU hides a path).
|
||||
Tiles of 1408 keep their central 1024 behind a 192 halo, exactly.
|
||||
|
||||
**Results.** PSNR after the display transform, held-out days, step 60 000:
|
||||
|
||||
| ISO | Network | Bilinear | Bilinear on a clean mosaic |
|
||||
|---|---|---|---|
|
||||
| 400 | 41.6 | 36.8 | 40.0 |
|
||||
| 1600 | 40.8 | 33.8 | 40.0 |
|
||||
| 6400 | 39.5 | 29.3 | 40.0 |
|
||||
| 25600 | 37.8 | 24.6 | 40.0 |
|
||||
|
||||
Unbiased in linear light on real frames (shadow level within 1 % of a heavily averaged bilinear).
|
||||
Checked against the app's own render for channel and axis order (`tools/check_against_app.py`).
|
||||
|
||||
**Precision (§8).** fp16: 0.00 dB at every ISO. int8 QDQ, calibrated on training tiles: −6 to −9 dB
|
||||
— the shadow steps §8 feared losing are lost. So the Hexagon refuses the role and the tablet runs f32
|
||||
on its CPU; the residual head of §8 is the route back.
|
||||
|
||||
**Noise for any Bayer body (§3.3).** Table, then `NoiseProfile`, then the frame itself: read, row and
|
||||
column noise from its masked border, the shot gain alone estimated from the quietest flat patches.
|
||||
On 130 6D frames the estimate is within ±10 % of the table from ISO 1000 up and scattered below. The
|
||||
network loses under 0.3 dB for σ off by 15–20 % and twice as much for under- as for over-estimating;
|
||||
the estimate leans high. Every Bayer body is offered the switch; develop says which source was used.
|
||||
|
||||
**Speed, a whole 6D frame (20 MP).** TensorRT fp16 3.1 s, ONNX Runtime CPU 14.4 s, on the laptop —
|
||||
measured while the GPU sat power-capped at an 810 MHz memory clock; uncapped is expected to be
|
||||
about four times faster. The Rust path reproduces the training repository's output to 2.5e-4 at
|
||||
worst; TensorRT fp16 is 75 dB from f32.
|
||||
|
||||
**Not yet:** the result is not cached across sessions (§7.1) — reopening recomputes; the tripod real
|
||||
pairs of §6.1; X-Trans (§9); the hand-written WGSL path, for which `export.py` already writes the
|
||||
weights blob and a manifest a shader can follow.
|
||||
|
||||
Reference in New Issue
Block a user