From 4b71ef0947972816ab55e9549439ec7b1bd5a2f8 Mon Sep 17 00:00:00 2001 From: Duncan Tourolle Date: Tue, 6 Oct 2026 23:14:11 -0400 Subject: [PATCH] Record whole-frame denoise in the spec and the model licences MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit denoise.md §14: why the tiles waste half of Best's work, where the any-size networks run and why only there, why the limit is the card's memory, and the measurement on _MG_8862 — 2.60 s in 1408 tiles, 1.37 s in two 4160 x 3248 tiles, the outputs within fp16's own spread. models/LICENCE.md lists the three re-exports. --- docs/dev/denoise.md | 45 +++++++++++++++++++++++++++++++++++++++++++++ models/LICENCE.md | 1 + 2 files changed, 46 insertions(+) diff --git a/docs/dev/denoise.md b/docs/dev/denoise.md index c01d6b5..4cb0985 100644 --- a/docs/dev/denoise.md +++ b/docs/dev/denoise.md @@ -526,3 +526,48 @@ reads the cache. **Packaging.** All six files in the APK (`BUNDLED`, 23 entries, +44.6 MB, ~41 MB compressed); the three f32 networks in the Arch package and the Windows installer, which stage `models/denoise` by directory. + +## 14. A whole frame, not 1408² tiles (after 0.23.0) + +A fixed 1408² tile is exact only past its halo, and Best's halo is 256: of every 1408² it computes +it keeps 896², 2.47 photosites of work for each one kept (Medium and Fast keep 1024², 1.89×). On a +GPU the network can instead run over the whole frame and its reflected border in one call, which is +exact by the same argument (§3.4) and wastes only the border. + +**The networks** are re-exported with any height and width (`mosaic-{best,medium,fast}.onnx` beside +the 1408 files; darkroom-denoise `tools/export_whole.py`), from the checkpoints the shipped files +came from. The tool refuses unless each matches its 1408 file at 1408² (max |Δ| = 0 for all three), +matches torch at 592 × 848, and equals tiled inference over the reflected frame in f64 (≤ 7e-16). +The APK leaves them out: the Hexagon takes fixed shapes. + +**Where they run.** `Role::WholeDenoiser` is served by TensorRT and the CUDA provider only, the rungs +where a new input size costs nothing at run time; MIGraphX, OpenVINO and CoreML compile per shape, +the Hexagon takes fixed shapes, and the CPU would hold gigabytes of f32 activations. Everywhere else +`whole_frame_limit()` is `None` and the 1408² tiles run as before. TensorRT gets an optimisation +profile up to `WHOLE_FRAME_MAX` (4608 × 3328) — without one a dynamic input compiles a new engine per +size at run time — through the runtime's V2 options, since `ort`'s builder has none, and keeps the +engine in a directory per model and profile (ONNX Runtime's cache key leaves the shape out). + +**The limit is the card's memory.** TensorRT plans its memory for the profile's largest shape. A +profile up to a whole 6D frame with Best's border (4608 × 6656) asked for 4.9–5.9 GB and would not +build on the 6 GB RTX 3050. At 15 MP the tiler (`tile::plan`) cuts the frame into the fewest equal +tiles under the limit: a 6D frame is two of 4160 × 3248, 27 MP of work for 20 MP kept, against 49 MP +in 1408² tiles. If a plan's first call fails, as a GPU out of memory does, its kept centre is halved +and the frame planned again. + +**Measured** 2026-10-06 on `_MG_8862` (6D, ISO 8000, 20 MP), RTX 3050 Laptop, TensorRT fp16, P3 / +5001 MHz, another session's paused training holding 1.3 GB: + +| Best | Network time | Against the tiles | +|---|---|---| +| 1408² tiles | 2.60 s | — | +| Whole frame, two 4160 × 3248 tiles | **1.37 s** | max \|Δ\| 0.0029, mean 1.1e-5 — fp16's own spread (GPU tiles against CPU tiles: 0.0025) | + +The first build of the whole-frame engine took 28 minutes, in the background at first launch, with +the 1408² tiles serving meanwhile — against about 3 minutes for the fixed one; the profile's range +is what it tunes across. A cached engine loads in about a second. + +In PyTorch fp16 the same network over the whole 20 MP frame in one call took 3.5× less than in +tiles, so a card that holds a whole frame gains more than the 6 GB one does; `WHOLE_FRAME_MAX` is a +constant sized for 6 GB until the limit follows the card's memory. + diff --git a/models/LICENCE.md b/models/LICENCE.md index bef223b..fb29745 100644 --- a/models/LICENCE.md +++ b/models/LICENCE.md @@ -137,6 +137,7 @@ declined, and the InsightFace grant of D13). | `denoise/mosaic-best-1408.onnx` | trained in the `darkroom-denoise` repository (2026-10-04, run `final`, 30 000 steps, from the experts of runs `m2` and `edges-100`) | 1,701 of the maintainer's own base-ISO raws and 6,000 synthetic scenes the repository draws itself, with the Canon EOS 6D's measured noise added | the learned demosaic and denoise, Best (FR-DEV-3g) | | `denoise/mosaic-medium-1408.onnx` | distilled from `final` in the same repository (2026-10-04, run `student-m`, 20 000 steps, from `m2`) | the same | Medium | | `denoise/mosaic-fast-1408.onnx` | distilled from `final` (2026-10-04, run `student-s`, 30 000 steps, from scratch) | the same | Fast | +| `denoise/mosaic-{best,medium,fast}.onnx` | the three networks above, re-exported with any height and width by `tools/export_whole.py` in the same repository (2026-10-06) from the same checkpoints; identical to the 1408 files at 1408² | the same | the same methods, a whole frame at a time on a GPU (denoise.md §14) | U-Nets of plain 3×3 convolutions, ReLU, strided and transposed convolutions and additive skips — no third-party architecture code or