Record whole-frame denoise in the spec and the model licences
denoise.md §14: why the tiles waste half of Best's work, where the any-size networks run and why only there, why the limit is the card's memory, and the measurement on _MG_8862 — 2.60 s in 1408 tiles, 1.37 s in two 4160 x 3248 tiles, the outputs within fp16's own spread. models/LICENCE.md lists the three re-exports.
This commit is contained in:
@@ -526,3 +526,48 @@ reads the cache.
|
||||
**Packaging.** All six files in the APK (`BUNDLED`, 23 entries, +44.6 MB, ~41 MB compressed); the
|
||||
three f32 networks in the Arch package and the Windows installer, which stage `models/denoise` by
|
||||
directory.
|
||||
|
||||
## 14. A whole frame, not 1408² tiles (after 0.23.0)
|
||||
|
||||
A fixed 1408² tile is exact only past its halo, and Best's halo is 256: of every 1408² it computes
|
||||
it keeps 896², 2.47 photosites of work for each one kept (Medium and Fast keep 1024², 1.89×). On a
|
||||
GPU the network can instead run over the whole frame and its reflected border in one call, which is
|
||||
exact by the same argument (§3.4) and wastes only the border.
|
||||
|
||||
**The networks** are re-exported with any height and width (`mosaic-{best,medium,fast}.onnx` beside
|
||||
the 1408 files; darkroom-denoise `tools/export_whole.py`), from the checkpoints the shipped files
|
||||
came from. The tool refuses unless each matches its 1408 file at 1408² (max |Δ| = 0 for all three),
|
||||
matches torch at 592 × 848, and equals tiled inference over the reflected frame in f64 (≤ 7e-16).
|
||||
The APK leaves them out: the Hexagon takes fixed shapes.
|
||||
|
||||
**Where they run.** `Role::WholeDenoiser` is served by TensorRT and the CUDA provider only, the rungs
|
||||
where a new input size costs nothing at run time; MIGraphX, OpenVINO and CoreML compile per shape,
|
||||
the Hexagon takes fixed shapes, and the CPU would hold gigabytes of f32 activations. Everywhere else
|
||||
`whole_frame_limit()` is `None` and the 1408² tiles run as before. TensorRT gets an optimisation
|
||||
profile up to `WHOLE_FRAME_MAX` (4608 × 3328) — without one a dynamic input compiles a new engine per
|
||||
size at run time — through the runtime's V2 options, since `ort`'s builder has none, and keeps the
|
||||
engine in a directory per model and profile (ONNX Runtime's cache key leaves the shape out).
|
||||
|
||||
**The limit is the card's memory.** TensorRT plans its memory for the profile's largest shape. A
|
||||
profile up to a whole 6D frame with Best's border (4608 × 6656) asked for 4.9–5.9 GB and would not
|
||||
build on the 6 GB RTX 3050. At 15 MP the tiler (`tile::plan`) cuts the frame into the fewest equal
|
||||
tiles under the limit: a 6D frame is two of 4160 × 3248, 27 MP of work for 20 MP kept, against 49 MP
|
||||
in 1408² tiles. If a plan's first call fails, as a GPU out of memory does, its kept centre is halved
|
||||
and the frame planned again.
|
||||
|
||||
**Measured** 2026-10-06 on `_MG_8862` (6D, ISO 8000, 20 MP), RTX 3050 Laptop, TensorRT fp16, P3 /
|
||||
5001 MHz, another session's paused training holding 1.3 GB:
|
||||
|
||||
| Best | Network time | Against the tiles |
|
||||
|---|---|---|
|
||||
| 1408² tiles | 2.60 s | — |
|
||||
| Whole frame, two 4160 × 3248 tiles | **1.37 s** | max \|Δ\| 0.0029, mean 1.1e-5 — fp16's own spread (GPU tiles against CPU tiles: 0.0025) |
|
||||
|
||||
The first build of the whole-frame engine took 28 minutes, in the background at first launch, with
|
||||
the 1408² tiles serving meanwhile — against about 3 minutes for the fixed one; the profile's range
|
||||
is what it tunes across. A cached engine loads in about a second.
|
||||
|
||||
In PyTorch fp16 the same network over the whole 20 MP frame in one call took 3.5× less than in
|
||||
tiles, so a card that holds a whole frame gains more than the 6 GB one does; `WHOLE_FRAME_MAX` is a
|
||||
constant sized for 6 GB until the limit follows the card's memory.
|
||||
|
||||
|
||||
@@ -137,6 +137,7 @@ declined, and the InsightFace grant of D13).
|
||||
| `denoise/mosaic-best-1408.onnx` | trained in the `darkroom-denoise` repository (2026-10-04, run `final`, 30 000 steps, from the experts of runs `m2` and `edges-100`) | 1,701 of the maintainer's own base-ISO raws and 6,000 synthetic scenes the repository draws itself, with the Canon EOS 6D's measured noise added | the learned demosaic and denoise, Best (FR-DEV-3g) |
|
||||
| `denoise/mosaic-medium-1408.onnx` | distilled from `final` in the same repository (2026-10-04, run `student-m`, 20 000 steps, from `m2`) | the same | Medium |
|
||||
| `denoise/mosaic-fast-1408.onnx` | distilled from `final` (2026-10-04, run `student-s`, 30 000 steps, from scratch) | the same | Fast |
|
||||
| `denoise/mosaic-{best,medium,fast}.onnx` | the three networks above, re-exported with any height and width by `tools/export_whole.py` in the same repository (2026-10-06) from the same checkpoints; identical to the 1408 files at 1408² | the same | the same methods, a whole frame at a time on a GPU (denoise.md §14) |
|
||||
|
||||
U-Nets of plain 3×3 convolutions, ReLU, strided and transposed
|
||||
convolutions and additive skips — no third-party architecture code or
|
||||
|
||||
Reference in New Issue
Block a user