Offer three denoise networks and a method to choose between them

AI Denoise's Apply switch becomes Method: Bilinear, Fast, Medium, Best,
default Best, so an untouched raw writes nothing and develops through the
mixture. `apply` is still read and never written: 0 is Bilinear, 1 keeps
a network already chosen.

- Best is the mixture of a flat and an edge expert with a learned gate;
  Medium and Fast are students distilled from it. 2.48 s, 0.79 s and
  0.57 s for a 20 MP frame on TensorRT fp16.
- Each network carries its own tile border (256 for the mixture, 192 for
  the students) through `dr_denoise::Shipped` and `TileNet::halo`.
- The file is hashed once at open and each network keys its own cached
  result; Bilinear keeps the result in memory for the way back.
- Each has an .a16w16 sibling for the Hexagon: 0.00 dB on the 6D gate,
  at most 0.11 dB with the noise scaled x0.5 to x4.
- APK BUNDLED 19 -> 23; the PKGBUILD installs all three.
This commit is contained in:
2026-10-04 08:02:25 -04:00
parent 14f08a565f
commit 06422a07db
26 changed files with 540 additions and 141 deletions
+69
View File
@@ -457,3 +457,72 @@ under Detail. So:
What it costs: every raw opened runs the network once, with the classical demosaic shown until the
result lands, and a first export of an unopened raw runs it too. Every raw renders differently from
0.21.0 unless switched off.
## 13. Three networks and a method (after 0.22.0)
The photographer asked for a choice between quality and time. `Method` replaces the Apply switch:
`Bilinear`, `Fast`, `Medium`, `Best`, by index in that order, default `Best`. An untouched raw
writes nothing and develops through `Best`. `apply` is still read and never written: 0 is
`Bilinear`, 1 keeps a network already chosen or is the default. A number past the list, from a newer
build, reads as the default. A build before this one ignores `method` and develops through its own
network, which is the most an older peer can do.
**The networks** (darkroom-denoise, every one trained on the same data and noise as §11, plus 1,201
further frames cropped from the library and 6,000 drawn scenes — polygons, lines of one to four
photosites, text, gratings — rendered at 4× through a random affine and smooth displacement, so
edges fall off the photosite grid):
| Method | File | Network | Parameters | GMAC / MP | Halo |
|---|---|---|---|---|---|
| Best | `mosaic-best-1408.onnx` | two U-Nets of §11's shape (a flat expert from `m2`, an edge expert from the ×100 edge-weighted run) and a 128 k-parameter gate that blends them per photosite | 6.4 M | 110 | 256 |
| Medium | `mosaic-medium-1408.onnx` | §11's U-Net, distilled from Best (75 % its output, 25 % the truth) | 3.2 M | 48 | 192 |
| Fast | `mosaic-fast-1408.onnx` | widths 16-32-64-128, blocks 1-1-1-2, distilled the same way | 0.93 M | 11 | 192 |
The gate learned on its own to trust the edge expert at 0.77–0.88 on edges and not at all on flat
areas. The mixture's receptive field is the experts' plus the gate's, so it keeps the centre of a
1408 tile past a 256 halo, where the single networks keep 1024 past 192. `dr_denoise::Shipped`
carries each file's halo, and `TileNet::halo` hands it to the tiler.
**Quality.** PSNR after the display transform on 1,842 held-out crops, and the width of a hard
edge on the drawn chart at ISO 6400 (truth 0.80 photosites; lower is sharper):
| | ISO 400 | 1600 | 6400 | 25600 | Edge width |
|---|---|---|---|---|---|
| §11's network | 40.61 | 39.77 | 38.43 | 36.67 | 1.77 |
| Best | 40.69 | 39.84 | 38.47 | 36.71 | 0.82 |
| Medium | 40.59 | 39.75 | 38.40 | 36.65 | 1.30 |
| Fast | 39.90 | 39.06 | 37.54 | 35.33 | 1.84 |
| Bilinear | 36.15 | 33.26 | 28.72 | 23.83 | 2.15 |
On photographs the three are close; on hard edges Best is half as wide as §11's network and
Medium most of the way there. Fast costs a dB at high ISO and edges as soft as §11's.
**Speed**, a whole 20 MP 6D frame, the network alone, TensorRT fp16 on the laptop's RTX 3050
(uncapped: memory at 5 GHz), engine already built: Best 2.48 s, Medium 0.79 s, Fast 0.57 s. Decode
and the hot-pixel pass add about 0.5 s. The first build of each TensorRT engine takes 80 s (Fast) to
190 s (Best), in the background at first launch, cached after.
**Before the network, two passes changed since §11.**
- *A noise-aware repair* (`dr_denoise::repair`) after the app's hot-pixel pass: a photosite more
than 8σ beyond every same-colour neighbour *and* every adjacent photosite, and more than twice
each adjacent one, is clamped to the brightest of its same-colour neighbours; a dead one, to the
darkest. The ratio test is what spares a point of light, whose neighbours are lit too. The networks
were trained behind the same pass (the Python and Rust agree: 935 repairs on an ISO 25600 frame).
- *The tiler feeds the network without waiting*: tiles are gathered on every core by a producer
thread one tile ahead, and the output is written back in parallel from the runtime's own buffer.
0.14 s of tiler for a frame, which is what keeps Fast under a second.
**The Hexagon.** Each network has an `.a16w16.onnx` sibling made by `tools/quantise-models.sh
--ranges`, the ranges from darkroom-3e's gate (96 training tiles, a third at noise ×2 and ×4). On
the 6D gate A16W16 loses 0.00 dB for all three; with the noise scaled ×0.5–×4 at most 0.11 dB.
A16W8 holds the gate (≤ 0.27 dB) but loses 0.63 dB on Medium at ×4, so A16W16 stays the form.
**Cache.** Each network keys its own results (§7.1 keys on the model's file name), and the file is
hashed once at open, so changing the method never re-reads it. Choosing `Bilinear` keeps the
network's result in memory for the way back; changing to another network drops it, and coming back
reads the cache.
**Packaging.** All six files in the APK (`BUNDLED`, 23 entries, +44.6 MB, ~41 MB compressed); the
three f32 networks in the Arch package and the Windows installer, which stage `models/denoise` by
directory.
File diff suppressed because one or more lines are too long