AI Denoise's Apply switch becomes Method: Bilinear, Fast, Medium, Best,
default Best, so an untouched raw writes nothing and develops through the
mixture. `apply` is still read and never written: 0 is Bilinear, 1 keeps
a network already chosen.
- Best is the mixture of a flat and an edge expert with a learned gate;
Medium and Fast are students distilled from it. 2.48 s, 0.79 s and
0.57 s for a 20 MP frame on TensorRT fp16.
- Each network carries its own tile border (256 for the mixture, 192 for
the students) through `dr_denoise::Shipped` and `TileNet::halo`.
- The file is hashed once at open and each network keys its own cached
result; Bilinear keeps the result in memory for the way back.
- Each has an .a16w16 sibling for the Hexagon: 0.00 dB on the 6D gate,
at most 0.11 dB with the noise scaled x0.5 to x4.
- APK BUNDLED 19 -> 23; the PKGBUILD installs all three.
A 20 MP frame spent 0.32 s outside the network: each tile's mosaic and
sigma gathered on one thread, then its 24 MB output copied out of the
runtime and back into the frame, all in series with the device. Tiles are
now gathered on every core by a producer thread one tile ahead, so the
gather overlaps the run; the centre is written back across cores; and the
tile interface hands its inputs over and lends its output, so neither
side is copied. With a stand-in network that does nothing, the tiler's own
time falls to 0.14 s at the 1408 tile and 0.09 s at 2048. The exactness
and Bayer-phase tests are unchanged and pass.
The app's hot-pixel pass takes gross defects only; at ISO 6400-25600 a 6D
frame keeps 1000-2000 photosites more than 8 sigma beyond all their
same-colour and adjacent neighbours, which the network turned into specks.
The same two tests with the threshold in the photosite's own sigma, plus
the factor of two that keeps a bright point of light (where 8 sigma is a
sliver of the signal). The next model is trained behind exactly this; on
an ISO 25600 frame the Rust and training code both repair 935.
The engine knew f32 and int8, and gave the Hexagon int8 for every role it
served. Measured on the tablet itself (inference.md §1.5), int8 lost
5% of the detector's faces at 40-80 px, moved the landmarks 1.5 px,
emptied the segmenter's scores and cost the denoiser 5-9 dB; fp16 the HTP
refuses outright. `Form` gains A16W8 and A16W16, and `Rung::form` now
names one per role: detectors and landmarks A16W8, the segmenter, scene
model, border filler and denoiser A16W16, XFeat int8. The embedder and
the eye classifiers stay on the CPU.
Each loader resolves its `<stem>.<form>.onnx` sibling; the segmenter and
XFeat, compiled into the binary, embed their quantised forms on Android
only and pick through `choose_embedded`. The probe, the compile step and
the cache fingerprint follow the form instead of assuming int8. Detectors
on the new form write `scrfd_*_a16+w600k_mbf`, and `model_ids` answers
for all three spellings.
On the tablet (ORT 1.29 + QNN 2.42), each shipped file against f32 on the
same inputs, and against the CPU's f32 time:
SCRFD 500m/2.5g/10g A16W8 100% of faces in every band 4.2/5.1/9.0 ms vs 17/56/198
landmarks A16W8 0.25 px in the 192 crop 0.5 ms vs 2.8
YOLO26n-seg A16W16 98.2% found, mask IoU 0.994 12.9 ms vs 90
scene model A16W16 98.9% of cells agree 15 ms vs 151
MI-GAN A16W16 41 dB from f32 in the fill 87 ms vs 488
XFeat int8 pano alignment 0.45 px (f32's own spread 0.41) 6.5 ms vs 58
denoiser A16W16 0.00 dB at every ISO 95 ms vs 1510 a tile
Face numbers are over public COCO val2017 photographs, not a library.
The APK carries the siblings (BUNDLED 15 -> 19; the old int8 detectors
removed), about 43 MB more. The Windows installer and its CI count skip
them; the Arch and Flatpak packages list their files and never had them.
The ladder example takes a role per model, which is how the per-role
forms above were seen landing on the NPU from the real probe.
A Bayer photograph keeps its mosaic in the session and is offered the AI
Denoise switch. Asked for, the network runs on the decode executor from a
hot-pixel-repaired copy — the app's own pass — with the frame's noise from
its best source, and its progress in the activity bar; the classical
demosaic shows until the result lands, and the finished job says where the
noise figures came from. Keep grain is a GrainBlend of the two, made once
per value; the render draws it as its source and the adjust pass never
knows. demosaiced stays the classical result, so the raw histogram, the
white balance picker, masks and segmentation still read the sensor.
The develop view reconciles on a 250 ms poll rather than on each way an
edit can change (slider, undo, preset, version, a sidecar from another
device): two comparisons when nothing changed, and no path that can forget.
A failure is not retried until the switch is toggled. An export of a
photograph that asks for it waits for a running job or computes it.
The noise model takes the best source the frame has: the body's measured
table (the Canon EOS 6D's, from the library), the DNG's NoiseProfile, or
the frame itself — read, row and column noise from its masked border, and
only the shot gain estimated, from the quietest flat patches. Checked on
130 6D frames, the estimate is within 10 % from ISO 1000 up; the network
loses under 0.3 dB for a sigma off by 15-20 %, so every Bayer body is
eligible.
Tiles of 1408 keep their central 1024 behind a 192-photosite halo, past the
185-photosite receptive field, and the frame is extended by reflection,
which keeps every photosite's colour; a pattern that starts on another
colour is read from one photosite up or left so the network sees RGGB, and
nothing is cropped. The tests run every Bayer phase, tiled against whole,
with a stand-in network of known reach.
The model ships as models/denoise/mosaic-1408.onnx (LFS), trained in
darkroom-denoise on the maintainer's own photographs, GPL like the code.
denoise_raw runs a file end to end: on a 6D frame at ISO 8000 the result
matches the training repository's own path to 2.5e-4 at worst, and takes
3.1 s on TensorRT fp16 (75 dB from f32) or 14.4 s on the CPU.