Best was a mixture of two experts and a gate, 110 GMAC a megapixel;
Medium a single network at 48 that was softer on real edges. fb-combo
(darkroom-denoise, 20 000 steps from fb-edges2, taught by the mixture
with a quarter of its crops from the edge-rich parts of the frames) is
Medium's shape and holds the mixture's edges on real photographs:
edge PSNR within 0.04-0.06 dB at ISO 1600/6400/25600, more sharpness
kept at all three, the chart's edge 0.89 photosites wide against 0.82.
It is 0.27 dB short on smooth areas at ISO 25600. It becomes Best, and
the methods are Bilinear, Fast and Best.
Saved edits keep their numbers: 2, which was Medium, is now Best, and
3, which was Best, is past the end and reads as the default, Best.
The network ships as mosaic-hq, a new name: the result cache keys a
model by name and size, and this one is byte for byte the old Medium's
size. Its tablet form (A16W16) lost 0.00 dB in simulated QDQ at every
ISO and at most 0.09 dB across the noise bracket.
denoise.md §14: why the tiles waste half of Best's work, where the any-size
networks run and why only there, why the limit is the card's memory, and
the measurement on _MG_8862 — 2.60 s in 1408 tiles, 1.37 s in two
4160 x 3248 tiles, the outputs within fp16's own spread. models/LICENCE.md
lists the three re-exports.
AI Denoise's Apply switch becomes Method: Bilinear, Fast, Medium, Best,
default Best, so an untouched raw writes nothing and develops through the
mixture. `apply` is still read and never written: 0 is Bilinear, 1 keeps
a network already chosen.
- Best is the mixture of a flat and an edge expert with a learned gate;
Medium and Fast are students distilled from it. 2.48 s, 0.79 s and
0.57 s for a 20 MP frame on TensorRT fp16.
- Each network carries its own tile border (256 for the mixture, 192 for
the students) through `dr_denoise::Shipped` and `TileNet::halo`.
- The file is hashed once at open and each network keys its own cached
result; Bilinear keeps the result in memory for the way back.
- Each has an .a16w16 sibling for the Hexagon: 0.00 dB on the 6D gate,
at most 0.11 dB with the noise scaled x0.5 to x4.
- APK BUNDLED 19 -> 23; the PKGBUILD installs all three.
The engine knew f32 and int8, and gave the Hexagon int8 for every role it
served. Measured on the tablet itself (inference.md §1.5), int8 lost
5% of the detector's faces at 40-80 px, moved the landmarks 1.5 px,
emptied the segmenter's scores and cost the denoiser 5-9 dB; fp16 the HTP
refuses outright. `Form` gains A16W8 and A16W16, and `Rung::form` now
names one per role: detectors and landmarks A16W8, the segmenter, scene
model, border filler and denoiser A16W16, XFeat int8. The embedder and
the eye classifiers stay on the CPU.
Each loader resolves its `<stem>.<form>.onnx` sibling; the segmenter and
XFeat, compiled into the binary, embed their quantised forms on Android
only and pick through `choose_embedded`. The probe, the compile step and
the cache fingerprint follow the form instead of assuming int8. Detectors
on the new form write `scrfd_*_a16+w600k_mbf`, and `model_ids` answers
for all three spellings.
On the tablet (ORT 1.29 + QNN 2.42), each shipped file against f32 on the
same inputs, and against the CPU's f32 time:
SCRFD 500m/2.5g/10g A16W8 100% of faces in every band 4.2/5.1/9.0 ms vs 17/56/198
landmarks A16W8 0.25 px in the 192 crop 0.5 ms vs 2.8
YOLO26n-seg A16W16 98.2% found, mask IoU 0.994 12.9 ms vs 90
scene model A16W16 98.9% of cells agree 15 ms vs 151
MI-GAN A16W16 41 dB from f32 in the fill 87 ms vs 488
XFeat int8 pano alignment 0.45 px (f32's own spread 0.41) 6.5 ms vs 58
denoiser A16W16 0.00 dB at every ISO 95 ms vs 1510 a tile
Face numbers are over public COCO val2017 photographs, not a library.
The APK carries the siblings (BUNDLED 15 -> 19; the old int8 detectors
removed), about 43 MB more. The Windows installer and its CI count skip
them; the Arch and Flatpak packages list their files and never had them.
The ladder example takes a role per model, which is how the per-role
forms above were seen landing on the NPU from the real probe.
The noise model takes the best source the frame has: the body's measured
table (the Canon EOS 6D's, from the library), the DNG's NoiseProfile, or
the frame itself — read, row and column noise from its masked border, and
only the shot gain estimated, from the quietest flat patches. Checked on
130 6D frames, the estimate is within 10 % from ISO 1000 up; the network
loses under 0.3 dB for a sigma off by 15-20 %, so every Bayer body is
eligible.
Tiles of 1408 keep their central 1024 behind a 192-photosite halo, past the
185-photosite receptive field, and the frame is extended by reflection,
which keeps every photosite's colour; a pattern that starts on another
colour is read from one photosite up or left so the network sees RGGB, and
nothing is cropped. The tests run every Bayer phase, tiled against whole,
with a stand-in network of known reach.
The model ships as models/denoise/mosaic-1408.onnx (LFS), trained in
darkroom-denoise on the maintainer's own photographs, GPL like the code.
denoise_raw runs a file end to end: on a 6D frame at ISO 8000 the result
matches the training repository's own path to 2.5e-4 at worst, and takes
3.1 s on TensorRT fp16 (75 dB from f32) or 14.4 s on the CPU.
docs/ had 26 developer documents flat beside the manual, and the two
audiences are very differently sized: most readers want the manual and
the gesture reference, a few want the register, the designs and the
measurements. The manual and gestures.md stay at the top; everything for
someone changing the code moves to docs/dev/, and the two documents that
name their own successors — the v0.1 milestone and the UI-refinement plan
— go to docs/dev/archive/ rather than being deleted, since both are still
cited. docs/README.md is the index, users first.
Every reference follows: code comments, Cargo manifests, the workflows,
the pre-commit hook, the bench and traceability tools (which locate the
repo root by docs/dev/requirements.md now), packaging, the Docker READMEs,
CLAUDE.md, CONTRIBUTING.md and the README. The matrix links one level
deeper and is regenerated. Links out of the moved documents into the tree
gain a level; a link checker over every Markdown file finds none broken.
Sargsyan et al., ICCV 2023; MIT code and weights (models/LICENCE.md),
exported by tools/export-migan.sh at a fixed 1×4×512×512 from the
authors' checkpoint — six operator types, 28 MB, in LFS like the rest.
The package installs it beside the scene model and the APK unpacks it
with the others.
Twelve real frames from the fixture set now align in 4.5 s — 4.4 s of
matching, 118 ms of bundle adjustment — where the first run took 51 s and
left the first two frames out.
The matcher computes each pair's similarity matrix once, across the
cores, with a dot product written to vectorise; both nearest-neighbour
directions read it. The frames that failed were portrait: fitted into the
landscape input they used 512 of 1024 px, and their thin overlap did not
survive at half resolution. The same weights are now exported at 768×1024
as well and the detector picks the shape by aspect. The example aligns
from embedded previews and draws the set on a cylinder; on the fixture the
sweep is 152° at a fitted 47.9 mm against the EXIF's 50, RMS 1.5 px, and
the overlaps show no ghosting.
tools/export-xfeat.sh exports the convolutional network alone at 768×1024
grayscale, on the pattern of export-seg-model.sh: thirteen standard
operator types, no dynamic axes, the keypoint decoding left to Rust.
examples/onnx_probe loads it through the ort-over-tract backend the app
ships with nothing unsupported and runs it in ~300 ms on the desktop CPU.
The weights are Apache-2.0, read from the repository's LICENSE, with no
grant on the checkpoint — recorded in models/LICENCE.md before they land,
as FR-MRG-8 asks. The probe stays: the next model will need the same
check.
`models/LICENCE.md` recorded, on 2026-08-21, that no YOLO model trained on
ADE20K existed in usable form — the stuff classes photography cares about,
sky and vegetation and water, had no model to come from. Re-checked
2026-08-30: Ultralytics now ships a `semantic` task with ADE20K
checkpoints, so `models/scene/` holds `yolo26s-sem-ade20k`.
This is an addition, not a replacement. A semantic model labels every
pixel but merges same-class pixels into one region, so it cannot tell
three people apart — which is exactly what clicking a subject needs, and
exactly what `segment/`'s COCO instance model already does. The scene tab
grades per category and does not care that instances are merged. Keeping
both is the point.
## The export is truncated, deliberately
Ultralytics ends the graph with `Resize -> ArgMax -> Cast` and hands back
a `[1, 640, 640]` u8 label map. The script cuts that tail and exposes the
classifier's `[1, 150, 80, 80]` f32 logits instead, for two reasons.
Cost: the Resize materialises 150 x 640 x 640 x f32, 246 MB, and ArgMax
then reduces across the channel axis, striding 409,600 elements per
comparison. On one loaded machine the full graph ran ~1160 ms against
~500 ms truncated — roughly four fifths of the time spent on work the
application discards. Those numbers were measured under contention and
are upper bounds, but the ratio is structural.
Softness: ArgMax destroys the per-class scores, and the scene tab needs
them. Softmax over the 150 channels, summed within each photographic
category, yields per-category weights summing to 1 at every pixel.
Feathering a partition of unity cannot double-grade a boundary, whereas
feathering hard labels outward from two adjacent categories paints both
grades into the overlap and haloes every horizon.
The discarded upsample was never information: the graph's true spatial
resolution is the 80x80 logit grid, and the application can resample from
that itself.
The tail is matched by op type and asserted before cutting, so an
upstream graph change fails loudly in the exporter rather than quietly
shipping a differently-shaped model.
Nothing reads these weights yet — the decode path, the category
descriptor grouping 150 classes into ~8 photographic ones, and the scene
tab are still to come. At 24 MB this model also wants the runtime-asset
treatment `models/face/` already gets on Android rather than
`include_bytes!`; embedding it would put ~35 MB of weights in the binary.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The weights were in two places: face detection and recognition in
`models/face/`, segmentation in `core/dr-segment/models/`. Nothing was
wrong with either path, but between them there was nowhere to look to
answer "how much model does this application carry", and that number is
about to start growing.
So the crate-local copy moves up beside the other. `models/` now holds
`face/` and `segment/`, and a `du -sh` of one directory is the whole
answer.
No content changes: the .onnx and its vocabulary are byte-identical, and
`LICENCE.md` moves up a level to cover the tree rather than one crate.
The LFS pattern in `.gitattributes` is `*.onnx` and already matched both
locations, so only its comment needed the new path.
`include_bytes!` is relative to the source file and `build.rs` runs with
the crate root as its working directory, which is why the two paths climb
a different number of levels.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>