The engine knew f32 and int8, and gave the Hexagon int8 for every role it
served. Measured on the tablet itself (inference.md §1.5), int8 lost
5% of the detector's faces at 40-80 px, moved the landmarks 1.5 px,
emptied the segmenter's scores and cost the denoiser 5-9 dB; fp16 the HTP
refuses outright. `Form` gains A16W8 and A16W16, and `Rung::form` now
names one per role: detectors and landmarks A16W8, the segmenter, scene
model, border filler and denoiser A16W16, XFeat int8. The embedder and
the eye classifiers stay on the CPU.
Each loader resolves its `<stem>.<form>.onnx` sibling; the segmenter and
XFeat, compiled into the binary, embed their quantised forms on Android
only and pick through `choose_embedded`. The probe, the compile step and
the cache fingerprint follow the form instead of assuming int8. Detectors
on the new form write `scrfd_*_a16+w600k_mbf`, and `model_ids` answers
for all three spellings.
On the tablet (ORT 1.29 + QNN 2.42), each shipped file against f32 on the
same inputs, and against the CPU's f32 time:
SCRFD 500m/2.5g/10g A16W8 100% of faces in every band 4.2/5.1/9.0 ms vs 17/56/198
landmarks A16W8 0.25 px in the 192 crop 0.5 ms vs 2.8
YOLO26n-seg A16W16 98.2% found, mask IoU 0.994 12.9 ms vs 90
scene model A16W16 98.9% of cells agree 15 ms vs 151
MI-GAN A16W16 41 dB from f32 in the fill 87 ms vs 488
XFeat int8 pano alignment 0.45 px (f32's own spread 0.41) 6.5 ms vs 58
denoiser A16W16 0.00 dB at every ISO 95 ms vs 1510 a tile
Face numbers are over public COCO val2017 photographs, not a library.
The APK carries the siblings (BUNDLED 15 -> 19; the old int8 detectors
removed), about 43 MB more. The Windows installer and its CI count skip
them; the Arch and Flatpak packages list their files and never had them.
The ladder example takes a role per model, which is how the per-role
forms above were seen landing on the NPU from the real probe.
The noise model takes the best source the frame has: the body's measured
table (the Canon EOS 6D's, from the library), the DNG's NoiseProfile, or
the frame itself — read, row and column noise from its masked border, and
only the shot gain estimated, from the quietest flat patches. Checked on
130 6D frames, the estimate is within 10 % from ISO 1000 up; the network
loses under 0.3 dB for a sigma off by 15-20 %, so every Bayer body is
eligible.
Tiles of 1408 keep their central 1024 behind a 192-photosite halo, past the
185-photosite receptive field, and the frame is extended by reflection,
which keeps every photosite's colour; a pattern that starts on another
colour is read from one photosite up or left so the network sees RGGB, and
nothing is cropped. The tests run every Bayer phase, tiled against whole,
with a stand-in network of known reach.
The model ships as models/denoise/mosaic-1408.onnx (LFS), trained in
darkroom-denoise on the maintainer's own photographs, GPL like the code.
denoise_raw runs a file end to end: on a 6D frame at ISO 8000 the result
matches the training repository's own path to 2.5e-4 at worst, and takes
3.1 s on TensorRT fp16 (75 dB from f32) or 14.4 s on the CPU.
docs/ had 26 developer documents flat beside the manual, and the two
audiences are very differently sized: most readers want the manual and
the gesture reference, a few want the register, the designs and the
measurements. The manual and gestures.md stay at the top; everything for
someone changing the code moves to docs/dev/, and the two documents that
name their own successors — the v0.1 milestone and the UI-refinement plan
— go to docs/dev/archive/ rather than being deleted, since both are still
cited. docs/README.md is the index, users first.
Every reference follows: code comments, Cargo manifests, the workflows,
the pre-commit hook, the bench and traceability tools (which locate the
repo root by docs/dev/requirements.md now), packaging, the Docker READMEs,
CLAUDE.md, CONTRIBUTING.md and the README. The matrix links one level
deeper and is regenerated. Links out of the moved documents into the tree
gain a level; a link checker over every Markdown file finds none broken.
Sargsyan et al., ICCV 2023; MIT code and weights (models/LICENCE.md),
exported by tools/export-migan.sh at a fixed 1×4×512×512 from the
authors' checkpoint — six operator types, 28 MB, in LFS like the rest.
The package installs it beside the scene model and the APK unpacks it
with the others.
The first int8 files found no faces at all, and for two reasons the
tool now guards against. The calibration set was landscape photographs
with no faces in them, so the score head's ranges had never seen the
face regime; the set is now proxies from the library itself. And ONNX
Runtime's strided and moving-average calibration modes both degrade
these graphs measurably (a quarter of the faces at eight images, none
at ninety-six), while driving the calibrator in chunks by hand gives
ranges identical to a single pass — so the tool does that, four images
at a time, and feeds quantize_static through its range cache.
Measured against f32 over 400 proxies (docs/inference.md §10.1): the
10g form finds every face above 32 px the f32 form finds; 500m and
2.5g find 96%, and what they lose sits at a median confidence of 0.52
against the 0.50 threshold. Shipped with the number on record.
The Android unpack list gains the three int8 files; without that the
tablet never saw them. D13's runtime half records the reopening.
tools/quantise-models.sh writes the QDQ form QNN's HTP backend takes
whole: opset 17, per-channel int8 weights, uint8 activations, ranges
from running the f32 graph over photographs fed exactly as the app
feeds them. The calibration is strided, four images at a time, because
every ONNX Runtime calibrator holds each image's whole set of
activations until it folds them — a gigabyte an image on the 10g
detector, and an OOM kill with no message when folded once at the end.
Release-time, never on the device (docs/inference.md §5): it needs
real photographs and a person reading the recall measurement that
gates whether each file is offered.
Twelve real frames from the fixture set now align in 4.5 s — 4.4 s of
matching, 118 ms of bundle adjustment — where the first run took 51 s and
left the first two frames out.
The matcher computes each pair's similarity matrix once, across the
cores, with a dot product written to vectorise; both nearest-neighbour
directions read it. The frames that failed were portrait: fitted into the
landscape input they used 512 of 1024 px, and their thin overlap did not
survive at half resolution. The same weights are now exported at 768×1024
as well and the detector picks the shape by aspect. The example aligns
from embedded previews and draws the set on a cylinder; on the fixture the
sweep is 152° at a fitted 47.9 mm against the EXIF's 50, RMS 1.5 px, and
the overlaps show no ghosting.
tools/export-xfeat.sh exports the convolutional network alone at 768×1024
grayscale, on the pattern of export-seg-model.sh: thirteen standard
operator types, no dynamic axes, the keypoint decoding left to Rust.
examples/onnx_probe loads it through the ort-over-tract backend the app
ships with nothing unsupported and runs it in ~300 ms on the desktop CPU.
The weights are Apache-2.0, read from the repository's LICENSE, with no
grant on the checkpoint — recorded in models/LICENCE.md before they land,
as FR-MRG-8 asks. The probe stays: the next model will need the same
check.
2d106det for the eye contours, OCEC for open or closed, SGC for
sunglasses — all three pinned to a batch of one by the same script as
the pair, and installed by every packager so the eyes-open filter works
out of the box. The two classifiers are MIT, code and weights; the
README records their provenance, SGC's undocumented training set, and
the hashes as fetched and as shipped.
faces.md §12.3 measured what the cheapest detector costs: the small
faces in every group shot, and a dog embedded a dozen times. Which
trade is right depends on the machine doing the sweep — a desktop left
overnight and a tablet on a battery want different answers — so the
detector is now a per-device setting, Fast / Balanced / Thorough on
the settings page beside the indexing button, persisted with the rest
of the settings file.
A detector is half of a model id. Every face, marker, shard and
calibration is keyed on faces.model_id precisely so that a model change
is a new id and a re-index rather than a silent change under existing
data, and a detector change is a model change: it decides which faces
exist and where the landmarks that align them land. So each choice
names its own pipeline. 500M keeps the bare "w600k_mbf" every existing
library was written under, so an upgrade disturbs nothing; the others
are qualified. Choosing one restarts coverage from zero under the new
id, the sweep re-detects, confirmed names carry across by box overlap,
and the sync shards are keyed by the same id so a peer on another
setting neither adopts nor pollutes them. The library controller
carries the id into the sync the same way it carries the cache budget,
because the sync starts from places that have no settings in reach.
All three shape-fixed exports ship — APK, Arch, Flatpak — since a
tablet has no other way to obtain the one it was not installed with;
the APK grows by twenty megabytes for the choice.
The weights landed last commit with nothing to read them. This is the
decoder, and the shape of it follows from one property worth stating
before the code: the categories must partition the image.
## Why a partition, and not a mask per category
The scene tab applies one grade to every pixel of a category — lift the
sky, desaturate foliage — and both grades meet at the horizon. If each
category carried an independent mask, feathering them outward would make
the boundary band belong to both, so both grades would land there and
every horizon would acquire a visible seam. Feathering has to *blend*
there, not accumulate.
So `marginalise` takes one softmax over all 150 channels and sums within
each category. Grouping cannot change a total of one, so the listed
categories plus the unlisted remainder sum to one at every pixel, by
construction rather than by normalising afterwards. `parse_categories`
refuses a descriptor that claims a class twice, because that is the one
input that would quietly make the property untrue.
## The descriptor is data, and hand-written
`models/scene/categories.txt` groups ADE20K's 150 classes into the eight
a photographer would recognise. It is a file rather than a table in Rust
for the reason `models/LICENCE.md` predicted — a vocabulary is model
metadata — and it is line-oriented with comments rather than JSON like
the `.classes.json` beside it, because that file is generated and this
one is argued. Why `swimming pool` is water and not architecture belongs
next to the line that says so.
Classes are named, not indexed. An index is silently wrong after a
re-export; a name is loudly wrong, and the loader refuses one the model
does not have.
## Resolution, kept visible
`Scene` holds the native 80×80 logit grid and resamples on demand rather
than upsampling once at load. The coarseness is real — it is what the
graph produces — and a type that hides it behind an early resize invites
callers to expect detail that was never there. `rasterise` is where the
letterbox inverse lives, once.
`Letterbox` and `Window` become `pub(crate)` and `to_proto` generalises
to `to_grid`, because both dense outputs this crate reads are an even
fraction of the same letterboxed square and differ only in the divisor.
## Verified by looking, which is the only way this gets verified
`examples/scene.rs` writes the photograph dimmed outside each category. A
transposed axis or an off-by-one in the inverse produces perfectly
plausible weights over slightly the wrong pixels, and no unit test
catches that. On an indoor frame the person mask lands on the person,
including the outstretched arm, and sky reads ~5% against a bright
ceiling.
It doubles as the benchmark, because every timing quoted while this model
was chosen came off a laptop compiling other things and none of them
belong in a document.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`models/LICENCE.md` recorded, on 2026-08-21, that no YOLO model trained on
ADE20K existed in usable form — the stuff classes photography cares about,
sky and vegetation and water, had no model to come from. Re-checked
2026-08-30: Ultralytics now ships a `semantic` task with ADE20K
checkpoints, so `models/scene/` holds `yolo26s-sem-ade20k`.
This is an addition, not a replacement. A semantic model labels every
pixel but merges same-class pixels into one region, so it cannot tell
three people apart — which is exactly what clicking a subject needs, and
exactly what `segment/`'s COCO instance model already does. The scene tab
grades per category and does not care that instances are merged. Keeping
both is the point.
## The export is truncated, deliberately
Ultralytics ends the graph with `Resize -> ArgMax -> Cast` and hands back
a `[1, 640, 640]` u8 label map. The script cuts that tail and exposes the
classifier's `[1, 150, 80, 80]` f32 logits instead, for two reasons.
Cost: the Resize materialises 150 x 640 x 640 x f32, 246 MB, and ArgMax
then reduces across the channel axis, striding 409,600 elements per
comparison. On one loaded machine the full graph ran ~1160 ms against
~500 ms truncated — roughly four fifths of the time spent on work the
application discards. Those numbers were measured under contention and
are upper bounds, but the ratio is structural.
Softness: ArgMax destroys the per-class scores, and the scene tab needs
them. Softmax over the 150 channels, summed within each photographic
category, yields per-category weights summing to 1 at every pixel.
Feathering a partition of unity cannot double-grade a boundary, whereas
feathering hard labels outward from two adjacent categories paints both
grades into the overlap and haloes every horizon.
The discarded upsample was never information: the graph's true spatial
resolution is the 80x80 logit grid, and the application can resample from
that itself.
The tail is matched by op type and asserted before cutting, so an
upstream graph change fails loudly in the exporter rather than quietly
shipping a differently-shaped model.
Nothing reads these weights yet — the decode path, the category
descriptor grouping 150 classes into ~8 photographic ones, and the scene
tab are still to come. At 24 MB this model also wants the runtime-asset
treatment `models/face/` already gets on Android rather than
`include_bytes!`; embedding it would put ~35 MB of weights in the binary.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The weights were in two places: face detection and recognition in
`models/face/`, segmentation in `core/dr-segment/models/`. Nothing was
wrong with either path, but between them there was nowhere to look to
answer "how much model does this application carry", and that number is
about to start growing.
So the crate-local copy moves up beside the other. `models/` now holds
`face/` and `segment/`, and a `du -sh` of one directory is the whole
answer.
No content changes: the .onnx and its vocabulary are byte-identical, and
`LICENCE.md` moves up a level to cover the tree rather than one crate.
The LFS pattern in `.gitattributes` is `*.onnx` and already matched both
locations, so only its comment needed the new path.
`include_bytes!` is relative to the source file and `build.rs` runs with
the crate root as its working directory, which is why the two paths climb
a different number of levels.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The Android bundling landed the weights under that platform's asset directory,
which was the wrong home the moment a second packager wanted them. `makepkg -si`
produced a desktop install with no model at all — the same "no face model is
installed" the phone used to show, for the same reason: nothing put the files
anywhere the app looks.
So `models/face/` at the root is the one copy, and both packagers read it:
assemble-apk.sh bundles it as APK assets, and the PKGBUILD installs it to
/usr/share/darkroom/models. Both refuse an LFS pointer rather than shipping a
130-byte file that fails inside the graph loader on a user's machine.
`face_models` now searches three places, most specific first: the account's own
directory, the shared user directory, then $XDG_DATA_DIRS. So a packaged pair is
found automatically and a pair the user placed by hand still outranks it — which
is what keeps a deliberate choice of weights from being overridden by an
upgrade.
$XDG_DATA_DIRS rather than a hard-coded /usr/share: that is the variable a
distribution, a prefix install or a Nix-style store already sets to say where
its data went, and its documented default is exactly the two paths that would
otherwise have been hard-coded. Empty on Android, which has no such directories
— there the APK's copy is unpacked into the shared user directory instead,
because an asset inside a package is not a path anything can read from.
Verified: the APK still carries both models at assets/models/, the PKGBUILD
parses and installs from the new path, 467 tests pass.
Includes the pkgver 0.6.0 → 0.7.0 bump that was already sitting uncommitted in
the working tree.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>