The engine knew f32 and int8, and gave the Hexagon int8 for every role it served. Measured on the tablet itself (inference.md §1.5), int8 lost 5% of the detector's faces at 40-80 px, moved the landmarks 1.5 px, emptied the segmenter's scores and cost the denoiser 5-9 dB; fp16 the HTP refuses outright. `Form` gains A16W8 and A16W16, and `Rung::form` now names one per role: detectors and landmarks A16W8, the segmenter, scene model, border filler and denoiser A16W16, XFeat int8. The embedder and the eye classifiers stay on the CPU. Each loader resolves its `<stem>.<form>.onnx` sibling; the segmenter and XFeat, compiled into the binary, embed their quantised forms on Android only and pick through `choose_embedded`. The probe, the compile step and the cache fingerprint follow the form instead of assuming int8. Detectors on the new form write `scrfd_*_a16+w600k_mbf`, and `model_ids` answers for all three spellings. On the tablet (ORT 1.29 + QNN 2.42), each shipped file against f32 on the same inputs, and against the CPU's f32 time: SCRFD 500m/2.5g/10g A16W8 100% of faces in every band 4.2/5.1/9.0 ms vs 17/56/198 landmarks A16W8 0.25 px in the 192 crop 0.5 ms vs 2.8 YOLO26n-seg A16W16 98.2% found, mask IoU 0.994 12.9 ms vs 90 scene model A16W16 98.9% of cells agree 15 ms vs 151 MI-GAN A16W16 41 dB from f32 in the fill 87 ms vs 488 XFeat int8 pano alignment 0.45 px (f32's own spread 0.41) 6.5 ms vs 58 denoiser A16W16 0.00 dB at every ISO 95 ms vs 1510 a tile Face numbers are over public COCO val2017 photographs, not a library. The APK carries the siblings (BUNDLED 15 -> 19; the old int8 detectors removed), about 43 MB more. The Windows installer and its CI count skip them; the Arch and Flatpak packages list their files and never had them. The ladder example takes a role per model, which is how the per-role forms above were seen landing on the NPU from the real probe.
Face models
The shape-fixed SCRFD detectors and the ArcFace/MobileFaceNet embedder, in LFS. One copy, packaged by every platform:
| Platform | How it ships | Where it lands |
|---|---|---|
| Android | assemble-apk.sh bundles them as APK assets; android_main unpacks on first launch |
the shared user data directory |
| Arch | PKGBUILD installs them |
/usr/share/darkroom/models/ |
library::face_models searches the account's own directory, then the shared user directory, then
$XDG_DATA_DIRS — so a pair the user placed by hand always outranks the packaged one.
scrfd_500m_640.onnx 2.5 MB "Fast" in settings — what every library was indexed with until 0.12
scrfd_2.5g_640.onnx 3.3 MB "Balanced" — 14% more faces for 12% more time (faces.md §12.3)
scrfd_10g_640.onnx 17 MB "Thorough" — a further 12% for 3× the time
arcface_mbf_b1.onnx 13 MB
2d106det_b1.onnx 4.8 MB 106 landmarks, for the eye boxes (faces.md §17)
ocec_s_b1.onnx 483 KB eyes open or closed, per eye
sgc_l_48_b1.onnx 6.1 MB sunglasses or not, per head
Which detector runs is a per-device setting (FaceSettings::detector); all three are installed so
the choice exists on every platform. Each is its own faces.model_id, so changing it re-indexes.
The last three are optional to the app: library::face_models reports them beside the pair when
all three are there, and a library without them indexes faces and simply has no eye readings —
every reader treats "never read" as unknown, never as closed. They are found in the same
directory as the pair, so a hand-placed pair does not pick up a package's eye models from a
directory it otherwise outranks.
scrfd_500m_640.a16w8.onnx the same three, in the form the Hexagon NPU takes
scrfd_2.5g_640.a16w8.onnx (docs/dev/inference.md §1.5) — 16-bit activations,
scrfd_10g_640.a16w8.onnx per-channel 8-bit weights, calibrated on 300 photographs
2d106det_b1.a16w8.onnx the landmarks, likewise
These are derived by tools/quantise-models.sh from the f32 files beside them and travel with
them: the engine loads the .a16w8.onnx sibling when the device's backend wants it and the
canonical file otherwise, and a library indexed on that form records it as a different detector
(scrfd_500m_a16+w600k_mbf), because it finds a different set of faces. Every other platform
ignores them, and the Windows installer leaves them out. The embedder and the eye classifiers have
no quantised form: the embedder's vectors must compare across devices (inference.md §7), and the
classifiers cost a millisecond on the CPU.
The detectors were int8 until §1.5 measured them on the tablet: int8 found 94–95% of f32's faces at 40–80 px, A16W8 all of them. A tablet that indexed under the int8 ids keeps those rows.
A clone without git-lfs gets a ~130-byte pointer where each model should be. Both packagers check
for exactly that and refuse, rather than shipping the pointer and failing inside tract on the user's
machine. Fix it with git lfs pull.
These are not what InsightFace ships. They came from buffalo_sc.zip, buffalo_m.zip, buffalo_l.zip
and buffalo_s.zip on the InsightFace v0.7 release with their input dimensions pinned, because tract
cannot parse any of the graphs while they are dynamic:
./tools/fix-face-model-shapes.sh det_500m.onnx models/face/scrfd_500m_640.onnx --input input.1=1,3,640,640
./tools/fix-face-model-shapes.sh det_2.5g.onnx models/face/scrfd_2.5g_640.onnx --input input.1=1,3,640,640
./tools/fix-face-model-shapes.sh det_10g.onnx models/face/scrfd_10g_640.onnx --input input.1=1,3,640,640
./tools/fix-face-model-shapes.sh w600k_mbf.onnx models/face/arcface_mbf_b1.onnx --dim None=1
./tools/fix-face-model-shapes.sh 2d106det.onnx models/face/2d106det_b1.onnx --dim None=1
2d106det_b1.onnx is from the same buffalo_l.zip as det_10g.onnx and under the same grant: the
106-point landmark model whose lid contours the eye boxes are cut from (docs/dev/faces.md §17.2).
sha256 as fetched f001b856…a7109dbf, as shipped afc2984c…03368ef26.
The weights carry a non-commercial research-only grant. They are here because this is a private repository and self-installed builds; they come back out before anything is published, and the restriction binds whoever uses the app, not only the project. docs/dev/faces.md §2.2a is the decision.
The two classifiers are a different matter
ocec_s_b1.onnx and sgc_l_48_b1.onnx are MIT, code and weights, from Katsuya Hyodo's
ultra-lightweight classifier series — the same author as the whole-body detector the reference
pipeline uses. They are not under the InsightFace grant and do not come out when the project
publishes. Read 2026-09-19:
| OCEC — open/closed eyes | SGC — sunglasses | |
|---|---|---|
| Source | github.com/PINTO0309/OCEC, release onnx, ocec_s.onnx |
github.com/PINTO0309/SGC, release onnx, sgc_is_l_48x48.onnx |
| Licence | MIT (repository LICENSE; no separate grant on the weights) |
MIT, likewise |
| Training data | Open and Closed Eyes (Młodawski 2024, HF, ODC-By 1.0 — attribution only), crops cut by DEIMv2-Wholebody34 (Apache 2.0) | Not stated. The README names no dataset and carries no acknowledgement; data/ holds a class-ratio plot and nothing else. |
| Input | one eye, 40×24, RGB, x/255 |
one head, 48×48, RGB, x/255 |
| Output | prob_open, a sigmoid |
prob_sunglasses, a sigmoid |
| sha256 as fetched | 9a346a08…c604ba9b64 |
9c13d937…a4063c73 |
| sha256 as shipped | c848d34c…4b29d04c7c |
d49b6206…7e9538525c6 |
Both shipped with a dynamic batch dimension that tract loads but the project pins anyway, with the same script as the pair:
./tools/fix-face-model-shapes.sh ocec_s.onnx models/face/ocec_s_b1.onnx --dim batch=1
./tools/fix-face-model-shapes.sh sgc_is_l_48x48.onnx models/face/sgc_l_48_b1.onnx --dim batch=1
The one thing D13's discipline turns up here is SGC's undocumented training set. The weights' grant is MIT and that is what binds a redistributor; but "trained on what" is the question this project reads first, and for SGC it has no answer. Recorded so it is a known gap rather than an assumption, and so that whoever finds a sunglasses classifier with a stated dataset knows what to replace.
Attribution, as ODC-By asks for the eye dataset:
Michał Młodawski, Open and Closed Eyes Dataset, July 2024, https://huggingface.co/datasets/MichalMlodawski/closed-open-eyes
The variant choice was measured, not taken from the F1 column: on family snapshots the S variant of OCEC read more open eyes as open than M or L did, which overfit their own domain (docs/dev/faces.md §17).