Files
DarkRoom/models/face/README.md
T
dtourolle 4ed29b9d81 Add the int8 detectors for the Hexagon, calibrated on real photographs
tools/quantise-models.sh writes the QDQ form QNN's HTP backend takes
whole: opset 17, per-channel int8 weights, uint8 activations, ranges
from running the f32 graph over photographs fed exactly as the app
feeds them. The calibration is strided, four images at a time, because
every ONNX Runtime calibrator holds each image's whole set of
activations until it folds them — a gigabyte an image on the 10g
detector, and an OOM kill with no message when folded once at the end.

Release-time, never on the device (docs/inference.md §5): it needs
real photographs and a person reading the recall measurement that
gates whether each file is offered.
2026-09-19 16:02:37 +02:00

6.3 KiB
Raw Blame History

Face models

The shape-fixed SCRFD detectors and the ArcFace/MobileFaceNet embedder, in LFS. One copy, packaged by every platform:

Platform How it ships Where it lands
Android assemble-apk.sh bundles them as APK assets; android_main unpacks on first launch the shared user data directory
Arch PKGBUILD installs them /usr/share/darkroom/models/

library::face_models searches the account's own directory, then the shared user directory, then $XDG_DATA_DIRS — so a pair the user placed by hand always outranks the packaged one.

scrfd_500m_640.onnx    2.5 MB   "Fast" in settings — what every library was indexed with until 0.12
scrfd_2.5g_640.onnx    3.3 MB   "Balanced" — 14% more faces for 12% more time (faces.md §12.3)
scrfd_10g_640.onnx      17 MB   "Thorough" — a further 12% for 3× the time
arcface_mbf_b1.onnx     13 MB
2d106det_b1.onnx       4.8 MB   106 landmarks, for the eye boxes (faces.md §17)
ocec_s_b1.onnx         483 KB   eyes open or closed, per eye
sgc_l_48_b1.onnx       6.1 MB   sunglasses or not, per head

Which detector runs is a per-device setting (FaceSettings::detector); all three are installed so the choice exists on every platform. Each is its own faces.model_id, so changing it re-indexes.

The last three are optional to the app: library::face_models reports them beside the pair when all three are there, and a library without them indexes faces and simply has no eye readings — every reader treats "never read" as unknown, never as closed. They are found in the same directory as the pair, so a hand-placed pair does not pick up a package's eye models from a directory it otherwise outranks.

scrfd_500m_640.int8.onnx   0.8 MB   the same three, in the form the Hexagon NPU takes
scrfd_2.5g_640.int8.onnx   0.9 MB   (docs/inference.md §5) — opset 17, per-channel int8
scrfd_10g_640.int8.onnx    4.3 MB   weights, uint8 activations, calibrated on 96 photographs

The int8 files are derived by tools/quantise-models.sh from the f32 ones beside them and travel with them: the engine loads the .int8.onnx sibling when the device's backend wants it and the canonical file otherwise, and a library indexed on the int8 form records it as a different detector (scrfd_500m_i8+w600k_mbf), because it finds a different set of faces. Every other platform ignores them. The embedder has no int8 form and never will (§7 of the same document).

A clone without git-lfs gets a ~130-byte pointer where each model should be. Both packagers check for exactly that and refuse, rather than shipping the pointer and failing inside tract on the user's machine. Fix it with git lfs pull.

These are not what InsightFace ships. They came from buffalo_sc.zip, buffalo_m.zip, buffalo_l.zip and buffalo_s.zip on the InsightFace v0.7 release with their input dimensions pinned, because tract cannot parse any of the graphs while they are dynamic:

./tools/fix-face-model-shapes.sh det_500m.onnx  models/face/scrfd_500m_640.onnx --input input.1=1,3,640,640
./tools/fix-face-model-shapes.sh det_2.5g.onnx  models/face/scrfd_2.5g_640.onnx --input input.1=1,3,640,640
./tools/fix-face-model-shapes.sh det_10g.onnx   models/face/scrfd_10g_640.onnx  --input input.1=1,3,640,640
./tools/fix-face-model-shapes.sh w600k_mbf.onnx models/face/arcface_mbf_b1.onnx --dim None=1
./tools/fix-face-model-shapes.sh 2d106det.onnx  models/face/2d106det_b1.onnx    --dim None=1

2d106det_b1.onnx is from the same buffalo_l.zip as det_10g.onnx and under the same grant: the 106-point landmark model whose lid contours the eye boxes are cut from (docs/faces.md §17.2). sha256 as fetched f001b856…a7109dbf, as shipped afc2984c…03368ef26.

The weights carry a non-commercial research-only grant. They are here because this is a private repository and self-installed builds; they come back out before anything is published, and the restriction binds whoever uses the app, not only the project. docs/faces.md §2.2a is the decision.

The two classifiers are a different matter

ocec_s_b1.onnx and sgc_l_48_b1.onnx are MIT, code and weights, from Katsuya Hyodo's ultra-lightweight classifier series — the same author as the whole-body detector the reference pipeline uses. They are not under the InsightFace grant and do not come out when the project publishes. Read 2026-09-19:

OCEC — open/closed eyes SGC — sunglasses
Source github.com/PINTO0309/OCEC, release onnx, ocec_s.onnx github.com/PINTO0309/SGC, release onnx, sgc_is_l_48x48.onnx
Licence MIT (repository LICENSE; no separate grant on the weights) MIT, likewise
Training data Open and Closed Eyes (Młodawski 2024, HF, ODC-By 1.0 — attribution only), crops cut by DEIMv2-Wholebody34 (Apache 2.0) Not stated. The README names no dataset and carries no acknowledgement; data/ holds a class-ratio plot and nothing else.
Input one eye, 40×24, RGB, x/255 one head, 48×48, RGB, x/255
Output prob_open, a sigmoid prob_sunglasses, a sigmoid
sha256 as fetched 9a346a08…c604ba9b64 9c13d937…a4063c73
sha256 as shipped c848d34c…4b29d04c7c d49b6206…7e9538525c6

Both shipped with a dynamic batch dimension that tract loads but the project pins anyway, with the same script as the pair:

./tools/fix-face-model-shapes.sh ocec_s.onnx         models/face/ocec_s_b1.onnx   --dim batch=1
./tools/fix-face-model-shapes.sh sgc_is_l_48x48.onnx models/face/sgc_l_48_b1.onnx --dim batch=1

The one thing D13's discipline turns up here is SGC's undocumented training set. The weights' grant is MIT and that is what binds a redistributor; but "trained on what" is the question this project reads first, and for SGC it has no answer. Recorded so it is a known gap rather than an assumption, and so that whoever finds a sunglasses classifier with a stated dataset knows what to replace.

Attribution, as ODC-By asks for the eye dataset:

Michał Młodawski, Open and Closed Eyes Dataset, July 2024, https://huggingface.co/datasets/MichalMlodawski/closed-open-eyes

The variant choice was measured, not taken from the F1 column: on family snapshots the S variant of OCEC read more open eyes as open than M or L did, which overfit their own domain (docs/faces.md §17).