Run each model on the Hexagon in the form measured to hold it

The engine knew f32 and int8, and gave the Hexagon int8 for every role it
served. Measured on the tablet itself (inference.md §1.5), int8 lost
5% of the detector's faces at 40-80 px, moved the landmarks 1.5 px,
emptied the segmenter's scores and cost the denoiser 5-9 dB; fp16 the HTP
refuses outright. `Form` gains A16W8 and A16W16, and `Rung::form` now
names one per role: detectors and landmarks A16W8, the segmenter, scene
model, border filler and denoiser A16W16, XFeat int8. The embedder and
the eye classifiers stay on the CPU.

Each loader resolves its `<stem>.<form>.onnx` sibling; the segmenter and
XFeat, compiled into the binary, embed their quantised forms on Android
only and pick through `choose_embedded`. The probe, the compile step and
the cache fingerprint follow the form instead of assuming int8. Detectors
on the new form write `scrfd_*_a16+w600k_mbf`, and `model_ids` answers
for all three spellings.

On the tablet (ORT 1.29 + QNN 2.42), each shipped file against f32 on the
same inputs, and against the CPU's f32 time:
  SCRFD 500m/2.5g/10g  A16W8   100% of faces in every band   4.2/5.1/9.0 ms vs 17/56/198
  landmarks            A16W8   0.25 px in the 192 crop        0.5 ms vs 2.8
  YOLO26n-seg          A16W16  98.2% found, mask IoU 0.994    12.9 ms vs 90
  scene model          A16W16  98.9% of cells agree           15 ms vs 151
  MI-GAN               A16W16  41 dB from f32 in the fill     87 ms vs 488
  XFeat                int8    pano alignment 0.45 px (f32's own spread 0.41)  6.5 ms vs 58
  denoiser             A16W16  0.00 dB at every ISO            95 ms vs 1510 a tile
Face numbers are over public COCO val2017 photographs, not a library.

The APK carries the siblings (BUNDLED 15 -> 19; the old int8 detectors
removed), about 43 MB more. The Windows installer and its CI count skip
them; the Arch and Flatpak packages list their files and never had them.
The ladder example takes a role per model, which is how the per-role
forms above were seen landing on the NPU from the real probe.
This commit is contained in:
2026-10-04 03:45:46 -04:00
parent 0e6ac09fd5
commit 5a8c3e4c40
39 changed files with 550 additions and 208 deletions
+7
View File
@@ -14,6 +14,13 @@ Both come from `https://huggingface.co/Ultralytics/YOLO26`. The face weights in
The keypoint weights in `keypoints/` and the border filler in `inpaint/` are
the other two, and the easiest — see the last two sections.
Every quantised sibling — `*.int8.onnx`, `*.a16w8.onnx`, `*.a16w16.onnx`, the
forms the tablet's Hexagon runs (`tools/quantise-models.sh`) — is the same
weights rounded, and carries exactly the grant of the file it was made from.
What calibration adds is one minimum and maximum per tensor: from public COCO
val2017 photographs (CC-BY 4.0) for the image models, and for the denoiser from
the same training tiles its weights were learned from. No image is in the files.
## The grant
**Ultralytics releases YOLO under AGPL-3.0**, and the weights carry the same
Binary file not shown.
Binary file not shown.
+14 -8
View File
@@ -28,15 +28,21 @@ every reader treats "never read" as unknown, never as closed. They are found in
directory as the pair, so a hand-placed pair does not pick up a package's eye models from a
directory it otherwise outranks.
scrfd_500m_640.int8.onnx 0.8 MB the same three, in the form the Hexagon NPU takes
scrfd_2.5g_640.int8.onnx 0.9 MB (docs/dev/inference.md §5) — opset 17, per-channel int8
scrfd_10g_640.int8.onnx 4.3 MB weights, uint8 activations, calibrated on 96 photographs
scrfd_500m_640.a16w8.onnx the same three, in the form the Hexagon NPU takes
scrfd_2.5g_640.a16w8.onnx (docs/dev/inference.md §1.5) — 16-bit activations,
scrfd_10g_640.a16w8.onnx per-channel 8-bit weights, calibrated on 300 photographs
2d106det_b1.a16w8.onnx the landmarks, likewise
The int8 files are **derived** by `tools/quantise-models.sh` from the f32 ones beside them and
travel with them: the engine loads the `.int8.onnx` sibling when the device's backend wants it and
the canonical file otherwise, and a library indexed on the int8 form records it as a different
detector (`scrfd_500m_i8+w600k_mbf`), because it finds a different set of faces. Every other
platform ignores them. The embedder has no int8 form and never will (§7 of the same document).
These are **derived** by `tools/quantise-models.sh` from the f32 files beside them and travel with
them: the engine loads the `.a16w8.onnx` sibling when the device's backend wants it and the
canonical file otherwise, and a library indexed on that form records it as a different detector
(`scrfd_500m_a16+w600k_mbf`), because it finds a different set of faces. Every other platform
ignores them, and the Windows installer leaves them out. The embedder and the eye classifiers have
no quantised form: the embedder's vectors must compare across devices (inference.md §7), and the
classifiers cost a millisecond on the CPU.
The detectors were int8 until §1.5 measured them on the tablet: int8 found 94–95% of f32's faces
at 40–80 px, A16W8 all of them. A tablet that indexed under the int8 ids keeps those rows.
**A clone without git-lfs gets a ~130-byte pointer where each model should be.** Both packagers check
for exactly that and refuse, rather than shipping the pointer and failing inside tract on the user's
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.