Run each model on the Hexagon in the form measured to hold it
The engine knew f32 and int8, and gave the Hexagon int8 for every role it served. Measured on the tablet itself (inference.md §1.5), int8 lost 5% of the detector's faces at 40-80 px, moved the landmarks 1.5 px, emptied the segmenter's scores and cost the denoiser 5-9 dB; fp16 the HTP refuses outright. `Form` gains A16W8 and A16W16, and `Rung::form` now names one per role: detectors and landmarks A16W8, the segmenter, scene model, border filler and denoiser A16W16, XFeat int8. The embedder and the eye classifiers stay on the CPU. Each loader resolves its `<stem>.<form>.onnx` sibling; the segmenter and XFeat, compiled into the binary, embed their quantised forms on Android only and pick through `choose_embedded`. The probe, the compile step and the cache fingerprint follow the form instead of assuming int8. Detectors on the new form write `scrfd_*_a16+w600k_mbf`, and `model_ids` answers for all three spellings. On the tablet (ORT 1.29 + QNN 2.42), each shipped file against f32 on the same inputs, and against the CPU's f32 time: SCRFD 500m/2.5g/10g A16W8 100% of faces in every band 4.2/5.1/9.0 ms vs 17/56/198 landmarks A16W8 0.25 px in the 192 crop 0.5 ms vs 2.8 YOLO26n-seg A16W16 98.2% found, mask IoU 0.994 12.9 ms vs 90 scene model A16W16 98.9% of cells agree 15 ms vs 151 MI-GAN A16W16 41 dB from f32 in the fill 87 ms vs 488 XFeat int8 pano alignment 0.45 px (f32's own spread 0.41) 6.5 ms vs 58 denoiser A16W16 0.00 dB at every ISO 95 ms vs 1510 a tile Face numbers are over public COCO val2017 photographs, not a library. The APK carries the siblings (BUNDLED 15 -> 19; the old int8 detectors removed), about 43 MB more. The Windows installer and its CI count skip them; the Arch and Flatpak packages list their files and never had them. The ladder example takes a role per model, which is how the per-role forms above were seen landing on the NPU from the real probe.
This commit is contained in:
@@ -14,6 +14,13 @@ Both come from `https://huggingface.co/Ultralytics/YOLO26`. The face weights in
|
||||
The keypoint weights in `keypoints/` and the border filler in `inpaint/` are
|
||||
the other two, and the easiest — see the last two sections.
|
||||
|
||||
Every quantised sibling — `*.int8.onnx`, `*.a16w8.onnx`, `*.a16w16.onnx`, the
|
||||
forms the tablet's Hexagon runs (`tools/quantise-models.sh`) — is the same
|
||||
weights rounded, and carries exactly the grant of the file it was made from.
|
||||
What calibration adds is one minimum and maximum per tensor: from public COCO
|
||||
val2017 photographs (CC-BY 4.0) for the image models, and for the denoiser from
|
||||
the same training tiles its weights were learned from. No image is in the files.
|
||||
|
||||
## The grant
|
||||
|
||||
**Ultralytics releases YOLO under AGPL-3.0**, and the weights carry the same
|
||||
|
||||
Binary file not shown.
Binary file not shown.
+14
-8
@@ -28,15 +28,21 @@ every reader treats "never read" as unknown, never as closed. They are found in
|
||||
directory as the pair, so a hand-placed pair does not pick up a package's eye models from a
|
||||
directory it otherwise outranks.
|
||||
|
||||
scrfd_500m_640.int8.onnx 0.8 MB the same three, in the form the Hexagon NPU takes
|
||||
scrfd_2.5g_640.int8.onnx 0.9 MB (docs/dev/inference.md §5) — opset 17, per-channel int8
|
||||
scrfd_10g_640.int8.onnx 4.3 MB weights, uint8 activations, calibrated on 96 photographs
|
||||
scrfd_500m_640.a16w8.onnx the same three, in the form the Hexagon NPU takes
|
||||
scrfd_2.5g_640.a16w8.onnx (docs/dev/inference.md §1.5) — 16-bit activations,
|
||||
scrfd_10g_640.a16w8.onnx per-channel 8-bit weights, calibrated on 300 photographs
|
||||
2d106det_b1.a16w8.onnx the landmarks, likewise
|
||||
|
||||
The int8 files are **derived** by `tools/quantise-models.sh` from the f32 ones beside them and
|
||||
travel with them: the engine loads the `.int8.onnx` sibling when the device's backend wants it and
|
||||
the canonical file otherwise, and a library indexed on the int8 form records it as a different
|
||||
detector (`scrfd_500m_i8+w600k_mbf`), because it finds a different set of faces. Every other
|
||||
platform ignores them. The embedder has no int8 form and never will (§7 of the same document).
|
||||
These are **derived** by `tools/quantise-models.sh` from the f32 files beside them and travel with
|
||||
them: the engine loads the `.a16w8.onnx` sibling when the device's backend wants it and the
|
||||
canonical file otherwise, and a library indexed on that form records it as a different detector
|
||||
(`scrfd_500m_a16+w600k_mbf`), because it finds a different set of faces. Every other platform
|
||||
ignores them, and the Windows installer leaves them out. The embedder and the eye classifiers have
|
||||
no quantised form: the embedder's vectors must compare across devices (inference.md §7), and the
|
||||
classifiers cost a millisecond on the CPU.
|
||||
|
||||
The detectors were int8 until §1.5 measured them on the tablet: int8 found 94–95% of f32's faces
|
||||
at 40–80 px, A16W8 all of them. A tablet that indexed under the int8 ids keeps those rows.
|
||||
|
||||
**A clone without git-lfs gets a ~130-byte pointer where each model should be.** Both packagers check
|
||||
for exactly that and refuse, rather than shipping the pointer and failing inside tract on the user's
|
||||
|
||||
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Reference in New Issue
Block a user