Run each model on the Hexagon in the form measured to hold it

The engine knew f32 and int8, and gave the Hexagon int8 for every role it
served. Measured on the tablet itself (inference.md §1.5), int8 lost
5% of the detector's faces at 40-80 px, moved the landmarks 1.5 px,
emptied the segmenter's scores and cost the denoiser 5-9 dB; fp16 the HTP
refuses outright. `Form` gains A16W8 and A16W16, and `Rung::form` now
names one per role: detectors and landmarks A16W8, the segmenter, scene
model, border filler and denoiser A16W16, XFeat int8. The embedder and
the eye classifiers stay on the CPU.

Each loader resolves its `<stem>.<form>.onnx` sibling; the segmenter and
XFeat, compiled into the binary, embed their quantised forms on Android
only and pick through `choose_embedded`. The probe, the compile step and
the cache fingerprint follow the form instead of assuming int8. Detectors
on the new form write `scrfd_*_a16+w600k_mbf`, and `model_ids` answers
for all three spellings.

On the tablet (ORT 1.29 + QNN 2.42), each shipped file against f32 on the
same inputs, and against the CPU's f32 time:
  SCRFD 500m/2.5g/10g  A16W8   100% of faces in every band   4.2/5.1/9.0 ms vs 17/56/198
  landmarks            A16W8   0.25 px in the 192 crop        0.5 ms vs 2.8
  YOLO26n-seg          A16W16  98.2% found, mask IoU 0.994    12.9 ms vs 90
  scene model          A16W16  98.9% of cells agree           15 ms vs 151
  MI-GAN               A16W16  41 dB from f32 in the fill     87 ms vs 488
  XFeat                int8    pano alignment 0.45 px (f32's own spread 0.41)  6.5 ms vs 58
  denoiser             A16W16  0.00 dB at every ISO            95 ms vs 1510 a tile
Face numbers are over public COCO val2017 photographs, not a library.

The APK carries the siblings (BUNDLED 15 -> 19; the old int8 detectors
removed), about 43 MB more. The Windows installer and its CI count skip
them; the Arch and Flatpak packages list their files and never had them.
The ladder example takes a role per model, which is how the per-role
forms above were seen landing on the NPU from the real probe.
This commit is contained in:
2026-10-04 03:45:46 -04:00
parent 0e6ac09fd5
commit 5a8c3e4c40
39 changed files with 550 additions and 208 deletions
+64 -4
View File
@@ -127,7 +127,8 @@ Three things the table settles.
five of six models. A whole-library face index on the tablet goes from ~100 ms + 39 ms per face
to ~1.4 ms + 12 ms per face, and the "Thorough" detector — 3× the cost of "Fast" today — becomes
free. Its price is that the models must be **quantised to int8**, which is an accuracy question
§5 has to answer before it is believed.
§5 has to answer before it is believed. (§1.5 answered it: int8 lost faces, and every model but
XFeat ships with 16-bit activations, at about three times these timings.)
- **The embedder does not gain from either accelerator.** 112×112 input, per-op overhead
dominates; it is 9 ms on the tablet's CPU and 12 ms on its NPU. It stays float, which §7 turns
from a performance footnote into a correctness rule.
@@ -137,6 +138,60 @@ Three things the table settles.
- **On AMD, MIGraphX fp16 is 4–17× the CPU provider** on the detectors and 60× on the
inpainter, with the same first-run compile cost as TensorRT and no rung between it and the CPU.
### 1.5 The Hexagon at every bit width · 2026-10-04
§1.1's Hexagon column is int8 calibrated on noise: timing only. This is the follow-up — every
model, every bit width the HTP offers, calibrated on real photographs and **scored on the tablet
itself** (ORT 1.29 + QNN 2.42, `htp_arch` 73), against the f32 model on the same inputs. The
"Form shipped" column is the files in `models/`, re-scored on the tablet after
`tools/quantise-models.sh` wrote them. The
calibration and scoring photographs are 800 from the public COCO val2017 set (CC-BY); the face
models' numbers are over the faces in them of at least 32 px. The tools are `tools/quantise-models.sh`
and the scratch harness described with it.
**What the HTP accepts.** fp16: nothing — every fp16 operator fails validation (3110), on QNN
2.42 and 2.50, with `htp_arch` and every `soc_model` tried; the fp16 rung stays off the table until
someone has Qualcomm's own SDK to say why. 4-bit weights (A8W4, A16W4): load, and wreck accuracy
(SCRFD finds 25–35% of f32's faces). What is left: **A8W8 (int8), A16W8 and A16W16**, all running
the whole graph. 16-bit activations cost about 3× int8's time, A16W16 about 4×.
| Model | ORT CPU f32 | Form shipped | Hexagon | On the tablet, against f32 | int8 for comparison |
|---|---|---|---|---|---|
| scrfd_500m / 2.5g / 10g | 17 / 56 / 198 ms | **A16W8** | 4.2 / 5.1 / 9.0 ms | 100% of faces found in every size band; keypoints 0.3–0.6% of the box | 94–95% of faces at 40–80 px |
| 2d106det (landmarks) | 2.8 ms | **A16W8** | 0.5 ms | 0.25 px in the 192 crop (eye points 0.20) | 1.5 px, and 29 partitions at 7.3 ms |
| yolo26n-seg | 90 ms | **A16W16**, tail in float | 12.9 ms | 98.2% of objects, mask IoU 0.994 | 74% (simulated) |
| yolo26s-sem-ade20k | 151 ms | **A16W16**, attention in float | 15 ms | 98.9% of cells agree on the class, TV 0.009 | 67% |
| migan-512 | 488 ms | **A16W16** | 87 ms | 41 dB from f32 in the fill (worst 1%: 30 dB) | 16 dB (simulated) |
| xfeat-1024 / 768 | 58 ms | **int8**, rewritten graph | 6.5 ms | panorama alignment 0.45 px from f32's — f32's own refit on 90% of its matches is 0.41 | — |
| mosaic-1408 (denoiser) | 1510 ms a tile | **A16W16**, rewritten graph | 95 ms a tile | 0.00 dB at every ISO; ≤ 0.09 dB with the noise scaled ×0.5–×4 | −4.7 to −9.2 dB |
| arcface_mbf (embedder) | 8.5 ms | f32, CPU | — | A16W16: cosine 0.9995, p1 0.9967 — misses §7's 0.999 gate | — |
| ocec, sgc (eyes) | 1, 1.7 ms | f32, CPU | — | sgc flips 1.45% of views even at A16W16; not worth a millisecond | — |
Four things the table needed that the f32 graphs did not have, all in `tools/htp_graph.py` and all
checked exact against the f32 graph before they are used:
- **Rank ≤ 5.** QNN's tensors stop at rank 5, and the denoiser packs the mosaic through a 6-D
reshape (6007 at compose). For one channel that reshape is `SpaceToDepth(2)`. XFeat's 8×8 unfold
is 224 Slices and 6-D Concats; it is `SpaceToDepth(8)` (736 nodes to 60).
- **No bilinear Resize at XFeat's sizes** (3110). A half-pixel bilinear resize between fixed sizes
is two constant matrices, so it is two MatMuls.
- **One scale per tensor.** The segmenter's output rows carry boxes in pixels beside scores in
0..1; quantised as one tensor the scores vanish. Everything from the Concat that builds the rows
stays float, on the CPU, where the top-300 selection costs nothing.
- **Float where the HTP's 16-bit arithmetic drifts.** The scene model's one attention block
(two MatMuls and a Softmax over 400 tokens) moved its agreement from 98.7% to 96.7%; it stays
float.
**ORT's CPU simulation of a QDQ graph is not the tablet.** It matched to the hundredth of a dB for
the denoiser and to rounding for the detectors, landmarks and XFeat, and it overstated MI-GAN by
27 dB and the scene model by three points. Every number above is the device's; a new form is not
measured until it has run there.
**XFeat's int8 loses keypoints and not the panorama.** 83% of f32's keypoints come back within
1.5 px; but over the twelve-frame `fixtures/pano/2025-08-05` sweep, the homographies fitted from
int8's matches land 0.45 px from f32's in the overlaps — the spread f32 shows against itself
(0.41).
---
## 2. The shape of the answer
@@ -146,7 +201,7 @@ winning:
| Platform | 1st | 2nd | 3rd | Floor |
|---|---|---|---|---|
| Android, Qualcomm with a Hexagon the shipped QNN skel covers (V68–V81) | QNN HTP, int8 model | ORT CPU, f32 model | — | tract |
| Android, Qualcomm with a Hexagon the shipped QNN skel covers (V68–V81) | QNN HTP, each model's quantised form (§1.5) | ORT CPU, f32 model | — | tract |
| Android, any other SoC | ORT CPU, f32 | — | — | tract |
| Linux / Windows, NVIDIA GPU | TensorRT, f32 model, fp16 engine | CUDA provider, f32 | ORT CPU, f32 | tract |
| Linux, AMD GPU with ROCm | MIGraphX, f32 model, fp16 program | ORT CPU, f32 | — | tract |
@@ -298,10 +353,10 @@ of which form they load:
| Form | Who produces it | When | Needed by |
|---|---|---|---|
| f32 ONNX, shape-fixed, **opset ≥ 13** | `tools/fix-face-model-shapes.sh`, `tools/export-seg-model.sh` | Release time, once | Every rung except Hexagon |
| int8 QDQ ONNX, per-channel, uint8 activations | `tools/quantise-models.sh` (new) | Release time, once, **calibrated on real photographs** | Hexagon |
| QDQ ONNX, per-channel — int8, A16W8 or A16W16 per model (§1.5) | `tools/quantise-models.sh` | Release time, once, **calibrated on real photographs**, scored on the tablet | Hexagon |
| TensorRT engine (`.engine`, per GPU architecture and TensorRT version) | The app, from the f32 file | First run on that device, in the background | TensorRT rung |
| MIGraphX program (`.mxr`, per GPU architecture, MIGraphX version and precision) | The app, from the f32 file | First run on that device, in the background | MIGraphX rung |
| QNN context binary | The app, from the int8 file | First run on that device, in the background | Hexagon rung |
| QNN context binary | The app, from the quantised file | First run on that device, in the background | Hexagon rung |
Two rules.
@@ -380,6 +435,11 @@ already the rule for the detector and because §5 is the gate on whether the int
enough to be *offered* at all. f32 on tract, ORT CPU, CUDA and TensorRT-f32 are one identity: the
same graph, the same arithmetic, differences at the last bit.
After §1.5 the Hexagon runs the detectors in **A16W8**, and that is a third spelling:
`scrfd_500m_a16+w600k_mbf` and its two siblings. Same rule, same reconciliation; a tablet that
indexed under `_i8` keeps those rows, and `FaceDetector::model_ids` answers "has this detector been
over this image" for all three forms.
**The embedder** is where comparability across devices is the whole point, and it is the one
model that no accelerator helps (§1.4). So: **the embedder runs in f32 on every rung.** On TensorRT
that means the embedder's engine is built without fp16 while the detector's is built with it; on