The strict flag refused the Hexagon over the ten quantise/dequantise
nodes at the graph's edges that QNN declines by policy, which cost
microseconds. A provider that hands real work to the CPU is slower than
the CPU floor and the timing already rejects it; the tablet measured
2.3 ms on the NPU against a 29.7 ms floor.
The first int8 files found no faces at all, and for two reasons the
tool now guards against. The calibration set was landscape photographs
with no faces in them, so the score head's ranges had never seen the
face regime; the set is now proxies from the library itself. And ONNX
Runtime's strided and moving-average calibration modes both degrade
these graphs measurably (a quarter of the faces at eight images, none
at ninety-six), while driving the calibrator in chunks by hand gives
ranges identical to a single pass — so the tool does that, four images
at a time, and feeds quantize_static through its range cache.
Measured against f32 over 400 proxies (docs/inference.md §10.1): the
10g form finds every face above 32 px the f32 form finds; 500m and
2.5g find 96%, and what they lose sits at a median confidence of 0.52
against the 0.50 threshold. Shipped with the number on record.
The Android unpack list gains the three int8 files; without that the
tablet never saw them. D13's runtime half records the reopening.
tract runs every model on one core on every platform. Measured against
ONNX Runtime's providers on the MagicPad 2 and the reference desktop:
ORT CPU alone is 3-10x, the Hexagon at int8 runs the detectors in
1-3 ms, TensorRT is ~2x the CUDA provider. NNAPI, XNNPACK, WebGPU and
CUDA int8 were tried and excluded with the numbers that excluded them.
The spec keeps the build C-free: ort::set_api takes a table from a
dlopened runtime or from ort-tract, chosen once per process. Rungs
are chosen by building a real session, cached until an input changes,
and compiled engines are built in the background after the first
frame. The embedder stays f32 everywhere; int8 detectors are a
distinct model_id and are gated on a recall measurement.