Commit Graph
4 Commits
Author SHA1 Message Date
dtourolle cbbe67fbd7 Let the probe's clock be its proof, not disable_cpu_ep_fallback
The strict flag refused the Hexagon over the ten quantise/dequantise
nodes at the graph's edges that QNN declines by policy, which cost
microseconds. A provider that hands real work to the CPU is slower than
the CPU floor and the timing already rejects it; the tablet measured
2.3 ms on the NPU against a 29.7 ms floor.
2026-09-19 16:13:19 +02:00
dtourolle 76bc5652d7 Calibrate the int8 detectors on library proxies, in chunks, and measure them
The first int8 files found no faces at all, and for two reasons the
tool now guards against. The calibration set was landscape photographs
with no faces in them, so the score head's ranges had never seen the
face regime; the set is now proxies from the library itself. And ONNX
Runtime's strided and moving-average calibration modes both degrade
these graphs measurably (a quarter of the faces at eight images, none
at ninety-six), while driving the calibrator in chunks by hand gives
ranges identical to a single pass — so the tool does that, four images
at a time, and feeds quantize_static through its range cache.

Measured against f32 over 400 proxies (docs/inference.md §10.1): the
10g form finds every face above 32 px the f32 form finds; 500m and
2.5g find 96%, and what they lose sits at a median confidence of 0.52
against the 0.50 threshold. Shipped with the number on record.

The Android unpack list gains the three int8 files; without that the
tablet never saw them. D13's runtime half records the reopening.
2026-09-19 16:02:44 +02:00
dtourolle caf21bea64 Name the crate dr-inference-engine 2026-09-19 16:02:37 +02:00
dtourolle 6739fdf908 Specify per-device inference backends, with the 2026-09-19 measurements
tract runs every model on one core on every platform. Measured against
ONNX Runtime's providers on the MagicPad 2 and the reference desktop:
ORT CPU alone is 3-10x, the Hexagon at int8 runs the detectors in
1-3 ms, TensorRT is ~2x the CUDA provider. NNAPI, XNNPACK, WebGPU and
CUDA int8 were tried and excluded with the numbers that excluded them.

The spec keeps the build C-free: ort::set_api takes a table from a
dlopened runtime or from ort-tract, chosen once per process. Rungs
are chosen by building a real session, cached until an input changes,
and compiled engines are built in the background after the first
frame. The embedder stays f32 everywhere; int8 detectors are a
distinct model_id and are gated on a recall measurement.
2026-09-19 16:02:37 +02:00