§1.6 is the Iris Xe measurement: OpenVINO fp16 1.3–5.8× the CPU provider on every shipped model, WebGPU behind it everywhere but the denoiser and MI-GAN. §2's ladder gains the Intel and generic rows, the generic one footnoted as unmeasured where it is meant to help. §3.1 lists the two bundled runtimes' licences; §3.2 is how one runtime of several is chosen per process. D13 notes the bundling.
699 lines
48 KiB
Markdown
699 lines
48 KiB
Markdown
# Inference backends — the runtime and the model, chosen per device
|
||
|
||
Spec for **S16**, the build that puts the neural models on the hardware each device actually has.
|
||
|
||
Every model DarkRoom runs today — the three SCRFD detectors, the ArcFace embedder, YOLO26n-seg and
|
||
the ADE20K scene model — runs through `tract`, on one CPU core, on every platform. That was the
|
||
right first answer: D13's runtime half chose it because it costs no C dependency, and
|
||
[faces.md](faces.md) and [segmentation.md](segmentation.md) were written against it. It is also
|
||
between 20× and 300× slower than what the same devices can do, and this document is the record of
|
||
having measured that and the specification of what replaces it.
|
||
|
||
**It does not reopen D13's licensing half.** The weights are the same files under the same grant.
|
||
It does reopen the *runtime* half, and §3 is where it says how far.
|
||
|
||
---
|
||
|
||
## 1. What was measured · 2026-09-19
|
||
|
||
One benchmark, two builds of it: `ort`'s API over `tract` (exactly what the app links) and `ort`'s
|
||
API over a dynamically loaded ONNX Runtime with each execution provider in turn. Random 640×640
|
||
input, three warm-ups, the median of 15–30 timed runs, milliseconds. The same input every run, so
|
||
the numbers are compute cost and nothing else.
|
||
|
||
### 1.1 The tablet — Honor MagicPad 2, Snapdragon 8s Gen 3
|
||
|
||
SM8635: 1× Cortex-X4, 4× A720, 3× A520, Adreno 735, Hexagon V73. Android 16.
|
||
|
||
| Model | **tract** (today) | ORT CPU f32 | ORT CPU int8 | Adreno f32 ¹ | **Hexagon int8** ² |
|
||
|---|---|---|---|---|---|
|
||
| scrfd_500m (Fast) | 98 | 16 | 8 | 21 | **1.4** |
|
||
| scrfd_2.5g (Balanced) | 161 | 59 | 19 | ✗ | **1.8** |
|
||
| scrfd_10g (Thorough) | 489 | 204 | 48 | ✗ | **3.2** |
|
||
| arcface_mbf (per face) | 39 | 9 | 13 | 24 | 12 |
|
||
| yolo26n-seg | 287 | 94 | 38 | 54 | **5.1** |
|
||
| yolo26s-sem-ade20k | 408 | 154 | 46 | 47 | **3.7** |
|
||
|
||
¹ Qualcomm's own GPU backend (`libQnnGpu.so`, OpenCL). Fails on the two larger SCRFD graphs at an
|
||
`AveragePool` the layout transformer cannot place. ONNX Runtime's WebGPU provider also runs on this
|
||
GPU and was slower than the CPU on every model; it is not in the table because it is not a
|
||
candidate.
|
||
² QNN's HTP backend. The Hexagon **refuses float32 and float16 tensors** in this ORT 1.29 + QNN
|
||
2.42 pairing (error 3110 on every node, with `enable_htp_fp16_precision` set or not); int8 QDQ
|
||
graphs run with 99.6% of nodes on the NPU — 1718 of 1725 for SCRFD-500m, the remainder being the
|
||
quantise/dequantise at the graph's edges — verified from the partition log, not inferred from the
|
||
timing.
|
||
|
||
Also tried and rejected: **NNAPI** — the device registers no neural-networks HAL at all, so the
|
||
provider has nothing to talk to; Google deprecated it in Android 15 and Qualcomm stopped shipping
|
||
drivers for it. **XNNPACK** — slower than ORT's default CPU kernels on every model that loaded, and
|
||
aborts inside its partitioner on the SCRFD graphs.
|
||
|
||
### 1.2 The desktop — RTX 3050 Laptop, Raptor Lake, 20 threads
|
||
|
||
| Model | **tract** (today) | ORT CPU f32 | ORT CPU int8 | CUDA f32 | CUDA fp16 | TensorRT f32 | TensorRT fp16 | TensorRT int8 |
|
||
|---|---|---|---|---|---|---|---|---|
|
||
| scrfd_500m | 104 | 12 | 8 | 5.6 | 3.7 | 2.5 | **1.8** | ✗ ⁴ |
|
||
| scrfd_2.5g | 162 | 27 | 12 | 6.2 | 5.3 | 2.9 | **1.9** | ✗ ⁴ |
|
||
| scrfd_10g | 514 | 99 | 32 | 14.5 | 9.3 | 7.6 | **3.3** | ✗ ⁴ |
|
||
| arcface_mbf | 45 | 15 | 18 | 1.1 | 0.8 | 0.9 | 0.7 | ✗ ⁴ |
|
||
| yolo26n-seg | 307 | 60 | 38 | 9.3 | ✗ ³ | 6.8 | **5.5** | ✗ ⁴ |
|
||
| yolo26s-sem-ade20k | 395 | 61 | 31 | 10.7 | ✗ ³ | 7.9 | **3.4** | ✗ ⁴ |
|
||
|
||
³ The offline fp16 conversion (`onnxconverter-common`) left a mixed-type node the CUDA provider
|
||
rejects. TensorRT converts to fp16 itself at engine build and does not have this problem, which is
|
||
one reason it is the target and the CUDA provider is the fallback.
|
||
⁴ TensorRT refuses the QDQ form ONNX Runtime's quantiser writes for the Hexagon (uint8
|
||
activations); it wants symmetric int8. Not pursued: fp16 needs no quantisation, no calibration and
|
||
no accuracy gate, and it is already 30–60× tract.
|
||
|
||
CUDA int8 is deliberately absent: the CUDA provider has no int8 kernels and runs a QDQ graph by
|
||
dequantising it, which measured *slower* than f32 (7.0 vs 5.6 ms on scrfd_500m). Int8 on NVIDIA is
|
||
TensorRT's job.
|
||
|
||
**TensorRT's first load is 16–116 s per model in f32 and 35–290 s in fp16** (yolo26n-seg the
|
||
worst: nearly five minutes), because it is compiling an engine for this exact GPU. The engine caches to disk and the second load is milliseconds. That
|
||
number is what §6 is designed around.
|
||
|
||
### 1.3 The desktop — Radeon RX 7900 XT, Threadripper 2920X, 24 threads · 2026-09-20
|
||
|
||
Arch's `onnxruntime-rocm` 1.29.0 against ROCm 7.2.4 and MIGraphX 7.2.3 (gfx1100), through the
|
||
same `ep_probe` harness (`core/dr-inference-engine/examples/ep_probe.rs`). Zero input, three
|
||
warm-ups, the median of 15 runs, on a machine doing nothing else. The build columns are the wall
|
||
clock of `Session` construction: cold is a MIGraphX compile of the graph for this GPU, cached is the
|
||
same session loading the program the cold build wrote.
|
||
|
||
| Model | ORT CPU f32 | MIGraphX f32 | **MIGraphX fp16** | Compile f32 / fp16 (s) | Cached load (s) |
|
||
|---|---|---|---|---|---|
|
||
| scrfd_500m (Fast) | 10.4 | 2.8 | **2.4** | 40 / 58 | 0.3 |
|
||
| scrfd_2.5g (Balanced) | 20.7 | 3.3 | **2.8** | 37 / 40 | 0.3 |
|
||
| scrfd_10g (Thorough) | 57.9 | 4.5 | **3.4** | 40 / 48 | 0.4 |
|
||
| arcface_mbf (per face) | 12.8 | 1.8 | 1.6 | 15 / 21 | 0.4 |
|
||
| yolo26n-seg | 49.3 | 8.4 | **7.5** | 110 / 136 | 0.9 |
|
||
| yolo26s-sem-ade20k | 55.8 | 4.8 | **3.8** | 50 / 60 | 0.5 |
|
||
| 2d106det (landmarks) | 9.9 | 1.2 | 1.0 | 16 / 21 | 0.2 |
|
||
| ocec_s (eye state) | 2.5 | 0.5 | 0.4 | 17 / 17 | 0.1 |
|
||
| xfeat-1024 | 26.8 | 10.5 | 9.9 | 37 / 53 | 0.3 |
|
||
| migan-512 (per tile) | 514 | 12.7 | **8.3** | 102 / 132 | 0.8 |
|
||
|
||
Three things the table settles.
|
||
|
||
- **The ROCm execution provider does not exist any more.** It was ONNX Runtime's CUDA-provider twin
|
||
for AMD, removed in the 1.23 release (AMD's builds dropped it from ROCm 7.1); 1.29's ROCm build ships
|
||
`libonnxruntime_providers_migraphx.so` and nothing else for AMD, and asking for `ROCm` answers
|
||
"not enabled in this build". So there is no non-compiling AMD rung to sit under MIGraphX the way
|
||
the CUDA provider sits under TensorRT: the AMD ladder is MIGraphX, then the CPU.
|
||
- **MIGraphX is a compiling provider, and its cache has to be asked for by name.** 15–135 s per
|
||
graph cold, under a second from its cache — TensorRT's shape exactly, and §6's design covers it.
|
||
Two things the provider does that the code has to know: `ort`'s builder fills the legacy options
|
||
struct, which 1.29 reads for the precision flags only, so the cache directory
|
||
(`migraphx_model_cache_dir`) reaches it only through the generic key/value registration; and
|
||
the cache key is the graph, the GPU and the MIGraphX version *without the precision*, so an fp16
|
||
session pointed at the f32 program's directory silently loads the f32 program (the first fp16
|
||
row measured here was that, before the directories were split).
|
||
- **fp16 is worth 10–35% over f32 on this card, not the 3× it is worth on TensorRT**, because
|
||
MIGraphX f32 is already 3–8× the CPU provider and the small graphs are launch-bound. The
|
||
detectors at 2.4–3.4 ms sit beside TensorRT fp16's 1.8–3.3 ms on the RTX 3050; the embedder is
|
||
1.6–1.8 ms on either precision and stays f32 (§7). The whole face pipeline for one image
|
||
(detector + landmarks + eyes + embedder) is under 6 ms.
|
||
|
||
### 1.4 What the numbers say
|
||
|
||
- **`tract` is single-threaded.** The tablet's one X4 core and one Raptor Lake core give the same
|
||
tract numbers. Replacing it with ONNX Runtime's CPU provider, *no accelerator involved*, is 3–6×
|
||
on the tablet and 8–10× on the desktop. That is the floor, and it is available on every platform
|
||
the app builds for.
|
||
- **The Hexagon is the standout.** A 6 W NPU running int8 beats a discrete RTX 3050 running f32 on
|
||
five of six models. A whole-library face index on the tablet goes from ~100 ms + 39 ms per face
|
||
to ~1.4 ms + 12 ms per face, and the "Thorough" detector — 3× the cost of "Fast" today — becomes
|
||
free. Its price is that the models must be **quantised to int8**, which is an accuracy question
|
||
§5 has to answer before it is believed. (§1.5 answered it: int8 lost faces, and every model but
|
||
XFeat ships with 16-bit activations, at about three times these timings.)
|
||
- **The embedder does not gain from either accelerator.** 112×112 input, per-op overhead
|
||
dominates; it is 9 ms on the tablet's CPU and 12 ms on its NPU. It stays float, which §7 turns
|
||
from a performance footnote into a correctness rule.
|
||
- **On NVIDIA, TensorRT fp16 ≈ 3× the CUDA provider**, and the CUDA provider ≈ 2× the
|
||
multi-threaded CPU; at fp16 the detectors are 1.8–3.3 ms with no quantisation at all. Both leave the twenty cores free for decoding during a batch index, which the table does not
|
||
show and which matters more than the ratio.
|
||
- **On AMD, MIGraphX fp16 is 4–17× the CPU provider** on the detectors and 60× on the
|
||
inpainter, with the same first-run compile cost as TensorRT and no rung between it and the CPU.
|
||
|
||
### 1.5 The Hexagon at every bit width · 2026-10-04
|
||
|
||
§1.1's Hexagon column is int8 calibrated on noise: timing only. This is the follow-up — every
|
||
model, every bit width the HTP offers, calibrated on real photographs and **scored on the tablet
|
||
itself** (ORT 1.29 + QNN 2.42, `htp_arch` 73), against the f32 model on the same inputs. The
|
||
"Form shipped" column is the files in `models/`, re-scored on the tablet after
|
||
`tools/quantise-models.sh` wrote them. The
|
||
calibration and scoring photographs are 800 from the public COCO val2017 set (CC-BY); the face
|
||
models' numbers are over the faces in them of at least 32 px. The tools are `tools/quantise-models.sh`
|
||
and the scratch harness described with it.
|
||
|
||
**What the HTP accepts.** fp16: nothing — every fp16 operator fails validation (3110), on QNN
|
||
2.42 and 2.50, with `htp_arch` and every `soc_model` tried; the fp16 rung stays off the table until
|
||
someone has Qualcomm's own SDK to say why. 4-bit weights (A8W4, A16W4): load, and wreck accuracy
|
||
(SCRFD finds 25–35% of f32's faces). What is left: **A8W8 (int8), A16W8 and A16W16**, all running
|
||
the whole graph. 16-bit activations cost about 3× int8's time, A16W16 about 4×.
|
||
|
||
| Model | ORT CPU f32 | Form shipped | Hexagon | On the tablet, against f32 | int8 for comparison |
|
||
|---|---|---|---|---|---|
|
||
| scrfd_500m / 2.5g / 10g | 17 / 56 / 198 ms | **A16W8** | 4.2 / 5.1 / 9.0 ms | 100% of faces found in every size band; keypoints 0.3–0.6% of the box | 94–95% of faces at 40–80 px |
|
||
| 2d106det (landmarks) | 2.8 ms | **A16W8** | 0.5 ms | 0.25 px in the 192 crop (eye points 0.20) | 1.5 px, and 29 partitions at 7.3 ms |
|
||
| yolo26n-seg | 90 ms | **A16W16**, tail in float | 12.9 ms | 98.2% of objects, mask IoU 0.994 | 74% (simulated) |
|
||
| yolo26s-sem-ade20k | 151 ms | **A16W16**, attention in float | 15 ms | 98.9% of cells agree on the class, TV 0.009 | 67% |
|
||
| migan-512 | 488 ms | **A16W16** | 87 ms | 41 dB from f32 in the fill (worst 1%: 30 dB) | 16 dB (simulated) |
|
||
| xfeat-1024 / 768 | 58 ms | **int8**, rewritten graph | 6.5 ms | panorama alignment 0.45 px from f32's — f32's own refit on 90% of its matches is 0.41 | — |
|
||
| mosaic-1408 (denoiser) | 1510 ms a tile | **A16W16**, rewritten graph | 95 ms a tile | 0.00 dB at every ISO; ≤ 0.09 dB with the noise scaled ×0.5–×4 | −4.7 to −9.2 dB |
|
||
| arcface_mbf (embedder) | 8.5 ms | f32, CPU | — | A16W16: cosine 0.9995, p1 0.9967 — misses §7's 0.999 gate | — |
|
||
| ocec, sgc (eyes) | 1, 1.7 ms | f32, CPU | — | sgc flips 1.45% of views even at A16W16; not worth a millisecond | — |
|
||
|
||
Four things the table needed that the f32 graphs did not have, all in `tools/htp_graph.py` and all
|
||
checked exact against the f32 graph before they are used:
|
||
|
||
- **Rank ≤ 5.** QNN's tensors stop at rank 5, and the denoiser packs the mosaic through a 6-D
|
||
reshape (6007 at compose). For one channel that reshape is `SpaceToDepth(2)`. XFeat's 8×8 unfold
|
||
is 224 Slices and 6-D Concats; it is `SpaceToDepth(8)` (736 nodes to 60).
|
||
- **No bilinear Resize at XFeat's sizes** (3110). A half-pixel bilinear resize between fixed sizes
|
||
is two constant matrices, so it is two MatMuls.
|
||
- **One scale per tensor.** The segmenter's output rows carry boxes in pixels beside scores in
|
||
0..1; quantised as one tensor the scores vanish. Everything from the Concat that builds the rows
|
||
stays float, on the CPU, where the top-300 selection costs nothing.
|
||
- **Float where the HTP's 16-bit arithmetic drifts.** The scene model's one attention block
|
||
(two MatMuls and a Softmax over 400 tokens) moved its agreement from 98.7% to 96.7%; it stays
|
||
float.
|
||
|
||
**ORT's CPU simulation of a QDQ graph is not the tablet.** It matched to the hundredth of a dB for
|
||
the denoiser and to rounding for the detectors, landmarks and XFeat, and it overstated MI-GAN by
|
||
27 dB and the scene model by three points. Every number above is the device's; a new form is not
|
||
measured until it has run there.
|
||
|
||
**XFeat's int8 loses keypoints and not the panorama.** 83% of f32's keypoints come back within
|
||
1.5 px; but over the twelve-frame `fixtures/pano/2025-08-05` sweep, the homographies fitted from
|
||
int8's matches land 0.45 px from f32's in the overlaps — the spread f32 shows against itself
|
||
(0.41).
|
||
|
||
---
|
||
|
||
### 1.6 Intel Iris Xe, and the generic rung · 2026-10-04
|
||
|
||
The RTX 3050 laptop's other GPU: Raptor Lake-P's Iris Xe (96 EU), Intel's `onnxruntime-openvino`
|
||
1.24.1 (OpenVINO 2025.4.1) and Microsoft's `onnxruntime-webgpu` 1.27.0, both PyPI wheels, through
|
||
`ep_probe`. Three warm-ups, the median of 15 runs. Another build shared the CPU during the run, so
|
||
the CPU columns are a little pessimistic; the GPU columns are not.
|
||
|
||
| Model | ORT CPU f32 | OpenVINO CPU | OpenVINO GPU f32 | **OpenVINO GPU fp16** | WebGPU (Iris Xe) |
|
||
|---|---|---|---|---|---|
|
||
| scrfd_500m (Fast) | 9.5 | 10.8 | 7.3 | **5.8** | 24.2 |
|
||
| scrfd_2.5g (Balanced) | 18.5 | 16.2 | 15.5 | **11.1** | 40.8 |
|
||
| scrfd_10g (Thorough) | 58.5 | 72.4 | 38.7 | **23.0** | 84.5 |
|
||
| arcface_mbf (per face) | 9.5 | 11.7 | **3.0** | 2.4 | 56.6 |
|
||
| 2d106det (landmarks) | 10.8 | 2.0 | 2.4 | **1.9** | 42.7 |
|
||
| yolo26s-sem-ade20k | 57.0 | 47.4 | 26.0 | **16.7** | 53.0 |
|
||
| xfeat-1024 | 23.5 | 18.5 | 19.9 | **17.5** | 34.2 |
|
||
| migan-512 (per tile) | 330 | ✗ ¹ | 89.7 | **57.2** | 275 |
|
||
| mosaic-fast-1408 (per tile) | 159 | 107 | 68.1 | **40.0** | 188 ² |
|
||
| mosaic-best-1408 (per tile) | 1109 | 1670 | 947 | **604** | 1034 ² |
|
||
|
||
¹ OpenVINO's CPU plugin refuses the graph at initialisation. Not shipped (§3.2), so moot.
|
||
² A later run, after `ep_probe` learned to feed the denoiser's two inputs, under heavier load: the
|
||
CPU provider took 256 and 1034 ms in that run, so WebGPU beat it by a quarter on mosaic-fast and tied
|
||
on mosaic-best — the only rows where it is not well behind.
|
||
|
||
- **OpenVINO on the Iris Xe beats ONNX Runtime's CPU provider on every model**, 1.3× on XFeat to
|
||
5.8× on MI-GAN, with a 1–3 s compile per graph and 0.1–0.4 s from its cache. It is the Intel rung.
|
||
fp16 is worth 1.3–1.7× over f32 here, against 1.1–1.35× on MIGraphX.
|
||
- **Its "GPU" is OpenCL's first GPU, not Intel's.** Before `intel-compute-runtime` was installed
|
||
the only OpenCL driver was NVIDIA's, and `device_type=GPU` ran on the RTX 3050 — slower than the
|
||
CPU, which is the probe's to catch. Read the process's maps for `libigdrcl` before believing a
|
||
number is the iGPU's.
|
||
- **WebGPU is slower than the CPU on the Iris Xe** on everything but MI-GAN, as it was on the
|
||
Adreno (§1.1), and on the RTX 3050 through Vulkan too. It is on the ladder anyway, as the generic
|
||
rung (§2): for GPUs no vendor rung covers — an AMD card on Windows or without ROCm, a Mali — where
|
||
it is unmeasured, and the probe's clock decides.
|
||
- **OpenVINO's CPU plugin is not a better floor.** It wins on some graphs and loses on scrfd_10g,
|
||
the embedder and mosaic-best, and refuses MI-GAN.
|
||
|
||
## 2. The shape of the answer
|
||
|
||
A **ladder per platform**, walked at start-up, with the first rung that builds a real session
|
||
winning:
|
||
|
||
| Platform | 1st | 2nd | 3rd | Floor |
|
||
|---|---|---|---|---|
|
||
| Android, Qualcomm with a Hexagon the shipped QNN skel covers (V68–V81) | QNN HTP, each model's quantised form (§1.5) | ORT CPU, f32 model | — | tract |
|
||
| Android, any other SoC ⁶ | WebGPU (Vulkan), f32 | ORT CPU, f32 | — | tract |
|
||
| Linux / Windows, NVIDIA GPU | TensorRT, f32 model, fp16 engine | CUDA provider, f32 | ORT CPU, f32 | tract |
|
||
| Linux, AMD GPU with ROCm | MIGraphX, f32 model, fp16 program | ORT CPU, f32 | — | tract |
|
||
| Linux / Windows, Intel GPU | OpenVINO, f32 model, fp16 program (§1.6) | ORT CPU, f32 | — | tract |
|
||
| Linux / Windows, any other GPU ⁶ | WebGPU (Vulkan / D3D12), f32 | ORT CPU, f32 | — | tract |
|
||
| macOS ⁵ | CoreML, f32 model, ML Program | ORT CPU, f32 | — | tract |
|
||
|
||
⁵ **Unmeasured**, and the one exception to the rule below: nobody here has a Mac. The rung is on
|
||
the ladder because the probe makes a wrong guess cheap — a CoreML that is slower than the CPU is
|
||
rejected by §4's clock, one that errors is recorded as failed, and one that takes the process
|
||
down is refused on the third launch (§4, `attempt`). The embedder stays on the CPU (§7). The first
|
||
macOS log that shows a probe line is this row's measurement; [macos.md](macos.md) says what to
|
||
ask for.
|
||
|
||
⁶ **The generic rung, unmeasured where it is meant to help.** WebGPU lost to the CPU on every GPU
|
||
it has been timed on — the Adreno, the Iris Xe, the RTX 3050 (§1.1, §1.6) — none of which it serves
|
||
here, since each has its own rung. It is on the ladder for the GPUs that have none, on the same
|
||
terms as CoreML: a WebGPU that is slower than the CPU is rejected by §4's clock, one that errors is
|
||
recorded as failed. Its first measurement on an AMD card without ROCm, or a Mali, is this row's.
|
||
|
||
Deliberately **not** on any ladder, with the measurement that excluded each: NNAPI (no driver),
|
||
XNNPACK (slower than CPU, aborts on SCRFD), the Adreno through QNN (works, but never where the
|
||
Hexagon does not also), CUDA int8 (slower than CUDA f32), the ROCm provider (gone: §1.3),
|
||
OpenVINO's CPU plugin as a floor (§1.6). A rung is added to this table by a measurement on this
|
||
page, not by a provider existing — the two footnoted rows are the exceptions, and say so.
|
||
|
||
The AMD ladder has no middle rung. TensorRT falls back to the CUDA provider while its engines
|
||
compile; MIGraphX has no such twin, so its fallback is the CPU provider, and the minute or two of
|
||
compiling on first run (§6) is spent at the floor's speed rather than at half the GPU's.
|
||
|
||
Two things the ladder is *not*: it is not a per-model choice — one backend serves every model on a
|
||
device, because §7's identity rule needs the detector and embedder on the same runtime for the
|
||
same reason `faces.model_id` pairs them; and it is not a per-account choice — it is a property of
|
||
the hardware, like `shared_face_models_dir` is, and it lives beside it.
|
||
|
||
---
|
||
|
||
## 3. The dependency policy, and how far this reopens it
|
||
|
||
D13 chose `ort` over `tract` because `alternative-backend` made ONNX Runtime's *API* available with
|
||
none of its *C*. Every rung above the floor needs the C++ ONNX Runtime and, for the two that matter
|
||
most, vendor libraries on top: Qualcomm's QNN runtime (~60 MB for one Hexagon generation; it is
|
||
per-SoC) and NVIDIA's TensorRT plus cuDNN (~600 MB with the CUDA libraries, and cuDNN's major
|
||
version must match what ONNX Runtime was built against — the Arch package on the reference desktop
|
||
was unusable for exactly that reason).
|
||
|
||
The policy protected the **build**: no C to cross-compile under the NDK, no toolchain to keep in
|
||
step. This document keeps that intact, and the mechanism is the one thing about `ort` that makes it
|
||
possible:
|
||
|
||
**`ort::set_api` accepts any `OrtApi` table.** With `alternative-backend` on, `ort` links nothing
|
||
and asks for the table once per process. The application can `dlopen` a `libonnxruntime.so` it
|
||
finds on disk, call `OrtGetApiBase()->GetApi(version)` and hand that table over; or, if there is no
|
||
such file, hand over `ort_tract::api()`. The Rust build is identical in both cases — pure Rust,
|
||
`cargo build --target aarch64-linux-android` sees the same dependency graph it sees today. What
|
||
changes is that the runtime is a **file the package installs**, next to the models, and the app
|
||
looks for it at start-up.
|
||
|
||
Consequences that follow and are accepted:
|
||
|
||
- **The runtime is chosen once per process**, because `set_api` is once per process. The ladder
|
||
in §2 is walked at start-up and the result is what every session in that process uses. There is
|
||
no "tract for this model, ORT for that one", and there is no falling back to tract *after* ONNX
|
||
Runtime has loaded — but there does not need to be: once the library loads, its CPU provider is
|
||
always there, and every fallback the ladder needs is between providers *inside* it.
|
||
- **Feature flags stay as they are.** `dr-face`'s `inference` and `dr-segment`'s `semantic`
|
||
continue to mean "compiled against `ort`'s API"; nothing at build time knows or cares which table
|
||
will be supplied. The one addition is a `native-probe` feature on the new crate (§8) that pulls in
|
||
`libloading`, which is pure Rust and already in the tree via `wgpu`.
|
||
- **The packagers ship the runtime, not the build.** The Arch package, the Flatpak manifest, the
|
||
NSIS installer and `assemble-apk.sh` each gain the ONNX Runtime library for their platform, and
|
||
the Android and NVIDIA variants gain the vendor libraries — each under the licence the packager
|
||
reads first (§3.1). A package without them is not broken; it is the tract build, and it says so
|
||
on the about screen.
|
||
- **The NDK problem does not come back.** `libonnxruntime.so` for Android is a prebuilt from
|
||
Maven (`com.microsoft.onnxruntime:onnxruntime-android-qnn`), extracted by `assemble-apk.sh` into
|
||
`jniLibs/` the way the models are bundled as assets today. Nothing compiles it.
|
||
|
||
### 3.1 Licences the packagers read before shipping a runtime
|
||
|
||
Written down now, because [segmentation.md §7](segmentation.md) established that reading the grant
|
||
is cheaper than discovering it at packaging time.
|
||
|
||
| Component | Licence | Redistributable in a self-distributed package? |
|
||
|---|---|---|
|
||
| ONNX Runtime | MIT | Yes |
|
||
| Qualcomm QNN runtime (`com.qualcomm.qti:qnn-runtime` on Maven) | Qualcomm AI Engine Direct SDK licence — proprietary, redistribution permitted for applications using it | Yes for the APK, with the licence text shipped; not for a source distribution. **To be read in full, not summarised from memory, before the APK gains it.** |
|
||
| CUDA runtime, cuDNN, TensorRT | NVIDIA EULAs — redistributable with an application, with the licence text, not modifiable | Yes for a package that bundles them. 600 MB. The alternative is to load them from the user's system install if present and skip the rung otherwise — which is what §4's probe does anyway. |
|
||
| ROCm (HIP, MIOpen, rocBLAS), MIGraphX | MIT | Yes, but the HIP SDK MIGraphX needs is ~15 GB installed. Same answer as NVIDIA: the user's system install, or the rung is skipped. |
|
||
|
||
The position this takes: the **GPU vendors' libraries are not bundled**. The desktop package probes for a
|
||
system CUDA/TensorRT or ROCm/MIGraphX install and uses it if it is version-compatible; a desktop without one runs on
|
||
ORT CPU, which is still 8–10× today. Bundling 600 MB for a rung that is 2× again is not a trade
|
||
worth making unmeasured, and it can be revisited by a measurement on a batch index. The **QNN
|
||
runtime is bundled** in the APK, because the Hexagon is the difference between a tablet that
|
||
indexes a library overnight and one that does it over lunch, and the package is 60 MB larger for
|
||
it.
|
||
|
||
Both positions are D13 territory and are recorded there (§12).
|
||
|
||
Two more runtimes ship in every desktop package since 0.23 (§3.2), and their parts are all
|
||
redistributable:
|
||
|
||
| Component | Licence | Shipped |
|
||
|---|---|---|
|
||
| Intel's `onnxruntime-openvino` build, with OpenVINO 2025.4.1 and oneTBB | MIT; Apache-2.0; Apache-2.0 | Linux and Windows packages, texts beside the libraries |
|
||
| Microsoft's WebGPU build (Dawn inside); on Windows the DirectX shader compiler | MIT; LLVM / MIT | Linux and Windows packages |
|
||
| Microsoft's stock `onnxruntime-android` (WebGPU) | MIT | The APK, as `libonnxruntime_generic.so` |
|
||
|
||
### 3.2 Several runtimes, one per process
|
||
|
||
A runtime carries one vendor's providers: Intel's build has OpenVINO, the `onnxruntime-gpu` wheel
|
||
CUDA and TensorRT, a ROCm build MIGraphX, Microsoft's WebGPU build the generic rung, the APK's QNN
|
||
build the Hexagon. No prebuilt carries two vendors, and `set_api` takes one table per process.
|
||
So a device that may hold several — the package's OpenVINO and WebGPU builds, a CUDA build the user
|
||
fetched, the distribution's ROCm build — has to choose which to load *before* the probe, and
|
||
cannot choose by trying.
|
||
|
||
`api::install` opens every runtime on the search list, asks each for `GetAvailableProviders`, and
|
||
loads the one that scores highest against the GPUs `hardware::detect` reads from files: a vendor
|
||
rung on its own vendor's GPU (NVIDIA driver, `/dev/kfd`, a Qualcomm SoC, macOS) above OpenVINO on
|
||
an Intel GPU (PCI vendor `0x8086`; on Windows Intel's DCH driver package) above WebGPU above a
|
||
CPU-only build. Equal scores keep the search order, a perfect fit ends the search — the APK's QNN
|
||
build is listed first, so on a Qualcomm device the generic build is never opened — and
|
||
`DARKROOM_ORT_DIR` wins outright. The losers stay mapped: unloading a C++ runtime whose static
|
||
constructors ran is a crash at exit waiting to happen.
|
||
|
||
The desktop packages install the two bundled builds under `runtimes/openvino` and
|
||
`runtimes/webgpu` beside each place a package installs to, from
|
||
`tools/fetch-bundled-runtimes.sh` (PyPI wheels pinned by SHA-256, pruned to the native libraries:
|
||
81 + 31 MB on Linux, 67 + 42 MB on Windows). On Windows the chosen runtime's directory is put on
|
||
`PATH`, because Intel's build leaves OpenVINO's DLLs for the loader to find there. The Flatpak has
|
||
no Intel OpenCL driver in its sandbox, so an Intel machine there settles on the CPU.
|
||
|
||
---
|
||
|
||
## 4. Selection — the probe, its cache, and what it may not do
|
||
|
||
**A rung is chosen by building a real session on it, not by asking whether it exists.** Both
|
||
failure modes that are not "the provider is absent" were hit on 2026-09-19: a driver in a wedged
|
||
state where the provider registered and the session then failed, and a provider that registered,
|
||
took the graph, and rejected every node at partition time. The probe therefore:
|
||
|
||
1. Loads the runtime library (§3), or falls to tract and stops.
|
||
2. Times the **smallest detector** on the CPU provider first — the floor. Then, for each rung
|
||
in this platform's ladder, in order: builds a session for the same model on that provider
|
||
with `error_on_failure`, runs it once on a fixed input, and times three more runs. **The rung
|
||
is taken only if its median beats the floor.** That one measurement is the proof the provider
|
||
took the graph: one that silently hands the work to the CPU is the CPU rung with extra
|
||
overhead, slower than the floor, and rejected. (ONNX Runtime's
|
||
`session.disable_cpu_ep_fallback` was the first draft of this proof and refuses the Hexagon
|
||
over the ten quantise/dequantise nodes at the graph's edges that QNN declines by policy.)
|
||
3. Records the outcome — rung, runtime version, provider version, device identity (GPU name and
|
||
compute capability; SoC model and Hexagon arch), and the models' content hashes — to a small
|
||
file beside `shared_face_models_dir`. The next start-up trusts the file **unless** any of those
|
||
inputs changed, in which case it probes again. A driver update, a runtime update, a new model
|
||
file: each invalidates the cache by construction, and none needs a "reset backend" button.
|
||
|
||
What the probe may not do:
|
||
|
||
- **Block the first frame.** It runs on the same background as `install_bundled_models` and for
|
||
the same reason: a TensorRT probe can take thirty seconds cold, and a tablet that stalls that long
|
||
is an ANR. Until it reports, every model request is answered by the floor the runtime supports
|
||
(ORT CPU if the library loaded, tract otherwise), and a job that started on the floor finishes
|
||
on it — a backend does not change under a running index.
|
||
- **Retry a rung that failed within a session.** A failed probe is cached as a failure with the
|
||
same inputs; the rung is tried again when an input changes. Otherwise a wedged driver means a
|
||
thirty-second stall on every launch.
|
||
- **Crash the app twice for the same reason.** The probe runs in the app's process, and a provider
|
||
can fail by aborting rather than by returning an error (XNNPACK on SCRFD, §2). Every session build
|
||
on a rung above the CPU — the probe's, and each background compile of §6 — writes what it is
|
||
attempting to `attempt` in the cache directory first and removes it after. A launch that finds the
|
||
file knows the last one died inside that attempt; after two such launches in a row the attempt is
|
||
refused and recorded like any other failure (a rung in `failed`, an engine in `refused`), until
|
||
the fingerprint changes. Two, not one, because quitting during a forty-second TensorRT compile
|
||
leaves the same file.
|
||
- **Choose for the user without saying so.** Settings gains one row, *Inference backend*, showing
|
||
what was chosen and why in one line ("Hexagon NPU · int8 · QNN 2.42"; "CPU · ONNX Runtime 1.30 ·
|
||
TensorRT probe failed: cuDNN 8 required"), with an override to force any lower rung. The about
|
||
screen carries the same line beside the model names NFR-SEC-5 already puts there.
|
||
|
||
---
|
||
|
||
## 5. Model variants, and who makes them
|
||
|
||
Every model exists in one **canonical** form — the f32 ONNX file the app ships or the user supplies
|
||
today — and, where a rung needs it, a **derived** form. The ladder's rungs are specified in terms
|
||
of which form they load:
|
||
|
||
| Form | Who produces it | When | Needed by |
|
||
|---|---|---|---|
|
||
| f32 ONNX, shape-fixed, **opset ≥ 13** | `tools/fix-face-model-shapes.sh`, `tools/export-seg-model.sh` | Release time, once | Every rung except Hexagon |
|
||
| QDQ ONNX, per-channel — int8, A16W8 or A16W16 per model (§1.5) | `tools/quantise-models.sh` | Release time, once, **calibrated on real photographs**, scored on the tablet | Hexagon |
|
||
| TensorRT engine (`.engine`, per GPU architecture and TensorRT version) | The app, from the f32 file | First run on that device, in the background | TensorRT rung |
|
||
| MIGraphX program (`.mxr`, per GPU architecture, MIGraphX version and precision) | The app, from the f32 file | First run on that device, in the background | MIGraphX rung |
|
||
| QNN context binary | The app, from the quantised file | First run on that device, in the background | Hexagon rung |
|
||
|
||
Two rules.
|
||
|
||
**Quantisation is a release-time step, not a device-time one.** The int8 files that produced §1's
|
||
numbers were calibrated on random noise, which is enough to time and worthless to trust. A real
|
||
int8 detector is calibrated on a few hundred real photographs and then measured against the f32
|
||
detector on the reference library by [faces.md §12.3](faces.md)'s method — faces found, per size
|
||
band, per detector — before it ships. That needs the reference library and a person reading the
|
||
result, and it happens once per model release, in `tools/`, beside the shape-fixing it already
|
||
depends on. The device never quantises anything.
|
||
|
||
The SCRFD and ArcFace files are **opset 11** as InsightFace exported them, and per-channel QDQ needs
|
||
13; `tools/fix-face-model-shapes.sh` gains an opset upgrade to 17 (`onnx.version_converter`,
|
||
`ir_version` 8), which `tract` has been verified to load and which every provider on this page
|
||
prefers. That is a change to the canonical file and so a change to the shipped models, and it
|
||
happens in the same model release as the int8 files.
|
||
|
||
**Compilation is a device-time step, and it is cached.** A TensorRT engine is specific to the GPU
|
||
it was built on and the TensorRT that built it; a MIGraphX program to the GPU and the MIGraphX
|
||
that built it; a QNN context binary to the Hexagon generation. None can ship. All are built by
|
||
the app the first time that rung is selected, in the background (§6), and written beside the
|
||
probe cache keyed by the same inputs. MIGraphX's own key leaves out the precision, so the app
|
||
gives its f32 and fp16 programs separate directories — otherwise the embedder's f32 build would
|
||
be served the detector's fp16 program, or the reverse. They are
|
||
**derived, disposable, and regenerable**: deleting the cache directory costs the next launch a
|
||
rebuild and nothing else, and the directory is excluded from anything that syncs (it is a peer of
|
||
`thumbs`, not of the catalog).
|
||
|
||
---
|
||
|
||
## 6. First run — building engines without the user waiting for them
|
||
|
||
The sequence on a device where a compiling rung (TensorRT, MIGraphX, Hexagon) is selected:
|
||
|
||
1. **Launch.** The runtime loads; the probe (§4) starts in the background; the app serves every
|
||
model request from the floor. Face indexing, segmentation and scene grading all work, at
|
||
today's speed or better (ORT CPU).
|
||
2. **Probe reports** — say, TensorRT. The compiling rung is now *selected* but has **no engines**.
|
||
Model requests continue on the fallback rung below it (CUDA provider for TensorRT; ORT CPU for
|
||
MIGraphX and Hexagon), which needs no compilation and is already faster than the floor.
|
||
3. **Engines build**, one model at a time, on a single low-priority background thread, smallest
|
||
model first so the detector — the one that runs per image — is ready soonest. On the reference
|
||
desktop that is ~1 minute for the first detector and ~10 minutes for all six at fp16; on the AMD
|
||
desktop 40 s for the first detector and ~8 minutes for the set; on the tablet the QNN
|
||
context binaries take 0.8–1.7 s each and the whole set is ready before the user has opened a
|
||
library. Each engine is written to a temporary name and renamed into place, so a request never
|
||
sees a half-written file.
|
||
4. **Requests move up as engines land.** A model whose engine exists loads it on the selected
|
||
rung; one whose engine is still building loads on the fallback. **A running job does not
|
||
switch** — an index that started on the CUDA provider finishes on it — because §7 needs one
|
||
`model_id` per job, and because a job is the wrong granularity for surprise.
|
||
5. On Android, the build runs only while the app is in the foreground and the device is not in
|
||
battery saver (NFR-RES-3): a context binary takes a second, so this costs nothing, and the rule
|
||
exists for the day a model takes longer.
|
||
|
||
Settings shows a one-line progress row while engines build ("Preparing GPU engines · 3 of 6") and
|
||
nothing when they are done. A build failure demotes the rung — it is recorded in the probe cache
|
||
as a failure with the model hash as an input, so a corrected model file retries it — and the app
|
||
carries on one rung down, saying so in the same row.
|
||
|
||
---
|
||
|
||
## 7. Identity — what changes `model_id` and what may not
|
||
|
||
`faces.model_id` exists so that two libraries indexed with different networks are never compared
|
||
as if they were one ([catalog.md §10.1](catalog.md); the trap is written up in
|
||
[faces.md §14](faces.md)). A backend that changes what a network *computes* is a different network
|
||
and must be a different `model_id`; one that changes only *where* it computes it must not be.
|
||
|
||
**The detector.** An int8 SCRFD finds a different set of faces from the f32 one — that is what
|
||
§5's acceptance measures — so **the numeric form is part of the detector's identity**:
|
||
`scrfd_500m` and `scrfd_500m_i8` are two detectors in `model_id`, and a library indexed on the
|
||
tablet's Hexagon and continued on the desktop is two populations, which a re-index on either side
|
||
reconciles the same way a switch from Fast to Thorough does today. That is acceptable because it is
|
||
already the rule for the detector and because §5 is the gate on whether the int8 form is close
|
||
enough to be *offered* at all. f32 on tract, ORT CPU, CUDA and TensorRT-f32 are one identity: the
|
||
same graph, the same arithmetic, differences at the last bit.
|
||
|
||
After §1.5 the Hexagon runs the detectors in **A16W8**, and that is a third spelling:
|
||
`scrfd_500m_a16+w600k_mbf` and its two siblings. Same rule, same reconciliation; a tablet that
|
||
indexed under `_i8` keeps those rows, and `FaceDetector::model_ids` answers "has this detector been
|
||
over this image" for all three forms.
|
||
|
||
**The embedder** is where comparability across devices is the whole point, and it is the one
|
||
model that no accelerator helps (§1.4). So: **the embedder runs in f32 on every rung.** On TensorRT
|
||
that means the embedder's engine is built without fp16 while the detector's is built with it; on
|
||
the Hexagon it means the embedder is not on the NPU at all — it runs on the ORT CPU rung at 9 ms,
|
||
and the ladder's "one backend per device" is, precisely, one backend *per model role*, with the
|
||
embedder pinned. A `w600k_mbf` embedding from any device is comparable with one from any other,
|
||
which is the property the identity system, the calibration and the cross-device merge all rest
|
||
on, and it is not for sale for 3 ms.
|
||
|
||
If S16 wants fp16 for the embedder later, the gate is written now: over the reference library's
|
||
faces, the cosine between the f32 and fp16 embedding of the same crop exceeds 0.999 for 99.9% of
|
||
faces and the calibration's fitted threshold moves by less than its own confidence interval
|
||
([faces.md §8.3](faces.md)). Until measured, f32.
|
||
|
||
**Segmentation and the scene model** carry no identity across devices — their outputs are
|
||
recomputed per image and never stored beyond the cache — so they take whatever the rung offers,
|
||
int8 included, subject to §10's own acceptance.
|
||
|
||
---
|
||
|
||
## 8. Crate shape — `core/dr-inference-engine`
|
||
|
||
The seam is the same shape as [storage.md](storage.md)'s: a small crate below the consumers that is
|
||
the **only** place naming a provider, a library file or a vendor, with the consumers reduced to
|
||
"give me a session for these bytes in this role".
|
||
|
||
```
|
||
core/dr-inference-engine
|
||
src/lib.rs Runtime (Tract | Onnx { lib, version }), Backend (rung), Role (Detector | Embedder | Segmenter)
|
||
src/probe.rs §4 — the ladder per platform, the session-build probe, the cache file
|
||
src/engines.rs §6 — background compilation, the cache directory, progress
|
||
src/session.rs open(role, bytes) -> ort::Session, applying the rung and the role's precision rule
|
||
src/api.rs the unsafe block that matters: dlopen libonnxruntime, fetch OrtApi, ort::set_api — or ort_tract::api()
|
||
```
|
||
|
||
- `dr-face` and `dr-segment` **delete** their private `install_backend` and their direct
|
||
`Session::builder()` calls and take an `&dr_inference_engine::Sessions` where they take model bytes today.
|
||
Their tests keep `tract` — `dr_inference_engine::Sessions::tract()` is a constructor and the test-only path.
|
||
- `dr-inference-engine` depends on `ort` with the same workspace features as today plus `cuda`, `tensorrt`,
|
||
`qnn`: those features add option builders, not linking, under `alternative-backend`. **Verified
|
||
for the QNN, CUDA and TensorRT builders on 2026-09-19** — they go through the API table's generic
|
||
`SessionOptionsAppendExecutionProvider*`. The NNAPI builder resolves a symbol directly and would
|
||
not; it is not needed and is not enabled. MIGraphX uses no `ort` feature at all: `ort`'s builder
|
||
fills the legacy options struct, which ONNX Runtime 1.29 reads for the precision flags and
|
||
nothing else, and the compiled-program cache directory only travels through the generic
|
||
key/value entry point (`migraphx_model_cache_dir`). `session::migraphx` makes that one call
|
||
on the API table itself.
|
||
- `dr-ui` owns the settings row, the about-screen line and the progress row; it holds one
|
||
`Sessions` per process, created at launch, and passes it down. `dr_ui::library` gains
|
||
`inference_cache_dir()` beside `shared_face_models_dir()`, on the same account-independent
|
||
footing and for the same reason.
|
||
- The Android entry point's `install_bundled_models` also extracts nothing new: `jniLibs/` is
|
||
loaded by the system loader, and `dr-inference-engine` on Android looks for `libonnxruntime.so` through
|
||
`dlopen` by bare name first, which resolves to the APK's copy, before any directory.
|
||
|
||
---
|
||
|
||
## 9. Threads and memory
|
||
|
||
- ONNX Runtime's intra-op pool is sized to the physical cores minus two on desktop and to the
|
||
performance cores on Android (the X4 and the A720s; the A520s are for the compositor). tract's
|
||
single thread today is the reason a batch index leaves nineteen cores idle; ORT CPU with the
|
||
pool is the reason it will not. One session per model per process; `Session::run` is
|
||
`&mut self`-free in `ort` and internally serialised, and the index job is the only caller.
|
||
- A TensorRT session pins GPU memory for its workspace; the builder is capped at 512 MB on the
|
||
reference 6 GB card and the cap is a setting, because the develop view's tiles share the card
|
||
(NFR-RES-2). The engine cache on disk is bounded by the model set — six engines, ~80 MB — and
|
||
needs no LRU.
|
||
- The Hexagon rung sets QNN's performance mode to `Burst` for the duration of an index job and
|
||
`Default` otherwise; a 5 ms detector does not need the NPU clocked up between images.
|
||
- The probe (§4) and the engine build (§6) run on one dedicated low-priority thread. They never
|
||
share the index job's pool: a probe that competes with the job it is meant to speed up is the
|
||
frame-budget trap in a new coat.
|
||
|
||
---
|
||
|
||
## 10. What S16 measures
|
||
|
||
In order, with the gate each is:
|
||
|
||
| # | Question | Gate |
|
||
|---|---|---|
|
||
| M1 | Does one binary carry both tables? `dlopen` + `set_api` on Linux, Windows and Android; `ort_tract::api()` when the file is absent. | Go / no-go for §3. If `set_api` cannot take a dynamically fetched table on some platform, that platform ships two binaries, and the cost is stated. |
|
||
| M2 | Do the int8 SCRFD detectors, **calibrated on real photographs**, find the faces? faces.md §12.3's method over the reference library, per size band, against f32. | Ship the int8 form for a detector only if it finds ≥ 97% of the f32 detector's faces above 40 px and the difference is not concentrated in one band. Otherwise that detector's Hexagon rung is ORT CPU int8-free, and the table in §1.1 says what that costs. |
|
||
| M3 | Does the embedder on ORT CPU beside a detector on the Hexagon (§7) produce embeddings within the f32 gate? | It must — same graph, same arithmetic. This is a check that the plumbing did not quantise it by accident. |
|
||
| M4 | Is a batch index on the tablet and on the desktop faster by the ratio §1 predicts, end to end, decode included? | The face index over the reference library (18,143 faces): report wall-clock on tract, on the floor and on the selected rung, and where the time went. The prediction is that decode becomes the bottleneck on both; if it does not, say why. |
|
||
| M5 | Does the first-run sequence (§6) hold: nothing blocks the first frame, engines land, requests move up, a running job does not switch? | Observed on both devices with the app's own progress row, and with the cache directory deleted between runs. |
|
||
| M6 | What does the APK weigh with the QNN runtime, and does a non-Qualcomm Android device (any one) still launch and index on the floor? | Size reported; launch verified on one non-Qualcomm device or an emulator. |
|
||
| M7 | The segmentation and scene models at int8 on the Hexagon: does the mask boundary move? segmentation.md §6's IoU against f32 over its corpus. | Ship int8 for a model only above the IoU floor that document set for arm B. |
|
||
|
||
M1 and M2 are the ones the rest is conditional on, and M2 is the one that needs a person.
|
||
|
||
### 10.1 M2 result · 2026-09-19
|
||
|
||
Each int8 detector against its own f32 form, over 400 proxies evenly spaced through the reference
|
||
library, on ONNX Runtime's CPU provider (the int8 graph is the same file the Hexagon loads;
|
||
`ui/dr-ui/examples/face_detectors`). Calibrated on 64 proxies from the same library, disjoint
|
||
from the 400.
|
||
|
||
| Detector | f32 faces | int8 faces | both | int8 only | f32 only | found ≥ 32 px | found, all sizes |
|
||
|---|---|---|---|---|---|---|---|
|
||
| scrfd_500m | 1342 | 1319 | 1287 | 32 | 55 | 95.6% | 95.9% |
|
||
| scrfd_2.5g | 1525 | 1482 | 1478 | 4 | 47 | 96.1% | 96.9% |
|
||
| scrfd_10g | 1769 | 1760 | 1749 | 11 | 20 | 100% | 98.9% |
|
||
|
||
The 10g form clears the 97% gate; 500m and 2.5g sit one point under it. What they lose is
|
||
specific: the faces in the "f32 only" column have a **median confidence of 0.52** against a
|
||
threshold of 0.50 — detections the f32 graph itself barely made, that int8 rounding drops to the
|
||
other side of the line — and the extra faces int8 finds are the same kind (median 0.51–0.52).
|
||
Not a size-band failure: the losses are spread across bands in proportion. Shipped as they are,
|
||
with the number on record; a threshold of 0.48 for the int8 forms would recover most of the
|
||
margin, and is the first thing to try if a library's count on the tablet reads low.
|
||
|
||
Two things the calibration taught, both in `tools/quantise-models.py`: the calibration set has
|
||
to contain faces (a first attempt on landscape photographs produced a graph that found nothing —
|
||
the score head's ranges had never seen the face regime), and ONNX Runtime's own strided and
|
||
moving-average calibration modes both measurably degrade the result on these graphs, while
|
||
driving the calibrator in chunks by hand reproduces the plain min/max ranges exactly.
|
||
|
||
---
|
||
|
||
## 11. Order
|
||
|
||
1. **`dr-inference-engine` with the two tables and the floor** — `set_api` from a dlopened runtime, tract
|
||
otherwise, ORT CPU as the only rung. Consumers moved over; tests unchanged. This alone is the
|
||
3–10× and is the build most of the value sits in. M1.
|
||
2. **The probe and its cache** (§4), with the settings row and the about line. Still CPU-only;
|
||
the ladder has one rung. M5's first half.
|
||
3. **`tools/quantise-models.sh`** and the opset upgrade; the int8 detectors calibrated and
|
||
measured. M2, M7. This is the step with a person in it and it runs in parallel with 4.
|
||
4. **The Hexagon rung**, the QNN runtime in the APK, the context-binary cache. M3, M6.
|
||
5. **The TensorRT and CUDA rungs** on desktop, the engine cache, the first-run sequence. M5's
|
||
second half.
|
||
6. **M4** last, on both devices, and the number goes in this document.
|
||
|
||
The Windows installer and the Flatpak manifest are touched in steps 1 and 5 only, and only to add
|
||
a file each; the Arch package likewise.
|
||
|
||
---
|
||
|
||
## 12. Register entries
|
||
|
||
**FR-INF-1 — Runtime selection.** On launch the application shall determine, per device and
|
||
without blocking the first frame, the fastest inference backend that can build and run a session
|
||
for the shipped models, by attempting it; shall record and reuse that determination until the
|
||
runtime, driver, hardware or models change; and shall display the backend in use in Settings and
|
||
on the about screen. *Acceptance:* §10 M1 and M5.
|
||
|
||
**FR-INF-2 — Derived engines.** Backends that require device-specific compilation shall compile in
|
||
the background after selection, shall serve requests from the next lower backend until each engine
|
||
is ready, and shall not change the backend of a job in progress. *Acceptance:* M5.
|
||
|
||
**FR-INF-3 — Model forms.** Quantised model forms are produced at release time from real
|
||
calibration data and are shipped only when they meet §10's accuracy gates against the canonical
|
||
form; the application never quantises on the device. *Acceptance:* M2, M7.
|
||
|
||
**NFR-INF-1 — Embedding comparability.** Face embeddings shall be computed at a precision whose
|
||
deviation from the f32 reference is within §7's gate, on every backend, so that embeddings from any
|
||
device are comparable. *Acceptance:* M3.
|
||
|
||
**D13 — updated.** The runtime half is reopened to the extent of §3: the Rust build stays C-free
|
||
under `alternative-backend`; packages may install a dynamically loaded ONNX Runtime and, per §3.1,
|
||
the Qualcomm QNN runtime; the NVIDIA libraries are not bundled. The licensing half is unchanged.
|
||
Since 0.23 every desktop package bundles two runtimes — Intel's OpenVINO build and the WebGPU
|
||
build — and the APK a second, generic one; the engine loads the one that fits the GPU (§3.2).
|
||
|
||
---
|
||
|
||
## 13. Requirements touched
|
||
|
||
FR-CULL-8 (indexing time is what this exists to change), FR-CULL-9 (NFR-INF-1 is the guard on its
|
||
calibration), NFR-RES-2 (TensorRT workspace against the develop view's tiles), NFR-RES-3
|
||
(background compilation on Android), NFR-SEC-5 (the about screen's line), NFR-COMPAT-2 (each
|
||
channel gains a runtime file), FR-PLAT-AND-1 (unchanged; the runtime lives in the APK, not in
|
||
storage), ARCH §6.1 (the once-per-image budget this was sized against no longer binds; what could
|
||
run per frame is a separate question this document does not open).
|