Files
dtourolle b3999bbc0c Record the Intel and generic rungs in the inference spec
§1.6 is the Iris Xe measurement: OpenVINO fp16 1.3–5.8× the CPU
provider on every shipped model, WebGPU behind it everywhere but the
denoiser and MI-GAN. §2's ladder gains the Intel and generic rows, the
generic one footnoted as unmeasured where it is meant to help. §3.1
lists the two bundled runtimes' licences; §3.2 is how one runtime of
several is chosen per process. D13 notes the bundling.
2026-10-04 21:00:15 -04:00

48 KiB
Raw Permalink Blame History

Inference backends — the runtime and the model, chosen per device

Spec for S16, the build that puts the neural models on the hardware each device actually has.

Every model DarkRoom runs today — the three SCRFD detectors, the ArcFace embedder, YOLO26n-seg and the ADE20K scene model — runs through tract, on one CPU core, on every platform. That was the right first answer: D13's runtime half chose it because it costs no C dependency, and faces.md and segmentation.md were written against it. It is also between 20× and 300× slower than what the same devices can do, and this document is the record of having measured that and the specification of what replaces it.

It does not reopen D13's licensing half. The weights are the same files under the same grant. It does reopen the runtime half, and §3 is where it says how far.


1. What was measured · 2026-09-19

One benchmark, two builds of it: ort's API over tract (exactly what the app links) and ort's API over a dynamically loaded ONNX Runtime with each execution provider in turn. Random 640×640 input, three warm-ups, the median of 15–30 timed runs, milliseconds. The same input every run, so the numbers are compute cost and nothing else.

1.1 The tablet — Honor MagicPad 2, Snapdragon 8s Gen 3

SM8635: 1× Cortex-X4, 4× A720, 3× A520, Adreno 735, Hexagon V73. Android 16.

Model tract (today) ORT CPU f32 ORT CPU int8 Adreno f32 ¹ Hexagon int8 ²
scrfd_500m (Fast) 98 16 8 21 1.4
scrfd_2.5g (Balanced) 161 59 19 ✗ 1.8
scrfd_10g (Thorough) 489 204 48 ✗ 3.2
arcface_mbf (per face) 39 9 13 24 12
yolo26n-seg 287 94 38 54 5.1
yolo26s-sem-ade20k 408 154 46 47 3.7

¹ Qualcomm's own GPU backend (libQnnGpu.so, OpenCL). Fails on the two larger SCRFD graphs at an AveragePool the layout transformer cannot place. ONNX Runtime's WebGPU provider also runs on this GPU and was slower than the CPU on every model; it is not in the table because it is not a candidate. ² QNN's HTP backend. The Hexagon refuses float32 and float16 tensors in this ORT 1.29 + QNN 2.42 pairing (error 3110 on every node, with enable_htp_fp16_precision set or not); int8 QDQ graphs run with 99.6% of nodes on the NPU — 1718 of 1725 for SCRFD-500m, the remainder being the quantise/dequantise at the graph's edges — verified from the partition log, not inferred from the timing.

Also tried and rejected: NNAPI — the device registers no neural-networks HAL at all, so the provider has nothing to talk to; Google deprecated it in Android 15 and Qualcomm stopped shipping drivers for it. XNNPACK — slower than ORT's default CPU kernels on every model that loaded, and aborts inside its partitioner on the SCRFD graphs.

1.2 The desktop — RTX 3050 Laptop, Raptor Lake, 20 threads

Model tract (today) ORT CPU f32 ORT CPU int8 CUDA f32 CUDA fp16 TensorRT f32 TensorRT fp16 TensorRT int8
scrfd_500m 104 12 8 5.6 3.7 2.5 1.8 ✗ ⁴
scrfd_2.5g 162 27 12 6.2 5.3 2.9 1.9 ✗ ⁴
scrfd_10g 514 99 32 14.5 9.3 7.6 3.3 ✗ ⁴
arcface_mbf 45 15 18 1.1 0.8 0.9 0.7 ✗ ⁴
yolo26n-seg 307 60 38 9.3 ✗ ³ 6.8 5.5 ✗ ⁴
yolo26s-sem-ade20k 395 61 31 10.7 ✗ ³ 7.9 3.4 ✗ ⁴

³ The offline fp16 conversion (onnxconverter-common) left a mixed-type node the CUDA provider rejects. TensorRT converts to fp16 itself at engine build and does not have this problem, which is one reason it is the target and the CUDA provider is the fallback. ⁴ TensorRT refuses the QDQ form ONNX Runtime's quantiser writes for the Hexagon (uint8 activations); it wants symmetric int8. Not pursued: fp16 needs no quantisation, no calibration and no accuracy gate, and it is already 30–60× tract.

CUDA int8 is deliberately absent: the CUDA provider has no int8 kernels and runs a QDQ graph by dequantising it, which measured slower than f32 (7.0 vs 5.6 ms on scrfd_500m). Int8 on NVIDIA is TensorRT's job.

TensorRT's first load is 16–116 s per model in f32 and 35–290 s in fp16 (yolo26n-seg the worst: nearly five minutes), because it is compiling an engine for this exact GPU. The engine caches to disk and the second load is milliseconds. That number is what §6 is designed around.

1.3 The desktop — Radeon RX 7900 XT, Threadripper 2920X, 24 threads · 2026-09-20

Arch's onnxruntime-rocm 1.29.0 against ROCm 7.2.4 and MIGraphX 7.2.3 (gfx1100), through the same ep_probe harness (core/dr-inference-engine/examples/ep_probe.rs). Zero input, three warm-ups, the median of 15 runs, on a machine doing nothing else. The build columns are the wall clock of Session construction: cold is a MIGraphX compile of the graph for this GPU, cached is the same session loading the program the cold build wrote.

Model ORT CPU f32 MIGraphX f32 MIGraphX fp16 Compile f32 / fp16 (s) Cached load (s)
scrfd_500m (Fast) 10.4 2.8 2.4 40 / 58 0.3
scrfd_2.5g (Balanced) 20.7 3.3 2.8 37 / 40 0.3
scrfd_10g (Thorough) 57.9 4.5 3.4 40 / 48 0.4
arcface_mbf (per face) 12.8 1.8 1.6 15 / 21 0.4
yolo26n-seg 49.3 8.4 7.5 110 / 136 0.9
yolo26s-sem-ade20k 55.8 4.8 3.8 50 / 60 0.5
2d106det (landmarks) 9.9 1.2 1.0 16 / 21 0.2
ocec_s (eye state) 2.5 0.5 0.4 17 / 17 0.1
xfeat-1024 26.8 10.5 9.9 37 / 53 0.3
migan-512 (per tile) 514 12.7 8.3 102 / 132 0.8

Three things the table settles.

  • The ROCm execution provider does not exist any more. It was ONNX Runtime's CUDA-provider twin for AMD, removed in the 1.23 release (AMD's builds dropped it from ROCm 7.1); 1.29's ROCm build ships libonnxruntime_providers_migraphx.so and nothing else for AMD, and asking for ROCm answers "not enabled in this build". So there is no non-compiling AMD rung to sit under MIGraphX the way the CUDA provider sits under TensorRT: the AMD ladder is MIGraphX, then the CPU.
  • MIGraphX is a compiling provider, and its cache has to be asked for by name. 15–135 s per graph cold, under a second from its cache — TensorRT's shape exactly, and §6's design covers it. Two things the provider does that the code has to know: ort's builder fills the legacy options struct, which 1.29 reads for the precision flags only, so the cache directory (migraphx_model_cache_dir) reaches it only through the generic key/value registration; and the cache key is the graph, the GPU and the MIGraphX version without the precision, so an fp16 session pointed at the f32 program's directory silently loads the f32 program (the first fp16 row measured here was that, before the directories were split).
  • fp16 is worth 10–35% over f32 on this card, not the 3× it is worth on TensorRT, because MIGraphX f32 is already 3–8× the CPU provider and the small graphs are launch-bound. The detectors at 2.4–3.4 ms sit beside TensorRT fp16's 1.8–3.3 ms on the RTX 3050; the embedder is 1.6–1.8 ms on either precision and stays f32 (§7). The whole face pipeline for one image (detector + landmarks + eyes + embedder) is under 6 ms.

1.4 What the numbers say

  • tract is single-threaded. The tablet's one X4 core and one Raptor Lake core give the same tract numbers. Replacing it with ONNX Runtime's CPU provider, no accelerator involved, is 3–6× on the tablet and 8–10× on the desktop. That is the floor, and it is available on every platform the app builds for.
  • The Hexagon is the standout. A 6 W NPU running int8 beats a discrete RTX 3050 running f32 on five of six models. A whole-library face index on the tablet goes from ~100 ms + 39 ms per face to ~1.4 ms + 12 ms per face, and the "Thorough" detector — 3× the cost of "Fast" today — becomes free. Its price is that the models must be quantised to int8, which is an accuracy question §5 has to answer before it is believed. (§1.5 answered it: int8 lost faces, and every model but XFeat ships with 16-bit activations, at about three times these timings.)
  • The embedder does not gain from either accelerator. 112×112 input, per-op overhead dominates; it is 9 ms on the tablet's CPU and 12 ms on its NPU. It stays float, which §7 turns from a performance footnote into a correctness rule.
  • On NVIDIA, TensorRT fp16 ≈ 3× the CUDA provider, and the CUDA provider ≈ 2× the multi-threaded CPU; at fp16 the detectors are 1.8–3.3 ms with no quantisation at all. Both leave the twenty cores free for decoding during a batch index, which the table does not show and which matters more than the ratio.
  • On AMD, MIGraphX fp16 is 4–17× the CPU provider on the detectors and 60× on the inpainter, with the same first-run compile cost as TensorRT and no rung between it and the CPU.

1.5 The Hexagon at every bit width · 2026-10-04

§1.1's Hexagon column is int8 calibrated on noise: timing only. This is the follow-up — every model, every bit width the HTP offers, calibrated on real photographs and scored on the tablet itself (ORT 1.29 + QNN 2.42, htp_arch 73), against the f32 model on the same inputs. The "Form shipped" column is the files in models/, re-scored on the tablet after tools/quantise-models.sh wrote them. The calibration and scoring photographs are 800 from the public COCO val2017 set (CC-BY); the face models' numbers are over the faces in them of at least 32 px. The tools are tools/quantise-models.sh and the scratch harness described with it.

What the HTP accepts. fp16: nothing — every fp16 operator fails validation (3110), on QNN 2.42 and 2.50, with htp_arch and every soc_model tried; the fp16 rung stays off the table until someone has Qualcomm's own SDK to say why. 4-bit weights (A8W4, A16W4): load, and wreck accuracy (SCRFD finds 25–35% of f32's faces). What is left: A8W8 (int8), A16W8 and A16W16, all running the whole graph. 16-bit activations cost about 3× int8's time, A16W16 about 4×.

Model ORT CPU f32 Form shipped Hexagon On the tablet, against f32 int8 for comparison
scrfd_500m / 2.5g / 10g 17 / 56 / 198 ms A16W8 4.2 / 5.1 / 9.0 ms 100% of faces found in every size band; keypoints 0.3–0.6% of the box 94–95% of faces at 40–80 px
2d106det (landmarks) 2.8 ms A16W8 0.5 ms 0.25 px in the 192 crop (eye points 0.20) 1.5 px, and 29 partitions at 7.3 ms
yolo26n-seg 90 ms A16W16, tail in float 12.9 ms 98.2% of objects, mask IoU 0.994 74% (simulated)
yolo26s-sem-ade20k 151 ms A16W16, attention in float 15 ms 98.9% of cells agree on the class, TV 0.009 67%
migan-512 488 ms A16W16 87 ms 41 dB from f32 in the fill (worst 1%: 30 dB) 16 dB (simulated)
xfeat-1024 / 768 58 ms int8, rewritten graph 6.5 ms panorama alignment 0.45 px from f32's — f32's own refit on 90% of its matches is 0.41 —
mosaic-1408 (denoiser) 1510 ms a tile A16W16, rewritten graph 95 ms a tile 0.00 dB at every ISO; ≤ 0.09 dB with the noise scaled ×0.5–×4 −4.7 to −9.2 dB
arcface_mbf (embedder) 8.5 ms f32, CPU — A16W16: cosine 0.9995, p1 0.9967 — misses §7's 0.999 gate —
ocec, sgc (eyes) 1, 1.7 ms f32, CPU — sgc flips 1.45% of views even at A16W16; not worth a millisecond —

Four things the table needed that the f32 graphs did not have, all in tools/htp_graph.py and all checked exact against the f32 graph before they are used:

  • Rank ≤ 5. QNN's tensors stop at rank 5, and the denoiser packs the mosaic through a 6-D reshape (6007 at compose). For one channel that reshape is SpaceToDepth(2). XFeat's 8×8 unfold is 224 Slices and 6-D Concats; it is SpaceToDepth(8) (736 nodes to 60).
  • No bilinear Resize at XFeat's sizes (3110). A half-pixel bilinear resize between fixed sizes is two constant matrices, so it is two MatMuls.
  • One scale per tensor. The segmenter's output rows carry boxes in pixels beside scores in 0..1; quantised as one tensor the scores vanish. Everything from the Concat that builds the rows stays float, on the CPU, where the top-300 selection costs nothing.
  • Float where the HTP's 16-bit arithmetic drifts. The scene model's one attention block (two MatMuls and a Softmax over 400 tokens) moved its agreement from 98.7% to 96.7%; it stays float.

ORT's CPU simulation of a QDQ graph is not the tablet. It matched to the hundredth of a dB for the denoiser and to rounding for the detectors, landmarks and XFeat, and it overstated MI-GAN by 27 dB and the scene model by three points. Every number above is the device's; a new form is not measured until it has run there.

XFeat's int8 loses keypoints and not the panorama. 83% of f32's keypoints come back within 1.5 px; but over the twelve-frame fixtures/pano/2025-08-05 sweep, the homographies fitted from int8's matches land 0.45 px from f32's in the overlaps — the spread f32 shows against itself (0.41).


1.6 Intel Iris Xe, and the generic rung · 2026-10-04

The RTX 3050 laptop's other GPU: Raptor Lake-P's Iris Xe (96 EU), Intel's onnxruntime-openvino 1.24.1 (OpenVINO 2025.4.1) and Microsoft's onnxruntime-webgpu 1.27.0, both PyPI wheels, through ep_probe. Three warm-ups, the median of 15 runs. Another build shared the CPU during the run, so the CPU columns are a little pessimistic; the GPU columns are not.

Model ORT CPU f32 OpenVINO CPU OpenVINO GPU f32 OpenVINO GPU fp16 WebGPU (Iris Xe)
scrfd_500m (Fast) 9.5 10.8 7.3 5.8 24.2
scrfd_2.5g (Balanced) 18.5 16.2 15.5 11.1 40.8
scrfd_10g (Thorough) 58.5 72.4 38.7 23.0 84.5
arcface_mbf (per face) 9.5 11.7 3.0 2.4 56.6
2d106det (landmarks) 10.8 2.0 2.4 1.9 42.7
yolo26s-sem-ade20k 57.0 47.4 26.0 16.7 53.0
xfeat-1024 23.5 18.5 19.9 17.5 34.2
migan-512 (per tile) 330 ✗ ¹ 89.7 57.2 275
mosaic-fast-1408 (per tile) 159 107 68.1 40.0 188 ²
mosaic-best-1408 (per tile) 1109 1670 947 604 1034 ²

¹ OpenVINO's CPU plugin refuses the graph at initialisation. Not shipped (§3.2), so moot. ² A later run, after ep_probe learned to feed the denoiser's two inputs, under heavier load: the CPU provider took 256 and 1034 ms in that run, so WebGPU beat it by a quarter on mosaic-fast and tied on mosaic-best — the only rows where it is not well behind.

  • OpenVINO on the Iris Xe beats ONNX Runtime's CPU provider on every model, 1.3× on XFeat to 5.8× on MI-GAN, with a 1–3 s compile per graph and 0.1–0.4 s from its cache. It is the Intel rung. fp16 is worth 1.3–1.7× over f32 here, against 1.1–1.35× on MIGraphX.
  • Its "GPU" is OpenCL's first GPU, not Intel's. Before intel-compute-runtime was installed the only OpenCL driver was NVIDIA's, and device_type=GPU ran on the RTX 3050 — slower than the CPU, which is the probe's to catch. Read the process's maps for libigdrcl before believing a number is the iGPU's.
  • WebGPU is slower than the CPU on the Iris Xe on everything but MI-GAN, as it was on the Adreno (§1.1), and on the RTX 3050 through Vulkan too. It is on the ladder anyway, as the generic rung (§2): for GPUs no vendor rung covers — an AMD card on Windows or without ROCm, a Mali — where it is unmeasured, and the probe's clock decides.
  • OpenVINO's CPU plugin is not a better floor. It wins on some graphs and loses on scrfd_10g, the embedder and mosaic-best, and refuses MI-GAN.

2. The shape of the answer

A ladder per platform, walked at start-up, with the first rung that builds a real session winning:

Platform 1st 2nd 3rd Floor
Android, Qualcomm with a Hexagon the shipped QNN skel covers (V68–V81) QNN HTP, each model's quantised form (§1.5) ORT CPU, f32 model — tract
Android, any other SoC ⁶ WebGPU (Vulkan), f32 ORT CPU, f32 — tract
Linux / Windows, NVIDIA GPU TensorRT, f32 model, fp16 engine CUDA provider, f32 ORT CPU, f32 tract
Linux, AMD GPU with ROCm MIGraphX, f32 model, fp16 program ORT CPU, f32 — tract
Linux / Windows, Intel GPU OpenVINO, f32 model, fp16 program (§1.6) ORT CPU, f32 — tract
Linux / Windows, any other GPU ⁶ WebGPU (Vulkan / D3D12), f32 ORT CPU, f32 — tract
macOS ⁵ CoreML, f32 model, ML Program ORT CPU, f32 — tract

⁵ Unmeasured, and the one exception to the rule below: nobody here has a Mac. The rung is on the ladder because the probe makes a wrong guess cheap — a CoreML that is slower than the CPU is rejected by §4's clock, one that errors is recorded as failed, and one that takes the process down is refused on the third launch (§4, attempt). The embedder stays on the CPU (§7). The first macOS log that shows a probe line is this row's measurement; macos.md says what to ask for.

⁶ The generic rung, unmeasured where it is meant to help. WebGPU lost to the CPU on every GPU it has been timed on — the Adreno, the Iris Xe, the RTX 3050 (§1.1, §1.6) — none of which it serves here, since each has its own rung. It is on the ladder for the GPUs that have none, on the same terms as CoreML: a WebGPU that is slower than the CPU is rejected by §4's clock, one that errors is recorded as failed. Its first measurement on an AMD card without ROCm, or a Mali, is this row's.

Deliberately not on any ladder, with the measurement that excluded each: NNAPI (no driver), XNNPACK (slower than CPU, aborts on SCRFD), the Adreno through QNN (works, but never where the Hexagon does not also), CUDA int8 (slower than CUDA f32), the ROCm provider (gone: §1.3), OpenVINO's CPU plugin as a floor (§1.6). A rung is added to this table by a measurement on this page, not by a provider existing — the two footnoted rows are the exceptions, and say so.

The AMD ladder has no middle rung. TensorRT falls back to the CUDA provider while its engines compile; MIGraphX has no such twin, so its fallback is the CPU provider, and the minute or two of compiling on first run (§6) is spent at the floor's speed rather than at half the GPU's.

Two things the ladder is not: it is not a per-model choice — one backend serves every model on a device, because §7's identity rule needs the detector and embedder on the same runtime for the same reason faces.model_id pairs them; and it is not a per-account choice — it is a property of the hardware, like shared_face_models_dir is, and it lives beside it.


3. The dependency policy, and how far this reopens it

D13 chose ort over tract because alternative-backend made ONNX Runtime's API available with none of its C. Every rung above the floor needs the C++ ONNX Runtime and, for the two that matter most, vendor libraries on top: Qualcomm's QNN runtime (~60 MB for one Hexagon generation; it is per-SoC) and NVIDIA's TensorRT plus cuDNN (~600 MB with the CUDA libraries, and cuDNN's major version must match what ONNX Runtime was built against — the Arch package on the reference desktop was unusable for exactly that reason).

The policy protected the build: no C to cross-compile under the NDK, no toolchain to keep in step. This document keeps that intact, and the mechanism is the one thing about ort that makes it possible:

ort::set_api accepts any OrtApi table. With alternative-backend on, ort links nothing and asks for the table once per process. The application can dlopen a libonnxruntime.so it finds on disk, call OrtGetApiBase()->GetApi(version) and hand that table over; or, if there is no such file, hand over ort_tract::api(). The Rust build is identical in both cases — pure Rust, cargo build --target aarch64-linux-android sees the same dependency graph it sees today. What changes is that the runtime is a file the package installs, next to the models, and the app looks for it at start-up.

Consequences that follow and are accepted:

  • The runtime is chosen once per process, because set_api is once per process. The ladder in §2 is walked at start-up and the result is what every session in that process uses. There is no "tract for this model, ORT for that one", and there is no falling back to tract after ONNX Runtime has loaded — but there does not need to be: once the library loads, its CPU provider is always there, and every fallback the ladder needs is between providers inside it.
  • Feature flags stay as they are. dr-face's inference and dr-segment's semantic continue to mean "compiled against ort's API"; nothing at build time knows or cares which table will be supplied. The one addition is a native-probe feature on the new crate (§8) that pulls in libloading, which is pure Rust and already in the tree via wgpu.
  • The packagers ship the runtime, not the build. The Arch package, the Flatpak manifest, the NSIS installer and assemble-apk.sh each gain the ONNX Runtime library for their platform, and the Android and NVIDIA variants gain the vendor libraries — each under the licence the packager reads first (§3.1). A package without them is not broken; it is the tract build, and it says so on the about screen.
  • The NDK problem does not come back. libonnxruntime.so for Android is a prebuilt from Maven (com.microsoft.onnxruntime:onnxruntime-android-qnn), extracted by assemble-apk.sh into jniLibs/ the way the models are bundled as assets today. Nothing compiles it.

3.1 Licences the packagers read before shipping a runtime

Written down now, because segmentation.md §7 established that reading the grant is cheaper than discovering it at packaging time.

Component Licence Redistributable in a self-distributed package?
ONNX Runtime MIT Yes
Qualcomm QNN runtime (com.qualcomm.qti:qnn-runtime on Maven) Qualcomm AI Engine Direct SDK licence — proprietary, redistribution permitted for applications using it Yes for the APK, with the licence text shipped; not for a source distribution. To be read in full, not summarised from memory, before the APK gains it.
CUDA runtime, cuDNN, TensorRT NVIDIA EULAs — redistributable with an application, with the licence text, not modifiable Yes for a package that bundles them. 600 MB. The alternative is to load them from the user's system install if present and skip the rung otherwise — which is what §4's probe does anyway.
ROCm (HIP, MIOpen, rocBLAS), MIGraphX MIT Yes, but the HIP SDK MIGraphX needs is ~15 GB installed. Same answer as NVIDIA: the user's system install, or the rung is skipped.

The position this takes: the GPU vendors' libraries are not bundled. The desktop package probes for a system CUDA/TensorRT or ROCm/MIGraphX install and uses it if it is version-compatible; a desktop without one runs on ORT CPU, which is still 8–10× today. Bundling 600 MB for a rung that is 2× again is not a trade worth making unmeasured, and it can be revisited by a measurement on a batch index. The QNN runtime is bundled in the APK, because the Hexagon is the difference between a tablet that indexes a library overnight and one that does it over lunch, and the package is 60 MB larger for it.

Both positions are D13 territory and are recorded there (§12).

Two more runtimes ship in every desktop package since 0.23 (§3.2), and their parts are all redistributable:

Component Licence Shipped
Intel's onnxruntime-openvino build, with OpenVINO 2025.4.1 and oneTBB MIT; Apache-2.0; Apache-2.0 Linux and Windows packages, texts beside the libraries
Microsoft's WebGPU build (Dawn inside); on Windows the DirectX shader compiler MIT; LLVM / MIT Linux and Windows packages
Microsoft's stock onnxruntime-android (WebGPU) MIT The APK, as libonnxruntime_generic.so

3.2 Several runtimes, one per process

A runtime carries one vendor's providers: Intel's build has OpenVINO, the onnxruntime-gpu wheel CUDA and TensorRT, a ROCm build MIGraphX, Microsoft's WebGPU build the generic rung, the APK's QNN build the Hexagon. No prebuilt carries two vendors, and set_api takes one table per process. So a device that may hold several — the package's OpenVINO and WebGPU builds, a CUDA build the user fetched, the distribution's ROCm build — has to choose which to load before the probe, and cannot choose by trying.

api::install opens every runtime on the search list, asks each for GetAvailableProviders, and loads the one that scores highest against the GPUs hardware::detect reads from files: a vendor rung on its own vendor's GPU (NVIDIA driver, /dev/kfd, a Qualcomm SoC, macOS) above OpenVINO on an Intel GPU (PCI vendor 0x8086; on Windows Intel's DCH driver package) above WebGPU above a CPU-only build. Equal scores keep the search order, a perfect fit ends the search — the APK's QNN build is listed first, so on a Qualcomm device the generic build is never opened — and DARKROOM_ORT_DIR wins outright. The losers stay mapped: unloading a C++ runtime whose static constructors ran is a crash at exit waiting to happen.

The desktop packages install the two bundled builds under runtimes/openvino and runtimes/webgpu beside each place a package installs to, from tools/fetch-bundled-runtimes.sh (PyPI wheels pinned by SHA-256, pruned to the native libraries: 81 + 31 MB on Linux, 67 + 42 MB on Windows). On Windows the chosen runtime's directory is put on PATH, because Intel's build leaves OpenVINO's DLLs for the loader to find there. The Flatpak has no Intel OpenCL driver in its sandbox, so an Intel machine there settles on the CPU.


4. Selection — the probe, its cache, and what it may not do

A rung is chosen by building a real session on it, not by asking whether it exists. Both failure modes that are not "the provider is absent" were hit on 2026-09-19: a driver in a wedged state where the provider registered and the session then failed, and a provider that registered, took the graph, and rejected every node at partition time. The probe therefore:

  1. Loads the runtime library (§3), or falls to tract and stops.
  2. Times the smallest detector on the CPU provider first — the floor. Then, for each rung in this platform's ladder, in order: builds a session for the same model on that provider with error_on_failure, runs it once on a fixed input, and times three more runs. The rung is taken only if its median beats the floor. That one measurement is the proof the provider took the graph: one that silently hands the work to the CPU is the CPU rung with extra overhead, slower than the floor, and rejected. (ONNX Runtime's session.disable_cpu_ep_fallback was the first draft of this proof and refuses the Hexagon over the ten quantise/dequantise nodes at the graph's edges that QNN declines by policy.)
  3. Records the outcome — rung, runtime version, provider version, device identity (GPU name and compute capability; SoC model and Hexagon arch), and the models' content hashes — to a small file beside shared_face_models_dir. The next start-up trusts the file unless any of those inputs changed, in which case it probes again. A driver update, a runtime update, a new model file: each invalidates the cache by construction, and none needs a "reset backend" button.

What the probe may not do:

  • Block the first frame. It runs on the same background as install_bundled_models and for the same reason: a TensorRT probe can take thirty seconds cold, and a tablet that stalls that long is an ANR. Until it reports, every model request is answered by the floor the runtime supports (ORT CPU if the library loaded, tract otherwise), and a job that started on the floor finishes on it — a backend does not change under a running index.
  • Retry a rung that failed within a session. A failed probe is cached as a failure with the same inputs; the rung is tried again when an input changes. Otherwise a wedged driver means a thirty-second stall on every launch.
  • Crash the app twice for the same reason. The probe runs in the app's process, and a provider can fail by aborting rather than by returning an error (XNNPACK on SCRFD, §2). Every session build on a rung above the CPU — the probe's, and each background compile of §6 — writes what it is attempting to attempt in the cache directory first and removes it after. A launch that finds the file knows the last one died inside that attempt; after two such launches in a row the attempt is refused and recorded like any other failure (a rung in failed, an engine in refused), until the fingerprint changes. Two, not one, because quitting during a forty-second TensorRT compile leaves the same file.
  • Choose for the user without saying so. Settings gains one row, Inference backend, showing what was chosen and why in one line ("Hexagon NPU · int8 · QNN 2.42"; "CPU · ONNX Runtime 1.30 · TensorRT probe failed: cuDNN 8 required"), with an override to force any lower rung. The about screen carries the same line beside the model names NFR-SEC-5 already puts there.

5. Model variants, and who makes them

Every model exists in one canonical form — the f32 ONNX file the app ships or the user supplies today — and, where a rung needs it, a derived form. The ladder's rungs are specified in terms of which form they load:

Form Who produces it When Needed by
f32 ONNX, shape-fixed, opset ≥ 13 tools/fix-face-model-shapes.sh, tools/export-seg-model.sh Release time, once Every rung except Hexagon
QDQ ONNX, per-channel — int8, A16W8 or A16W16 per model (§1.5) tools/quantise-models.sh Release time, once, calibrated on real photographs, scored on the tablet Hexagon
TensorRT engine (.engine, per GPU architecture and TensorRT version) The app, from the f32 file First run on that device, in the background TensorRT rung
MIGraphX program (.mxr, per GPU architecture, MIGraphX version and precision) The app, from the f32 file First run on that device, in the background MIGraphX rung
QNN context binary The app, from the quantised file First run on that device, in the background Hexagon rung

Two rules.

Quantisation is a release-time step, not a device-time one. The int8 files that produced §1's numbers were calibrated on random noise, which is enough to time and worthless to trust. A real int8 detector is calibrated on a few hundred real photographs and then measured against the f32 detector on the reference library by faces.md §12.3's method — faces found, per size band, per detector — before it ships. That needs the reference library and a person reading the result, and it happens once per model release, in tools/, beside the shape-fixing it already depends on. The device never quantises anything.

The SCRFD and ArcFace files are opset 11 as InsightFace exported them, and per-channel QDQ needs 13; tools/fix-face-model-shapes.sh gains an opset upgrade to 17 (onnx.version_converter, ir_version 8), which tract has been verified to load and which every provider on this page prefers. That is a change to the canonical file and so a change to the shipped models, and it happens in the same model release as the int8 files.

Compilation is a device-time step, and it is cached. A TensorRT engine is specific to the GPU it was built on and the TensorRT that built it; a MIGraphX program to the GPU and the MIGraphX that built it; a QNN context binary to the Hexagon generation. None can ship. All are built by the app the first time that rung is selected, in the background (§6), and written beside the probe cache keyed by the same inputs. MIGraphX's own key leaves out the precision, so the app gives its f32 and fp16 programs separate directories — otherwise the embedder's f32 build would be served the detector's fp16 program, or the reverse. They are derived, disposable, and regenerable: deleting the cache directory costs the next launch a rebuild and nothing else, and the directory is excluded from anything that syncs (it is a peer of thumbs, not of the catalog).


6. First run — building engines without the user waiting for them

The sequence on a device where a compiling rung (TensorRT, MIGraphX, Hexagon) is selected:

  1. Launch. The runtime loads; the probe (§4) starts in the background; the app serves every model request from the floor. Face indexing, segmentation and scene grading all work, at today's speed or better (ORT CPU).
  2. Probe reports — say, TensorRT. The compiling rung is now selected but has no engines. Model requests continue on the fallback rung below it (CUDA provider for TensorRT; ORT CPU for MIGraphX and Hexagon), which needs no compilation and is already faster than the floor.
  3. Engines build, one model at a time, on a single low-priority background thread, smallest model first so the detector — the one that runs per image — is ready soonest. On the reference desktop that is ~1 minute for the first detector and ~10 minutes for all six at fp16; on the AMD desktop 40 s for the first detector and ~8 minutes for the set; on the tablet the QNN context binaries take 0.8–1.7 s each and the whole set is ready before the user has opened a library. Each engine is written to a temporary name and renamed into place, so a request never sees a half-written file.
  4. Requests move up as engines land. A model whose engine exists loads it on the selected rung; one whose engine is still building loads on the fallback. A running job does not switch — an index that started on the CUDA provider finishes on it — because §7 needs one model_id per job, and because a job is the wrong granularity for surprise.
  5. On Android, the build runs only while the app is in the foreground and the device is not in battery saver (NFR-RES-3): a context binary takes a second, so this costs nothing, and the rule exists for the day a model takes longer.

Settings shows a one-line progress row while engines build ("Preparing GPU engines · 3 of 6") and nothing when they are done. A build failure demotes the rung — it is recorded in the probe cache as a failure with the model hash as an input, so a corrected model file retries it — and the app carries on one rung down, saying so in the same row.


7. Identity — what changes model_id and what may not

faces.model_id exists so that two libraries indexed with different networks are never compared as if they were one (catalog.md §10.1; the trap is written up in faces.md §14). A backend that changes what a network computes is a different network and must be a different model_id; one that changes only where it computes it must not be.

The detector. An int8 SCRFD finds a different set of faces from the f32 one — that is what §5's acceptance measures — so the numeric form is part of the detector's identity: scrfd_500m and scrfd_500m_i8 are two detectors in model_id, and a library indexed on the tablet's Hexagon and continued on the desktop is two populations, which a re-index on either side reconciles the same way a switch from Fast to Thorough does today. That is acceptable because it is already the rule for the detector and because §5 is the gate on whether the int8 form is close enough to be offered at all. f32 on tract, ORT CPU, CUDA and TensorRT-f32 are one identity: the same graph, the same arithmetic, differences at the last bit.

After §1.5 the Hexagon runs the detectors in A16W8, and that is a third spelling: scrfd_500m_a16+w600k_mbf and its two siblings. Same rule, same reconciliation; a tablet that indexed under _i8 keeps those rows, and FaceDetector::model_ids answers "has this detector been over this image" for all three forms.

The embedder is where comparability across devices is the whole point, and it is the one model that no accelerator helps (§1.4). So: the embedder runs in f32 on every rung. On TensorRT that means the embedder's engine is built without fp16 while the detector's is built with it; on the Hexagon it means the embedder is not on the NPU at all — it runs on the ORT CPU rung at 9 ms, and the ladder's "one backend per device" is, precisely, one backend per model role, with the embedder pinned. A w600k_mbf embedding from any device is comparable with one from any other, which is the property the identity system, the calibration and the cross-device merge all rest on, and it is not for sale for 3 ms.

If S16 wants fp16 for the embedder later, the gate is written now: over the reference library's faces, the cosine between the f32 and fp16 embedding of the same crop exceeds 0.999 for 99.9% of faces and the calibration's fitted threshold moves by less than its own confidence interval (faces.md §8.3). Until measured, f32.

Segmentation and the scene model carry no identity across devices — their outputs are recomputed per image and never stored beyond the cache — so they take whatever the rung offers, int8 included, subject to §10's own acceptance.


8. Crate shape — core/dr-inference-engine

The seam is the same shape as storage.md's: a small crate below the consumers that is the only place naming a provider, a library file or a vendor, with the consumers reduced to "give me a session for these bytes in this role".

core/dr-inference-engine
  src/lib.rs        Runtime (Tract | Onnx { lib, version }), Backend (rung), Role (Detector | Embedder | Segmenter)
  src/probe.rs      §4 — the ladder per platform, the session-build probe, the cache file
  src/engines.rs    §6 — background compilation, the cache directory, progress
  src/session.rs    open(role, bytes) -> ort::Session, applying the rung and the role's precision rule
  src/api.rs        the unsafe block that matters: dlopen libonnxruntime, fetch OrtApi, ort::set_api — or ort_tract::api()
  • dr-face and dr-segment delete their private install_backend and their direct Session::builder() calls and take an &dr_inference_engine::Sessions where they take model bytes today. Their tests keep tract — dr_inference_engine::Sessions::tract() is a constructor and the test-only path.
  • dr-inference-engine depends on ort with the same workspace features as today plus cuda, tensorrt, qnn: those features add option builders, not linking, under alternative-backend. Verified for the QNN, CUDA and TensorRT builders on 2026-09-19 — they go through the API table's generic SessionOptionsAppendExecutionProvider*. The NNAPI builder resolves a symbol directly and would not; it is not needed and is not enabled. MIGraphX uses no ort feature at all: ort's builder fills the legacy options struct, which ONNX Runtime 1.29 reads for the precision flags and nothing else, and the compiled-program cache directory only travels through the generic key/value entry point (migraphx_model_cache_dir). session::migraphx makes that one call on the API table itself.
  • dr-ui owns the settings row, the about-screen line and the progress row; it holds one Sessions per process, created at launch, and passes it down. dr_ui::library gains inference_cache_dir() beside shared_face_models_dir(), on the same account-independent footing and for the same reason.
  • The Android entry point's install_bundled_models also extracts nothing new: jniLibs/ is loaded by the system loader, and dr-inference-engine on Android looks for libonnxruntime.so through dlopen by bare name first, which resolves to the APK's copy, before any directory.

9. Threads and memory

  • ONNX Runtime's intra-op pool is sized to the physical cores minus two on desktop and to the performance cores on Android (the X4 and the A720s; the A520s are for the compositor). tract's single thread today is the reason a batch index leaves nineteen cores idle; ORT CPU with the pool is the reason it will not. One session per model per process; Session::run is &mut self-free in ort and internally serialised, and the index job is the only caller.
  • A TensorRT session pins GPU memory for its workspace; the builder is capped at 512 MB on the reference 6 GB card and the cap is a setting, because the develop view's tiles share the card (NFR-RES-2). The engine cache on disk is bounded by the model set — six engines, ~80 MB — and needs no LRU.
  • The Hexagon rung sets QNN's performance mode to Burst for the duration of an index job and Default otherwise; a 5 ms detector does not need the NPU clocked up between images.
  • The probe (§4) and the engine build (§6) run on one dedicated low-priority thread. They never share the index job's pool: a probe that competes with the job it is meant to speed up is the frame-budget trap in a new coat.

10. What S16 measures

In order, with the gate each is:

# Question Gate
M1 Does one binary carry both tables? dlopen + set_api on Linux, Windows and Android; ort_tract::api() when the file is absent. Go / no-go for §3. If set_api cannot take a dynamically fetched table on some platform, that platform ships two binaries, and the cost is stated.
M2 Do the int8 SCRFD detectors, calibrated on real photographs, find the faces? faces.md §12.3's method over the reference library, per size band, against f32. Ship the int8 form for a detector only if it finds ≥ 97% of the f32 detector's faces above 40 px and the difference is not concentrated in one band. Otherwise that detector's Hexagon rung is ORT CPU int8-free, and the table in §1.1 says what that costs.
M3 Does the embedder on ORT CPU beside a detector on the Hexagon (§7) produce embeddings within the f32 gate? It must — same graph, same arithmetic. This is a check that the plumbing did not quantise it by accident.
M4 Is a batch index on the tablet and on the desktop faster by the ratio §1 predicts, end to end, decode included? The face index over the reference library (18,143 faces): report wall-clock on tract, on the floor and on the selected rung, and where the time went. The prediction is that decode becomes the bottleneck on both; if it does not, say why.
M5 Does the first-run sequence (§6) hold: nothing blocks the first frame, engines land, requests move up, a running job does not switch? Observed on both devices with the app's own progress row, and with the cache directory deleted between runs.
M6 What does the APK weigh with the QNN runtime, and does a non-Qualcomm Android device (any one) still launch and index on the floor? Size reported; launch verified on one non-Qualcomm device or an emulator.
M7 The segmentation and scene models at int8 on the Hexagon: does the mask boundary move? segmentation.md §6's IoU against f32 over its corpus. Ship int8 for a model only above the IoU floor that document set for arm B.

M1 and M2 are the ones the rest is conditional on, and M2 is the one that needs a person.

10.1 M2 result · 2026-09-19

Each int8 detector against its own f32 form, over 400 proxies evenly spaced through the reference library, on ONNX Runtime's CPU provider (the int8 graph is the same file the Hexagon loads; ui/dr-ui/examples/face_detectors). Calibrated on 64 proxies from the same library, disjoint from the 400.

Detector f32 faces int8 faces both int8 only f32 only found ≥ 32 px found, all sizes
scrfd_500m 1342 1319 1287 32 55 95.6% 95.9%
scrfd_2.5g 1525 1482 1478 4 47 96.1% 96.9%
scrfd_10g 1769 1760 1749 11 20 100% 98.9%

The 10g form clears the 97% gate; 500m and 2.5g sit one point under it. What they lose is specific: the faces in the "f32 only" column have a median confidence of 0.52 against a threshold of 0.50 — detections the f32 graph itself barely made, that int8 rounding drops to the other side of the line — and the extra faces int8 finds are the same kind (median 0.51–0.52). Not a size-band failure: the losses are spread across bands in proportion. Shipped as they are, with the number on record; a threshold of 0.48 for the int8 forms would recover most of the margin, and is the first thing to try if a library's count on the tablet reads low.

Two things the calibration taught, both in tools/quantise-models.py: the calibration set has to contain faces (a first attempt on landscape photographs produced a graph that found nothing — the score head's ranges had never seen the face regime), and ONNX Runtime's own strided and moving-average calibration modes both measurably degrade the result on these graphs, while driving the calibrator in chunks by hand reproduces the plain min/max ranges exactly.


11. Order

  1. dr-inference-engine with the two tables and the floor — set_api from a dlopened runtime, tract otherwise, ORT CPU as the only rung. Consumers moved over; tests unchanged. This alone is the 3–10× and is the build most of the value sits in. M1.
  2. The probe and its cache (§4), with the settings row and the about line. Still CPU-only; the ladder has one rung. M5's first half.
  3. tools/quantise-models.sh and the opset upgrade; the int8 detectors calibrated and measured. M2, M7. This is the step with a person in it and it runs in parallel with 4.
  4. The Hexagon rung, the QNN runtime in the APK, the context-binary cache. M3, M6.
  5. The TensorRT and CUDA rungs on desktop, the engine cache, the first-run sequence. M5's second half.
  6. M4 last, on both devices, and the number goes in this document.

The Windows installer and the Flatpak manifest are touched in steps 1 and 5 only, and only to add a file each; the Arch package likewise.


12. Register entries

FR-INF-1 — Runtime selection. On launch the application shall determine, per device and without blocking the first frame, the fastest inference backend that can build and run a session for the shipped models, by attempting it; shall record and reuse that determination until the runtime, driver, hardware or models change; and shall display the backend in use in Settings and on the about screen. Acceptance: §10 M1 and M5.

FR-INF-2 — Derived engines. Backends that require device-specific compilation shall compile in the background after selection, shall serve requests from the next lower backend until each engine is ready, and shall not change the backend of a job in progress. Acceptance: M5.

FR-INF-3 — Model forms. Quantised model forms are produced at release time from real calibration data and are shipped only when they meet §10's accuracy gates against the canonical form; the application never quantises on the device. Acceptance: M2, M7.

NFR-INF-1 — Embedding comparability. Face embeddings shall be computed at a precision whose deviation from the f32 reference is within §7's gate, on every backend, so that embeddings from any device are comparable. Acceptance: M3.

D13 — updated. The runtime half is reopened to the extent of §3: the Rust build stays C-free under alternative-backend; packages may install a dynamically loaded ONNX Runtime and, per §3.1, the Qualcomm QNN runtime; the NVIDIA libraries are not bundled. The licensing half is unchanged. Since 0.23 every desktop package bundles two runtimes — Intel's OpenVINO build and the WebGPU build — and the APK a second, generic one; the engine loads the one that fits the GPU (§3.2).


13. Requirements touched

FR-CULL-8 (indexing time is what this exists to change), FR-CULL-9 (NFR-INF-1 is the guard on its calibration), NFR-RES-2 (TensorRT workspace against the develop view's tiles), NFR-RES-3 (background compilation on Android), NFR-SEC-5 (the about screen's line), NFR-COMPAT-2 (each channel gains a runtime file), FR-PLAT-AND-1 (unchanged; the runtime lives in the APK, not in storage), ARCH §6.1 (the once-per-image budget this was sized against no longer binds; what could run per frame is a separate question this document does not open).