The probe runs in the app's process, and a provider can fail by aborting rather than by returning an error — XNNPACK did on SCRFD. A rung that does that once would do it on every launch, before the first photograph is on screen. Every session build above the CPU, the probe's and each background compile's, now writes what it is attempting to `attempt` in the cache directory first and removes it after. After two launches in a row that died inside the same attempt it is refused and recorded — a rung in `failed`, an engine in the new `refused` — until the fingerprint changes. Two, not one, because quitting during a TensorRT compile leaves the same file.
37 KiB
Inference backends — the runtime and the model, chosen per device
Spec for S16, the build that puts the neural models on the hardware each device actually has.
Every model DarkRoom runs today — the three SCRFD detectors, the ArcFace embedder, YOLO26n-seg and
the ADE20K scene model — runs through tract, on one CPU core, on every platform. That was the
right first answer: D13's runtime half chose it because it costs no C dependency, and
faces.md and segmentation.md were written against it. It is also
between 20× and 300× slower than what the same devices can do, and this document is the record of
having measured that and the specification of what replaces it.
It does not reopen D13's licensing half. The weights are the same files under the same grant. It does reopen the runtime half, and §3 is where it says how far.
1. What was measured · 2026-09-19
One benchmark, two builds of it: ort's API over tract (exactly what the app links) and ort's
API over a dynamically loaded ONNX Runtime with each execution provider in turn. Random 640×640
input, three warm-ups, the median of 15–30 timed runs, milliseconds. The same input every run, so
the numbers are compute cost and nothing else.
1.1 The tablet — Honor MagicPad 2, Snapdragon 8s Gen 3
SM8635: 1× Cortex-X4, 4× A720, 3× A520, Adreno 735, Hexagon V73. Android 16.
| Model | tract (today) | ORT CPU f32 | ORT CPU int8 | Adreno f32 ¹ | Hexagon int8 ² |
|---|---|---|---|---|---|
| scrfd_500m (Fast) | 98 | 16 | 8 | 21 | 1.4 |
| scrfd_2.5g (Balanced) | 161 | 59 | 19 | ✗ | 1.8 |
| scrfd_10g (Thorough) | 489 | 204 | 48 | ✗ | 3.2 |
| arcface_mbf (per face) | 39 | 9 | 13 | 24 | 12 |
| yolo26n-seg | 287 | 94 | 38 | 54 | 5.1 |
| yolo26s-sem-ade20k | 408 | 154 | 46 | 47 | 3.7 |
¹ Qualcomm's own GPU backend (libQnnGpu.so, OpenCL). Fails on the two larger SCRFD graphs at an
AveragePool the layout transformer cannot place. ONNX Runtime's WebGPU provider also runs on this
GPU and was slower than the CPU on every model; it is not in the table because it is not a
candidate.
² QNN's HTP backend. The Hexagon refuses float32 and float16 tensors in this ORT 1.29 + QNN
2.42 pairing (error 3110 on every node, with enable_htp_fp16_precision set or not); int8 QDQ
graphs run with 99.6% of nodes on the NPU — 1718 of 1725 for SCRFD-500m, the remainder being the
quantise/dequantise at the graph's edges — verified from the partition log, not inferred from the
timing.
Also tried and rejected: NNAPI — the device registers no neural-networks HAL at all, so the provider has nothing to talk to; Google deprecated it in Android 15 and Qualcomm stopped shipping drivers for it. XNNPACK — slower than ORT's default CPU kernels on every model that loaded, and aborts inside its partitioner on the SCRFD graphs.
1.2 The desktop — RTX 3050 Laptop, Raptor Lake, 20 threads
| Model | tract (today) | ORT CPU f32 | ORT CPU int8 | CUDA f32 | CUDA fp16 | TensorRT f32 | TensorRT fp16 | TensorRT int8 |
|---|---|---|---|---|---|---|---|---|
| scrfd_500m | 104 | 12 | 8 | 5.6 | 3.7 | 2.5 | 1.8 | ✗ ⁴ |
| scrfd_2.5g | 162 | 27 | 12 | 6.2 | 5.3 | 2.9 | 1.9 | ✗ ⁴ |
| scrfd_10g | 514 | 99 | 32 | 14.5 | 9.3 | 7.6 | 3.3 | ✗ ⁴ |
| arcface_mbf | 45 | 15 | 18 | 1.1 | 0.8 | 0.9 | 0.7 | ✗ ⁴ |
| yolo26n-seg | 307 | 60 | 38 | 9.3 | ✗ ³ | 6.8 | 5.5 | ✗ ⁴ |
| yolo26s-sem-ade20k | 395 | 61 | 31 | 10.7 | ✗ ³ | 7.9 | 3.4 | ✗ ⁴ |
³ The offline fp16 conversion (onnxconverter-common) left a mixed-type node the CUDA provider
rejects. TensorRT converts to fp16 itself at engine build and does not have this problem, which is
one reason it is the target and the CUDA provider is the fallback.
⁴ TensorRT refuses the QDQ form ONNX Runtime's quantiser writes for the Hexagon (uint8
activations); it wants symmetric int8. Not pursued: fp16 needs no quantisation, no calibration and
no accuracy gate, and it is already 30–60× tract.
CUDA int8 is deliberately absent: the CUDA provider has no int8 kernels and runs a QDQ graph by dequantising it, which measured slower than f32 (7.0 vs 5.6 ms on scrfd_500m). Int8 on NVIDIA is TensorRT's job.
TensorRT's first load is 16–116 s per model in f32 and 35–290 s in fp16 (yolo26n-seg the worst: nearly five minutes), because it is compiling an engine for this exact GPU. The engine caches to disk and the second load is milliseconds. That number is what §6 is designed around.
1.3 The desktop — Radeon RX 7900 XT, Threadripper 2920X, 24 threads · 2026-09-20
Arch's onnxruntime-rocm 1.29.0 against ROCm 7.2.4 and MIGraphX 7.2.3 (gfx1100), through the
same ep_probe harness (core/dr-inference-engine/examples/ep_probe.rs). Zero input, three
warm-ups, the median of 15 runs, on a machine doing nothing else. The build columns are the wall
clock of Session construction: cold is a MIGraphX compile of the graph for this GPU, cached is the
same session loading the program the cold build wrote.
| Model | ORT CPU f32 | MIGraphX f32 | MIGraphX fp16 | Compile f32 / fp16 (s) | Cached load (s) |
|---|---|---|---|---|---|
| scrfd_500m (Fast) | 10.4 | 2.8 | 2.4 | 40 / 58 | 0.3 |
| scrfd_2.5g (Balanced) | 20.7 | 3.3 | 2.8 | 37 / 40 | 0.3 |
| scrfd_10g (Thorough) | 57.9 | 4.5 | 3.4 | 40 / 48 | 0.4 |
| arcface_mbf (per face) | 12.8 | 1.8 | 1.6 | 15 / 21 | 0.4 |
| yolo26n-seg | 49.3 | 8.4 | 7.5 | 110 / 136 | 0.9 |
| yolo26s-sem-ade20k | 55.8 | 4.8 | 3.8 | 50 / 60 | 0.5 |
| 2d106det (landmarks) | 9.9 | 1.2 | 1.0 | 16 / 21 | 0.2 |
| ocec_s (eye state) | 2.5 | 0.5 | 0.4 | 17 / 17 | 0.1 |
| xfeat-1024 | 26.8 | 10.5 | 9.9 | 37 / 53 | 0.3 |
| migan-512 (per tile) | 514 | 12.7 | 8.3 | 102 / 132 | 0.8 |
Three things the table settles.
- The ROCm execution provider does not exist any more. It was ONNX Runtime's CUDA-provider twin
for AMD, removed in the 1.23 release (AMD's builds dropped it from ROCm 7.1); 1.29's ROCm build ships
libonnxruntime_providers_migraphx.soand nothing else for AMD, and asking forROCmanswers "not enabled in this build". So there is no non-compiling AMD rung to sit under MIGraphX the way the CUDA provider sits under TensorRT: the AMD ladder is MIGraphX, then the CPU. - MIGraphX is a compiling provider, and its cache has to be asked for by name. 15–135 s per
graph cold, under a second from its cache — TensorRT's shape exactly, and §6's design covers it.
Two things the provider does that the code has to know:
ort's builder fills the legacy options struct, which 1.29 reads for the precision flags only, so the cache directory (migraphx_model_cache_dir) reaches it only through the generic key/value registration; and the cache key is the graph, the GPU and the MIGraphX version without the precision, so an fp16 session pointed at the f32 program's directory silently loads the f32 program (the first fp16 row measured here was that, before the directories were split). - fp16 is worth 10–35% over f32 on this card, not the 3× it is worth on TensorRT, because MIGraphX f32 is already 3–8× the CPU provider and the small graphs are launch-bound. The detectors at 2.4–3.4 ms sit beside TensorRT fp16's 1.8–3.3 ms on the RTX 3050; the embedder is 1.6–1.8 ms on either precision and stays f32 (§7). The whole face pipeline for one image (detector + landmarks + eyes + embedder) is under 6 ms.
1.4 What the numbers say
tractis single-threaded. The tablet's one X4 core and one Raptor Lake core give the same tract numbers. Replacing it with ONNX Runtime's CPU provider, no accelerator involved, is 3–6× on the tablet and 8–10× on the desktop. That is the floor, and it is available on every platform the app builds for.- The Hexagon is the standout. A 6 W NPU running int8 beats a discrete RTX 3050 running f32 on five of six models. A whole-library face index on the tablet goes from ~100 ms + 39 ms per face to ~1.4 ms + 12 ms per face, and the "Thorough" detector — 3× the cost of "Fast" today — becomes free. Its price is that the models must be quantised to int8, which is an accuracy question §5 has to answer before it is believed.
- The embedder does not gain from either accelerator. 112×112 input, per-op overhead dominates; it is 9 ms on the tablet's CPU and 12 ms on its NPU. It stays float, which §7 turns from a performance footnote into a correctness rule.
- On NVIDIA, TensorRT fp16 ≈ 3× the CUDA provider, and the CUDA provider ≈ 2× the multi-threaded CPU; at fp16 the detectors are 1.8–3.3 ms with no quantisation at all. Both leave the twenty cores free for decoding during a batch index, which the table does not show and which matters more than the ratio.
- On AMD, MIGraphX fp16 is 4–17× the CPU provider on the detectors and 60× on the inpainter, with the same first-run compile cost as TensorRT and no rung between it and the CPU.
2. The shape of the answer
A ladder per platform, walked at start-up, with the first rung that builds a real session winning:
| Platform | 1st | 2nd | 3rd | Floor |
|---|---|---|---|---|
| Android, Qualcomm with a Hexagon the shipped QNN skel covers (V68–V81) | QNN HTP, int8 model | ORT CPU, f32 model | — | tract |
| Android, any other SoC | ORT CPU, f32 | — | — | tract |
| Linux / Windows, NVIDIA GPU | TensorRT, f32 model, fp16 engine | CUDA provider, f32 | ORT CPU, f32 | tract |
| Linux, AMD GPU with ROCm | MIGraphX, f32 model, fp16 program | ORT CPU, f32 | — | tract |
| Linux / Windows, no GPU stack | ORT CPU, f32 | — | — | tract |
| macOS ⁵ | ORT CPU, f32 | — | — | tract |
⁵ CoreML is the obvious rung and is unmeasured; it is listed so its absence is a gap and not an oversight.
Deliberately not on any ladder, with the measurement that excluded each: NNAPI (no driver), XNNPACK (slower than CPU, aborts on SCRFD), WebGPU (slower than CPU), the Adreno through QNN (works, but never where the Hexagon does not also), CUDA int8 (slower than CUDA f32), the ROCm provider (gone: §1.3). A rung is added to this table by a measurement on this page, not by a provider existing.
The AMD ladder has no middle rung. TensorRT falls back to the CUDA provider while its engines compile; MIGraphX has no such twin, so its fallback is the CPU provider, and the minute or two of compiling on first run (§6) is spent at the floor's speed rather than at half the GPU's.
Two things the ladder is not: it is not a per-model choice — one backend serves every model on a
device, because §7's identity rule needs the detector and embedder on the same runtime for the
same reason faces.model_id pairs them; and it is not a per-account choice — it is a property of
the hardware, like shared_face_models_dir is, and it lives beside it.
3. The dependency policy, and how far this reopens it
D13 chose ort over tract because alternative-backend made ONNX Runtime's API available with
none of its C. Every rung above the floor needs the C++ ONNX Runtime and, for the two that matter
most, vendor libraries on top: Qualcomm's QNN runtime (~60 MB for one Hexagon generation; it is
per-SoC) and NVIDIA's TensorRT plus cuDNN (~600 MB with the CUDA libraries, and cuDNN's major
version must match what ONNX Runtime was built against — the Arch package on the reference desktop
was unusable for exactly that reason).
The policy protected the build: no C to cross-compile under the NDK, no toolchain to keep in
step. This document keeps that intact, and the mechanism is the one thing about ort that makes it
possible:
ort::set_api accepts any OrtApi table. With alternative-backend on, ort links nothing
and asks for the table once per process. The application can dlopen a libonnxruntime.so it
finds on disk, call OrtGetApiBase()->GetApi(version) and hand that table over; or, if there is no
such file, hand over ort_tract::api(). The Rust build is identical in both cases — pure Rust,
cargo build --target aarch64-linux-android sees the same dependency graph it sees today. What
changes is that the runtime is a file the package installs, next to the models, and the app
looks for it at start-up.
Consequences that follow and are accepted:
- The runtime is chosen once per process, because
set_apiis once per process. The ladder in §2 is walked at start-up and the result is what every session in that process uses. There is no "tract for this model, ORT for that one", and there is no falling back to tract after ONNX Runtime has loaded — but there does not need to be: once the library loads, its CPU provider is always there, and every fallback the ladder needs is between providers inside it. - Feature flags stay as they are.
dr-face'sinferenceanddr-segment'ssemanticcontinue to mean "compiled againstort's API"; nothing at build time knows or cares which table will be supplied. The one addition is anative-probefeature on the new crate (§8) that pulls inlibloading, which is pure Rust and already in the tree viawgpu. - The packagers ship the runtime, not the build. The Arch package, the Flatpak manifest, the
NSIS installer and
assemble-apk.sheach gain the ONNX Runtime library for their platform, and the Android and NVIDIA variants gain the vendor libraries — each under the licence the packager reads first (§3.1). A package without them is not broken; it is the tract build, and it says so on the about screen. - The NDK problem does not come back.
libonnxruntime.sofor Android is a prebuilt from Maven (com.microsoft.onnxruntime:onnxruntime-android-qnn), extracted byassemble-apk.shintojniLibs/the way the models are bundled as assets today. Nothing compiles it.
3.1 Licences the packagers read before shipping a runtime
Written down now, because segmentation.md §7 established that reading the grant is cheaper than discovering it at packaging time.
| Component | Licence | Redistributable in a self-distributed package? |
|---|---|---|
| ONNX Runtime | MIT | Yes |
Qualcomm QNN runtime (com.qualcomm.qti:qnn-runtime on Maven) |
Qualcomm AI Engine Direct SDK licence — proprietary, redistribution permitted for applications using it | Yes for the APK, with the licence text shipped; not for a source distribution. To be read in full, not summarised from memory, before the APK gains it. |
| CUDA runtime, cuDNN, TensorRT | NVIDIA EULAs — redistributable with an application, with the licence text, not modifiable | Yes for a package that bundles them. 600 MB. The alternative is to load them from the user's system install if present and skip the rung otherwise — which is what §4's probe does anyway. |
| ROCm (HIP, MIOpen, rocBLAS), MIGraphX | MIT | Yes, but the HIP SDK MIGraphX needs is ~15 GB installed. Same answer as NVIDIA: the user's system install, or the rung is skipped. |
The position this takes: the GPU vendors' libraries are not bundled. The desktop package probes for a system CUDA/TensorRT or ROCm/MIGraphX install and uses it if it is version-compatible; a desktop without one runs on ORT CPU, which is still 8–10× today. Bundling 600 MB for a rung that is 2× again is not a trade worth making unmeasured, and it can be revisited by a measurement on a batch index. The QNN runtime is bundled in the APK, because the Hexagon is the difference between a tablet that indexes a library overnight and one that does it over lunch, and the package is 60 MB larger for it.
Both positions are D13 territory and are recorded there (§12).
4. Selection — the probe, its cache, and what it may not do
A rung is chosen by building a real session on it, not by asking whether it exists. Both failure modes that are not "the provider is absent" were hit on 2026-09-19: a driver in a wedged state where the provider registered and the session then failed, and a provider that registered, took the graph, and rejected every node at partition time. The probe therefore:
- Loads the runtime library (§3), or falls to tract and stops.
- Times the smallest detector on the CPU provider first — the floor. Then, for each rung
in this platform's ladder, in order: builds a session for the same model on that provider
with
error_on_failure, runs it once on a fixed input, and times three more runs. The rung is taken only if its median beats the floor. That one measurement is the proof the provider took the graph: one that silently hands the work to the CPU is the CPU rung with extra overhead, slower than the floor, and rejected. (ONNX Runtime'ssession.disable_cpu_ep_fallbackwas the first draft of this proof and refuses the Hexagon over the ten quantise/dequantise nodes at the graph's edges that QNN declines by policy.) - Records the outcome — rung, runtime version, provider version, device identity (GPU name and
compute capability; SoC model and Hexagon arch), and the models' content hashes — to a small
file beside
shared_face_models_dir. The next start-up trusts the file unless any of those inputs changed, in which case it probes again. A driver update, a runtime update, a new model file: each invalidates the cache by construction, and none needs a "reset backend" button.
What the probe may not do:
- Block the first frame. It runs on the same background as
install_bundled_modelsand for the same reason: a TensorRT probe can take thirty seconds cold, and a tablet that stalls that long is an ANR. Until it reports, every model request is answered by the floor the runtime supports (ORT CPU if the library loaded, tract otherwise), and a job that started on the floor finishes on it — a backend does not change under a running index. - Retry a rung that failed within a session. A failed probe is cached as a failure with the same inputs; the rung is tried again when an input changes. Otherwise a wedged driver means a thirty-second stall on every launch.
- Crash the app twice for the same reason. The probe runs in the app's process, and a provider
can fail by aborting rather than by returning an error (XNNPACK on SCRFD, §2). Every session build
on a rung above the CPU — the probe's, and each background compile of §6 — writes what it is
attempting to
attemptin the cache directory first and removes it after. A launch that finds the file knows the last one died inside that attempt; after two such launches in a row the attempt is refused and recorded like any other failure (a rung infailed, an engine inrefused), until the fingerprint changes. Two, not one, because quitting during a forty-second TensorRT compile leaves the same file. - Choose for the user without saying so. Settings gains one row, Inference backend, showing what was chosen and why in one line ("Hexagon NPU · int8 · QNN 2.42"; "CPU · ONNX Runtime 1.30 · TensorRT probe failed: cuDNN 8 required"), with an override to force any lower rung. The about screen carries the same line beside the model names NFR-SEC-5 already puts there.
5. Model variants, and who makes them
Every model exists in one canonical form — the f32 ONNX file the app ships or the user supplies today — and, where a rung needs it, a derived form. The ladder's rungs are specified in terms of which form they load:
| Form | Who produces it | When | Needed by |
|---|---|---|---|
| f32 ONNX, shape-fixed, opset ≥ 13 | tools/fix-face-model-shapes.sh, tools/export-seg-model.sh |
Release time, once | Every rung except Hexagon |
| int8 QDQ ONNX, per-channel, uint8 activations | tools/quantise-models.sh (new) |
Release time, once, calibrated on real photographs | Hexagon |
TensorRT engine (.engine, per GPU architecture and TensorRT version) |
The app, from the f32 file | First run on that device, in the background | TensorRT rung |
MIGraphX program (.mxr, per GPU architecture, MIGraphX version and precision) |
The app, from the f32 file | First run on that device, in the background | MIGraphX rung |
| QNN context binary | The app, from the int8 file | First run on that device, in the background | Hexagon rung |
Two rules.
Quantisation is a release-time step, not a device-time one. The int8 files that produced §1's
numbers were calibrated on random noise, which is enough to time and worthless to trust. A real
int8 detector is calibrated on a few hundred real photographs and then measured against the f32
detector on the reference library by faces.md §12.3's method — faces found, per size
band, per detector — before it ships. That needs the reference library and a person reading the
result, and it happens once per model release, in tools/, beside the shape-fixing it already
depends on. The device never quantises anything.
The SCRFD and ArcFace files are opset 11 as InsightFace exported them, and per-channel QDQ needs
13; tools/fix-face-model-shapes.sh gains an opset upgrade to 17 (onnx.version_converter,
ir_version 8), which tract has been verified to load and which every provider on this page
prefers. That is a change to the canonical file and so a change to the shipped models, and it
happens in the same model release as the int8 files.
Compilation is a device-time step, and it is cached. A TensorRT engine is specific to the GPU
it was built on and the TensorRT that built it; a MIGraphX program to the GPU and the MIGraphX
that built it; a QNN context binary to the Hexagon generation. None can ship. All are built by
the app the first time that rung is selected, in the background (§6), and written beside the
probe cache keyed by the same inputs. MIGraphX's own key leaves out the precision, so the app
gives its f32 and fp16 programs separate directories — otherwise the embedder's f32 build would
be served the detector's fp16 program, or the reverse. They are
derived, disposable, and regenerable: deleting the cache directory costs the next launch a
rebuild and nothing else, and the directory is excluded from anything that syncs (it is a peer of
thumbs, not of the catalog).
6. First run — building engines without the user waiting for them
The sequence on a device where a compiling rung (TensorRT, MIGraphX, Hexagon) is selected:
- Launch. The runtime loads; the probe (§4) starts in the background; the app serves every model request from the floor. Face indexing, segmentation and scene grading all work, at today's speed or better (ORT CPU).
- Probe reports — say, TensorRT. The compiling rung is now selected but has no engines. Model requests continue on the fallback rung below it (CUDA provider for TensorRT; ORT CPU for MIGraphX and Hexagon), which needs no compilation and is already faster than the floor.
- Engines build, one model at a time, on a single low-priority background thread, smallest model first so the detector — the one that runs per image — is ready soonest. On the reference desktop that is ~1 minute for the first detector and ~10 minutes for all six at fp16; on the AMD desktop 40 s for the first detector and ~8 minutes for the set; on the tablet the QNN context binaries take 0.8–1.7 s each and the whole set is ready before the user has opened a library. Each engine is written to a temporary name and renamed into place, so a request never sees a half-written file.
- Requests move up as engines land. A model whose engine exists loads it on the selected
rung; one whose engine is still building loads on the fallback. A running job does not
switch — an index that started on the CUDA provider finishes on it — because §7 needs one
model_idper job, and because a job is the wrong granularity for surprise. - On Android, the build runs only while the app is in the foreground and the device is not in battery saver (NFR-RES-3): a context binary takes a second, so this costs nothing, and the rule exists for the day a model takes longer.
Settings shows a one-line progress row while engines build ("Preparing GPU engines · 3 of 6") and nothing when they are done. A build failure demotes the rung — it is recorded in the probe cache as a failure with the model hash as an input, so a corrected model file retries it — and the app carries on one rung down, saying so in the same row.
7. Identity — what changes model_id and what may not
faces.model_id exists so that two libraries indexed with different networks are never compared
as if they were one (catalog.md §10.1; the trap is written up in
faces.md §14). A backend that changes what a network computes is a different network
and must be a different model_id; one that changes only where it computes it must not be.
The detector. An int8 SCRFD finds a different set of faces from the f32 one — that is what
§5's acceptance measures — so the numeric form is part of the detector's identity:
scrfd_500m and scrfd_500m_i8 are two detectors in model_id, and a library indexed on the
tablet's Hexagon and continued on the desktop is two populations, which a re-index on either side
reconciles the same way a switch from Fast to Thorough does today. That is acceptable because it is
already the rule for the detector and because §5 is the gate on whether the int8 form is close
enough to be offered at all. f32 on tract, ORT CPU, CUDA and TensorRT-f32 are one identity: the
same graph, the same arithmetic, differences at the last bit.
The embedder is where comparability across devices is the whole point, and it is the one
model that no accelerator helps (§1.4). So: the embedder runs in f32 on every rung. On TensorRT
that means the embedder's engine is built without fp16 while the detector's is built with it; on
the Hexagon it means the embedder is not on the NPU at all — it runs on the ORT CPU rung at 9 ms,
and the ladder's "one backend per device" is, precisely, one backend per model role, with the
embedder pinned. A w600k_mbf embedding from any device is comparable with one from any other,
which is the property the identity system, the calibration and the cross-device merge all rest
on, and it is not for sale for 3 ms.
If S16 wants fp16 for the embedder later, the gate is written now: over the reference library's faces, the cosine between the f32 and fp16 embedding of the same crop exceeds 0.999 for 99.9% of faces and the calibration's fitted threshold moves by less than its own confidence interval (faces.md §8.3). Until measured, f32.
Segmentation and the scene model carry no identity across devices — their outputs are recomputed per image and never stored beyond the cache — so they take whatever the rung offers, int8 included, subject to §10's own acceptance.
8. Crate shape — core/dr-inference-engine
The seam is the same shape as storage.md's: a small crate below the consumers that is the only place naming a provider, a library file or a vendor, with the consumers reduced to "give me a session for these bytes in this role".
core/dr-inference-engine
src/lib.rs Runtime (Tract | Onnx { lib, version }), Backend (rung), Role (Detector | Embedder | Segmenter)
src/probe.rs §4 — the ladder per platform, the session-build probe, the cache file
src/engines.rs §6 — background compilation, the cache directory, progress
src/session.rs open(role, bytes) -> ort::Session, applying the rung and the role's precision rule
src/api.rs the unsafe block that matters: dlopen libonnxruntime, fetch OrtApi, ort::set_api — or ort_tract::api()
dr-faceanddr-segmentdelete their privateinstall_backendand their directSession::builder()calls and take an&dr_inference_engine::Sessionswhere they take model bytes today. Their tests keeptract—dr_inference_engine::Sessions::tract()is a constructor and the test-only path.dr-inference-enginedepends onortwith the same workspace features as today pluscuda,tensorrt,qnn: those features add option builders, not linking, underalternative-backend. Verified for the QNN, CUDA and TensorRT builders on 2026-09-19 — they go through the API table's genericSessionOptionsAppendExecutionProvider*. The NNAPI builder resolves a symbol directly and would not; it is not needed and is not enabled. MIGraphX uses noortfeature at all:ort's builder fills the legacy options struct, which ONNX Runtime 1.29 reads for the precision flags and nothing else, and the compiled-program cache directory only travels through the generic key/value entry point (migraphx_model_cache_dir).session::migraphxmakes that one call on the API table itself.dr-uiowns the settings row, the about-screen line and the progress row; it holds oneSessionsper process, created at launch, and passes it down.dr_ui::librarygainsinference_cache_dir()besideshared_face_models_dir(), on the same account-independent footing and for the same reason.- The Android entry point's
install_bundled_modelsalso extracts nothing new:jniLibs/is loaded by the system loader, anddr-inference-engineon Android looks forlibonnxruntime.sothroughdlopenby bare name first, which resolves to the APK's copy, before any directory.
9. Threads and memory
- ONNX Runtime's intra-op pool is sized to the physical cores minus two on desktop and to the
performance cores on Android (the X4 and the A720s; the A520s are for the compositor). tract's
single thread today is the reason a batch index leaves nineteen cores idle; ORT CPU with the
pool is the reason it will not. One session per model per process;
Session::runis&mut self-free inortand internally serialised, and the index job is the only caller. - A TensorRT session pins GPU memory for its workspace; the builder is capped at 512 MB on the reference 6 GB card and the cap is a setting, because the develop view's tiles share the card (NFR-RES-2). The engine cache on disk is bounded by the model set — six engines, ~80 MB — and needs no LRU.
- The Hexagon rung sets QNN's performance mode to
Burstfor the duration of an index job andDefaultotherwise; a 5 ms detector does not need the NPU clocked up between images. - The probe (§4) and the engine build (§6) run on one dedicated low-priority thread. They never share the index job's pool: a probe that competes with the job it is meant to speed up is the frame-budget trap in a new coat.
10. What S16 measures
In order, with the gate each is:
| # | Question | Gate |
|---|---|---|
| M1 | Does one binary carry both tables? dlopen + set_api on Linux, Windows and Android; ort_tract::api() when the file is absent. |
Go / no-go for §3. If set_api cannot take a dynamically fetched table on some platform, that platform ships two binaries, and the cost is stated. |
| M2 | Do the int8 SCRFD detectors, calibrated on real photographs, find the faces? faces.md §12.3's method over the reference library, per size band, against f32. | Ship the int8 form for a detector only if it finds ≥ 97% of the f32 detector's faces above 40 px and the difference is not concentrated in one band. Otherwise that detector's Hexagon rung is ORT CPU int8-free, and the table in §1.1 says what that costs. |
| M3 | Does the embedder on ORT CPU beside a detector on the Hexagon (§7) produce embeddings within the f32 gate? | It must — same graph, same arithmetic. This is a check that the plumbing did not quantise it by accident. |
| M4 | Is a batch index on the tablet and on the desktop faster by the ratio §1 predicts, end to end, decode included? | The face index over the reference library (18,143 faces): report wall-clock on tract, on the floor and on the selected rung, and where the time went. The prediction is that decode becomes the bottleneck on both; if it does not, say why. |
| M5 | Does the first-run sequence (§6) hold: nothing blocks the first frame, engines land, requests move up, a running job does not switch? | Observed on both devices with the app's own progress row, and with the cache directory deleted between runs. |
| M6 | What does the APK weigh with the QNN runtime, and does a non-Qualcomm Android device (any one) still launch and index on the floor? | Size reported; launch verified on one non-Qualcomm device or an emulator. |
| M7 | The segmentation and scene models at int8 on the Hexagon: does the mask boundary move? segmentation.md §6's IoU against f32 over its corpus. | Ship int8 for a model only above the IoU floor that document set for arm B. |
M1 and M2 are the ones the rest is conditional on, and M2 is the one that needs a person.
10.1 M2 result · 2026-09-19
Each int8 detector against its own f32 form, over 400 proxies evenly spaced through the reference
library, on ONNX Runtime's CPU provider (the int8 graph is the same file the Hexagon loads;
ui/dr-ui/examples/face_detectors). Calibrated on 64 proxies from the same library, disjoint
from the 400.
| Detector | f32 faces | int8 faces | both | int8 only | f32 only | found ≥ 32 px | found, all sizes |
|---|---|---|---|---|---|---|---|
| scrfd_500m | 1342 | 1319 | 1287 | 32 | 55 | 95.6% | 95.9% |
| scrfd_2.5g | 1525 | 1482 | 1478 | 4 | 47 | 96.1% | 96.9% |
| scrfd_10g | 1769 | 1760 | 1749 | 11 | 20 | 100% | 98.9% |
The 10g form clears the 97% gate; 500m and 2.5g sit one point under it. What they lose is specific: the faces in the "f32 only" column have a median confidence of 0.52 against a threshold of 0.50 — detections the f32 graph itself barely made, that int8 rounding drops to the other side of the line — and the extra faces int8 finds are the same kind (median 0.51–0.52). Not a size-band failure: the losses are spread across bands in proportion. Shipped as they are, with the number on record; a threshold of 0.48 for the int8 forms would recover most of the margin, and is the first thing to try if a library's count on the tablet reads low.
Two things the calibration taught, both in tools/quantise-models.py: the calibration set has
to contain faces (a first attempt on landscape photographs produced a graph that found nothing —
the score head's ranges had never seen the face regime), and ONNX Runtime's own strided and
moving-average calibration modes both measurably degrade the result on these graphs, while
driving the calibrator in chunks by hand reproduces the plain min/max ranges exactly.
11. Order
dr-inference-enginewith the two tables and the floor —set_apifrom a dlopened runtime, tract otherwise, ORT CPU as the only rung. Consumers moved over; tests unchanged. This alone is the 3–10× and is the build most of the value sits in. M1.- The probe and its cache (§4), with the settings row and the about line. Still CPU-only; the ladder has one rung. M5's first half.
tools/quantise-models.shand the opset upgrade; the int8 detectors calibrated and measured. M2, M7. This is the step with a person in it and it runs in parallel with 4.- The Hexagon rung, the QNN runtime in the APK, the context-binary cache. M3, M6.
- The TensorRT and CUDA rungs on desktop, the engine cache, the first-run sequence. M5's second half.
- M4 last, on both devices, and the number goes in this document.
The Windows installer and the Flatpak manifest are touched in steps 1 and 5 only, and only to add a file each; the Arch package likewise.
12. Register entries
FR-INF-1 — Runtime selection. On launch the application shall determine, per device and without blocking the first frame, the fastest inference backend that can build and run a session for the shipped models, by attempting it; shall record and reuse that determination until the runtime, driver, hardware or models change; and shall display the backend in use in Settings and on the about screen. Acceptance: §10 M1 and M5.
FR-INF-2 — Derived engines. Backends that require device-specific compilation shall compile in the background after selection, shall serve requests from the next lower backend until each engine is ready, and shall not change the backend of a job in progress. Acceptance: M5.
FR-INF-3 — Model forms. Quantised model forms are produced at release time from real calibration data and are shipped only when they meet §10's accuracy gates against the canonical form; the application never quantises on the device. Acceptance: M2, M7.
NFR-INF-1 — Embedding comparability. Face embeddings shall be computed at a precision whose deviation from the f32 reference is within §7's gate, on every backend, so that embeddings from any device are comparable. Acceptance: M3.
D13 — updated. The runtime half is reopened to the extent of §3: the Rust build stays C-free
under alternative-backend; packages may install a dynamically loaded ONNX Runtime and, per §3.1,
the Qualcomm QNN runtime; the NVIDIA libraries are not bundled. The licensing half is unchanged.
13. Requirements touched
FR-CULL-8 (indexing time is what this exists to change), FR-CULL-9 (NFR-INF-1 is the guard on its calibration), NFR-RES-2 (TensorRT workspace against the develop view's tiles), NFR-RES-3 (background compilation on Android), NFR-SEC-5 (the about screen's line), NFR-COMPAT-2 (each channel gains a runtime file), FR-PLAT-AND-1 (unchanged; the runtime lives in the APK, not in storage), ARCH §6.1 (the once-per-image budget this was sized against no longer binds; what could run per frame is a separate question this document does not open).