docs/ had 26 developer documents flat beside the manual, and the two audiences are very differently sized: most readers want the manual and the gesture reference, a few want the register, the designs and the measurements. The manual and gestures.md stay at the top; everything for someone changing the code moves to docs/dev/, and the two documents that name their own successors — the v0.1 milestone and the UI-refinement plan — go to docs/dev/archive/ rather than being deleted, since both are still cited. docs/README.md is the index, users first. Every reference follows: code comments, Cargo manifests, the workflows, the pre-commit hook, the bench and traceability tools (which locate the repo root by docs/dev/requirements.md now), packaging, the Docker READMEs, CLAUDE.md, CONTRIBUTING.md and the README. The matrix links one level deeper and is regenerated. Links out of the moved documents into the tree gain a level; a link checker over every Markdown file finds none broken.
485 lines
32 KiB
Markdown
485 lines
32 KiB
Markdown
# Inference backends — the runtime and the model, chosen per device
|
||
|
||
Spec for **S16**, the build that puts the neural models on the hardware each device actually has.
|
||
|
||
Every model DarkRoom runs today — the three SCRFD detectors, the ArcFace embedder, YOLO26n-seg and
|
||
the ADE20K scene model — runs through `tract`, on one CPU core, on every platform. That was the
|
||
right first answer: D13's runtime half chose it because it costs no C dependency, and
|
||
[faces.md](faces.md) and [segmentation.md](segmentation.md) were written against it. It is also
|
||
between 20× and 300× slower than what the same devices can do, and this document is the record of
|
||
having measured that and the specification of what replaces it.
|
||
|
||
**It does not reopen D13's licensing half.** The weights are the same files under the same grant.
|
||
It does reopen the *runtime* half, and §3 is where it says how far.
|
||
|
||
---
|
||
|
||
## 1. What was measured · 2026-09-19
|
||
|
||
One benchmark, two builds of it: `ort`'s API over `tract` (exactly what the app links) and `ort`'s
|
||
API over a dynamically loaded ONNX Runtime with each execution provider in turn. Random 640×640
|
||
input, three warm-ups, the median of 15–30 timed runs, milliseconds. The same input every run, so
|
||
the numbers are compute cost and nothing else.
|
||
|
||
### 1.1 The tablet — Honor MagicPad 2, Snapdragon 8s Gen 3
|
||
|
||
SM8635: 1× Cortex-X4, 4× A720, 3× A520, Adreno 735, Hexagon V73. Android 16.
|
||
|
||
| Model | **tract** (today) | ORT CPU f32 | ORT CPU int8 | Adreno f32 ¹ | **Hexagon int8** ² |
|
||
|---|---|---|---|---|---|
|
||
| scrfd_500m (Fast) | 98 | 16 | 8 | 21 | **1.4** |
|
||
| scrfd_2.5g (Balanced) | 161 | 59 | 19 | ✗ | **1.8** |
|
||
| scrfd_10g (Thorough) | 489 | 204 | 48 | ✗ | **3.2** |
|
||
| arcface_mbf (per face) | 39 | 9 | 13 | 24 | 12 |
|
||
| yolo26n-seg | 287 | 94 | 38 | 54 | **5.1** |
|
||
| yolo26s-sem-ade20k | 408 | 154 | 46 | 47 | **3.7** |
|
||
|
||
¹ Qualcomm's own GPU backend (`libQnnGpu.so`, OpenCL). Fails on the two larger SCRFD graphs at an
|
||
`AveragePool` the layout transformer cannot place. ONNX Runtime's WebGPU provider also runs on this
|
||
GPU and was slower than the CPU on every model; it is not in the table because it is not a
|
||
candidate.
|
||
² QNN's HTP backend. The Hexagon **refuses float32 and float16 tensors** in this ORT 1.29 + QNN
|
||
2.42 pairing (error 3110 on every node, with `enable_htp_fp16_precision` set or not); int8 QDQ
|
||
graphs run with 99.6% of nodes on the NPU — 1718 of 1725 for SCRFD-500m, the remainder being the
|
||
quantise/dequantise at the graph's edges — verified from the partition log, not inferred from the
|
||
timing.
|
||
|
||
Also tried and rejected: **NNAPI** — the device registers no neural-networks HAL at all, so the
|
||
provider has nothing to talk to; Google deprecated it in Android 15 and Qualcomm stopped shipping
|
||
drivers for it. **XNNPACK** — slower than ORT's default CPU kernels on every model that loaded, and
|
||
aborts inside its partitioner on the SCRFD graphs.
|
||
|
||
### 1.2 The desktop — RTX 3050 Laptop, Raptor Lake, 20 threads
|
||
|
||
| Model | **tract** (today) | ORT CPU f32 | ORT CPU int8 | CUDA f32 | CUDA fp16 | TensorRT f32 | TensorRT fp16 | TensorRT int8 |
|
||
|---|---|---|---|---|---|---|---|---|
|
||
| scrfd_500m | 104 | 12 | 8 | 5.6 | 3.7 | 2.5 | **1.8** | ✗ ⁴ |
|
||
| scrfd_2.5g | 162 | 27 | 12 | 6.2 | 5.3 | 2.9 | **1.9** | ✗ ⁴ |
|
||
| scrfd_10g | 514 | 99 | 32 | 14.5 | 9.3 | 7.6 | **3.3** | ✗ ⁴ |
|
||
| arcface_mbf | 45 | 15 | 18 | 1.1 | 0.8 | 0.9 | 0.7 | ✗ ⁴ |
|
||
| yolo26n-seg | 307 | 60 | 38 | 9.3 | ✗ ³ | 6.8 | **5.5** | ✗ ⁴ |
|
||
| yolo26s-sem-ade20k | 395 | 61 | 31 | 10.7 | ✗ ³ | 7.9 | **3.4** | ✗ ⁴ |
|
||
|
||
³ The offline fp16 conversion (`onnxconverter-common`) left a mixed-type node the CUDA provider
|
||
rejects. TensorRT converts to fp16 itself at engine build and does not have this problem, which is
|
||
one reason it is the target and the CUDA provider is the fallback.
|
||
⁴ TensorRT refuses the QDQ form ONNX Runtime's quantiser writes for the Hexagon (uint8
|
||
activations); it wants symmetric int8. Not pursued: fp16 needs no quantisation, no calibration and
|
||
no accuracy gate, and it is already 30–60× tract.
|
||
|
||
CUDA int8 is deliberately absent: the CUDA provider has no int8 kernels and runs a QDQ graph by
|
||
dequantising it, which measured *slower* than f32 (7.0 vs 5.6 ms on scrfd_500m). Int8 on NVIDIA is
|
||
TensorRT's job.
|
||
|
||
**TensorRT's first load is 16–116 s per model in f32 and 35–290 s in fp16** (yolo26n-seg the
|
||
worst: nearly five minutes), because it is compiling an engine for this exact GPU. The engine caches to disk and the second load is milliseconds. That
|
||
number is what §6 is designed around.
|
||
|
||
### 1.3 What the numbers say
|
||
|
||
- **`tract` is single-threaded.** The tablet's one X4 core and one Raptor Lake core give the same
|
||
tract numbers. Replacing it with ONNX Runtime's CPU provider, *no accelerator involved*, is 3–6×
|
||
on the tablet and 8–10× on the desktop. That is the floor, and it is available on every platform
|
||
the app builds for.
|
||
- **The Hexagon is the standout.** A 6 W NPU running int8 beats a discrete RTX 3050 running f32 on
|
||
five of six models. A whole-library face index on the tablet goes from ~100 ms + 39 ms per face
|
||
to ~1.4 ms + 12 ms per face, and the "Thorough" detector — 3× the cost of "Fast" today — becomes
|
||
free. Its price is that the models must be **quantised to int8**, which is an accuracy question
|
||
§5 has to answer before it is believed.
|
||
- **The embedder does not gain from either accelerator.** 112×112 input, per-op overhead
|
||
dominates; it is 9 ms on the tablet's CPU and 12 ms on its NPU. It stays float, which §7 turns
|
||
from a performance footnote into a correctness rule.
|
||
- **On NVIDIA, TensorRT fp16 ≈ 3× the CUDA provider**, and the CUDA provider ≈ 2× the
|
||
multi-threaded CPU; at fp16 the detectors are 1.8–3.3 ms with no quantisation at all. Both leave the twenty cores free for decoding during a batch index, which the table does not
|
||
show and which matters more than the ratio.
|
||
|
||
---
|
||
|
||
## 2. The shape of the answer
|
||
|
||
A **ladder per platform**, walked at start-up, with the first rung that builds a real session
|
||
winning:
|
||
|
||
| Platform | 1st | 2nd | 3rd | Floor |
|
||
|---|---|---|---|---|
|
||
| Android, Qualcomm with a Hexagon the shipped QNN skel covers (V68–V81) | QNN HTP, int8 model | ORT CPU, f32 model | — | tract |
|
||
| Android, any other SoC | ORT CPU, f32 | — | — | tract |
|
||
| Linux / Windows, NVIDIA GPU | TensorRT, f32 model, fp16 engine | CUDA provider, f32 | ORT CPU, f32 | tract |
|
||
| Linux / Windows, no NVIDIA | ORT CPU, f32 | — | — | tract |
|
||
| macOS ⁵ | ORT CPU, f32 | — | — | tract |
|
||
|
||
⁵ CoreML is the obvious rung and is unmeasured; it is listed so its absence is a gap and not an
|
||
oversight.
|
||
|
||
Deliberately **not** on any ladder, with the measurement that excluded each: NNAPI (no driver),
|
||
XNNPACK (slower than CPU, aborts on SCRFD), WebGPU (slower than CPU), the Adreno through QNN (works,
|
||
but never where the Hexagon does not also), CUDA int8 (slower than CUDA f32). A rung is added to this
|
||
table by a measurement on this page, not by a provider existing.
|
||
|
||
Two things the ladder is *not*: it is not a per-model choice — one backend serves every model on a
|
||
device, because §7's identity rule needs the detector and embedder on the same runtime for the
|
||
same reason `faces.model_id` pairs them; and it is not a per-account choice — it is a property of
|
||
the hardware, like `shared_face_models_dir` is, and it lives beside it.
|
||
|
||
---
|
||
|
||
## 3. The dependency policy, and how far this reopens it
|
||
|
||
D13 chose `ort` over `tract` because `alternative-backend` made ONNX Runtime's *API* available with
|
||
none of its *C*. Every rung above the floor needs the C++ ONNX Runtime and, for the two that matter
|
||
most, vendor libraries on top: Qualcomm's QNN runtime (~60 MB for one Hexagon generation; it is
|
||
per-SoC) and NVIDIA's TensorRT plus cuDNN (~600 MB with the CUDA libraries, and cuDNN's major
|
||
version must match what ONNX Runtime was built against — the Arch package on the reference desktop
|
||
was unusable for exactly that reason).
|
||
|
||
The policy protected the **build**: no C to cross-compile under the NDK, no toolchain to keep in
|
||
step. This document keeps that intact, and the mechanism is the one thing about `ort` that makes it
|
||
possible:
|
||
|
||
**`ort::set_api` accepts any `OrtApi` table.** With `alternative-backend` on, `ort` links nothing
|
||
and asks for the table once per process. The application can `dlopen` a `libonnxruntime.so` it
|
||
finds on disk, call `OrtGetApiBase()->GetApi(version)` and hand that table over; or, if there is no
|
||
such file, hand over `ort_tract::api()`. The Rust build is identical in both cases — pure Rust,
|
||
`cargo build --target aarch64-linux-android` sees the same dependency graph it sees today. What
|
||
changes is that the runtime is a **file the package installs**, next to the models, and the app
|
||
looks for it at start-up.
|
||
|
||
Consequences that follow and are accepted:
|
||
|
||
- **The runtime is chosen once per process**, because `set_api` is once per process. The ladder
|
||
in §2 is walked at start-up and the result is what every session in that process uses. There is
|
||
no "tract for this model, ORT for that one", and there is no falling back to tract *after* ONNX
|
||
Runtime has loaded — but there does not need to be: once the library loads, its CPU provider is
|
||
always there, and every fallback the ladder needs is between providers *inside* it.
|
||
- **Feature flags stay as they are.** `dr-face`'s `inference` and `dr-segment`'s `semantic`
|
||
continue to mean "compiled against `ort`'s API"; nothing at build time knows or cares which table
|
||
will be supplied. The one addition is a `native-probe` feature on the new crate (§8) that pulls in
|
||
`libloading`, which is pure Rust and already in the tree via `wgpu`.
|
||
- **The packagers ship the runtime, not the build.** The Arch package, the Flatpak manifest, the
|
||
NSIS installer and `assemble-apk.sh` each gain the ONNX Runtime library for their platform, and
|
||
the Android and NVIDIA variants gain the vendor libraries — each under the licence the packager
|
||
reads first (§3.1). A package without them is not broken; it is the tract build, and it says so
|
||
on the about screen.
|
||
- **The NDK problem does not come back.** `libonnxruntime.so` for Android is a prebuilt from
|
||
Maven (`com.microsoft.onnxruntime:onnxruntime-android-qnn`), extracted by `assemble-apk.sh` into
|
||
`jniLibs/` the way the models are bundled as assets today. Nothing compiles it.
|
||
|
||
### 3.1 Licences the packagers read before shipping a runtime
|
||
|
||
Written down now, because [segmentation.md §7](segmentation.md) established that reading the grant
|
||
is cheaper than discovering it at packaging time.
|
||
|
||
| Component | Licence | Redistributable in a self-distributed package? |
|
||
|---|---|---|
|
||
| ONNX Runtime | MIT | Yes |
|
||
| Qualcomm QNN runtime (`com.qualcomm.qti:qnn-runtime` on Maven) | Qualcomm AI Engine Direct SDK licence — proprietary, redistribution permitted for applications using it | Yes for the APK, with the licence text shipped; not for a source distribution. **To be read in full, not summarised from memory, before the APK gains it.** |
|
||
| CUDA runtime, cuDNN, TensorRT | NVIDIA EULAs — redistributable with an application, with the licence text, not modifiable | Yes for a package that bundles them. 600 MB. The alternative is to load them from the user's system install if present and skip the rung otherwise — which is what §4's probe does anyway. |
|
||
|
||
The position this takes: the **NVIDIA libraries are not bundled**. The desktop package probes for a
|
||
system CUDA/TensorRT install and uses it if it is version-compatible; a desktop without one runs on
|
||
ORT CPU, which is still 8–10× today. Bundling 600 MB for a rung that is 2× again is not a trade
|
||
worth making unmeasured, and it can be revisited by a measurement on a batch index. The **QNN
|
||
runtime is bundled** in the APK, because the Hexagon is the difference between a tablet that
|
||
indexes a library overnight and one that does it over lunch, and the package is 60 MB larger for
|
||
it.
|
||
|
||
Both positions are D13 territory and are recorded there (§12).
|
||
|
||
---
|
||
|
||
## 4. Selection — the probe, its cache, and what it may not do
|
||
|
||
**A rung is chosen by building a real session on it, not by asking whether it exists.** Both
|
||
failure modes that are not "the provider is absent" were hit on 2026-09-19: a driver in a wedged
|
||
state where the provider registered and the session then failed, and a provider that registered,
|
||
took the graph, and rejected every node at partition time. The probe therefore:
|
||
|
||
1. Loads the runtime library (§3), or falls to tract and stops.
|
||
2. Times the **smallest detector** on the CPU provider first — the floor. Then, for each rung
|
||
in this platform's ladder, in order: builds a session for the same model on that provider
|
||
with `error_on_failure`, runs it once on a fixed input, and times three more runs. **The rung
|
||
is taken only if its median beats the floor.** That one measurement is the proof the provider
|
||
took the graph: one that silently hands the work to the CPU is the CPU rung with extra
|
||
overhead, slower than the floor, and rejected. (ONNX Runtime's
|
||
`session.disable_cpu_ep_fallback` was the first draft of this proof and refuses the Hexagon
|
||
over the ten quantise/dequantise nodes at the graph's edges that QNN declines by policy.)
|
||
3. Records the outcome — rung, runtime version, provider version, device identity (GPU name and
|
||
compute capability; SoC model and Hexagon arch), and the models' content hashes — to a small
|
||
file beside `shared_face_models_dir`. The next start-up trusts the file **unless** any of those
|
||
inputs changed, in which case it probes again. A driver update, a runtime update, a new model
|
||
file: each invalidates the cache by construction, and none needs a "reset backend" button.
|
||
|
||
What the probe may not do:
|
||
|
||
- **Block the first frame.** It runs on the same background as `install_bundled_models` and for
|
||
the same reason: a TensorRT probe can take thirty seconds cold, and a tablet that stalls that long
|
||
is an ANR. Until it reports, every model request is answered by the floor the runtime supports
|
||
(ORT CPU if the library loaded, tract otherwise), and a job that started on the floor finishes
|
||
on it — a backend does not change under a running index.
|
||
- **Retry a rung that failed within a session.** A failed probe is cached as a failure with the
|
||
same inputs; the rung is tried again when an input changes. Otherwise a wedged driver means a
|
||
thirty-second stall on every launch.
|
||
- **Choose for the user without saying so.** Settings gains one row, *Inference backend*, showing
|
||
what was chosen and why in one line ("Hexagon NPU · int8 · QNN 2.42"; "CPU · ONNX Runtime 1.30 ·
|
||
TensorRT probe failed: cuDNN 8 required"), with an override to force any lower rung. The about
|
||
screen carries the same line beside the model names NFR-SEC-5 already puts there.
|
||
|
||
---
|
||
|
||
## 5. Model variants, and who makes them
|
||
|
||
Every model exists in one **canonical** form — the f32 ONNX file the app ships or the user supplies
|
||
today — and, where a rung needs it, a **derived** form. The ladder's rungs are specified in terms
|
||
of which form they load:
|
||
|
||
| Form | Who produces it | When | Needed by |
|
||
|---|---|---|---|
|
||
| f32 ONNX, shape-fixed, **opset ≥ 13** | `tools/fix-face-model-shapes.sh`, `tools/export-seg-model.sh` | Release time, once | Every rung except Hexagon |
|
||
| int8 QDQ ONNX, per-channel, uint8 activations | `tools/quantise-models.sh` (new) | Release time, once, **calibrated on real photographs** | Hexagon |
|
||
| TensorRT engine (`.engine`, per GPU architecture and TensorRT version) | The app, from the f32 file | First run on that device, in the background | TensorRT rung |
|
||
| QNN context binary | The app, from the int8 file | First run on that device, in the background | Hexagon rung |
|
||
|
||
Two rules.
|
||
|
||
**Quantisation is a release-time step, not a device-time one.** The int8 files that produced §1's
|
||
numbers were calibrated on random noise, which is enough to time and worthless to trust. A real
|
||
int8 detector is calibrated on a few hundred real photographs and then measured against the f32
|
||
detector on the reference library by [faces.md §12.3](faces.md)'s method — faces found, per size
|
||
band, per detector — before it ships. That needs the reference library and a person reading the
|
||
result, and it happens once per model release, in `tools/`, beside the shape-fixing it already
|
||
depends on. The device never quantises anything.
|
||
|
||
The SCRFD and ArcFace files are **opset 11** as InsightFace exported them, and per-channel QDQ needs
|
||
13; `tools/fix-face-model-shapes.sh` gains an opset upgrade to 17 (`onnx.version_converter`,
|
||
`ir_version` 8), which `tract` has been verified to load and which every provider on this page
|
||
prefers. That is a change to the canonical file and so a change to the shipped models, and it
|
||
happens in the same model release as the int8 files.
|
||
|
||
**Compilation is a device-time step, and it is cached.** A TensorRT engine is specific to the GPU
|
||
it was built on and the TensorRT that built it; a QNN context binary is specific to the Hexagon
|
||
generation. Neither can ship. Both are built by the app the first time that rung is selected, in
|
||
the background (§6), and written beside the probe cache keyed by the same inputs. They are
|
||
**derived, disposable, and regenerable**: deleting the cache directory costs the next launch a
|
||
rebuild and nothing else, and the directory is excluded from anything that syncs (it is a peer of
|
||
`thumbs`, not of the catalog).
|
||
|
||
---
|
||
|
||
## 6. First run — building engines without the user waiting for them
|
||
|
||
The sequence on a device where a compiling rung (TensorRT, Hexagon) is selected:
|
||
|
||
1. **Launch.** The runtime loads; the probe (§4) starts in the background; the app serves every
|
||
model request from the floor. Face indexing, segmentation and scene grading all work, at
|
||
today's speed or better (ORT CPU).
|
||
2. **Probe reports** — say, TensorRT. The compiling rung is now *selected* but has **no engines**.
|
||
Model requests continue on the fallback rung below it (CUDA provider for TensorRT; ORT CPU for
|
||
Hexagon), which needs no compilation and is already faster than the floor.
|
||
3. **Engines build**, one model at a time, on a single low-priority background thread, smallest
|
||
model first so the detector — the one that runs per image — is ready soonest. On the reference
|
||
desktop that is ~1 minute for the first detector and ~10 minutes for all six at fp16; on the tablet the QNN
|
||
context binaries take 0.8–1.7 s each and the whole set is ready before the user has opened a
|
||
library. Each engine is written to a temporary name and renamed into place, so a request never
|
||
sees a half-written file.
|
||
4. **Requests move up as engines land.** A model whose engine exists loads it on the selected
|
||
rung; one whose engine is still building loads on the fallback. **A running job does not
|
||
switch** — an index that started on the CUDA provider finishes on it — because §7 needs one
|
||
`model_id` per job, and because a job is the wrong granularity for surprise.
|
||
5. On Android, the build runs only while the app is in the foreground and the device is not in
|
||
battery saver (NFR-RES-3): a context binary takes a second, so this costs nothing, and the rule
|
||
exists for the day a model takes longer.
|
||
|
||
Settings shows a one-line progress row while engines build ("Preparing GPU engines · 3 of 6") and
|
||
nothing when they are done. A build failure demotes the rung — it is recorded in the probe cache
|
||
as a failure with the model hash as an input, so a corrected model file retries it — and the app
|
||
carries on one rung down, saying so in the same row.
|
||
|
||
---
|
||
|
||
## 7. Identity — what changes `model_id` and what may not
|
||
|
||
`faces.model_id` exists so that two libraries indexed with different networks are never compared
|
||
as if they were one ([catalog.md §10.1](catalog.md); the trap is written up in
|
||
[faces.md §14](faces.md)). A backend that changes what a network *computes* is a different network
|
||
and must be a different `model_id`; one that changes only *where* it computes it must not be.
|
||
|
||
**The detector.** An int8 SCRFD finds a different set of faces from the f32 one — that is what
|
||
§5's acceptance measures — so **the numeric form is part of the detector's identity**:
|
||
`scrfd_500m` and `scrfd_500m_i8` are two detectors in `model_id`, and a library indexed on the
|
||
tablet's Hexagon and continued on the desktop is two populations, which a re-index on either side
|
||
reconciles the same way a switch from Fast to Thorough does today. That is acceptable because it is
|
||
already the rule for the detector and because §5 is the gate on whether the int8 form is close
|
||
enough to be *offered* at all. f32 on tract, ORT CPU, CUDA and TensorRT-f32 are one identity: the
|
||
same graph, the same arithmetic, differences at the last bit.
|
||
|
||
**The embedder** is where comparability across devices is the whole point, and it is the one
|
||
model that no accelerator helps (§1.3). So: **the embedder runs in f32 on every rung.** On TensorRT
|
||
that means the embedder's engine is built without fp16 while the detector's is built with it; on
|
||
the Hexagon it means the embedder is not on the NPU at all — it runs on the ORT CPU rung at 9 ms,
|
||
and the ladder's "one backend per device" is, precisely, one backend *per model role*, with the
|
||
embedder pinned. A `w600k_mbf` embedding from any device is comparable with one from any other,
|
||
which is the property the identity system, the calibration and the cross-device merge all rest
|
||
on, and it is not for sale for 3 ms.
|
||
|
||
If S16 wants fp16 for the embedder later, the gate is written now: over the reference library's
|
||
faces, the cosine between the f32 and fp16 embedding of the same crop exceeds 0.999 for 99.9% of
|
||
faces and the calibration's fitted threshold moves by less than its own confidence interval
|
||
([faces.md §8.3](faces.md)). Until measured, f32.
|
||
|
||
**Segmentation and the scene model** carry no identity across devices — their outputs are
|
||
recomputed per image and never stored beyond the cache — so they take whatever the rung offers,
|
||
int8 included, subject to §10's own acceptance.
|
||
|
||
---
|
||
|
||
## 8. Crate shape — `core/dr-inference-engine`
|
||
|
||
The seam is the same shape as [storage.md](storage.md)'s: a small crate below the consumers that is
|
||
the **only** place naming a provider, a library file or a vendor, with the consumers reduced to
|
||
"give me a session for these bytes in this role".
|
||
|
||
```
|
||
core/dr-inference-engine
|
||
src/lib.rs Runtime (Tract | Onnx { lib, version }), Backend (rung), Role (Detector | Embedder | Segmenter)
|
||
src/probe.rs §4 — the ladder per platform, the session-build probe, the cache file
|
||
src/engines.rs §6 — background compilation, the cache directory, progress
|
||
src/session.rs open(role, bytes) -> ort::Session, applying the rung and the role's precision rule
|
||
src/api.rs the one unsafe block: dlopen libonnxruntime, fetch OrtApi, ort::set_api — or ort_tract::api()
|
||
```
|
||
|
||
- `dr-face` and `dr-segment` **delete** their private `install_backend` and their direct
|
||
`Session::builder()` calls and take an `&dr_inference_engine::Sessions` where they take model bytes today.
|
||
Their tests keep `tract` — `dr_inference_engine::Sessions::tract()` is a constructor and the test-only path.
|
||
- `dr-inference-engine` depends on `ort` with the same workspace features as today plus `cuda`, `tensorrt`,
|
||
`qnn`: those features add option builders, not linking, under `alternative-backend`. **Verified
|
||
for the QNN, CUDA and TensorRT builders on 2026-09-19** — they go through the API table's generic
|
||
`SessionOptionsAppendExecutionProvider*`. The NNAPI builder resolves a symbol directly and would
|
||
not; it is not needed and is not enabled.
|
||
- `dr-ui` owns the settings row, the about-screen line and the progress row; it holds one
|
||
`Sessions` per process, created at launch, and passes it down. `dr_ui::library` gains
|
||
`inference_cache_dir()` beside `shared_face_models_dir()`, on the same account-independent
|
||
footing and for the same reason.
|
||
- The Android entry point's `install_bundled_models` also extracts nothing new: `jniLibs/` is
|
||
loaded by the system loader, and `dr-inference-engine` on Android looks for `libonnxruntime.so` through
|
||
`dlopen` by bare name first, which resolves to the APK's copy, before any directory.
|
||
|
||
---
|
||
|
||
## 9. Threads and memory
|
||
|
||
- ONNX Runtime's intra-op pool is sized to the physical cores minus two on desktop and to the
|
||
performance cores on Android (the X4 and the A720s; the A520s are for the compositor). tract's
|
||
single thread today is the reason a batch index leaves nineteen cores idle; ORT CPU with the
|
||
pool is the reason it will not. One session per model per process; `Session::run` is
|
||
`&mut self`-free in `ort` and internally serialised, and the index job is the only caller.
|
||
- A TensorRT session pins GPU memory for its workspace; the builder is capped at 512 MB on the
|
||
reference 6 GB card and the cap is a setting, because the develop view's tiles share the card
|
||
(NFR-RES-2). The engine cache on disk is bounded by the model set — six engines, ~80 MB — and
|
||
needs no LRU.
|
||
- The Hexagon rung sets QNN's performance mode to `Burst` for the duration of an index job and
|
||
`Default` otherwise; a 5 ms detector does not need the NPU clocked up between images.
|
||
- The probe (§4) and the engine build (§6) run on one dedicated low-priority thread. They never
|
||
share the index job's pool: a probe that competes with the job it is meant to speed up is the
|
||
frame-budget trap in a new coat.
|
||
|
||
---
|
||
|
||
## 10. What S16 measures
|
||
|
||
In order, with the gate each is:
|
||
|
||
| # | Question | Gate |
|
||
|---|---|---|
|
||
| M1 | Does one binary carry both tables? `dlopen` + `set_api` on Linux, Windows and Android; `ort_tract::api()` when the file is absent. | Go / no-go for §3. If `set_api` cannot take a dynamically fetched table on some platform, that platform ships two binaries, and the cost is stated. |
|
||
| M2 | Do the int8 SCRFD detectors, **calibrated on real photographs**, find the faces? faces.md §12.3's method over the reference library, per size band, against f32. | Ship the int8 form for a detector only if it finds ≥ 97% of the f32 detector's faces above 40 px and the difference is not concentrated in one band. Otherwise that detector's Hexagon rung is ORT CPU int8-free, and the table in §1.1 says what that costs. |
|
||
| M3 | Does the embedder on ORT CPU beside a detector on the Hexagon (§7) produce embeddings within the f32 gate? | It must — same graph, same arithmetic. This is a check that the plumbing did not quantise it by accident. |
|
||
| M4 | Is a batch index on the tablet and on the desktop faster by the ratio §1 predicts, end to end, decode included? | The face index over the reference library (18,143 faces): report wall-clock on tract, on the floor and on the selected rung, and where the time went. The prediction is that decode becomes the bottleneck on both; if it does not, say why. |
|
||
| M5 | Does the first-run sequence (§6) hold: nothing blocks the first frame, engines land, requests move up, a running job does not switch? | Observed on both devices with the app's own progress row, and with the cache directory deleted between runs. |
|
||
| M6 | What does the APK weigh with the QNN runtime, and does a non-Qualcomm Android device (any one) still launch and index on the floor? | Size reported; launch verified on one non-Qualcomm device or an emulator. |
|
||
| M7 | The segmentation and scene models at int8 on the Hexagon: does the mask boundary move? segmentation.md §6's IoU against f32 over its corpus. | Ship int8 for a model only above the IoU floor that document set for arm B. |
|
||
|
||
M1 and M2 are the ones the rest is conditional on, and M2 is the one that needs a person.
|
||
|
||
### 10.1 M2 result · 2026-09-19
|
||
|
||
Each int8 detector against its own f32 form, over 400 proxies evenly spaced through the reference
|
||
library, on ONNX Runtime's CPU provider (the int8 graph is the same file the Hexagon loads;
|
||
`ui/dr-ui/examples/face_detectors`). Calibrated on 64 proxies from the same library, disjoint
|
||
from the 400.
|
||
|
||
| Detector | f32 faces | int8 faces | both | int8 only | f32 only | found ≥ 32 px | found, all sizes |
|
||
|---|---|---|---|---|---|---|---|
|
||
| scrfd_500m | 1342 | 1319 | 1287 | 32 | 55 | 95.6% | 95.9% |
|
||
| scrfd_2.5g | 1525 | 1482 | 1478 | 4 | 47 | 96.1% | 96.9% |
|
||
| scrfd_10g | 1769 | 1760 | 1749 | 11 | 20 | 100% | 98.9% |
|
||
|
||
The 10g form clears the 97% gate; 500m and 2.5g sit one point under it. What they lose is
|
||
specific: the faces in the "f32 only" column have a **median confidence of 0.52** against a
|
||
threshold of 0.50 — detections the f32 graph itself barely made, that int8 rounding drops to the
|
||
other side of the line — and the extra faces int8 finds are the same kind (median 0.51–0.52).
|
||
Not a size-band failure: the losses are spread across bands in proportion. Shipped as they are,
|
||
with the number on record; a threshold of 0.48 for the int8 forms would recover most of the
|
||
margin, and is the first thing to try if a library's count on the tablet reads low.
|
||
|
||
Two things the calibration taught, both in `tools/quantise-models.py`: the calibration set has
|
||
to contain faces (a first attempt on landscape photographs produced a graph that found nothing —
|
||
the score head's ranges had never seen the face regime), and ONNX Runtime's own strided and
|
||
moving-average calibration modes both measurably degrade the result on these graphs, while
|
||
driving the calibrator in chunks by hand reproduces the plain min/max ranges exactly.
|
||
|
||
---
|
||
|
||
## 11. Order
|
||
|
||
1. **`dr-inference-engine` with the two tables and the floor** — `set_api` from a dlopened runtime, tract
|
||
otherwise, ORT CPU as the only rung. Consumers moved over; tests unchanged. This alone is the
|
||
3–10× and is the build most of the value sits in. M1.
|
||
2. **The probe and its cache** (§4), with the settings row and the about line. Still CPU-only;
|
||
the ladder has one rung. M5's first half.
|
||
3. **`tools/quantise-models.sh`** and the opset upgrade; the int8 detectors calibrated and
|
||
measured. M2, M7. This is the step with a person in it and it runs in parallel with 4.
|
||
4. **The Hexagon rung**, the QNN runtime in the APK, the context-binary cache. M3, M6.
|
||
5. **The TensorRT and CUDA rungs** on desktop, the engine cache, the first-run sequence. M5's
|
||
second half.
|
||
6. **M4** last, on both devices, and the number goes in this document.
|
||
|
||
The Windows installer and the Flatpak manifest are touched in steps 1 and 5 only, and only to add
|
||
a file each; the Arch package likewise.
|
||
|
||
---
|
||
|
||
## 12. Register entries
|
||
|
||
**FR-INF-1 — Runtime selection.** On launch the application shall determine, per device and
|
||
without blocking the first frame, the fastest inference backend that can build and run a session
|
||
for the shipped models, by attempting it; shall record and reuse that determination until the
|
||
runtime, driver, hardware or models change; and shall display the backend in use in Settings and
|
||
on the about screen. *Acceptance:* §10 M1 and M5.
|
||
|
||
**FR-INF-2 — Derived engines.** Backends that require device-specific compilation shall compile in
|
||
the background after selection, shall serve requests from the next lower backend until each engine
|
||
is ready, and shall not change the backend of a job in progress. *Acceptance:* M5.
|
||
|
||
**FR-INF-3 — Model forms.** Quantised model forms are produced at release time from real
|
||
calibration data and are shipped only when they meet §10's accuracy gates against the canonical
|
||
form; the application never quantises on the device. *Acceptance:* M2, M7.
|
||
|
||
**NFR-INF-1 — Embedding comparability.** Face embeddings shall be computed at a precision whose
|
||
deviation from the f32 reference is within §7's gate, on every backend, so that embeddings from any
|
||
device are comparable. *Acceptance:* M3.
|
||
|
||
**D13 — updated.** The runtime half is reopened to the extent of §3: the Rust build stays C-free
|
||
under `alternative-backend`; packages may install a dynamically loaded ONNX Runtime and, per §3.1,
|
||
the Qualcomm QNN runtime; the NVIDIA libraries are not bundled. The licensing half is unchanged.
|
||
|
||
---
|
||
|
||
## 13. Requirements touched
|
||
|
||
FR-CULL-8 (indexing time is what this exists to change), FR-CULL-9 (NFR-INF-1 is the guard on its
|
||
calibration), NFR-RES-2 (TensorRT workspace against the develop view's tiles), NFR-RES-3
|
||
(background compilation on Android), NFR-SEC-5 (the about screen's line), NFR-COMPAT-2 (each
|
||
channel gains a runtime file), FR-PLAT-AND-1 (unchanged; the runtime lives in the APK, not in
|
||
storage), ARCH §6.1 (the once-per-image budget this was sized against no longer binds; what could
|
||
run per frame is a separate question this document does not open).
|