Put the developer docs under docs/dev and index the folder for users first
docs/ had 26 developer documents flat beside the manual, and the two audiences are very differently sized: most readers want the manual and the gesture reference, a few want the register, the designs and the measurements. The manual and gestures.md stay at the top; everything for someone changing the code moves to docs/dev/, and the two documents that name their own successors — the v0.1 milestone and the UI-refinement plan — go to docs/dev/archive/ rather than being deleted, since both are still cited. docs/README.md is the index, users first. Every reference follows: code comments, Cargo manifests, the workflows, the pre-commit hook, the bench and traceability tools (which locate the repo root by docs/dev/requirements.md now), packaging, the Docker READMEs, CLAUDE.md, CONTRIBUTING.md and the README. The matrix links one level deeper and is regenerated. Links out of the moved documents into the tree gain a level; a link checker over every Markdown file finds none broken.
This commit is contained in:
@@ -0,0 +1,544 @@
|
||||
# Inference backends — the runtime and the model, chosen per device
|
||||
|
||||
Spec for **S16**, the build that puts the neural models on the hardware each device actually has.
|
||||
|
||||
Every model DarkRoom runs today — the three SCRFD detectors, the ArcFace embedder, YOLO26n-seg and
|
||||
the ADE20K scene model — runs through `tract`, on one CPU core, on every platform. That was the
|
||||
right first answer: D13's runtime half chose it because it costs no C dependency, and
|
||||
[faces.md](faces.md) and [segmentation.md](segmentation.md) were written against it. It is also
|
||||
between 20× and 300× slower than what the same devices can do, and this document is the record of
|
||||
having measured that and the specification of what replaces it.
|
||||
|
||||
**It does not reopen D13's licensing half.** The weights are the same files under the same grant.
|
||||
It does reopen the *runtime* half, and §3 is where it says how far.
|
||||
|
||||
---
|
||||
|
||||
## 1. What was measured · 2026-09-19
|
||||
|
||||
One benchmark, two builds of it: `ort`'s API over `tract` (exactly what the app links) and `ort`'s
|
||||
API over a dynamically loaded ONNX Runtime with each execution provider in turn. Random 640×640
|
||||
input, three warm-ups, the median of 15–30 timed runs, milliseconds. The same input every run, so
|
||||
the numbers are compute cost and nothing else.
|
||||
|
||||
### 1.1 The tablet — Honor MagicPad 2, Snapdragon 8s Gen 3
|
||||
|
||||
SM8635: 1× Cortex-X4, 4× A720, 3× A520, Adreno 735, Hexagon V73. Android 16.
|
||||
|
||||
| Model | **tract** (today) | ORT CPU f32 | ORT CPU int8 | Adreno f32 ¹ | **Hexagon int8** ² |
|
||||
|---|---|---|---|---|---|
|
||||
| scrfd_500m (Fast) | 98 | 16 | 8 | 21 | **1.4** |
|
||||
| scrfd_2.5g (Balanced) | 161 | 59 | 19 | ✗ | **1.8** |
|
||||
| scrfd_10g (Thorough) | 489 | 204 | 48 | ✗ | **3.2** |
|
||||
| arcface_mbf (per face) | 39 | 9 | 13 | 24 | 12 |
|
||||
| yolo26n-seg | 287 | 94 | 38 | 54 | **5.1** |
|
||||
| yolo26s-sem-ade20k | 408 | 154 | 46 | 47 | **3.7** |
|
||||
|
||||
¹ Qualcomm's own GPU backend (`libQnnGpu.so`, OpenCL). Fails on the two larger SCRFD graphs at an
|
||||
`AveragePool` the layout transformer cannot place. ONNX Runtime's WebGPU provider also runs on this
|
||||
GPU and was slower than the CPU on every model; it is not in the table because it is not a
|
||||
candidate.
|
||||
² QNN's HTP backend. The Hexagon **refuses float32 and float16 tensors** in this ORT 1.29 + QNN
|
||||
2.42 pairing (error 3110 on every node, with `enable_htp_fp16_precision` set or not); int8 QDQ
|
||||
graphs run with 99.6% of nodes on the NPU — 1718 of 1725 for SCRFD-500m, the remainder being the
|
||||
quantise/dequantise at the graph's edges — verified from the partition log, not inferred from the
|
||||
timing.
|
||||
|
||||
Also tried and rejected: **NNAPI** — the device registers no neural-networks HAL at all, so the
|
||||
provider has nothing to talk to; Google deprecated it in Android 15 and Qualcomm stopped shipping
|
||||
drivers for it. **XNNPACK** — slower than ORT's default CPU kernels on every model that loaded, and
|
||||
aborts inside its partitioner on the SCRFD graphs.
|
||||
|
||||
### 1.2 The desktop — RTX 3050 Laptop, Raptor Lake, 20 threads
|
||||
|
||||
| Model | **tract** (today) | ORT CPU f32 | ORT CPU int8 | CUDA f32 | CUDA fp16 | TensorRT f32 | TensorRT fp16 | TensorRT int8 |
|
||||
|---|---|---|---|---|---|---|---|---|
|
||||
| scrfd_500m | 104 | 12 | 8 | 5.6 | 3.7 | 2.5 | **1.8** | ✗ ⁴ |
|
||||
| scrfd_2.5g | 162 | 27 | 12 | 6.2 | 5.3 | 2.9 | **1.9** | ✗ ⁴ |
|
||||
| scrfd_10g | 514 | 99 | 32 | 14.5 | 9.3 | 7.6 | **3.3** | ✗ ⁴ |
|
||||
| arcface_mbf | 45 | 15 | 18 | 1.1 | 0.8 | 0.9 | 0.7 | ✗ ⁴ |
|
||||
| yolo26n-seg | 307 | 60 | 38 | 9.3 | ✗ ³ | 6.8 | **5.5** | ✗ ⁴ |
|
||||
| yolo26s-sem-ade20k | 395 | 61 | 31 | 10.7 | ✗ ³ | 7.9 | **3.4** | ✗ ⁴ |
|
||||
|
||||
³ The offline fp16 conversion (`onnxconverter-common`) left a mixed-type node the CUDA provider
|
||||
rejects. TensorRT converts to fp16 itself at engine build and does not have this problem, which is
|
||||
one reason it is the target and the CUDA provider is the fallback.
|
||||
⁴ TensorRT refuses the QDQ form ONNX Runtime's quantiser writes for the Hexagon (uint8
|
||||
activations); it wants symmetric int8. Not pursued: fp16 needs no quantisation, no calibration and
|
||||
no accuracy gate, and it is already 30–60× tract.
|
||||
|
||||
CUDA int8 is deliberately absent: the CUDA provider has no int8 kernels and runs a QDQ graph by
|
||||
dequantising it, which measured *slower* than f32 (7.0 vs 5.6 ms on scrfd_500m). Int8 on NVIDIA is
|
||||
TensorRT's job.
|
||||
|
||||
**TensorRT's first load is 16–116 s per model in f32 and 35–290 s in fp16** (yolo26n-seg the
|
||||
worst: nearly five minutes), because it is compiling an engine for this exact GPU. The engine caches to disk and the second load is milliseconds. That
|
||||
number is what §6 is designed around.
|
||||
|
||||
### 1.3 The desktop — Radeon RX 7900 XT, Threadripper 2920X, 24 threads · 2026-09-20
|
||||
|
||||
Arch's `onnxruntime-rocm` 1.29.0 against ROCm 7.2.4 and MIGraphX 7.2.3 (gfx1100), through the
|
||||
same `ep_probe` harness (`core/dr-inference-engine/examples/ep_probe.rs`). Zero input, three
|
||||
warm-ups, the median of 15 runs, on a machine doing nothing else. The build columns are the wall
|
||||
clock of `Session` construction: cold is a MIGraphX compile of the graph for this GPU, cached is the
|
||||
same session loading the program the cold build wrote.
|
||||
|
||||
| Model | ORT CPU f32 | MIGraphX f32 | **MIGraphX fp16** | Compile f32 / fp16 (s) | Cached load (s) |
|
||||
|---|---|---|---|---|---|
|
||||
| scrfd_500m (Fast) | 10.4 | 2.8 | **2.4** | 40 / 58 | 0.3 |
|
||||
| scrfd_2.5g (Balanced) | 20.7 | 3.3 | **2.8** | 37 / 40 | 0.3 |
|
||||
| scrfd_10g (Thorough) | 57.9 | 4.5 | **3.4** | 40 / 48 | 0.4 |
|
||||
| arcface_mbf (per face) | 12.8 | 1.8 | 1.6 | 15 / 21 | 0.4 |
|
||||
| yolo26n-seg | 49.3 | 8.4 | **7.5** | 110 / 136 | 0.9 |
|
||||
| yolo26s-sem-ade20k | 55.8 | 4.8 | **3.8** | 50 / 60 | 0.5 |
|
||||
| 2d106det (landmarks) | 9.9 | 1.2 | 1.0 | 16 / 21 | 0.2 |
|
||||
| ocec_s (eye state) | 2.5 | 0.5 | 0.4 | 17 / 17 | 0.1 |
|
||||
| xfeat-1024 | 26.8 | 10.5 | 9.9 | 37 / 53 | 0.3 |
|
||||
| migan-512 (per tile) | 514 | 12.7 | **8.3** | 102 / 132 | 0.8 |
|
||||
|
||||
Three things the table settles.
|
||||
|
||||
- **The ROCm execution provider does not exist any more.** It was ONNX Runtime's CUDA-provider twin
|
||||
for AMD, removed in the 1.23 release (AMD's builds dropped it from ROCm 7.1); 1.29's ROCm build ships
|
||||
`libonnxruntime_providers_migraphx.so` and nothing else for AMD, and asking for `ROCm` answers
|
||||
"not enabled in this build". So there is no non-compiling AMD rung to sit under MIGraphX the way
|
||||
the CUDA provider sits under TensorRT: the AMD ladder is MIGraphX, then the CPU.
|
||||
- **MIGraphX is a compiling provider, and its cache has to be asked for by name.** 15–135 s per
|
||||
graph cold, under a second from its cache — TensorRT's shape exactly, and §6's design covers it.
|
||||
Two things the provider does that the code has to know: `ort`'s builder fills the legacy options
|
||||
struct, which 1.29 reads for the precision flags only, so the cache directory
|
||||
(`migraphx_model_cache_dir`) reaches it only through the generic key/value registration; and
|
||||
the cache key is the graph, the GPU and the MIGraphX version *without the precision*, so an fp16
|
||||
session pointed at the f32 program's directory silently loads the f32 program (the first fp16
|
||||
row measured here was that, before the directories were split).
|
||||
- **fp16 is worth 10–35% over f32 on this card, not the 3× it is worth on TensorRT**, because
|
||||
MIGraphX f32 is already 3–8× the CPU provider and the small graphs are launch-bound. The
|
||||
detectors at 2.4–3.4 ms sit beside TensorRT fp16's 1.8–3.3 ms on the RTX 3050; the embedder is
|
||||
1.6–1.8 ms on either precision and stays f32 (§7). The whole face pipeline for one image
|
||||
(detector + landmarks + eyes + embedder) is under 6 ms.
|
||||
|
||||
### 1.4 What the numbers say
|
||||
|
||||
- **`tract` is single-threaded.** The tablet's one X4 core and one Raptor Lake core give the same
|
||||
tract numbers. Replacing it with ONNX Runtime's CPU provider, *no accelerator involved*, is 3–6×
|
||||
on the tablet and 8–10× on the desktop. That is the floor, and it is available on every platform
|
||||
the app builds for.
|
||||
- **The Hexagon is the standout.** A 6 W NPU running int8 beats a discrete RTX 3050 running f32 on
|
||||
five of six models. A whole-library face index on the tablet goes from ~100 ms + 39 ms per face
|
||||
to ~1.4 ms + 12 ms per face, and the "Thorough" detector — 3× the cost of "Fast" today — becomes
|
||||
free. Its price is that the models must be **quantised to int8**, which is an accuracy question
|
||||
§5 has to answer before it is believed.
|
||||
- **The embedder does not gain from either accelerator.** 112×112 input, per-op overhead
|
||||
dominates; it is 9 ms on the tablet's CPU and 12 ms on its NPU. It stays float, which §7 turns
|
||||
from a performance footnote into a correctness rule.
|
||||
- **On NVIDIA, TensorRT fp16 ≈ 3× the CUDA provider**, and the CUDA provider ≈ 2× the
|
||||
multi-threaded CPU; at fp16 the detectors are 1.8–3.3 ms with no quantisation at all. Both leave the twenty cores free for decoding during a batch index, which the table does not
|
||||
show and which matters more than the ratio.
|
||||
- **On AMD, MIGraphX fp16 is 4–17× the CPU provider** on the detectors and 60× on the
|
||||
inpainter, with the same first-run compile cost as TensorRT and no rung between it and the CPU.
|
||||
|
||||
---
|
||||
|
||||
## 2. The shape of the answer
|
||||
|
||||
A **ladder per platform**, walked at start-up, with the first rung that builds a real session
|
||||
winning:
|
||||
|
||||
| Platform | 1st | 2nd | 3rd | Floor |
|
||||
|---|---|---|---|---|
|
||||
| Android, Qualcomm with a Hexagon the shipped QNN skel covers (V68–V81) | QNN HTP, int8 model | ORT CPU, f32 model | — | tract |
|
||||
| Android, any other SoC | ORT CPU, f32 | — | — | tract |
|
||||
| Linux / Windows, NVIDIA GPU | TensorRT, f32 model, fp16 engine | CUDA provider, f32 | ORT CPU, f32 | tract |
|
||||
| Linux, AMD GPU with ROCm | MIGraphX, f32 model, fp16 program | ORT CPU, f32 | — | tract |
|
||||
| Linux / Windows, no GPU stack | ORT CPU, f32 | — | — | tract |
|
||||
| macOS ⁵ | ORT CPU, f32 | — | — | tract |
|
||||
|
||||
⁵ CoreML is the obvious rung and is unmeasured; it is listed so its absence is a gap and not an
|
||||
oversight.
|
||||
|
||||
Deliberately **not** on any ladder, with the measurement that excluded each: NNAPI (no driver),
|
||||
XNNPACK (slower than CPU, aborts on SCRFD), WebGPU (slower than CPU), the Adreno through QNN (works,
|
||||
but never where the Hexagon does not also), CUDA int8 (slower than CUDA f32), the ROCm provider
|
||||
(gone: §1.3). A rung is added to this table by a measurement on this page, not by a provider
|
||||
existing.
|
||||
|
||||
The AMD ladder has no middle rung. TensorRT falls back to the CUDA provider while its engines
|
||||
compile; MIGraphX has no such twin, so its fallback is the CPU provider, and the minute or two of
|
||||
compiling on first run (§6) is spent at the floor's speed rather than at half the GPU's.
|
||||
|
||||
Two things the ladder is *not*: it is not a per-model choice — one backend serves every model on a
|
||||
device, because §7's identity rule needs the detector and embedder on the same runtime for the
|
||||
same reason `faces.model_id` pairs them; and it is not a per-account choice — it is a property of
|
||||
the hardware, like `shared_face_models_dir` is, and it lives beside it.
|
||||
|
||||
---
|
||||
|
||||
## 3. The dependency policy, and how far this reopens it
|
||||
|
||||
D13 chose `ort` over `tract` because `alternative-backend` made ONNX Runtime's *API* available with
|
||||
none of its *C*. Every rung above the floor needs the C++ ONNX Runtime and, for the two that matter
|
||||
most, vendor libraries on top: Qualcomm's QNN runtime (~60 MB for one Hexagon generation; it is
|
||||
per-SoC) and NVIDIA's TensorRT plus cuDNN (~600 MB with the CUDA libraries, and cuDNN's major
|
||||
version must match what ONNX Runtime was built against — the Arch package on the reference desktop
|
||||
was unusable for exactly that reason).
|
||||
|
||||
The policy protected the **build**: no C to cross-compile under the NDK, no toolchain to keep in
|
||||
step. This document keeps that intact, and the mechanism is the one thing about `ort` that makes it
|
||||
possible:
|
||||
|
||||
**`ort::set_api` accepts any `OrtApi` table.** With `alternative-backend` on, `ort` links nothing
|
||||
and asks for the table once per process. The application can `dlopen` a `libonnxruntime.so` it
|
||||
finds on disk, call `OrtGetApiBase()->GetApi(version)` and hand that table over; or, if there is no
|
||||
such file, hand over `ort_tract::api()`. The Rust build is identical in both cases — pure Rust,
|
||||
`cargo build --target aarch64-linux-android` sees the same dependency graph it sees today. What
|
||||
changes is that the runtime is a **file the package installs**, next to the models, and the app
|
||||
looks for it at start-up.
|
||||
|
||||
Consequences that follow and are accepted:
|
||||
|
||||
- **The runtime is chosen once per process**, because `set_api` is once per process. The ladder
|
||||
in §2 is walked at start-up and the result is what every session in that process uses. There is
|
||||
no "tract for this model, ORT for that one", and there is no falling back to tract *after* ONNX
|
||||
Runtime has loaded — but there does not need to be: once the library loads, its CPU provider is
|
||||
always there, and every fallback the ladder needs is between providers *inside* it.
|
||||
- **Feature flags stay as they are.** `dr-face`'s `inference` and `dr-segment`'s `semantic`
|
||||
continue to mean "compiled against `ort`'s API"; nothing at build time knows or cares which table
|
||||
will be supplied. The one addition is a `native-probe` feature on the new crate (§8) that pulls in
|
||||
`libloading`, which is pure Rust and already in the tree via `wgpu`.
|
||||
- **The packagers ship the runtime, not the build.** The Arch package, the Flatpak manifest, the
|
||||
NSIS installer and `assemble-apk.sh` each gain the ONNX Runtime library for their platform, and
|
||||
the Android and NVIDIA variants gain the vendor libraries — each under the licence the packager
|
||||
reads first (§3.1). A package without them is not broken; it is the tract build, and it says so
|
||||
on the about screen.
|
||||
- **The NDK problem does not come back.** `libonnxruntime.so` for Android is a prebuilt from
|
||||
Maven (`com.microsoft.onnxruntime:onnxruntime-android-qnn`), extracted by `assemble-apk.sh` into
|
||||
`jniLibs/` the way the models are bundled as assets today. Nothing compiles it.
|
||||
|
||||
### 3.1 Licences the packagers read before shipping a runtime
|
||||
|
||||
Written down now, because [segmentation.md §7](segmentation.md) established that reading the grant
|
||||
is cheaper than discovering it at packaging time.
|
||||
|
||||
| Component | Licence | Redistributable in a self-distributed package? |
|
||||
|---|---|---|
|
||||
| ONNX Runtime | MIT | Yes |
|
||||
| Qualcomm QNN runtime (`com.qualcomm.qti:qnn-runtime` on Maven) | Qualcomm AI Engine Direct SDK licence — proprietary, redistribution permitted for applications using it | Yes for the APK, with the licence text shipped; not for a source distribution. **To be read in full, not summarised from memory, before the APK gains it.** |
|
||||
| CUDA runtime, cuDNN, TensorRT | NVIDIA EULAs — redistributable with an application, with the licence text, not modifiable | Yes for a package that bundles them. 600 MB. The alternative is to load them from the user's system install if present and skip the rung otherwise — which is what §4's probe does anyway. |
|
||||
| ROCm (HIP, MIOpen, rocBLAS), MIGraphX | MIT | Yes, but the HIP SDK MIGraphX needs is ~15 GB installed. Same answer as NVIDIA: the user's system install, or the rung is skipped. |
|
||||
|
||||
The position this takes: the **GPU vendors' libraries are not bundled**. The desktop package probes for a
|
||||
system CUDA/TensorRT or ROCm/MIGraphX install and uses it if it is version-compatible; a desktop without one runs on
|
||||
ORT CPU, which is still 8–10× today. Bundling 600 MB for a rung that is 2× again is not a trade
|
||||
worth making unmeasured, and it can be revisited by a measurement on a batch index. The **QNN
|
||||
runtime is bundled** in the APK, because the Hexagon is the difference between a tablet that
|
||||
indexes a library overnight and one that does it over lunch, and the package is 60 MB larger for
|
||||
it.
|
||||
|
||||
Both positions are D13 territory and are recorded there (§12).
|
||||
|
||||
---
|
||||
|
||||
## 4. Selection — the probe, its cache, and what it may not do
|
||||
|
||||
**A rung is chosen by building a real session on it, not by asking whether it exists.** Both
|
||||
failure modes that are not "the provider is absent" were hit on 2026-09-19: a driver in a wedged
|
||||
state where the provider registered and the session then failed, and a provider that registered,
|
||||
took the graph, and rejected every node at partition time. The probe therefore:
|
||||
|
||||
1. Loads the runtime library (§3), or falls to tract and stops.
|
||||
2. Times the **smallest detector** on the CPU provider first — the floor. Then, for each rung
|
||||
in this platform's ladder, in order: builds a session for the same model on that provider
|
||||
with `error_on_failure`, runs it once on a fixed input, and times three more runs. **The rung
|
||||
is taken only if its median beats the floor.** That one measurement is the proof the provider
|
||||
took the graph: one that silently hands the work to the CPU is the CPU rung with extra
|
||||
overhead, slower than the floor, and rejected. (ONNX Runtime's
|
||||
`session.disable_cpu_ep_fallback` was the first draft of this proof and refuses the Hexagon
|
||||
over the ten quantise/dequantise nodes at the graph's edges that QNN declines by policy.)
|
||||
3. Records the outcome — rung, runtime version, provider version, device identity (GPU name and
|
||||
compute capability; SoC model and Hexagon arch), and the models' content hashes — to a small
|
||||
file beside `shared_face_models_dir`. The next start-up trusts the file **unless** any of those
|
||||
inputs changed, in which case it probes again. A driver update, a runtime update, a new model
|
||||
file: each invalidates the cache by construction, and none needs a "reset backend" button.
|
||||
|
||||
What the probe may not do:
|
||||
|
||||
- **Block the first frame.** It runs on the same background as `install_bundled_models` and for
|
||||
the same reason: a TensorRT probe can take thirty seconds cold, and a tablet that stalls that long
|
||||
is an ANR. Until it reports, every model request is answered by the floor the runtime supports
|
||||
(ORT CPU if the library loaded, tract otherwise), and a job that started on the floor finishes
|
||||
on it — a backend does not change under a running index.
|
||||
- **Retry a rung that failed within a session.** A failed probe is cached as a failure with the
|
||||
same inputs; the rung is tried again when an input changes. Otherwise a wedged driver means a
|
||||
thirty-second stall on every launch.
|
||||
- **Choose for the user without saying so.** Settings gains one row, *Inference backend*, showing
|
||||
what was chosen and why in one line ("Hexagon NPU · int8 · QNN 2.42"; "CPU · ONNX Runtime 1.30 ·
|
||||
TensorRT probe failed: cuDNN 8 required"), with an override to force any lower rung. The about
|
||||
screen carries the same line beside the model names NFR-SEC-5 already puts there.
|
||||
|
||||
---
|
||||
|
||||
## 5. Model variants, and who makes them
|
||||
|
||||
Every model exists in one **canonical** form — the f32 ONNX file the app ships or the user supplies
|
||||
today — and, where a rung needs it, a **derived** form. The ladder's rungs are specified in terms
|
||||
of which form they load:
|
||||
|
||||
| Form | Who produces it | When | Needed by |
|
||||
|---|---|---|---|
|
||||
| f32 ONNX, shape-fixed, **opset ≥ 13** | `tools/fix-face-model-shapes.sh`, `tools/export-seg-model.sh` | Release time, once | Every rung except Hexagon |
|
||||
| int8 QDQ ONNX, per-channel, uint8 activations | `tools/quantise-models.sh` (new) | Release time, once, **calibrated on real photographs** | Hexagon |
|
||||
| TensorRT engine (`.engine`, per GPU architecture and TensorRT version) | The app, from the f32 file | First run on that device, in the background | TensorRT rung |
|
||||
| MIGraphX program (`.mxr`, per GPU architecture, MIGraphX version and precision) | The app, from the f32 file | First run on that device, in the background | MIGraphX rung |
|
||||
| QNN context binary | The app, from the int8 file | First run on that device, in the background | Hexagon rung |
|
||||
|
||||
Two rules.
|
||||
|
||||
**Quantisation is a release-time step, not a device-time one.** The int8 files that produced §1's
|
||||
numbers were calibrated on random noise, which is enough to time and worthless to trust. A real
|
||||
int8 detector is calibrated on a few hundred real photographs and then measured against the f32
|
||||
detector on the reference library by [faces.md §12.3](faces.md)'s method — faces found, per size
|
||||
band, per detector — before it ships. That needs the reference library and a person reading the
|
||||
result, and it happens once per model release, in `tools/`, beside the shape-fixing it already
|
||||
depends on. The device never quantises anything.
|
||||
|
||||
The SCRFD and ArcFace files are **opset 11** as InsightFace exported them, and per-channel QDQ needs
|
||||
13; `tools/fix-face-model-shapes.sh` gains an opset upgrade to 17 (`onnx.version_converter`,
|
||||
`ir_version` 8), which `tract` has been verified to load and which every provider on this page
|
||||
prefers. That is a change to the canonical file and so a change to the shipped models, and it
|
||||
happens in the same model release as the int8 files.
|
||||
|
||||
**Compilation is a device-time step, and it is cached.** A TensorRT engine is specific to the GPU
|
||||
it was built on and the TensorRT that built it; a MIGraphX program to the GPU and the MIGraphX
|
||||
that built it; a QNN context binary to the Hexagon generation. None can ship. All are built by
|
||||
the app the first time that rung is selected, in the background (§6), and written beside the
|
||||
probe cache keyed by the same inputs. MIGraphX's own key leaves out the precision, so the app
|
||||
gives its f32 and fp16 programs separate directories — otherwise the embedder's f32 build would
|
||||
be served the detector's fp16 program, or the reverse. They are
|
||||
**derived, disposable, and regenerable**: deleting the cache directory costs the next launch a
|
||||
rebuild and nothing else, and the directory is excluded from anything that syncs (it is a peer of
|
||||
`thumbs`, not of the catalog).
|
||||
|
||||
---
|
||||
|
||||
## 6. First run — building engines without the user waiting for them
|
||||
|
||||
The sequence on a device where a compiling rung (TensorRT, MIGraphX, Hexagon) is selected:
|
||||
|
||||
1. **Launch.** The runtime loads; the probe (§4) starts in the background; the app serves every
|
||||
model request from the floor. Face indexing, segmentation and scene grading all work, at
|
||||
today's speed or better (ORT CPU).
|
||||
2. **Probe reports** — say, TensorRT. The compiling rung is now *selected* but has **no engines**.
|
||||
Model requests continue on the fallback rung below it (CUDA provider for TensorRT; ORT CPU for
|
||||
MIGraphX and Hexagon), which needs no compilation and is already faster than the floor.
|
||||
3. **Engines build**, one model at a time, on a single low-priority background thread, smallest
|
||||
model first so the detector — the one that runs per image — is ready soonest. On the reference
|
||||
desktop that is ~1 minute for the first detector and ~10 minutes for all six at fp16; on the AMD
|
||||
desktop 40 s for the first detector and ~8 minutes for the set; on the tablet the QNN
|
||||
context binaries take 0.8–1.7 s each and the whole set is ready before the user has opened a
|
||||
library. Each engine is written to a temporary name and renamed into place, so a request never
|
||||
sees a half-written file.
|
||||
4. **Requests move up as engines land.** A model whose engine exists loads it on the selected
|
||||
rung; one whose engine is still building loads on the fallback. **A running job does not
|
||||
switch** — an index that started on the CUDA provider finishes on it — because §7 needs one
|
||||
`model_id` per job, and because a job is the wrong granularity for surprise.
|
||||
5. On Android, the build runs only while the app is in the foreground and the device is not in
|
||||
battery saver (NFR-RES-3): a context binary takes a second, so this costs nothing, and the rule
|
||||
exists for the day a model takes longer.
|
||||
|
||||
Settings shows a one-line progress row while engines build ("Preparing GPU engines · 3 of 6") and
|
||||
nothing when they are done. A build failure demotes the rung — it is recorded in the probe cache
|
||||
as a failure with the model hash as an input, so a corrected model file retries it — and the app
|
||||
carries on one rung down, saying so in the same row.
|
||||
|
||||
---
|
||||
|
||||
## 7. Identity — what changes `model_id` and what may not
|
||||
|
||||
`faces.model_id` exists so that two libraries indexed with different networks are never compared
|
||||
as if they were one ([catalog.md §10.1](catalog.md); the trap is written up in
|
||||
[faces.md §14](faces.md)). A backend that changes what a network *computes* is a different network
|
||||
and must be a different `model_id`; one that changes only *where* it computes it must not be.
|
||||
|
||||
**The detector.** An int8 SCRFD finds a different set of faces from the f32 one — that is what
|
||||
§5's acceptance measures — so **the numeric form is part of the detector's identity**:
|
||||
`scrfd_500m` and `scrfd_500m_i8` are two detectors in `model_id`, and a library indexed on the
|
||||
tablet's Hexagon and continued on the desktop is two populations, which a re-index on either side
|
||||
reconciles the same way a switch from Fast to Thorough does today. That is acceptable because it is
|
||||
already the rule for the detector and because §5 is the gate on whether the int8 form is close
|
||||
enough to be *offered* at all. f32 on tract, ORT CPU, CUDA and TensorRT-f32 are one identity: the
|
||||
same graph, the same arithmetic, differences at the last bit.
|
||||
|
||||
**The embedder** is where comparability across devices is the whole point, and it is the one
|
||||
model that no accelerator helps (§1.4). So: **the embedder runs in f32 on every rung.** On TensorRT
|
||||
that means the embedder's engine is built without fp16 while the detector's is built with it; on
|
||||
the Hexagon it means the embedder is not on the NPU at all — it runs on the ORT CPU rung at 9 ms,
|
||||
and the ladder's "one backend per device" is, precisely, one backend *per model role*, with the
|
||||
embedder pinned. A `w600k_mbf` embedding from any device is comparable with one from any other,
|
||||
which is the property the identity system, the calibration and the cross-device merge all rest
|
||||
on, and it is not for sale for 3 ms.
|
||||
|
||||
If S16 wants fp16 for the embedder later, the gate is written now: over the reference library's
|
||||
faces, the cosine between the f32 and fp16 embedding of the same crop exceeds 0.999 for 99.9% of
|
||||
faces and the calibration's fitted threshold moves by less than its own confidence interval
|
||||
([faces.md §8.3](faces.md)). Until measured, f32.
|
||||
|
||||
**Segmentation and the scene model** carry no identity across devices — their outputs are
|
||||
recomputed per image and never stored beyond the cache — so they take whatever the rung offers,
|
||||
int8 included, subject to §10's own acceptance.
|
||||
|
||||
---
|
||||
|
||||
## 8. Crate shape — `core/dr-inference-engine`
|
||||
|
||||
The seam is the same shape as [storage.md](storage.md)'s: a small crate below the consumers that is
|
||||
the **only** place naming a provider, a library file or a vendor, with the consumers reduced to
|
||||
"give me a session for these bytes in this role".
|
||||
|
||||
```
|
||||
core/dr-inference-engine
|
||||
src/lib.rs Runtime (Tract | Onnx { lib, version }), Backend (rung), Role (Detector | Embedder | Segmenter)
|
||||
src/probe.rs §4 — the ladder per platform, the session-build probe, the cache file
|
||||
src/engines.rs §6 — background compilation, the cache directory, progress
|
||||
src/session.rs open(role, bytes) -> ort::Session, applying the rung and the role's precision rule
|
||||
src/api.rs the unsafe block that matters: dlopen libonnxruntime, fetch OrtApi, ort::set_api — or ort_tract::api()
|
||||
```
|
||||
|
||||
- `dr-face` and `dr-segment` **delete** their private `install_backend` and their direct
|
||||
`Session::builder()` calls and take an `&dr_inference_engine::Sessions` where they take model bytes today.
|
||||
Their tests keep `tract` — `dr_inference_engine::Sessions::tract()` is a constructor and the test-only path.
|
||||
- `dr-inference-engine` depends on `ort` with the same workspace features as today plus `cuda`, `tensorrt`,
|
||||
`qnn`: those features add option builders, not linking, under `alternative-backend`. **Verified
|
||||
for the QNN, CUDA and TensorRT builders on 2026-09-19** — they go through the API table's generic
|
||||
`SessionOptionsAppendExecutionProvider*`. The NNAPI builder resolves a symbol directly and would
|
||||
not; it is not needed and is not enabled. MIGraphX uses no `ort` feature at all: `ort`'s builder
|
||||
fills the legacy options struct, which ONNX Runtime 1.29 reads for the precision flags and
|
||||
nothing else, and the compiled-program cache directory only travels through the generic
|
||||
key/value entry point (`migraphx_model_cache_dir`). `session::migraphx` makes that one call
|
||||
on the API table itself.
|
||||
- `dr-ui` owns the settings row, the about-screen line and the progress row; it holds one
|
||||
`Sessions` per process, created at launch, and passes it down. `dr_ui::library` gains
|
||||
`inference_cache_dir()` beside `shared_face_models_dir()`, on the same account-independent
|
||||
footing and for the same reason.
|
||||
- The Android entry point's `install_bundled_models` also extracts nothing new: `jniLibs/` is
|
||||
loaded by the system loader, and `dr-inference-engine` on Android looks for `libonnxruntime.so` through
|
||||
`dlopen` by bare name first, which resolves to the APK's copy, before any directory.
|
||||
|
||||
---
|
||||
|
||||
## 9. Threads and memory
|
||||
|
||||
- ONNX Runtime's intra-op pool is sized to the physical cores minus two on desktop and to the
|
||||
performance cores on Android (the X4 and the A720s; the A520s are for the compositor). tract's
|
||||
single thread today is the reason a batch index leaves nineteen cores idle; ORT CPU with the
|
||||
pool is the reason it will not. One session per model per process; `Session::run` is
|
||||
`&mut self`-free in `ort` and internally serialised, and the index job is the only caller.
|
||||
- A TensorRT session pins GPU memory for its workspace; the builder is capped at 512 MB on the
|
||||
reference 6 GB card and the cap is a setting, because the develop view's tiles share the card
|
||||
(NFR-RES-2). The engine cache on disk is bounded by the model set — six engines, ~80 MB — and
|
||||
needs no LRU.
|
||||
- The Hexagon rung sets QNN's performance mode to `Burst` for the duration of an index job and
|
||||
`Default` otherwise; a 5 ms detector does not need the NPU clocked up between images.
|
||||
- The probe (§4) and the engine build (§6) run on one dedicated low-priority thread. They never
|
||||
share the index job's pool: a probe that competes with the job it is meant to speed up is the
|
||||
frame-budget trap in a new coat.
|
||||
|
||||
---
|
||||
|
||||
## 10. What S16 measures
|
||||
|
||||
In order, with the gate each is:
|
||||
|
||||
| # | Question | Gate |
|
||||
|---|---|---|
|
||||
| M1 | Does one binary carry both tables? `dlopen` + `set_api` on Linux, Windows and Android; `ort_tract::api()` when the file is absent. | Go / no-go for §3. If `set_api` cannot take a dynamically fetched table on some platform, that platform ships two binaries, and the cost is stated. |
|
||||
| M2 | Do the int8 SCRFD detectors, **calibrated on real photographs**, find the faces? faces.md §12.3's method over the reference library, per size band, against f32. | Ship the int8 form for a detector only if it finds ≥ 97% of the f32 detector's faces above 40 px and the difference is not concentrated in one band. Otherwise that detector's Hexagon rung is ORT CPU int8-free, and the table in §1.1 says what that costs. |
|
||||
| M3 | Does the embedder on ORT CPU beside a detector on the Hexagon (§7) produce embeddings within the f32 gate? | It must — same graph, same arithmetic. This is a check that the plumbing did not quantise it by accident. |
|
||||
| M4 | Is a batch index on the tablet and on the desktop faster by the ratio §1 predicts, end to end, decode included? | The face index over the reference library (18,143 faces): report wall-clock on tract, on the floor and on the selected rung, and where the time went. The prediction is that decode becomes the bottleneck on both; if it does not, say why. |
|
||||
| M5 | Does the first-run sequence (§6) hold: nothing blocks the first frame, engines land, requests move up, a running job does not switch? | Observed on both devices with the app's own progress row, and with the cache directory deleted between runs. |
|
||||
| M6 | What does the APK weigh with the QNN runtime, and does a non-Qualcomm Android device (any one) still launch and index on the floor? | Size reported; launch verified on one non-Qualcomm device or an emulator. |
|
||||
| M7 | The segmentation and scene models at int8 on the Hexagon: does the mask boundary move? segmentation.md §6's IoU against f32 over its corpus. | Ship int8 for a model only above the IoU floor that document set for arm B. |
|
||||
|
||||
M1 and M2 are the ones the rest is conditional on, and M2 is the one that needs a person.
|
||||
|
||||
### 10.1 M2 result · 2026-09-19
|
||||
|
||||
Each int8 detector against its own f32 form, over 400 proxies evenly spaced through the reference
|
||||
library, on ONNX Runtime's CPU provider (the int8 graph is the same file the Hexagon loads;
|
||||
`ui/dr-ui/examples/face_detectors`). Calibrated on 64 proxies from the same library, disjoint
|
||||
from the 400.
|
||||
|
||||
| Detector | f32 faces | int8 faces | both | int8 only | f32 only | found ≥ 32 px | found, all sizes |
|
||||
|---|---|---|---|---|---|---|---|
|
||||
| scrfd_500m | 1342 | 1319 | 1287 | 32 | 55 | 95.6% | 95.9% |
|
||||
| scrfd_2.5g | 1525 | 1482 | 1478 | 4 | 47 | 96.1% | 96.9% |
|
||||
| scrfd_10g | 1769 | 1760 | 1749 | 11 | 20 | 100% | 98.9% |
|
||||
|
||||
The 10g form clears the 97% gate; 500m and 2.5g sit one point under it. What they lose is
|
||||
specific: the faces in the "f32 only" column have a **median confidence of 0.52** against a
|
||||
threshold of 0.50 — detections the f32 graph itself barely made, that int8 rounding drops to the
|
||||
other side of the line — and the extra faces int8 finds are the same kind (median 0.51–0.52).
|
||||
Not a size-band failure: the losses are spread across bands in proportion. Shipped as they are,
|
||||
with the number on record; a threshold of 0.48 for the int8 forms would recover most of the
|
||||
margin, and is the first thing to try if a library's count on the tablet reads low.
|
||||
|
||||
Two things the calibration taught, both in `tools/quantise-models.py`: the calibration set has
|
||||
to contain faces (a first attempt on landscape photographs produced a graph that found nothing —
|
||||
the score head's ranges had never seen the face regime), and ONNX Runtime's own strided and
|
||||
moving-average calibration modes both measurably degrade the result on these graphs, while
|
||||
driving the calibrator in chunks by hand reproduces the plain min/max ranges exactly.
|
||||
|
||||
---
|
||||
|
||||
## 11. Order
|
||||
|
||||
1. **`dr-inference-engine` with the two tables and the floor** — `set_api` from a dlopened runtime, tract
|
||||
otherwise, ORT CPU as the only rung. Consumers moved over; tests unchanged. This alone is the
|
||||
3–10× and is the build most of the value sits in. M1.
|
||||
2. **The probe and its cache** (§4), with the settings row and the about line. Still CPU-only;
|
||||
the ladder has one rung. M5's first half.
|
||||
3. **`tools/quantise-models.sh`** and the opset upgrade; the int8 detectors calibrated and
|
||||
measured. M2, M7. This is the step with a person in it and it runs in parallel with 4.
|
||||
4. **The Hexagon rung**, the QNN runtime in the APK, the context-binary cache. M3, M6.
|
||||
5. **The TensorRT and CUDA rungs** on desktop, the engine cache, the first-run sequence. M5's
|
||||
second half.
|
||||
6. **M4** last, on both devices, and the number goes in this document.
|
||||
|
||||
The Windows installer and the Flatpak manifest are touched in steps 1 and 5 only, and only to add
|
||||
a file each; the Arch package likewise.
|
||||
|
||||
---
|
||||
|
||||
## 12. Register entries
|
||||
|
||||
**FR-INF-1 — Runtime selection.** On launch the application shall determine, per device and
|
||||
without blocking the first frame, the fastest inference backend that can build and run a session
|
||||
for the shipped models, by attempting it; shall record and reuse that determination until the
|
||||
runtime, driver, hardware or models change; and shall display the backend in use in Settings and
|
||||
on the about screen. *Acceptance:* §10 M1 and M5.
|
||||
|
||||
**FR-INF-2 — Derived engines.** Backends that require device-specific compilation shall compile in
|
||||
the background after selection, shall serve requests from the next lower backend until each engine
|
||||
is ready, and shall not change the backend of a job in progress. *Acceptance:* M5.
|
||||
|
||||
**FR-INF-3 — Model forms.** Quantised model forms are produced at release time from real
|
||||
calibration data and are shipped only when they meet §10's accuracy gates against the canonical
|
||||
form; the application never quantises on the device. *Acceptance:* M2, M7.
|
||||
|
||||
**NFR-INF-1 — Embedding comparability.** Face embeddings shall be computed at a precision whose
|
||||
deviation from the f32 reference is within §7's gate, on every backend, so that embeddings from any
|
||||
device are comparable. *Acceptance:* M3.
|
||||
|
||||
**D13 — updated.** The runtime half is reopened to the extent of §3: the Rust build stays C-free
|
||||
under `alternative-backend`; packages may install a dynamically loaded ONNX Runtime and, per §3.1,
|
||||
the Qualcomm QNN runtime; the NVIDIA libraries are not bundled. The licensing half is unchanged.
|
||||
|
||||
---
|
||||
|
||||
## 13. Requirements touched
|
||||
|
||||
FR-CULL-8 (indexing time is what this exists to change), FR-CULL-9 (NFR-INF-1 is the guard on its
|
||||
calibration), NFR-RES-2 (TensorRT workspace against the develop view's tiles), NFR-RES-3
|
||||
(background compilation on Android), NFR-SEC-5 (the about screen's line), NFR-COMPAT-2 (each
|
||||
channel gains a runtime file), FR-PLAT-AND-1 (unchanged; the runtime lives in the APK, not in
|
||||
storage), ARCH §6.1 (the once-per-image budget this was sized against no longer binds; what could
|
||||
run per frame is a separate question this document does not open).
|
||||
Reference in New Issue
Block a user