Files
DarkRoom/docs/inference.md
T

455 lines
30 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Inference backends — the runtime and the model, chosen per device
Spec for **S16**, the build that puts the neural models on the hardware each device actually has.
Every model DarkRoom runs today — the three SCRFD detectors, the ArcFace embedder, YOLO26n-seg and
the ADE20K scene model — runs through `tract`, on one CPU core, on every platform. That was the
right first answer: D13's runtime half chose it because it costs no C dependency, and
[faces.md](faces.md) and [segmentation.md](segmentation.md) were written against it. It is also
between 20× and 300× slower than what the same devices can do, and this document is the record of
having measured that and the specification of what replaces it.
**It does not reopen D13's licensing half.** The weights are the same files under the same grant.
It does reopen the *runtime* half, and §3 is where it says how far.
---
## 1. What was measured · 2026-09-19
One benchmark, two builds of it: `ort`'s API over `tract` (exactly what the app links) and `ort`'s
API over a dynamically loaded ONNX Runtime with each execution provider in turn. Random 640×640
input, three warm-ups, the median of 15–30 timed runs, milliseconds. The same input every run, so
the numbers are compute cost and nothing else.
### 1.1 The tablet — Honor MagicPad 2, Snapdragon 8s Gen 3
SM8635: 1× Cortex-X4, 4× A720, 3× A520, Adreno 735, Hexagon V73. Android 16.
| Model | **tract** (today) | ORT CPU f32 | ORT CPU int8 | Adreno f32 ¹ | **Hexagon int8** ² |
|---|---|---|---|---|---|
| scrfd_500m (Fast) | 98 | 16 | 8 | 21 | **1.4** |
| scrfd_2.5g (Balanced) | 161 | 59 | 19 | ✗ | **1.8** |
| scrfd_10g (Thorough) | 489 | 204 | 48 | ✗ | **3.2** |
| arcface_mbf (per face) | 39 | 9 | 13 | 24 | 12 |
| yolo26n-seg | 287 | 94 | 38 | 54 | **5.1** |
| yolo26s-sem-ade20k | 408 | 154 | 46 | 47 | **3.7** |
¹ Qualcomm's own GPU backend (`libQnnGpu.so`, OpenCL). Fails on the two larger SCRFD graphs at an
`AveragePool` the layout transformer cannot place. ONNX Runtime's WebGPU provider also runs on this
GPU and was slower than the CPU on every model; it is not in the table because it is not a
candidate.
² QNN's HTP backend. The Hexagon **refuses float32 and float16 tensors** in this ORT 1.29 + QNN
2.42 pairing (error 3110 on every node, with `enable_htp_fp16_precision` set or not); int8 QDQ
graphs run with 99.6% of nodes on the NPU — 1718 of 1725 for SCRFD-500m, the remainder being the
quantise/dequantise at the graph's edges — verified from the partition log, not inferred from the
timing.
Also tried and rejected: **NNAPI** — the device registers no neural-networks HAL at all, so the
provider has nothing to talk to; Google deprecated it in Android 15 and Qualcomm stopped shipping
drivers for it. **XNNPACK** — slower than ORT's default CPU kernels on every model that loaded, and
aborts inside its partitioner on the SCRFD graphs.
### 1.2 The desktop — RTX 3050 Laptop, Raptor Lake, 20 threads
| Model | **tract** (today) | ORT CPU f32 | ORT CPU int8 | CUDA f32 | CUDA fp16 | TensorRT f32 | TensorRT fp16 | TensorRT int8 |
|---|---|---|---|---|---|---|---|---|
| scrfd_500m | 104 | 12 | 8 | 5.6 | 3.7 | 2.5 | **1.8** | ✗ ⁴ |
| scrfd_2.5g | 162 | 27 | 12 | 6.2 | 5.3 | 2.9 | **1.9** | ✗ ⁴ |
| scrfd_10g | 514 | 99 | 32 | 14.5 | 9.3 | 7.6 | **3.3** | ✗ ⁴ |
| arcface_mbf | 45 | 15 | 18 | 1.1 | 0.8 | 0.9 | 0.7 | ✗ ⁴ |
| yolo26n-seg | 307 | 60 | 38 | 9.3 | ✗ ³ | 6.8 | **5.5** | ✗ ⁴ |
| yolo26s-sem-ade20k | 395 | 61 | 31 | 10.7 | ✗ ³ | 7.9 | **3.4** | ✗ ⁴ |
³ The offline fp16 conversion (`onnxconverter-common`) left a mixed-type node the CUDA provider
rejects. TensorRT converts to fp16 itself at engine build and does not have this problem, which is
one reason it is the target and the CUDA provider is the fallback.
⁴ TensorRT refuses the QDQ form ONNX Runtime's quantiser writes for the Hexagon (uint8
activations); it wants symmetric int8. Not pursued: fp16 needs no quantisation, no calibration and
no accuracy gate, and it is already 30–60× tract.
CUDA int8 is deliberately absent: the CUDA provider has no int8 kernels and runs a QDQ graph by
dequantising it, which measured *slower* than f32 (7.0 vs 5.6 ms on scrfd_500m). Int8 on NVIDIA is
TensorRT's job.
**TensorRT's first load is 16–116 s per model in f32 and 35–290 s in fp16** (yolo26n-seg the
worst: nearly five minutes), because it is compiling an engine for this exact GPU. The engine caches to disk and the second load is milliseconds. That
number is what §6 is designed around.
### 1.3 What the numbers say
- **`tract` is single-threaded.** The tablet's one X4 core and one Raptor Lake core give the same
tract numbers. Replacing it with ONNX Runtime's CPU provider, *no accelerator involved*, is 3–6×
on the tablet and 8–10× on the desktop. That is the floor, and it is available on every platform
the app builds for.
- **The Hexagon is the standout.** A 6 W NPU running int8 beats a discrete RTX 3050 running f32 on
five of six models. A whole-library face index on the tablet goes from ~100 ms + 39 ms per face
to ~1.4 ms + 12 ms per face, and the "Thorough" detector — 3× the cost of "Fast" today — becomes
free. Its price is that the models must be **quantised to int8**, which is an accuracy question
§5 has to answer before it is believed.
- **The embedder does not gain from either accelerator.** 112×112 input, per-op overhead
dominates; it is 9 ms on the tablet's CPU and 12 ms on its NPU. It stays float, which §7 turns
from a performance footnote into a correctness rule.
- **On NVIDIA, TensorRT fp16 ≈ 3× the CUDA provider**, and the CUDA provider ≈ 2× the
multi-threaded CPU; at fp16 the detectors are 1.8–3.3 ms with no quantisation at all. Both leave the twenty cores free for decoding during a batch index, which the table does not
show and which matters more than the ratio.
---
## 2. The shape of the answer
A **ladder per platform**, walked at start-up, with the first rung that builds a real session
winning:
| Platform | 1st | 2nd | 3rd | Floor |
|---|---|---|---|---|
| Android, Qualcomm with a Hexagon the shipped QNN skel covers (V68–V81) | QNN HTP, int8 model | ORT CPU, f32 model | — | tract |
| Android, any other SoC | ORT CPU, f32 | — | — | tract |
| Linux / Windows, NVIDIA GPU | TensorRT, f32 model, fp16 engine | CUDA provider, f32 | ORT CPU, f32 | tract |
| Linux / Windows, no NVIDIA | ORT CPU, f32 | — | — | tract |
| macOS ⁵ | ORT CPU, f32 | — | — | tract |
⁵ CoreML is the obvious rung and is unmeasured; it is listed so its absence is a gap and not an
oversight.
Deliberately **not** on any ladder, with the measurement that excluded each: NNAPI (no driver),
XNNPACK (slower than CPU, aborts on SCRFD), WebGPU (slower than CPU), the Adreno through QNN (works,
but never where the Hexagon does not also), CUDA int8 (slower than CUDA f32). A rung is added to this
table by a measurement on this page, not by a provider existing.
Two things the ladder is *not*: it is not a per-model choice — one backend serves every model on a
device, because §7's identity rule needs the detector and embedder on the same runtime for the
same reason `faces.model_id` pairs them; and it is not a per-account choice — it is a property of
the hardware, like `shared_face_models_dir` is, and it lives beside it.
---
## 3. The dependency policy, and how far this reopens it
D13 chose `ort` over `tract` because `alternative-backend` made ONNX Runtime's *API* available with
none of its *C*. Every rung above the floor needs the C++ ONNX Runtime and, for the two that matter
most, vendor libraries on top: Qualcomm's QNN runtime (~60 MB for one Hexagon generation; it is
per-SoC) and NVIDIA's TensorRT plus cuDNN (~600 MB with the CUDA libraries, and cuDNN's major
version must match what ONNX Runtime was built against — the Arch package on the reference desktop
was unusable for exactly that reason).
The policy protected the **build**: no C to cross-compile under the NDK, no toolchain to keep in
step. This document keeps that intact, and the mechanism is the one thing about `ort` that makes it
possible:
**`ort::set_api` accepts any `OrtApi` table.** With `alternative-backend` on, `ort` links nothing
and asks for the table once per process. The application can `dlopen` a `libonnxruntime.so` it
finds on disk, call `OrtGetApiBase()->GetApi(version)` and hand that table over; or, if there is no
such file, hand over `ort_tract::api()`. The Rust build is identical in both cases — pure Rust,
`cargo build --target aarch64-linux-android` sees the same dependency graph it sees today. What
changes is that the runtime is a **file the package installs**, next to the models, and the app
looks for it at start-up.
Consequences that follow and are accepted:
- **The runtime is chosen once per process**, because `set_api` is once per process. The ladder
in §2 is walked at start-up and the result is what every session in that process uses. There is
no "tract for this model, ORT for that one", and there is no falling back to tract *after* ONNX
Runtime has loaded — but there does not need to be: once the library loads, its CPU provider is
always there, and every fallback the ladder needs is between providers *inside* it.
- **Feature flags stay as they are.** `dr-face`'s `inference` and `dr-segment`'s `semantic`
continue to mean "compiled against `ort`'s API"; nothing at build time knows or cares which table
will be supplied. The one addition is a `native-probe` feature on the new crate (§8) that pulls in
`libloading`, which is pure Rust and already in the tree via `wgpu`.
- **The packagers ship the runtime, not the build.** The Arch package, the Flatpak manifest, the
NSIS installer and `assemble-apk.sh` each gain the ONNX Runtime library for their platform, and
the Android and NVIDIA variants gain the vendor libraries — each under the licence the packager
reads first (§3.1). A package without them is not broken; it is the tract build, and it says so
on the about screen.
- **The NDK problem does not come back.** `libonnxruntime.so` for Android is a prebuilt from
Maven (`com.microsoft.onnxruntime:onnxruntime-android-qnn`), extracted by `assemble-apk.sh` into
`jniLibs/` the way the models are bundled as assets today. Nothing compiles it.
### 3.1 Licences the packagers read before shipping a runtime
Written down now, because [segmentation.md §7](segmentation.md) established that reading the grant
is cheaper than discovering it at packaging time.
| Component | Licence | Redistributable in a self-distributed package? |
|---|---|---|
| ONNX Runtime | MIT | Yes |
| Qualcomm QNN runtime (`com.qualcomm.qti:qnn-runtime` on Maven) | Qualcomm AI Engine Direct SDK licence — proprietary, redistribution permitted for applications using it | Yes for the APK, with the licence text shipped; not for a source distribution. **To be read in full, not summarised from memory, before the APK gains it.** |
| CUDA runtime, cuDNN, TensorRT | NVIDIA EULAs — redistributable with an application, with the licence text, not modifiable | Yes for a package that bundles them. 600 MB. The alternative is to load them from the user's system install if present and skip the rung otherwise — which is what §4's probe does anyway. |
The position this takes: the **NVIDIA libraries are not bundled**. The desktop package probes for a
system CUDA/TensorRT install and uses it if it is version-compatible; a desktop without one runs on
ORT CPU, which is still 8–10× today. Bundling 600 MB for a rung that is 2× again is not a trade
worth making unmeasured, and it can be revisited by a measurement on a batch index. The **QNN
runtime is bundled** in the APK, because the Hexagon is the difference between a tablet that
indexes a library overnight and one that does it over lunch, and the package is 60 MB larger for
it.
Both positions are D13 territory and are recorded there (§12).
---
## 4. Selection — the probe, its cache, and what it may not do
**A rung is chosen by building a real session on it, not by asking whether it exists.** Both
failure modes that are not "the provider is absent" were hit on 2026-09-19: a driver in a wedged
state where the provider registered and the session then failed, and a provider that registered,
took the graph, and rejected every node at partition time. The probe therefore:
1. Loads the runtime library (§3), or falls to tract and stops.
2. For each rung in this platform's ladder, in order: builds a session for the **smallest model
in the set** (`scrfd_500m`) on that provider with `error_on_failure`, runs it once on a fixed
input, and reads back the provider assignment from the session — the rung is taken only if the
provider ran **at least 95% of the graph's nodes**. A provider that silently hands the graph to
the CPU is the CPU rung with extra overhead, and the app should say "CPU".
3. Records the outcome — rung, runtime version, provider version, device identity (GPU name and
compute capability; SoC model and Hexagon arch), and the models' content hashes — to a small
file beside `shared_face_models_dir`. The next start-up trusts the file **unless** any of those
inputs changed, in which case it probes again. A driver update, a runtime update, a new model
file: each invalidates the cache by construction, and none needs a "reset backend" button.
What the probe may not do:
- **Block the first frame.** It runs on the same background as `install_bundled_models` and for
the same reason: a TensorRT probe can take thirty seconds cold, and a tablet that stalls that long
is an ANR. Until it reports, every model request is answered by the floor the runtime supports
(ORT CPU if the library loaded, tract otherwise), and a job that started on the floor finishes
on it — a backend does not change under a running index.
- **Retry a rung that failed within a session.** A failed probe is cached as a failure with the
same inputs; the rung is tried again when an input changes. Otherwise a wedged driver means a
thirty-second stall on every launch.
- **Choose for the user without saying so.** Settings gains one row, *Inference backend*, showing
what was chosen and why in one line ("Hexagon NPU · int8 · QNN 2.42"; "CPU · ONNX Runtime 1.30 ·
TensorRT probe failed: cuDNN 8 required"), with an override to force any lower rung. The about
screen carries the same line beside the model names NFR-SEC-5 already puts there.
---
## 5. Model variants, and who makes them
Every model exists in one **canonical** form — the f32 ONNX file the app ships or the user supplies
today — and, where a rung needs it, a **derived** form. The ladder's rungs are specified in terms
of which form they load:
| Form | Who produces it | When | Needed by |
|---|---|---|---|
| f32 ONNX, shape-fixed, **opset ≥ 13** | `tools/fix-face-model-shapes.sh`, `tools/export-seg-model.sh` | Release time, once | Every rung except Hexagon |
| int8 QDQ ONNX, per-channel, uint8 activations | `tools/quantise-models.sh` (new) | Release time, once, **calibrated on real photographs** | Hexagon |
| TensorRT engine (`.engine`, per GPU architecture and TensorRT version) | The app, from the f32 file | First run on that device, in the background | TensorRT rung |
| QNN context binary | The app, from the int8 file | First run on that device, in the background | Hexagon rung |
Two rules.
**Quantisation is a release-time step, not a device-time one.** The int8 files that produced §1's
numbers were calibrated on random noise, which is enough to time and worthless to trust. A real
int8 detector is calibrated on a few hundred real photographs and then measured against the f32
detector on the reference library by [faces.md §12.3](faces.md)'s method — faces found, per size
band, per detector — before it ships. That needs the reference library and a person reading the
result, and it happens once per model release, in `tools/`, beside the shape-fixing it already
depends on. The device never quantises anything.
The SCRFD and ArcFace files are **opset 11** as InsightFace exported them, and per-channel QDQ needs
13; `tools/fix-face-model-shapes.sh` gains an opset upgrade to 17 (`onnx.version_converter`,
`ir_version` 8), which `tract` has been verified to load and which every provider on this page
prefers. That is a change to the canonical file and so a change to the shipped models, and it
happens in the same model release as the int8 files.
**Compilation is a device-time step, and it is cached.** A TensorRT engine is specific to the GPU
it was built on and the TensorRT that built it; a QNN context binary is specific to the Hexagon
generation. Neither can ship. Both are built by the app the first time that rung is selected, in
the background (§6), and written beside the probe cache keyed by the same inputs. They are
**derived, disposable, and regenerable**: deleting the cache directory costs the next launch a
rebuild and nothing else, and the directory is excluded from anything that syncs (it is a peer of
`thumbs`, not of the catalog).
---
## 6. First run — building engines without the user waiting for them
The sequence on a device where a compiling rung (TensorRT, Hexagon) is selected:
1. **Launch.** The runtime loads; the probe (§4) starts in the background; the app serves every
model request from the floor. Face indexing, segmentation and scene grading all work, at
today's speed or better (ORT CPU).
2. **Probe reports** — say, TensorRT. The compiling rung is now *selected* but has **no engines**.
Model requests continue on the fallback rung below it (CUDA provider for TensorRT; ORT CPU for
Hexagon), which needs no compilation and is already faster than the floor.
3. **Engines build**, one model at a time, on a single low-priority background thread, smallest
model first so the detector — the one that runs per image — is ready soonest. On the reference
desktop that is ~1 minute for the first detector and ~10 minutes for all six at fp16; on the tablet the QNN
context binaries take 0.8–1.7 s each and the whole set is ready before the user has opened a
library. Each engine is written to a temporary name and renamed into place, so a request never
sees a half-written file.
4. **Requests move up as engines land.** A model whose engine exists loads it on the selected
rung; one whose engine is still building loads on the fallback. **A running job does not
switch** — an index that started on the CUDA provider finishes on it — because §7 needs one
`model_id` per job, and because a job is the wrong granularity for surprise.
5. On Android, the build runs only while the app is in the foreground and the device is not in
battery saver (NFR-RES-3): a context binary takes a second, so this costs nothing, and the rule
exists for the day a model takes longer.
Settings shows a one-line progress row while engines build ("Preparing GPU engines · 3 of 6") and
nothing when they are done. A build failure demotes the rung — it is recorded in the probe cache
as a failure with the model hash as an input, so a corrected model file retries it — and the app
carries on one rung down, saying so in the same row.
---
## 7. Identity — what changes `model_id` and what may not
`faces.model_id` exists so that two libraries indexed with different networks are never compared
as if they were one ([catalog.md §10.1](catalog.md); the trap is written up in
[faces.md §14](faces.md)). A backend that changes what a network *computes* is a different network
and must be a different `model_id`; one that changes only *where* it computes it must not be.
**The detector.** An int8 SCRFD finds a different set of faces from the f32 one — that is what
§5's acceptance measures — so **the numeric form is part of the detector's identity**:
`scrfd_500m` and `scrfd_500m_i8` are two detectors in `model_id`, and a library indexed on the
tablet's Hexagon and continued on the desktop is two populations, which a re-index on either side
reconciles the same way a switch from Fast to Thorough does today. That is acceptable because it is
already the rule for the detector and because §5 is the gate on whether the int8 form is close
enough to be *offered* at all. f32 on tract, ORT CPU, CUDA and TensorRT-f32 are one identity: the
same graph, the same arithmetic, differences at the last bit.
**The embedder** is where comparability across devices is the whole point, and it is the one
model that no accelerator helps (§1.3). So: **the embedder runs in f32 on every rung.** On TensorRT
that means the embedder's engine is built without fp16 while the detector's is built with it; on
the Hexagon it means the embedder is not on the NPU at all — it runs on the ORT CPU rung at 9 ms,
and the ladder's "one backend per device" is, precisely, one backend *per model role*, with the
embedder pinned. A `w600k_mbf` embedding from any device is comparable with one from any other,
which is the property the identity system, the calibration and the cross-device merge all rest
on, and it is not for sale for 3 ms.
If S16 wants fp16 for the embedder later, the gate is written now: over the reference library's
faces, the cosine between the f32 and fp16 embedding of the same crop exceeds 0.999 for 99.9% of
faces and the calibration's fitted threshold moves by less than its own confidence interval
([faces.md §8.3](faces.md)). Until measured, f32.
**Segmentation and the scene model** carry no identity across devices — their outputs are
recomputed per image and never stored beyond the cache — so they take whatever the rung offers,
int8 included, subject to §10's own acceptance.
---
## 8. Crate shape — `core/dr-inference-engine`
The seam is the same shape as [storage.md](storage.md)'s: a small crate below the consumers that is
the **only** place naming a provider, a library file or a vendor, with the consumers reduced to
"give me a session for these bytes in this role".
```
core/dr-inference-engine
src/lib.rs Runtime (Tract | Onnx { lib, version }), Backend (rung), Role (Detector | Embedder | Segmenter)
src/probe.rs §4 — the ladder per platform, the session-build probe, the cache file
src/engines.rs §6 — background compilation, the cache directory, progress
src/session.rs open(role, bytes) -> ort::Session, applying the rung and the role's precision rule
src/api.rs the one unsafe block: dlopen libonnxruntime, fetch OrtApi, ort::set_api — or ort_tract::api()
```
- `dr-face` and `dr-segment` **delete** their private `install_backend` and their direct
`Session::builder()` calls and take an `&dr_inference_engine::Sessions` where they take model bytes today.
Their tests keep `tract` — `dr_inference_engine::Sessions::tract()` is a constructor and the test-only path.
- `dr-inference-engine` depends on `ort` with the same workspace features as today plus `cuda`, `tensorrt`,
`qnn`: those features add option builders, not linking, under `alternative-backend`. **Verified
for the QNN, CUDA and TensorRT builders on 2026-09-19** — they go through the API table's generic
`SessionOptionsAppendExecutionProvider*`. The NNAPI builder resolves a symbol directly and would
not; it is not needed and is not enabled.
- `dr-ui` owns the settings row, the about-screen line and the progress row; it holds one
`Sessions` per process, created at launch, and passes it down. `dr_ui::library` gains
`inference_cache_dir()` beside `shared_face_models_dir()`, on the same account-independent
footing and for the same reason.
- The Android entry point's `install_bundled_models` also extracts nothing new: `jniLibs/` is
loaded by the system loader, and `dr-inference-engine` on Android looks for `libonnxruntime.so` through
`dlopen` by bare name first, which resolves to the APK's copy, before any directory.
---
## 9. Threads and memory
- ONNX Runtime's intra-op pool is sized to the physical cores minus two on desktop and to the
performance cores on Android (the X4 and the A720s; the A520s are for the compositor). tract's
single thread today is the reason a batch index leaves nineteen cores idle; ORT CPU with the
pool is the reason it will not. One session per model per process; `Session::run` is
`&mut self`-free in `ort` and internally serialised, and the index job is the only caller.
- A TensorRT session pins GPU memory for its workspace; the builder is capped at 512 MB on the
reference 6 GB card and the cap is a setting, because the develop view's tiles share the card
(NFR-RES-2). The engine cache on disk is bounded by the model set — six engines, ~80 MB — and
needs no LRU.
- The Hexagon rung sets QNN's performance mode to `Burst` for the duration of an index job and
`Default` otherwise; a 5 ms detector does not need the NPU clocked up between images.
- The probe (§4) and the engine build (§6) run on one dedicated low-priority thread. They never
share the index job's pool: a probe that competes with the job it is meant to speed up is the
frame-budget trap in a new coat.
---
## 10. What S16 measures
In order, with the gate each is:
| # | Question | Gate |
|---|---|---|
| M1 | Does one binary carry both tables? `dlopen` + `set_api` on Linux, Windows and Android; `ort_tract::api()` when the file is absent. | Go / no-go for §3. If `set_api` cannot take a dynamically fetched table on some platform, that platform ships two binaries, and the cost is stated. |
| M2 | Do the int8 SCRFD detectors, **calibrated on real photographs**, find the faces? faces.md §12.3's method over the reference library, per size band, against f32. | Ship the int8 form for a detector only if it finds ≥ 97% of the f32 detector's faces above 40 px and the difference is not concentrated in one band. Otherwise that detector's Hexagon rung is ORT CPU int8-free, and the table in §1.1 says what that costs. |
| M3 | Does the embedder on ORT CPU beside a detector on the Hexagon (§7) produce embeddings within the f32 gate? | It must — same graph, same arithmetic. This is a check that the plumbing did not quantise it by accident. |
| M4 | Is a batch index on the tablet and on the desktop faster by the ratio §1 predicts, end to end, decode included? | The face index over the reference library (18,143 faces): report wall-clock on tract, on the floor and on the selected rung, and where the time went. The prediction is that decode becomes the bottleneck on both; if it does not, say why. |
| M5 | Does the first-run sequence (§6) hold: nothing blocks the first frame, engines land, requests move up, a running job does not switch? | Observed on both devices with the app's own progress row, and with the cache directory deleted between runs. |
| M6 | What does the APK weigh with the QNN runtime, and does a non-Qualcomm Android device (any one) still launch and index on the floor? | Size reported; launch verified on one non-Qualcomm device or an emulator. |
| M7 | The segmentation and scene models at int8 on the Hexagon: does the mask boundary move? segmentation.md §6's IoU against f32 over its corpus. | Ship int8 for a model only above the IoU floor that document set for arm B. |
M1 and M2 are the ones the rest is conditional on, and M2 is the one that needs a person.
---
## 11. Order
1. **`dr-inference-engine` with the two tables and the floor** — `set_api` from a dlopened runtime, tract
otherwise, ORT CPU as the only rung. Consumers moved over; tests unchanged. This alone is the
3–10× and is the build most of the value sits in. M1.
2. **The probe and its cache** (§4), with the settings row and the about line. Still CPU-only;
the ladder has one rung. M5's first half.
3. **`tools/quantise-models.sh`** and the opset upgrade; the int8 detectors calibrated and
measured. M2, M7. This is the step with a person in it and it runs in parallel with 4.
4. **The Hexagon rung**, the QNN runtime in the APK, the context-binary cache. M3, M6.
5. **The TensorRT and CUDA rungs** on desktop, the engine cache, the first-run sequence. M5's
second half.
6. **M4** last, on both devices, and the number goes in this document.
The Windows installer and the Flatpak manifest are touched in steps 1 and 5 only, and only to add
a file each; the Arch package likewise.
---
## 12. Register entries
**FR-INF-1 — Runtime selection.** On launch the application shall determine, per device and
without blocking the first frame, the fastest inference backend that can build and run a session
for the shipped models, by attempting it; shall record and reuse that determination until the
runtime, driver, hardware or models change; and shall display the backend in use in Settings and
on the about screen. *Acceptance:* §10 M1 and M5.
**FR-INF-2 — Derived engines.** Backends that require device-specific compilation shall compile in
the background after selection, shall serve requests from the next lower backend until each engine
is ready, and shall not change the backend of a job in progress. *Acceptance:* M5.
**FR-INF-3 — Model forms.** Quantised model forms are produced at release time from real
calibration data and are shipped only when they meet §10's accuracy gates against the canonical
form; the application never quantises on the device. *Acceptance:* M2, M7.
**NFR-INF-1 — Embedding comparability.** Face embeddings shall be computed at a precision whose
deviation from the f32 reference is within §7's gate, on every backend, so that embeddings from any
device are comparable. *Acceptance:* M3.
**D13 — updated.** The runtime half is reopened to the extent of §3: the Rust build stays C-free
under `alternative-backend`; packages may install a dynamically loaded ONNX Runtime and, per §3.1,
the Qualcomm QNN runtime; the NVIDIA libraries are not bundled. The licensing half is unchanged.
---
## 13. Requirements touched
FR-CULL-8 (indexing time is what this exists to change), FR-CULL-9 (NFR-INF-1 is the guard on its
calibration), NFR-RES-2 (TensorRT workspace against the develop view's tiles), NFR-RES-3
(background compilation on Android), NFR-SEC-5 (the about screen's line), NFR-COMPAT-2 (each
channel gains a runtime file), FR-PLAT-AND-1 (unchanged; the runtime lives in the APK, not in
storage), ARCH §6.1 (the once-per-image budget this was sized against no longer binds; what could
run per frame is a separate question this document does not open).