From 6739fdf908bd0176fc97177c78615a1ef4b6513e Mon Sep 17 00:00:00 2001 From: Duncan Tourolle Date: Sat, 19 Sep 2026 13:53:06 +0200 Subject: [PATCH] Specify per-device inference backends, with the 2026-09-19 measurements tract runs every model on one core on every platform. Measured against ONNX Runtime's providers on the MagicPad 2 and the reference desktop: ORT CPU alone is 3-10x, the Hexagon at int8 runs the detectors in 1-3 ms, TensorRT is ~2x the CUDA provider. NNAPI, XNNPACK, WebGPU and CUDA int8 were tried and excluded with the numbers that excluded them. The spec keeps the build C-free: ort::set_api takes a table from a dlopened runtime or from ort-tract, chosen once per process. Rungs are chosen by building a real session, cached until an input changes, and compiled engines are built in the background after the first frame. The embedder stays f32 everywhere; int8 detectors are a distinct model_id and are gated on a recall measurement. --- docs/inference.md | 454 ++++++++++++++++++++++++++++++++++++++++++++++ 1 file changed, 454 insertions(+) create mode 100644 docs/inference.md diff --git a/docs/inference.md b/docs/inference.md new file mode 100644 index 0000000..84d9b30 --- /dev/null +++ b/docs/inference.md @@ -0,0 +1,454 @@ +# Inference backends — the runtime and the model, chosen per device + +Spec for **S16**, the build that puts the neural models on the hardware each device actually has. + +Every model DarkRoom runs today — the three SCRFD detectors, the ArcFace embedder, YOLO26n-seg and +the ADE20K scene model — runs through `tract`, on one CPU core, on every platform. That was the +right first answer: D13's runtime half chose it because it costs no C dependency, and +[faces.md](faces.md) and [segmentation.md](segmentation.md) were written against it. It is also +between 20× and 300× slower than what the same devices can do, and this document is the record of +having measured that and the specification of what replaces it. + +**It does not reopen D13's licensing half.** The weights are the same files under the same grant. +It does reopen the *runtime* half, and §3 is where it says how far. + +--- + +## 1. What was measured · 2026-09-19 + +One benchmark, two builds of it: `ort`'s API over `tract` (exactly what the app links) and `ort`'s +API over a dynamically loaded ONNX Runtime with each execution provider in turn. Random 640×640 +input, three warm-ups, the median of 15–30 timed runs, milliseconds. The same input every run, so +the numbers are compute cost and nothing else. + +### 1.1 The tablet — Honor MagicPad 2, Snapdragon 8s Gen 3 + +SM8635: 1× Cortex-X4, 4× A720, 3× A520, Adreno 735, Hexagon V73. Android 16. + +| Model | **tract** (today) | ORT CPU f32 | ORT CPU int8 | Adreno f32 ¹ | **Hexagon int8** ² | +|---|---|---|---|---|---| +| scrfd_500m (Fast) | 98 | 16 | 8 | 21 | **1.4** | +| scrfd_2.5g (Balanced) | 161 | 59 | 19 | ✗ | **1.8** | +| scrfd_10g (Thorough) | 489 | 204 | 48 | ✗ | **3.2** | +| arcface_mbf (per face) | 39 | 9 | 13 | 24 | 12 | +| yolo26n-seg | 287 | 94 | 38 | 54 | **5.1** | +| yolo26s-sem-ade20k | 408 | 154 | 46 | 47 | **3.7** | + +¹ Qualcomm's own GPU backend (`libQnnGpu.so`, OpenCL). Fails on the two larger SCRFD graphs at an +`AveragePool` the layout transformer cannot place. ONNX Runtime's WebGPU provider also runs on this +GPU and was slower than the CPU on every model; it is not in the table because it is not a +candidate. +² QNN's HTP backend. The Hexagon **refuses float32 and float16 tensors** in this ORT 1.29 + QNN +2.42 pairing (error 3110 on every node, with `enable_htp_fp16_precision` set or not); int8 QDQ +graphs run with 99.6% of nodes on the NPU — 1718 of 1725 for SCRFD-500m, the remainder being the +quantise/dequantise at the graph's edges — verified from the partition log, not inferred from the +timing. + +Also tried and rejected: **NNAPI** — the device registers no neural-networks HAL at all, so the +provider has nothing to talk to; Google deprecated it in Android 15 and Qualcomm stopped shipping +drivers for it. **XNNPACK** — slower than ORT's default CPU kernels on every model that loaded, and +aborts inside its partitioner on the SCRFD graphs. + +### 1.2 The desktop — RTX 3050 Laptop, Raptor Lake, 20 threads + +| Model | **tract** (today) | ORT CPU f32 | ORT CPU int8 | CUDA f32 | CUDA fp16 | TensorRT f32 | TensorRT fp16 | TensorRT int8 | +|---|---|---|---|---|---|---|---|---| +| scrfd_500m | 104 | 12 | 8 | 5.6 | 3.7 | 2.5 | **1.8** | ✗ ⁴ | +| scrfd_2.5g | 162 | 27 | 12 | 6.2 | 5.3 | 2.9 | **1.9** | ✗ ⁴ | +| scrfd_10g | 514 | 99 | 32 | 14.5 | 9.3 | 7.6 | **3.3** | ✗ ⁴ | +| arcface_mbf | 45 | 15 | 18 | 1.1 | 0.8 | 0.9 | 0.7 | ✗ ⁴ | +| yolo26n-seg | 307 | 60 | 38 | 9.3 | ✗ ³ | 6.8 | **5.5** | ✗ ⁴ | +| yolo26s-sem-ade20k | 395 | 61 | 31 | 10.7 | ✗ ³ | 7.9 | **3.4** | ✗ ⁴ | + +³ The offline fp16 conversion (`onnxconverter-common`) left a mixed-type node the CUDA provider +rejects. TensorRT converts to fp16 itself at engine build and does not have this problem, which is +one reason it is the target and the CUDA provider is the fallback. +⁴ TensorRT refuses the QDQ form ONNX Runtime's quantiser writes for the Hexagon (uint8 +activations); it wants symmetric int8. Not pursued: fp16 needs no quantisation, no calibration and +no accuracy gate, and it is already 30–60× tract. + +CUDA int8 is deliberately absent: the CUDA provider has no int8 kernels and runs a QDQ graph by +dequantising it, which measured *slower* than f32 (7.0 vs 5.6 ms on scrfd_500m). Int8 on NVIDIA is +TensorRT's job. + +**TensorRT's first load is 16–116 s per model in f32 and 35–290 s in fp16** (yolo26n-seg the +worst: nearly five minutes), because it is compiling an engine for this exact GPU. The engine caches to disk and the second load is milliseconds. That +number is what §6 is designed around. + +### 1.3 What the numbers say + +- **`tract` is single-threaded.** The tablet's one X4 core and one Raptor Lake core give the same + tract numbers. Replacing it with ONNX Runtime's CPU provider, *no accelerator involved*, is 3–6× + on the tablet and 8–10× on the desktop. That is the floor, and it is available on every platform + the app builds for. +- **The Hexagon is the standout.** A 6 W NPU running int8 beats a discrete RTX 3050 running f32 on + five of six models. A whole-library face index on the tablet goes from ~100 ms + 39 ms per face + to ~1.4 ms + 12 ms per face, and the "Thorough" detector — 3× the cost of "Fast" today — becomes + free. Its price is that the models must be **quantised to int8**, which is an accuracy question + §5 has to answer before it is believed. +- **The embedder does not gain from either accelerator.** 112×112 input, per-op overhead + dominates; it is 9 ms on the tablet's CPU and 12 ms on its NPU. It stays float, which §7 turns + from a performance footnote into a correctness rule. +- **On NVIDIA, TensorRT fp16 ≈ 3× the CUDA provider**, and the CUDA provider ≈ 2× the + multi-threaded CPU; at fp16 the detectors are 1.8–3.3 ms with no quantisation at all. Both leave the twenty cores free for decoding during a batch index, which the table does not + show and which matters more than the ratio. + +--- + +## 2. The shape of the answer + +A **ladder per platform**, walked at start-up, with the first rung that builds a real session +winning: + +| Platform | 1st | 2nd | 3rd | Floor | +|---|---|---|---|---| +| Android, Qualcomm with a Hexagon the shipped QNN skel covers (V68–V81) | QNN HTP, int8 model | ORT CPU, f32 model | — | tract | +| Android, any other SoC | ORT CPU, f32 | — | — | tract | +| Linux / Windows, NVIDIA GPU | TensorRT, f32 model, fp16 engine | CUDA provider, f32 | ORT CPU, f32 | tract | +| Linux / Windows, no NVIDIA | ORT CPU, f32 | — | — | tract | +| macOS ⁵ | ORT CPU, f32 | — | — | tract | + +⁵ CoreML is the obvious rung and is unmeasured; it is listed so its absence is a gap and not an +oversight. + +Deliberately **not** on any ladder, with the measurement that excluded each: NNAPI (no driver), +XNNPACK (slower than CPU, aborts on SCRFD), WebGPU (slower than CPU), the Adreno through QNN (works, +but never where the Hexagon does not also), CUDA int8 (slower than CUDA f32). A rung is added to this +table by a measurement on this page, not by a provider existing. + +Two things the ladder is *not*: it is not a per-model choice — one backend serves every model on a +device, because §7's identity rule needs the detector and embedder on the same runtime for the +same reason `faces.model_id` pairs them; and it is not a per-account choice — it is a property of +the hardware, like `shared_face_models_dir` is, and it lives beside it. + +--- + +## 3. The dependency policy, and how far this reopens it + +D13 chose `ort` over `tract` because `alternative-backend` made ONNX Runtime's *API* available with +none of its *C*. Every rung above the floor needs the C++ ONNX Runtime and, for the two that matter +most, vendor libraries on top: Qualcomm's QNN runtime (~60 MB for one Hexagon generation; it is +per-SoC) and NVIDIA's TensorRT plus cuDNN (~600 MB with the CUDA libraries, and cuDNN's major +version must match what ONNX Runtime was built against — the Arch package on the reference desktop +was unusable for exactly that reason). + +The policy protected the **build**: no C to cross-compile under the NDK, no toolchain to keep in +step. This document keeps that intact, and the mechanism is the one thing about `ort` that makes it +possible: + +**`ort::set_api` accepts any `OrtApi` table.** With `alternative-backend` on, `ort` links nothing +and asks for the table once per process. The application can `dlopen` a `libonnxruntime.so` it +finds on disk, call `OrtGetApiBase()->GetApi(version)` and hand that table over; or, if there is no +such file, hand over `ort_tract::api()`. The Rust build is identical in both cases — pure Rust, +`cargo build --target aarch64-linux-android` sees the same dependency graph it sees today. What +changes is that the runtime is a **file the package installs**, next to the models, and the app +looks for it at start-up. + +Consequences that follow and are accepted: + +- **The runtime is chosen once per process**, because `set_api` is once per process. The ladder + in §2 is walked at start-up and the result is what every session in that process uses. There is + no "tract for this model, ORT for that one", and there is no falling back to tract *after* ONNX + Runtime has loaded — but there does not need to be: once the library loads, its CPU provider is + always there, and every fallback the ladder needs is between providers *inside* it. +- **Feature flags stay as they are.** `dr-face`'s `inference` and `dr-segment`'s `semantic` + continue to mean "compiled against `ort`'s API"; nothing at build time knows or cares which table + will be supplied. The one addition is a `native-probe` feature on the new crate (§8) that pulls in + `libloading`, which is pure Rust and already in the tree via `wgpu`. +- **The packagers ship the runtime, not the build.** The Arch package, the Flatpak manifest, the + NSIS installer and `assemble-apk.sh` each gain the ONNX Runtime library for their platform, and + the Android and NVIDIA variants gain the vendor libraries — each under the licence the packager + reads first (§3.1). A package without them is not broken; it is the tract build, and it says so + on the about screen. +- **The NDK problem does not come back.** `libonnxruntime.so` for Android is a prebuilt from + Maven (`com.microsoft.onnxruntime:onnxruntime-android-qnn`), extracted by `assemble-apk.sh` into + `jniLibs/` the way the models are bundled as assets today. Nothing compiles it. + +### 3.1 Licences the packagers read before shipping a runtime + +Written down now, because [segmentation.md §7](segmentation.md) established that reading the grant +is cheaper than discovering it at packaging time. + +| Component | Licence | Redistributable in a self-distributed package? | +|---|---|---| +| ONNX Runtime | MIT | Yes | +| Qualcomm QNN runtime (`com.qualcomm.qti:qnn-runtime` on Maven) | Qualcomm AI Engine Direct SDK licence — proprietary, redistribution permitted for applications using it | Yes for the APK, with the licence text shipped; not for a source distribution. **To be read in full, not summarised from memory, before the APK gains it.** | +| CUDA runtime, cuDNN, TensorRT | NVIDIA EULAs — redistributable with an application, with the licence text, not modifiable | Yes for a package that bundles them. 600 MB. The alternative is to load them from the user's system install if present and skip the rung otherwise — which is what §4's probe does anyway. | + +The position this takes: the **NVIDIA libraries are not bundled**. The desktop package probes for a +system CUDA/TensorRT install and uses it if it is version-compatible; a desktop without one runs on +ORT CPU, which is still 8–10× today. Bundling 600 MB for a rung that is 2× again is not a trade +worth making unmeasured, and it can be revisited by a measurement on a batch index. The **QNN +runtime is bundled** in the APK, because the Hexagon is the difference between a tablet that +indexes a library overnight and one that does it over lunch, and the package is 60 MB larger for +it. + +Both positions are D13 territory and are recorded there (§12). + +--- + +## 4. Selection — the probe, its cache, and what it may not do + +**A rung is chosen by building a real session on it, not by asking whether it exists.** Both +failure modes that are not "the provider is absent" were hit on 2026-09-19: a driver in a wedged +state where the provider registered and the session then failed, and a provider that registered, +took the graph, and rejected every node at partition time. The probe therefore: + +1. Loads the runtime library (§3), or falls to tract and stops. +2. For each rung in this platform's ladder, in order: builds a session for the **smallest model + in the set** (`scrfd_500m`) on that provider with `error_on_failure`, runs it once on a fixed + input, and reads back the provider assignment from the session — the rung is taken only if the + provider ran **at least 95% of the graph's nodes**. A provider that silently hands the graph to + the CPU is the CPU rung with extra overhead, and the app should say "CPU". +3. Records the outcome — rung, runtime version, provider version, device identity (GPU name and + compute capability; SoC model and Hexagon arch), and the models' content hashes — to a small + file beside `shared_face_models_dir`. The next start-up trusts the file **unless** any of those + inputs changed, in which case it probes again. A driver update, a runtime update, a new model + file: each invalidates the cache by construction, and none needs a "reset backend" button. + +What the probe may not do: + +- **Block the first frame.** It runs on the same background as `install_bundled_models` and for + the same reason: a TensorRT probe can take thirty seconds cold, and a tablet that stalls that long + is an ANR. Until it reports, every model request is answered by the floor the runtime supports + (ORT CPU if the library loaded, tract otherwise), and a job that started on the floor finishes + on it — a backend does not change under a running index. +- **Retry a rung that failed within a session.** A failed probe is cached as a failure with the + same inputs; the rung is tried again when an input changes. Otherwise a wedged driver means a + thirty-second stall on every launch. +- **Choose for the user without saying so.** Settings gains one row, *Inference backend*, showing + what was chosen and why in one line ("Hexagon NPU · int8 · QNN 2.42"; "CPU · ONNX Runtime 1.30 · + TensorRT probe failed: cuDNN 8 required"), with an override to force any lower rung. The about + screen carries the same line beside the model names NFR-SEC-5 already puts there. + +--- + +## 5. Model variants, and who makes them + +Every model exists in one **canonical** form — the f32 ONNX file the app ships or the user supplies +today — and, where a rung needs it, a **derived** form. The ladder's rungs are specified in terms +of which form they load: + +| Form | Who produces it | When | Needed by | +|---|---|---|---| +| f32 ONNX, shape-fixed, **opset ≥ 13** | `tools/fix-face-model-shapes.sh`, `tools/export-seg-model.sh` | Release time, once | Every rung except Hexagon | +| int8 QDQ ONNX, per-channel, uint8 activations | `tools/quantise-models.sh` (new) | Release time, once, **calibrated on real photographs** | Hexagon | +| TensorRT engine (`.engine`, per GPU architecture and TensorRT version) | The app, from the f32 file | First run on that device, in the background | TensorRT rung | +| QNN context binary | The app, from the int8 file | First run on that device, in the background | Hexagon rung | + +Two rules. + +**Quantisation is a release-time step, not a device-time one.** The int8 files that produced §1's +numbers were calibrated on random noise, which is enough to time and worthless to trust. A real +int8 detector is calibrated on a few hundred real photographs and then measured against the f32 +detector on the reference library by [faces.md §12.3](faces.md)'s method — faces found, per size +band, per detector — before it ships. That needs the reference library and a person reading the +result, and it happens once per model release, in `tools/`, beside the shape-fixing it already +depends on. The device never quantises anything. + +The SCRFD and ArcFace files are **opset 11** as InsightFace exported them, and per-channel QDQ needs +13; `tools/fix-face-model-shapes.sh` gains an opset upgrade to 17 (`onnx.version_converter`, +`ir_version` 8), which `tract` has been verified to load and which every provider on this page +prefers. That is a change to the canonical file and so a change to the shipped models, and it +happens in the same model release as the int8 files. + +**Compilation is a device-time step, and it is cached.** A TensorRT engine is specific to the GPU +it was built on and the TensorRT that built it; a QNN context binary is specific to the Hexagon +generation. Neither can ship. Both are built by the app the first time that rung is selected, in +the background (§6), and written beside the probe cache keyed by the same inputs. They are +**derived, disposable, and regenerable**: deleting the cache directory costs the next launch a +rebuild and nothing else, and the directory is excluded from anything that syncs (it is a peer of +`thumbs`, not of the catalog). + +--- + +## 6. First run — building engines without the user waiting for them + +The sequence on a device where a compiling rung (TensorRT, Hexagon) is selected: + +1. **Launch.** The runtime loads; the probe (§4) starts in the background; the app serves every + model request from the floor. Face indexing, segmentation and scene grading all work, at + today's speed or better (ORT CPU). +2. **Probe reports** — say, TensorRT. The compiling rung is now *selected* but has **no engines**. + Model requests continue on the fallback rung below it (CUDA provider for TensorRT; ORT CPU for + Hexagon), which needs no compilation and is already faster than the floor. +3. **Engines build**, one model at a time, on a single low-priority background thread, smallest + model first so the detector — the one that runs per image — is ready soonest. On the reference + desktop that is ~1 minute for the first detector and ~10 minutes for all six at fp16; on the tablet the QNN + context binaries take 0.8–1.7 s each and the whole set is ready before the user has opened a + library. Each engine is written to a temporary name and renamed into place, so a request never + sees a half-written file. +4. **Requests move up as engines land.** A model whose engine exists loads it on the selected + rung; one whose engine is still building loads on the fallback. **A running job does not + switch** — an index that started on the CUDA provider finishes on it — because §7 needs one + `model_id` per job, and because a job is the wrong granularity for surprise. +5. On Android, the build runs only while the app is in the foreground and the device is not in + battery saver (NFR-RES-3): a context binary takes a second, so this costs nothing, and the rule + exists for the day a model takes longer. + +Settings shows a one-line progress row while engines build ("Preparing GPU engines · 3 of 6") and +nothing when they are done. A build failure demotes the rung — it is recorded in the probe cache +as a failure with the model hash as an input, so a corrected model file retries it — and the app +carries on one rung down, saying so in the same row. + +--- + +## 7. Identity — what changes `model_id` and what may not + +`faces.model_id` exists so that two libraries indexed with different networks are never compared +as if they were one ([catalog.md §10.1](catalog.md); the trap is written up in +[faces.md §14](faces.md)). A backend that changes what a network *computes* is a different network +and must be a different `model_id`; one that changes only *where* it computes it must not be. + +**The detector.** An int8 SCRFD finds a different set of faces from the f32 one — that is what +§5's acceptance measures — so **the numeric form is part of the detector's identity**: +`scrfd_500m` and `scrfd_500m_i8` are two detectors in `model_id`, and a library indexed on the +tablet's Hexagon and continued on the desktop is two populations, which a re-index on either side +reconciles the same way a switch from Fast to Thorough does today. That is acceptable because it is +already the rule for the detector and because §5 is the gate on whether the int8 form is close +enough to be *offered* at all. f32 on tract, ORT CPU, CUDA and TensorRT-f32 are one identity: the +same graph, the same arithmetic, differences at the last bit. + +**The embedder** is where comparability across devices is the whole point, and it is the one +model that no accelerator helps (§1.3). So: **the embedder runs in f32 on every rung.** On TensorRT +that means the embedder's engine is built without fp16 while the detector's is built with it; on +the Hexagon it means the embedder is not on the NPU at all — it runs on the ORT CPU rung at 9 ms, +and the ladder's "one backend per device" is, precisely, one backend *per model role*, with the +embedder pinned. A `w600k_mbf` embedding from any device is comparable with one from any other, +which is the property the identity system, the calibration and the cross-device merge all rest +on, and it is not for sale for 3 ms. + +If S16 wants fp16 for the embedder later, the gate is written now: over the reference library's +faces, the cosine between the f32 and fp16 embedding of the same crop exceeds 0.999 for 99.9% of +faces and the calibration's fitted threshold moves by less than its own confidence interval +([faces.md §8.3](faces.md)). Until measured, f32. + +**Segmentation and the scene model** carry no identity across devices — their outputs are +recomputed per image and never stored beyond the cache — so they take whatever the rung offers, +int8 included, subject to §10's own acceptance. + +--- + +## 8. Crate shape — `core/dr-infer` + +The seam is the same shape as [storage.md](storage.md)'s: a small crate below the consumers that is +the **only** place naming a provider, a library file or a vendor, with the consumers reduced to +"give me a session for these bytes in this role". + +``` +core/dr-infer + src/lib.rs Runtime (Tract | Onnx { lib, version }), Backend (rung), Role (Detector | Embedder | Segmenter) + src/probe.rs §4 — the ladder per platform, the session-build probe, the cache file + src/engines.rs §6 — background compilation, the cache directory, progress + src/session.rs open(role, bytes) -> ort::Session, applying the rung and the role's precision rule + src/api.rs the one unsafe block: dlopen libonnxruntime, fetch OrtApi, ort::set_api — or ort_tract::api() +``` + +- `dr-face` and `dr-segment` **delete** their private `install_backend` and their direct + `Session::builder()` calls and take an `&dr_infer::Sessions` where they take model bytes today. + Their tests keep `tract` — `dr_infer::Sessions::tract()` is a constructor and the test-only path. +- `dr-infer` depends on `ort` with the same workspace features as today plus `cuda`, `tensorrt`, + `qnn`: those features add option builders, not linking, under `alternative-backend`. **Verified + for the QNN, CUDA and TensorRT builders on 2026-09-19** — they go through the API table's generic + `SessionOptionsAppendExecutionProvider*`. The NNAPI builder resolves a symbol directly and would + not; it is not needed and is not enabled. +- `dr-ui` owns the settings row, the about-screen line and the progress row; it holds one + `Sessions` per process, created at launch, and passes it down. `dr_ui::library` gains + `inference_cache_dir()` beside `shared_face_models_dir()`, on the same account-independent + footing and for the same reason. +- The Android entry point's `install_bundled_models` also extracts nothing new: `jniLibs/` is + loaded by the system loader, and `dr-infer` on Android looks for `libonnxruntime.so` through + `dlopen` by bare name first, which resolves to the APK's copy, before any directory. + +--- + +## 9. Threads and memory + +- ONNX Runtime's intra-op pool is sized to the physical cores minus two on desktop and to the + performance cores on Android (the X4 and the A720s; the A520s are for the compositor). tract's + single thread today is the reason a batch index leaves nineteen cores idle; ORT CPU with the + pool is the reason it will not. One session per model per process; `Session::run` is + `&mut self`-free in `ort` and internally serialised, and the index job is the only caller. +- A TensorRT session pins GPU memory for its workspace; the builder is capped at 512 MB on the + reference 6 GB card and the cap is a setting, because the develop view's tiles share the card + (NFR-RES-2). The engine cache on disk is bounded by the model set — six engines, ~80 MB — and + needs no LRU. +- The Hexagon rung sets QNN's performance mode to `Burst` for the duration of an index job and + `Default` otherwise; a 5 ms detector does not need the NPU clocked up between images. +- The probe (§4) and the engine build (§6) run on one dedicated low-priority thread. They never + share the index job's pool: a probe that competes with the job it is meant to speed up is the + frame-budget trap in a new coat. + +--- + +## 10. What S16 measures + +In order, with the gate each is: + +| # | Question | Gate | +|---|---|---| +| M1 | Does one binary carry both tables? `dlopen` + `set_api` on Linux, Windows and Android; `ort_tract::api()` when the file is absent. | Go / no-go for §3. If `set_api` cannot take a dynamically fetched table on some platform, that platform ships two binaries, and the cost is stated. | +| M2 | Do the int8 SCRFD detectors, **calibrated on real photographs**, find the faces? faces.md §12.3's method over the reference library, per size band, against f32. | Ship the int8 form for a detector only if it finds ≥ 97% of the f32 detector's faces above 40 px and the difference is not concentrated in one band. Otherwise that detector's Hexagon rung is ORT CPU int8-free, and the table in §1.1 says what that costs. | +| M3 | Does the embedder on ORT CPU beside a detector on the Hexagon (§7) produce embeddings within the f32 gate? | It must — same graph, same arithmetic. This is a check that the plumbing did not quantise it by accident. | +| M4 | Is a batch index on the tablet and on the desktop faster by the ratio §1 predicts, end to end, decode included? | The face index over the reference library (18,143 faces): report wall-clock on tract, on the floor and on the selected rung, and where the time went. The prediction is that decode becomes the bottleneck on both; if it does not, say why. | +| M5 | Does the first-run sequence (§6) hold: nothing blocks the first frame, engines land, requests move up, a running job does not switch? | Observed on both devices with the app's own progress row, and with the cache directory deleted between runs. | +| M6 | What does the APK weigh with the QNN runtime, and does a non-Qualcomm Android device (any one) still launch and index on the floor? | Size reported; launch verified on one non-Qualcomm device or an emulator. | +| M7 | The segmentation and scene models at int8 on the Hexagon: does the mask boundary move? segmentation.md §6's IoU against f32 over its corpus. | Ship int8 for a model only above the IoU floor that document set for arm B. | + +M1 and M2 are the ones the rest is conditional on, and M2 is the one that needs a person. + +--- + +## 11. Order + +1. **`dr-infer` with the two tables and the floor** — `set_api` from a dlopened runtime, tract + otherwise, ORT CPU as the only rung. Consumers moved over; tests unchanged. This alone is the + 3–10× and is the build most of the value sits in. M1. +2. **The probe and its cache** (§4), with the settings row and the about line. Still CPU-only; + the ladder has one rung. M5's first half. +3. **`tools/quantise-models.sh`** and the opset upgrade; the int8 detectors calibrated and + measured. M2, M7. This is the step with a person in it and it runs in parallel with 4. +4. **The Hexagon rung**, the QNN runtime in the APK, the context-binary cache. M3, M6. +5. **The TensorRT and CUDA rungs** on desktop, the engine cache, the first-run sequence. M5's + second half. +6. **M4** last, on both devices, and the number goes in this document. + +The Windows installer and the Flatpak manifest are touched in steps 1 and 5 only, and only to add +a file each; the Arch package likewise. + +--- + +## 12. Register entries + +**FR-INF-1 — Runtime selection.** On launch the application shall determine, per device and +without blocking the first frame, the fastest inference backend that can build and run a session +for the shipped models, by attempting it; shall record and reuse that determination until the +runtime, driver, hardware or models change; and shall display the backend in use in Settings and +on the about screen. *Acceptance:* §10 M1 and M5. + +**FR-INF-2 — Derived engines.** Backends that require device-specific compilation shall compile in +the background after selection, shall serve requests from the next lower backend until each engine +is ready, and shall not change the backend of a job in progress. *Acceptance:* M5. + +**FR-INF-3 — Model forms.** Quantised model forms are produced at release time from real +calibration data and are shipped only when they meet §10's accuracy gates against the canonical +form; the application never quantises on the device. *Acceptance:* M2, M7. + +**NFR-INF-1 — Embedding comparability.** Face embeddings shall be computed at a precision whose +deviation from the f32 reference is within §7's gate, on every backend, so that embeddings from any +device are comparable. *Acceptance:* M3. + +**D13 — updated.** The runtime half is reopened to the extent of §3: the Rust build stays C-free +under `alternative-backend`; packages may install a dynamically loaded ONNX Runtime and, per §3.1, +the Qualcomm QNN runtime; the NVIDIA libraries are not bundled. The licensing half is unchanged. + +--- + +## 13. Requirements touched + +FR-CULL-8 (indexing time is what this exists to change), FR-CULL-9 (NFR-INF-1 is the guard on its +calibration), NFR-RES-2 (TensorRT workspace against the develop view's tiles), NFR-RES-3 +(background compilation on Android), NFR-SEC-5 (the about screen's line), NFR-COMPAT-2 (each +channel gains a runtime file), FR-PLAT-AND-1 (unchanged; the runtime lives in the APK, not in +storage), ARCH §6.1 (the once-per-image budget this was sized against no longer binds; what could +run per frame is a separate question this document does not open).