Files
DarkRoom/docs/inference.md
T

30 KiB
Raw Blame History

Inference backends — the runtime and the model, chosen per device

Spec for S16, the build that puts the neural models on the hardware each device actually has.

Every model DarkRoom runs today — the three SCRFD detectors, the ArcFace embedder, YOLO26n-seg and the ADE20K scene model — runs through tract, on one CPU core, on every platform. That was the right first answer: D13's runtime half chose it because it costs no C dependency, and faces.md and segmentation.md were written against it. It is also between 20× and 300× slower than what the same devices can do, and this document is the record of having measured that and the specification of what replaces it.

It does not reopen D13's licensing half. The weights are the same files under the same grant. It does reopen the runtime half, and §3 is where it says how far.


1. What was measured · 2026-09-19

One benchmark, two builds of it: ort's API over tract (exactly what the app links) and ort's API over a dynamically loaded ONNX Runtime with each execution provider in turn. Random 640×640 input, three warm-ups, the median of 15–30 timed runs, milliseconds. The same input every run, so the numbers are compute cost and nothing else.

1.1 The tablet — Honor MagicPad 2, Snapdragon 8s Gen 3

SM8635: 1× Cortex-X4, 4× A720, 3× A520, Adreno 735, Hexagon V73. Android 16.

Model tract (today) ORT CPU f32 ORT CPU int8 Adreno f32 ¹ Hexagon int8 ²
scrfd_500m (Fast) 98 16 8 21 1.4
scrfd_2.5g (Balanced) 161 59 19 ✗ 1.8
scrfd_10g (Thorough) 489 204 48 ✗ 3.2
arcface_mbf (per face) 39 9 13 24 12
yolo26n-seg 287 94 38 54 5.1
yolo26s-sem-ade20k 408 154 46 47 3.7

¹ Qualcomm's own GPU backend (libQnnGpu.so, OpenCL). Fails on the two larger SCRFD graphs at an AveragePool the layout transformer cannot place. ONNX Runtime's WebGPU provider also runs on this GPU and was slower than the CPU on every model; it is not in the table because it is not a candidate. ² QNN's HTP backend. The Hexagon refuses float32 and float16 tensors in this ORT 1.29 + QNN 2.42 pairing (error 3110 on every node, with enable_htp_fp16_precision set or not); int8 QDQ graphs run with 99.6% of nodes on the NPU — 1718 of 1725 for SCRFD-500m, the remainder being the quantise/dequantise at the graph's edges — verified from the partition log, not inferred from the timing.

Also tried and rejected: NNAPI — the device registers no neural-networks HAL at all, so the provider has nothing to talk to; Google deprecated it in Android 15 and Qualcomm stopped shipping drivers for it. XNNPACK — slower than ORT's default CPU kernels on every model that loaded, and aborts inside its partitioner on the SCRFD graphs.

1.2 The desktop — RTX 3050 Laptop, Raptor Lake, 20 threads

Model tract (today) ORT CPU f32 ORT CPU int8 CUDA f32 CUDA fp16 TensorRT f32 TensorRT fp16 TensorRT int8
scrfd_500m 104 12 8 5.6 3.7 2.5 1.8 ✗ ⁴
scrfd_2.5g 162 27 12 6.2 5.3 2.9 1.9 ✗ ⁴
scrfd_10g 514 99 32 14.5 9.3 7.6 3.3 ✗ ⁴
arcface_mbf 45 15 18 1.1 0.8 0.9 0.7 ✗ ⁴
yolo26n-seg 307 60 38 9.3 ✗ ³ 6.8 5.5 ✗ ⁴
yolo26s-sem-ade20k 395 61 31 10.7 ✗ ³ 7.9 3.4 ✗ ⁴

³ The offline fp16 conversion (onnxconverter-common) left a mixed-type node the CUDA provider rejects. TensorRT converts to fp16 itself at engine build and does not have this problem, which is one reason it is the target and the CUDA provider is the fallback. ⁴ TensorRT refuses the QDQ form ONNX Runtime's quantiser writes for the Hexagon (uint8 activations); it wants symmetric int8. Not pursued: fp16 needs no quantisation, no calibration and no accuracy gate, and it is already 30–60× tract.

CUDA int8 is deliberately absent: the CUDA provider has no int8 kernels and runs a QDQ graph by dequantising it, which measured slower than f32 (7.0 vs 5.6 ms on scrfd_500m). Int8 on NVIDIA is TensorRT's job.

TensorRT's first load is 16–116 s per model in f32 and 35–290 s in fp16 (yolo26n-seg the worst: nearly five minutes), because it is compiling an engine for this exact GPU. The engine caches to disk and the second load is milliseconds. That number is what §6 is designed around.

1.3 What the numbers say

  • tract is single-threaded. The tablet's one X4 core and one Raptor Lake core give the same tract numbers. Replacing it with ONNX Runtime's CPU provider, no accelerator involved, is 3–6× on the tablet and 8–10× on the desktop. That is the floor, and it is available on every platform the app builds for.
  • The Hexagon is the standout. A 6 W NPU running int8 beats a discrete RTX 3050 running f32 on five of six models. A whole-library face index on the tablet goes from ~100 ms + 39 ms per face to ~1.4 ms + 12 ms per face, and the "Thorough" detector — 3× the cost of "Fast" today — becomes free. Its price is that the models must be quantised to int8, which is an accuracy question §5 has to answer before it is believed.
  • The embedder does not gain from either accelerator. 112×112 input, per-op overhead dominates; it is 9 ms on the tablet's CPU and 12 ms on its NPU. It stays float, which §7 turns from a performance footnote into a correctness rule.
  • On NVIDIA, TensorRT fp16 ≈ 3× the CUDA provider, and the CUDA provider ≈ 2× the multi-threaded CPU; at fp16 the detectors are 1.8–3.3 ms with no quantisation at all. Both leave the twenty cores free for decoding during a batch index, which the table does not show and which matters more than the ratio.

2. The shape of the answer

A ladder per platform, walked at start-up, with the first rung that builds a real session winning:

Platform 1st 2nd 3rd Floor
Android, Qualcomm with a Hexagon the shipped QNN skel covers (V68–V81) QNN HTP, int8 model ORT CPU, f32 model — tract
Android, any other SoC ORT CPU, f32 — — tract
Linux / Windows, NVIDIA GPU TensorRT, f32 model, fp16 engine CUDA provider, f32 ORT CPU, f32 tract
Linux / Windows, no NVIDIA ORT CPU, f32 — — tract
macOS ⁵ ORT CPU, f32 — — tract

⁵ CoreML is the obvious rung and is unmeasured; it is listed so its absence is a gap and not an oversight.

Deliberately not on any ladder, with the measurement that excluded each: NNAPI (no driver), XNNPACK (slower than CPU, aborts on SCRFD), WebGPU (slower than CPU), the Adreno through QNN (works, but never where the Hexagon does not also), CUDA int8 (slower than CUDA f32). A rung is added to this table by a measurement on this page, not by a provider existing.

Two things the ladder is not: it is not a per-model choice — one backend serves every model on a device, because §7's identity rule needs the detector and embedder on the same runtime for the same reason faces.model_id pairs them; and it is not a per-account choice — it is a property of the hardware, like shared_face_models_dir is, and it lives beside it.


3. The dependency policy, and how far this reopens it

D13 chose ort over tract because alternative-backend made ONNX Runtime's API available with none of its C. Every rung above the floor needs the C++ ONNX Runtime and, for the two that matter most, vendor libraries on top: Qualcomm's QNN runtime (~60 MB for one Hexagon generation; it is per-SoC) and NVIDIA's TensorRT plus cuDNN (~600 MB with the CUDA libraries, and cuDNN's major version must match what ONNX Runtime was built against — the Arch package on the reference desktop was unusable for exactly that reason).

The policy protected the build: no C to cross-compile under the NDK, no toolchain to keep in step. This document keeps that intact, and the mechanism is the one thing about ort that makes it possible:

ort::set_api accepts any OrtApi table. With alternative-backend on, ort links nothing and asks for the table once per process. The application can dlopen a libonnxruntime.so it finds on disk, call OrtGetApiBase()->GetApi(version) and hand that table over; or, if there is no such file, hand over ort_tract::api(). The Rust build is identical in both cases — pure Rust, cargo build --target aarch64-linux-android sees the same dependency graph it sees today. What changes is that the runtime is a file the package installs, next to the models, and the app looks for it at start-up.

Consequences that follow and are accepted:

  • The runtime is chosen once per process, because set_api is once per process. The ladder in §2 is walked at start-up and the result is what every session in that process uses. There is no "tract for this model, ORT for that one", and there is no falling back to tract after ONNX Runtime has loaded — but there does not need to be: once the library loads, its CPU provider is always there, and every fallback the ladder needs is between providers inside it.
  • Feature flags stay as they are. dr-face's inference and dr-segment's semantic continue to mean "compiled against ort's API"; nothing at build time knows or cares which table will be supplied. The one addition is a native-probe feature on the new crate (§8) that pulls in libloading, which is pure Rust and already in the tree via wgpu.
  • The packagers ship the runtime, not the build. The Arch package, the Flatpak manifest, the NSIS installer and assemble-apk.sh each gain the ONNX Runtime library for their platform, and the Android and NVIDIA variants gain the vendor libraries — each under the licence the packager reads first (§3.1). A package without them is not broken; it is the tract build, and it says so on the about screen.
  • The NDK problem does not come back. libonnxruntime.so for Android is a prebuilt from Maven (com.microsoft.onnxruntime:onnxruntime-android-qnn), extracted by assemble-apk.sh into jniLibs/ the way the models are bundled as assets today. Nothing compiles it.

3.1 Licences the packagers read before shipping a runtime

Written down now, because segmentation.md §7 established that reading the grant is cheaper than discovering it at packaging time.

Component Licence Redistributable in a self-distributed package?
ONNX Runtime MIT Yes
Qualcomm QNN runtime (com.qualcomm.qti:qnn-runtime on Maven) Qualcomm AI Engine Direct SDK licence — proprietary, redistribution permitted for applications using it Yes for the APK, with the licence text shipped; not for a source distribution. To be read in full, not summarised from memory, before the APK gains it.
CUDA runtime, cuDNN, TensorRT NVIDIA EULAs — redistributable with an application, with the licence text, not modifiable Yes for a package that bundles them. 600 MB. The alternative is to load them from the user's system install if present and skip the rung otherwise — which is what §4's probe does anyway.

The position this takes: the NVIDIA libraries are not bundled. The desktop package probes for a system CUDA/TensorRT install and uses it if it is version-compatible; a desktop without one runs on ORT CPU, which is still 8–10× today. Bundling 600 MB for a rung that is 2× again is not a trade worth making unmeasured, and it can be revisited by a measurement on a batch index. The QNN runtime is bundled in the APK, because the Hexagon is the difference between a tablet that indexes a library overnight and one that does it over lunch, and the package is 60 MB larger for it.

Both positions are D13 territory and are recorded there (§12).


4. Selection — the probe, its cache, and what it may not do

A rung is chosen by building a real session on it, not by asking whether it exists. Both failure modes that are not "the provider is absent" were hit on 2026-09-19: a driver in a wedged state where the provider registered and the session then failed, and a provider that registered, took the graph, and rejected every node at partition time. The probe therefore:

  1. Loads the runtime library (§3), or falls to tract and stops.
  2. For each rung in this platform's ladder, in order: builds a session for the smallest model in the set (scrfd_500m) on that provider with error_on_failure, runs it once on a fixed input, and reads back the provider assignment from the session — the rung is taken only if the provider ran at least 95% of the graph's nodes. A provider that silently hands the graph to the CPU is the CPU rung with extra overhead, and the app should say "CPU".
  3. Records the outcome — rung, runtime version, provider version, device identity (GPU name and compute capability; SoC model and Hexagon arch), and the models' content hashes — to a small file beside shared_face_models_dir. The next start-up trusts the file unless any of those inputs changed, in which case it probes again. A driver update, a runtime update, a new model file: each invalidates the cache by construction, and none needs a "reset backend" button.

What the probe may not do:

  • Block the first frame. It runs on the same background as install_bundled_models and for the same reason: a TensorRT probe can take thirty seconds cold, and a tablet that stalls that long is an ANR. Until it reports, every model request is answered by the floor the runtime supports (ORT CPU if the library loaded, tract otherwise), and a job that started on the floor finishes on it — a backend does not change under a running index.
  • Retry a rung that failed within a session. A failed probe is cached as a failure with the same inputs; the rung is tried again when an input changes. Otherwise a wedged driver means a thirty-second stall on every launch.
  • Choose for the user without saying so. Settings gains one row, Inference backend, showing what was chosen and why in one line ("Hexagon NPU · int8 · QNN 2.42"; "CPU · ONNX Runtime 1.30 · TensorRT probe failed: cuDNN 8 required"), with an override to force any lower rung. The about screen carries the same line beside the model names NFR-SEC-5 already puts there.

5. Model variants, and who makes them

Every model exists in one canonical form — the f32 ONNX file the app ships or the user supplies today — and, where a rung needs it, a derived form. The ladder's rungs are specified in terms of which form they load:

Form Who produces it When Needed by
f32 ONNX, shape-fixed, opset ≥ 13 tools/fix-face-model-shapes.sh, tools/export-seg-model.sh Release time, once Every rung except Hexagon
int8 QDQ ONNX, per-channel, uint8 activations tools/quantise-models.sh (new) Release time, once, calibrated on real photographs Hexagon
TensorRT engine (.engine, per GPU architecture and TensorRT version) The app, from the f32 file First run on that device, in the background TensorRT rung
QNN context binary The app, from the int8 file First run on that device, in the background Hexagon rung

Two rules.

Quantisation is a release-time step, not a device-time one. The int8 files that produced §1's numbers were calibrated on random noise, which is enough to time and worthless to trust. A real int8 detector is calibrated on a few hundred real photographs and then measured against the f32 detector on the reference library by faces.md §12.3's method — faces found, per size band, per detector — before it ships. That needs the reference library and a person reading the result, and it happens once per model release, in tools/, beside the shape-fixing it already depends on. The device never quantises anything.

The SCRFD and ArcFace files are opset 11 as InsightFace exported them, and per-channel QDQ needs 13; tools/fix-face-model-shapes.sh gains an opset upgrade to 17 (onnx.version_converter, ir_version 8), which tract has been verified to load and which every provider on this page prefers. That is a change to the canonical file and so a change to the shipped models, and it happens in the same model release as the int8 files.

Compilation is a device-time step, and it is cached. A TensorRT engine is specific to the GPU it was built on and the TensorRT that built it; a QNN context binary is specific to the Hexagon generation. Neither can ship. Both are built by the app the first time that rung is selected, in the background (§6), and written beside the probe cache keyed by the same inputs. They are derived, disposable, and regenerable: deleting the cache directory costs the next launch a rebuild and nothing else, and the directory is excluded from anything that syncs (it is a peer of thumbs, not of the catalog).


6. First run — building engines without the user waiting for them

The sequence on a device where a compiling rung (TensorRT, Hexagon) is selected:

  1. Launch. The runtime loads; the probe (§4) starts in the background; the app serves every model request from the floor. Face indexing, segmentation and scene grading all work, at today's speed or better (ORT CPU).
  2. Probe reports — say, TensorRT. The compiling rung is now selected but has no engines. Model requests continue on the fallback rung below it (CUDA provider for TensorRT; ORT CPU for Hexagon), which needs no compilation and is already faster than the floor.
  3. Engines build, one model at a time, on a single low-priority background thread, smallest model first so the detector — the one that runs per image — is ready soonest. On the reference desktop that is ~1 minute for the first detector and ~10 minutes for all six at fp16; on the tablet the QNN context binaries take 0.8–1.7 s each and the whole set is ready before the user has opened a library. Each engine is written to a temporary name and renamed into place, so a request never sees a half-written file.
  4. Requests move up as engines land. A model whose engine exists loads it on the selected rung; one whose engine is still building loads on the fallback. A running job does not switch — an index that started on the CUDA provider finishes on it — because §7 needs one model_id per job, and because a job is the wrong granularity for surprise.
  5. On Android, the build runs only while the app is in the foreground and the device is not in battery saver (NFR-RES-3): a context binary takes a second, so this costs nothing, and the rule exists for the day a model takes longer.

Settings shows a one-line progress row while engines build ("Preparing GPU engines · 3 of 6") and nothing when they are done. A build failure demotes the rung — it is recorded in the probe cache as a failure with the model hash as an input, so a corrected model file retries it — and the app carries on one rung down, saying so in the same row.


7. Identity — what changes model_id and what may not

faces.model_id exists so that two libraries indexed with different networks are never compared as if they were one (catalog.md §10.1; the trap is written up in faces.md §14). A backend that changes what a network computes is a different network and must be a different model_id; one that changes only where it computes it must not be.

The detector. An int8 SCRFD finds a different set of faces from the f32 one — that is what §5's acceptance measures — so the numeric form is part of the detector's identity: scrfd_500m and scrfd_500m_i8 are two detectors in model_id, and a library indexed on the tablet's Hexagon and continued on the desktop is two populations, which a re-index on either side reconciles the same way a switch from Fast to Thorough does today. That is acceptable because it is already the rule for the detector and because §5 is the gate on whether the int8 form is close enough to be offered at all. f32 on tract, ORT CPU, CUDA and TensorRT-f32 are one identity: the same graph, the same arithmetic, differences at the last bit.

The embedder is where comparability across devices is the whole point, and it is the one model that no accelerator helps (§1.3). So: the embedder runs in f32 on every rung. On TensorRT that means the embedder's engine is built without fp16 while the detector's is built with it; on the Hexagon it means the embedder is not on the NPU at all — it runs on the ORT CPU rung at 9 ms, and the ladder's "one backend per device" is, precisely, one backend per model role, with the embedder pinned. A w600k_mbf embedding from any device is comparable with one from any other, which is the property the identity system, the calibration and the cross-device merge all rest on, and it is not for sale for 3 ms.

If S16 wants fp16 for the embedder later, the gate is written now: over the reference library's faces, the cosine between the f32 and fp16 embedding of the same crop exceeds 0.999 for 99.9% of faces and the calibration's fitted threshold moves by less than its own confidence interval (faces.md §8.3). Until measured, f32.

Segmentation and the scene model carry no identity across devices — their outputs are recomputed per image and never stored beyond the cache — so they take whatever the rung offers, int8 included, subject to §10's own acceptance.


8. Crate shape — core/dr-inference-engine

The seam is the same shape as storage.md's: a small crate below the consumers that is the only place naming a provider, a library file or a vendor, with the consumers reduced to "give me a session for these bytes in this role".

core/dr-inference-engine
  src/lib.rs        Runtime (Tract | Onnx { lib, version }), Backend (rung), Role (Detector | Embedder | Segmenter)
  src/probe.rs      §4 — the ladder per platform, the session-build probe, the cache file
  src/engines.rs    §6 — background compilation, the cache directory, progress
  src/session.rs    open(role, bytes) -> ort::Session, applying the rung and the role's precision rule
  src/api.rs        the one unsafe block: dlopen libonnxruntime, fetch OrtApi, ort::set_api — or ort_tract::api()
  • dr-face and dr-segment delete their private install_backend and their direct Session::builder() calls and take an &dr_inference_engine::Sessions where they take model bytes today. Their tests keep tract — dr_inference_engine::Sessions::tract() is a constructor and the test-only path.
  • dr-inference-engine depends on ort with the same workspace features as today plus cuda, tensorrt, qnn: those features add option builders, not linking, under alternative-backend. Verified for the QNN, CUDA and TensorRT builders on 2026-09-19 — they go through the API table's generic SessionOptionsAppendExecutionProvider*. The NNAPI builder resolves a symbol directly and would not; it is not needed and is not enabled.
  • dr-ui owns the settings row, the about-screen line and the progress row; it holds one Sessions per process, created at launch, and passes it down. dr_ui::library gains inference_cache_dir() beside shared_face_models_dir(), on the same account-independent footing and for the same reason.
  • The Android entry point's install_bundled_models also extracts nothing new: jniLibs/ is loaded by the system loader, and dr-inference-engine on Android looks for libonnxruntime.so through dlopen by bare name first, which resolves to the APK's copy, before any directory.

9. Threads and memory

  • ONNX Runtime's intra-op pool is sized to the physical cores minus two on desktop and to the performance cores on Android (the X4 and the A720s; the A520s are for the compositor). tract's single thread today is the reason a batch index leaves nineteen cores idle; ORT CPU with the pool is the reason it will not. One session per model per process; Session::run is &mut self-free in ort and internally serialised, and the index job is the only caller.
  • A TensorRT session pins GPU memory for its workspace; the builder is capped at 512 MB on the reference 6 GB card and the cap is a setting, because the develop view's tiles share the card (NFR-RES-2). The engine cache on disk is bounded by the model set — six engines, ~80 MB — and needs no LRU.
  • The Hexagon rung sets QNN's performance mode to Burst for the duration of an index job and Default otherwise; a 5 ms detector does not need the NPU clocked up between images.
  • The probe (§4) and the engine build (§6) run on one dedicated low-priority thread. They never share the index job's pool: a probe that competes with the job it is meant to speed up is the frame-budget trap in a new coat.

10. What S16 measures

In order, with the gate each is:

# Question Gate
M1 Does one binary carry both tables? dlopen + set_api on Linux, Windows and Android; ort_tract::api() when the file is absent. Go / no-go for §3. If set_api cannot take a dynamically fetched table on some platform, that platform ships two binaries, and the cost is stated.
M2 Do the int8 SCRFD detectors, calibrated on real photographs, find the faces? faces.md §12.3's method over the reference library, per size band, against f32. Ship the int8 form for a detector only if it finds ≥ 97% of the f32 detector's faces above 40 px and the difference is not concentrated in one band. Otherwise that detector's Hexagon rung is ORT CPU int8-free, and the table in §1.1 says what that costs.
M3 Does the embedder on ORT CPU beside a detector on the Hexagon (§7) produce embeddings within the f32 gate? It must — same graph, same arithmetic. This is a check that the plumbing did not quantise it by accident.
M4 Is a batch index on the tablet and on the desktop faster by the ratio §1 predicts, end to end, decode included? The face index over the reference library (18,143 faces): report wall-clock on tract, on the floor and on the selected rung, and where the time went. The prediction is that decode becomes the bottleneck on both; if it does not, say why.
M5 Does the first-run sequence (§6) hold: nothing blocks the first frame, engines land, requests move up, a running job does not switch? Observed on both devices with the app's own progress row, and with the cache directory deleted between runs.
M6 What does the APK weigh with the QNN runtime, and does a non-Qualcomm Android device (any one) still launch and index on the floor? Size reported; launch verified on one non-Qualcomm device or an emulator.
M7 The segmentation and scene models at int8 on the Hexagon: does the mask boundary move? segmentation.md §6's IoU against f32 over its corpus. Ship int8 for a model only above the IoU floor that document set for arm B.

M1 and M2 are the ones the rest is conditional on, and M2 is the one that needs a person.


11. Order

  1. dr-inference-engine with the two tables and the floor — set_api from a dlopened runtime, tract otherwise, ORT CPU as the only rung. Consumers moved over; tests unchanged. This alone is the 3–10× and is the build most of the value sits in. M1.
  2. The probe and its cache (§4), with the settings row and the about line. Still CPU-only; the ladder has one rung. M5's first half.
  3. tools/quantise-models.sh and the opset upgrade; the int8 detectors calibrated and measured. M2, M7. This is the step with a person in it and it runs in parallel with 4.
  4. The Hexagon rung, the QNN runtime in the APK, the context-binary cache. M3, M6.
  5. The TensorRT and CUDA rungs on desktop, the engine cache, the first-run sequence. M5's second half.
  6. M4 last, on both devices, and the number goes in this document.

The Windows installer and the Flatpak manifest are touched in steps 1 and 5 only, and only to add a file each; the Arch package likewise.


12. Register entries

FR-INF-1 — Runtime selection. On launch the application shall determine, per device and without blocking the first frame, the fastest inference backend that can build and run a session for the shipped models, by attempting it; shall record and reuse that determination until the runtime, driver, hardware or models change; and shall display the backend in use in Settings and on the about screen. Acceptance: §10 M1 and M5.

FR-INF-2 — Derived engines. Backends that require device-specific compilation shall compile in the background after selection, shall serve requests from the next lower backend until each engine is ready, and shall not change the backend of a job in progress. Acceptance: M5.

FR-INF-3 — Model forms. Quantised model forms are produced at release time from real calibration data and are shipped only when they meet §10's accuracy gates against the canonical form; the application never quantises on the device. Acceptance: M2, M7.

NFR-INF-1 — Embedding comparability. Face embeddings shall be computed at a precision whose deviation from the f32 reference is within §7's gate, on every backend, so that embeddings from any device are comparable. Acceptance: M3.

D13 — updated. The runtime half is reopened to the extent of §3: the Rust build stays C-free under alternative-backend; packages may install a dynamically loaded ONNX Runtime and, per §3.1, the Qualcomm QNN runtime; the NVIDIA libraries are not bundled. The licensing half is unchanged.


13. Requirements touched

FR-CULL-8 (indexing time is what this exists to change), FR-CULL-9 (NFR-INF-1 is the guard on its calibration), NFR-RES-2 (TensorRT workspace against the develop view's tiles), NFR-RES-3 (background compilation on Android), NFR-SEC-5 (the about screen's line), NFR-COMPAT-2 (each channel gains a runtime file), FR-PLAT-AND-1 (unchanged; the runtime lives in the APK, not in storage), ARCH §6.1 (the once-per-image budget this was sized against no longer binds; what could run per frame is a separate question this document does not open).