Merge: master at 0.13.6, with the drag-ghost file and the shared model lookup ported into the split modules
This commit is contained in:
+75
-15
@@ -75,7 +75,49 @@ TensorRT's job.
|
||||
worst: nearly five minutes), because it is compiling an engine for this exact GPU. The engine caches to disk and the second load is milliseconds. That
|
||||
number is what §6 is designed around.
|
||||
|
||||
### 1.3 What the numbers say
|
||||
### 1.3 The desktop — Radeon RX 7900 XT, Threadripper 2920X, 24 threads · 2026-09-20
|
||||
|
||||
Arch's `onnxruntime-rocm` 1.29.0 against ROCm 7.2.4 and MIGraphX 7.2.3 (gfx1100), through the
|
||||
same `ep_probe` harness (`core/dr-inference-engine/examples/ep_probe.rs`). Zero input, three
|
||||
warm-ups, the median of 15 runs, on a machine doing nothing else. The build columns are the wall
|
||||
clock of `Session` construction: cold is a MIGraphX compile of the graph for this GPU, cached is the
|
||||
same session loading the program the cold build wrote.
|
||||
|
||||
| Model | ORT CPU f32 | MIGraphX f32 | **MIGraphX fp16** | Compile f32 / fp16 (s) | Cached load (s) |
|
||||
|---|---|---|---|---|---|
|
||||
| scrfd_500m (Fast) | 10.4 | 2.8 | **2.4** | 40 / 58 | 0.3 |
|
||||
| scrfd_2.5g (Balanced) | 20.7 | 3.3 | **2.8** | 37 / 40 | 0.3 |
|
||||
| scrfd_10g (Thorough) | 57.9 | 4.5 | **3.4** | 40 / 48 | 0.4 |
|
||||
| arcface_mbf (per face) | 12.8 | 1.8 | 1.6 | 15 / 21 | 0.4 |
|
||||
| yolo26n-seg | 49.3 | 8.4 | **7.5** | 110 / 136 | 0.9 |
|
||||
| yolo26s-sem-ade20k | 55.8 | 4.8 | **3.8** | 50 / 60 | 0.5 |
|
||||
| 2d106det (landmarks) | 9.9 | 1.2 | 1.0 | 16 / 21 | 0.2 |
|
||||
| ocec_s (eye state) | 2.5 | 0.5 | 0.4 | 17 / 17 | 0.1 |
|
||||
| xfeat-1024 | 26.8 | 10.5 | 9.9 | 37 / 53 | 0.3 |
|
||||
| migan-512 (per tile) | 514 | 12.7 | **8.3** | 102 / 132 | 0.8 |
|
||||
|
||||
Three things the table settles.
|
||||
|
||||
- **The ROCm execution provider does not exist any more.** It was ONNX Runtime's CUDA-provider twin
|
||||
for AMD, removed in the 1.23 release (AMD's builds dropped it from ROCm 7.1); 1.29's ROCm build ships
|
||||
`libonnxruntime_providers_migraphx.so` and nothing else for AMD, and asking for `ROCm` answers
|
||||
"not enabled in this build". So there is no non-compiling AMD rung to sit under MIGraphX the way
|
||||
the CUDA provider sits under TensorRT: the AMD ladder is MIGraphX, then the CPU.
|
||||
- **MIGraphX is a compiling provider, and its cache has to be asked for by name.** 15–135 s per
|
||||
graph cold, under a second from its cache — TensorRT's shape exactly, and §6's design covers it.
|
||||
Two things the provider does that the code has to know: `ort`'s builder fills the legacy options
|
||||
struct, which 1.29 reads for the precision flags only, so the cache directory
|
||||
(`migraphx_model_cache_dir`) reaches it only through the generic key/value registration; and
|
||||
the cache key is the graph, the GPU and the MIGraphX version *without the precision*, so an fp16
|
||||
session pointed at the f32 program's directory silently loads the f32 program (the first fp16
|
||||
row measured here was that, before the directories were split).
|
||||
- **fp16 is worth 10–35% over f32 on this card, not the 3× it is worth on TensorRT**, because
|
||||
MIGraphX f32 is already 3–8× the CPU provider and the small graphs are launch-bound. The
|
||||
detectors at 2.4–3.4 ms sit beside TensorRT fp16's 1.8–3.3 ms on the RTX 3050; the embedder is
|
||||
1.6–1.8 ms on either precision and stays f32 (§7). The whole face pipeline for one image
|
||||
(detector + landmarks + eyes + embedder) is under 6 ms.
|
||||
|
||||
### 1.4 What the numbers say
|
||||
|
||||
- **`tract` is single-threaded.** The tablet's one X4 core and one Raptor Lake core give the same
|
||||
tract numbers. Replacing it with ONNX Runtime's CPU provider, *no accelerator involved*, is 3–6×
|
||||
@@ -92,6 +134,8 @@ number is what §6 is designed around.
|
||||
- **On NVIDIA, TensorRT fp16 ≈ 3× the CUDA provider**, and the CUDA provider ≈ 2× the
|
||||
multi-threaded CPU; at fp16 the detectors are 1.8–3.3 ms with no quantisation at all. Both leave the twenty cores free for decoding during a batch index, which the table does not
|
||||
show and which matters more than the ratio.
|
||||
- **On AMD, MIGraphX fp16 is 4–17× the CPU provider** on the detectors and 60× on the
|
||||
inpainter, with the same first-run compile cost as TensorRT and no rung between it and the CPU.
|
||||
|
||||
---
|
||||
|
||||
@@ -105,7 +149,8 @@ winning:
|
||||
| Android, Qualcomm with a Hexagon the shipped QNN skel covers (V68–V81) | QNN HTP, int8 model | ORT CPU, f32 model | — | tract |
|
||||
| Android, any other SoC | ORT CPU, f32 | — | — | tract |
|
||||
| Linux / Windows, NVIDIA GPU | TensorRT, f32 model, fp16 engine | CUDA provider, f32 | ORT CPU, f32 | tract |
|
||||
| Linux / Windows, no NVIDIA | ORT CPU, f32 | — | — | tract |
|
||||
| Linux, AMD GPU with ROCm | MIGraphX, f32 model, fp16 program | ORT CPU, f32 | — | tract |
|
||||
| Linux / Windows, no GPU stack | ORT CPU, f32 | — | — | tract |
|
||||
| macOS ⁵ | ORT CPU, f32 | — | — | tract |
|
||||
|
||||
⁵ CoreML is the obvious rung and is unmeasured; it is listed so its absence is a gap and not an
|
||||
@@ -113,8 +158,13 @@ oversight.
|
||||
|
||||
Deliberately **not** on any ladder, with the measurement that excluded each: NNAPI (no driver),
|
||||
XNNPACK (slower than CPU, aborts on SCRFD), WebGPU (slower than CPU), the Adreno through QNN (works,
|
||||
but never where the Hexagon does not also), CUDA int8 (slower than CUDA f32). A rung is added to this
|
||||
table by a measurement on this page, not by a provider existing.
|
||||
but never where the Hexagon does not also), CUDA int8 (slower than CUDA f32), the ROCm provider
|
||||
(gone: §1.3). A rung is added to this table by a measurement on this page, not by a provider
|
||||
existing.
|
||||
|
||||
The AMD ladder has no middle rung. TensorRT falls back to the CUDA provider while its engines
|
||||
compile; MIGraphX has no such twin, so its fallback is the CPU provider, and the minute or two of
|
||||
compiling on first run (§6) is spent at the floor's speed rather than at half the GPU's.
|
||||
|
||||
Two things the ladder is *not*: it is not a per-model choice — one backend serves every model on a
|
||||
device, because §7's identity rule needs the detector and embedder on the same runtime for the
|
||||
@@ -174,9 +224,10 @@ is cheaper than discovering it at packaging time.
|
||||
| ONNX Runtime | MIT | Yes |
|
||||
| Qualcomm QNN runtime (`com.qualcomm.qti:qnn-runtime` on Maven) | Qualcomm AI Engine Direct SDK licence — proprietary, redistribution permitted for applications using it | Yes for the APK, with the licence text shipped; not for a source distribution. **To be read in full, not summarised from memory, before the APK gains it.** |
|
||||
| CUDA runtime, cuDNN, TensorRT | NVIDIA EULAs — redistributable with an application, with the licence text, not modifiable | Yes for a package that bundles them. 600 MB. The alternative is to load them from the user's system install if present and skip the rung otherwise — which is what §4's probe does anyway. |
|
||||
| ROCm (HIP, MIOpen, rocBLAS), MIGraphX | MIT | Yes, but the HIP SDK MIGraphX needs is ~15 GB installed. Same answer as NVIDIA: the user's system install, or the rung is skipped. |
|
||||
|
||||
The position this takes: the **NVIDIA libraries are not bundled**. The desktop package probes for a
|
||||
system CUDA/TensorRT install and uses it if it is version-compatible; a desktop without one runs on
|
||||
The position this takes: the **GPU vendors' libraries are not bundled**. The desktop package probes for a
|
||||
system CUDA/TensorRT or ROCm/MIGraphX install and uses it if it is version-compatible; a desktop without one runs on
|
||||
ORT CPU, which is still 8–10× today. Bundling 600 MB for a rung that is 2× again is not a trade
|
||||
worth making unmeasured, and it can be revisited by a measurement on a batch index. The **QNN
|
||||
runtime is bundled** in the APK, because the Hexagon is the difference between a tablet that
|
||||
@@ -237,6 +288,7 @@ of which form they load:
|
||||
| f32 ONNX, shape-fixed, **opset ≥ 13** | `tools/fix-face-model-shapes.sh`, `tools/export-seg-model.sh` | Release time, once | Every rung except Hexagon |
|
||||
| int8 QDQ ONNX, per-channel, uint8 activations | `tools/quantise-models.sh` (new) | Release time, once, **calibrated on real photographs** | Hexagon |
|
||||
| TensorRT engine (`.engine`, per GPU architecture and TensorRT version) | The app, from the f32 file | First run on that device, in the background | TensorRT rung |
|
||||
| MIGraphX program (`.mxr`, per GPU architecture, MIGraphX version and precision) | The app, from the f32 file | First run on that device, in the background | MIGraphX rung |
|
||||
| QNN context binary | The app, from the int8 file | First run on that device, in the background | Hexagon rung |
|
||||
|
||||
Two rules.
|
||||
@@ -256,9 +308,12 @@ prefers. That is a change to the canonical file and so a change to the shipped m
|
||||
happens in the same model release as the int8 files.
|
||||
|
||||
**Compilation is a device-time step, and it is cached.** A TensorRT engine is specific to the GPU
|
||||
it was built on and the TensorRT that built it; a QNN context binary is specific to the Hexagon
|
||||
generation. Neither can ship. Both are built by the app the first time that rung is selected, in
|
||||
the background (§6), and written beside the probe cache keyed by the same inputs. They are
|
||||
it was built on and the TensorRT that built it; a MIGraphX program to the GPU and the MIGraphX
|
||||
that built it; a QNN context binary to the Hexagon generation. None can ship. All are built by
|
||||
the app the first time that rung is selected, in the background (§6), and written beside the
|
||||
probe cache keyed by the same inputs. MIGraphX's own key leaves out the precision, so the app
|
||||
gives its f32 and fp16 programs separate directories — otherwise the embedder's f32 build would
|
||||
be served the detector's fp16 program, or the reverse. They are
|
||||
**derived, disposable, and regenerable**: deleting the cache directory costs the next launch a
|
||||
rebuild and nothing else, and the directory is excluded from anything that syncs (it is a peer of
|
||||
`thumbs`, not of the catalog).
|
||||
@@ -267,17 +322,18 @@ rebuild and nothing else, and the directory is excluded from anything that syncs
|
||||
|
||||
## 6. First run — building engines without the user waiting for them
|
||||
|
||||
The sequence on a device where a compiling rung (TensorRT, Hexagon) is selected:
|
||||
The sequence on a device where a compiling rung (TensorRT, MIGraphX, Hexagon) is selected:
|
||||
|
||||
1. **Launch.** The runtime loads; the probe (§4) starts in the background; the app serves every
|
||||
model request from the floor. Face indexing, segmentation and scene grading all work, at
|
||||
today's speed or better (ORT CPU).
|
||||
2. **Probe reports** — say, TensorRT. The compiling rung is now *selected* but has **no engines**.
|
||||
Model requests continue on the fallback rung below it (CUDA provider for TensorRT; ORT CPU for
|
||||
Hexagon), which needs no compilation and is already faster than the floor.
|
||||
MIGraphX and Hexagon), which needs no compilation and is already faster than the floor.
|
||||
3. **Engines build**, one model at a time, on a single low-priority background thread, smallest
|
||||
model first so the detector — the one that runs per image — is ready soonest. On the reference
|
||||
desktop that is ~1 minute for the first detector and ~10 minutes for all six at fp16; on the tablet the QNN
|
||||
desktop that is ~1 minute for the first detector and ~10 minutes for all six at fp16; on the AMD
|
||||
desktop 40 s for the first detector and ~8 minutes for the set; on the tablet the QNN
|
||||
context binaries take 0.8–1.7 s each and the whole set is ready before the user has opened a
|
||||
library. Each engine is written to a temporary name and renamed into place, so a request never
|
||||
sees a half-written file.
|
||||
@@ -313,7 +369,7 @@ enough to be *offered* at all. f32 on tract, ORT CPU, CUDA and TensorRT-f32 are
|
||||
same graph, the same arithmetic, differences at the last bit.
|
||||
|
||||
**The embedder** is where comparability across devices is the whole point, and it is the one
|
||||
model that no accelerator helps (§1.3). So: **the embedder runs in f32 on every rung.** On TensorRT
|
||||
model that no accelerator helps (§1.4). So: **the embedder runs in f32 on every rung.** On TensorRT
|
||||
that means the embedder's engine is built without fp16 while the detector's is built with it; on
|
||||
the Hexagon it means the embedder is not on the NPU at all — it runs on the ORT CPU rung at 9 ms,
|
||||
and the ladder's "one backend per device" is, precisely, one backend *per model role*, with the
|
||||
@@ -344,7 +400,7 @@ core/dr-inference-engine
|
||||
src/probe.rs §4 — the ladder per platform, the session-build probe, the cache file
|
||||
src/engines.rs §6 — background compilation, the cache directory, progress
|
||||
src/session.rs open(role, bytes) -> ort::Session, applying the rung and the role's precision rule
|
||||
src/api.rs the one unsafe block: dlopen libonnxruntime, fetch OrtApi, ort::set_api — or ort_tract::api()
|
||||
src/api.rs the unsafe block that matters: dlopen libonnxruntime, fetch OrtApi, ort::set_api — or ort_tract::api()
|
||||
```
|
||||
|
||||
- `dr-face` and `dr-segment` **delete** their private `install_backend` and their direct
|
||||
@@ -354,7 +410,11 @@ core/dr-inference-engine
|
||||
`qnn`: those features add option builders, not linking, under `alternative-backend`. **Verified
|
||||
for the QNN, CUDA and TensorRT builders on 2026-09-19** — they go through the API table's generic
|
||||
`SessionOptionsAppendExecutionProvider*`. The NNAPI builder resolves a symbol directly and would
|
||||
not; it is not needed and is not enabled.
|
||||
not; it is not needed and is not enabled. MIGraphX uses no `ort` feature at all: `ort`'s builder
|
||||
fills the legacy options struct, which ONNX Runtime 1.29 reads for the precision flags and
|
||||
nothing else, and the compiled-program cache directory only travels through the generic
|
||||
key/value entry point (`migraphx_model_cache_dir`). `session::migraphx` makes that one call
|
||||
on the API table itself.
|
||||
- `dr-ui` owns the settings row, the about-screen line and the progress row; it holds one
|
||||
`Sessions` per process, created at launch, and passes it down. `dr_ui::library` gains
|
||||
`inference_cache_dir()` beside `shared_face_models_dir()`, on the same account-independent
|
||||
|
||||
+21
-6
@@ -538,9 +538,24 @@ each loss weighting did, and the two runs abandoned (blur under L1 in
|
||||
the hole; a brick pattern under a strong adversarial term against a
|
||||
discriminator that had not learned) — is `runs/` in `darkroom-infill`.
|
||||
|
||||
**What is still wrong.** The ground fill is softer than its context —
|
||||
texture, not structure, is what a night on a laptop GPU could not finish.
|
||||
The levers, in order: a discriminator that learns (a pretrained one —
|
||||
MI-GAN's own from the unfused checkpoint — instead of a PatchGAN from
|
||||
scratch), feature matching, and more steps at 512. FR-MRG-4's
|
||||
*experimental* stays.
|
||||
**Second model, the same day.** The morning's fill was soft in the deep
|
||||
ground bands. The afternoon's run trained the generator against MI-GAN's
|
||||
own pretrained discriminator (non-saturating loss, lazy R1, feature
|
||||
matching), with fresh noise inputs while training and flip/translation
|
||||
augmentation of the discriminator's input — both needed, or the generator
|
||||
settles into a periodic texture the discriminator cannot see. The shipped
|
||||
weights (step 4 750 of `runs/border-v6`) are level with the stock model on
|
||||
LPIPS (edge 0.125 / corner 0.187 against 0.121 / 0.183) while keeping the
|
||||
PSNR gain (edge 18.0 / corner 15.7 against 16.9 / 14.6). On the fixture
|
||||
the ground bands now carry texture at the right tone; at 1:1 a faint
|
||||
regular hatch is visible in the deepest part.
|
||||
|
||||
**What is still wrong.** The hatch, and any deep textured void the
|
||||
generator must invent. The better answer for those is not generative:
|
||||
seed the void with the picture's own texture in hexagonal cells, let the
|
||||
discriminator rank the candidates, and let the generator heal only the
|
||||
gaps between cells — built and measured in `darkroom-infill`
|
||||
(`infill/hexfill.py`), the most convincing scree corner produced so far,
|
||||
and the next thing to port into `dr_pano::fill` (it needs the
|
||||
discriminator as a second model, ~80 MB fp16). FR-MRG-4's *experimental*
|
||||
stays.
|
||||
|
||||
+12
-12
File diff suppressed because one or more lines are too long
Reference in New Issue
Block a user