Add the MIGraphX rung for AMD GPUs
Benchmarks / CPU and I/O (per commit) (push) Failing after 6m20s
Benchmarks / Frame budget (on demand) (push) Skipped
Build and test / Desktop (Linux) (push) Failing after 45s
Build and test / Layer separation (push) Successful in 26s
Traceability / Requirement traces (push) Failing after 46s
🐳 Android image / Build and push (push) Successful in 1s
Build and test / android-image (push) Successful in 1s
🐳 Windows image / Build and push (push) Successful in 1s
Build and test / windows-image (push) Successful in 1s
Build and test / Android (aarch64) (push) Failing after 2m19s
Build and test / Windows (x86_64, cross) (push) Failing after 3m2s

Measured on a Radeon RX 7900 XT against Arch's onnxruntime-rocm 1.29
(docs/inference.md §1.3): MIGraphX fp16 runs the detectors at 2.4–3.4 ms
against 10–58 ms on the CPU provider, the inpainter at 8 ms against 514,
with a 15–135 s compile per graph the first time and under a second from
its cache after. A compiling rung on TensorRT's terms, wired the same way.

The ROCm execution provider is gone (removed in ONNX Runtime 1.23), so the
AMD ladder is MIGraphX then the CPU, with no non-compiling rung between.

MIGraphX is registered through the runtime's generic key/value entry
point rather than ort's builder: 1.29 reads the legacy options struct for
its precision flags only, and the compiled-program cache directory
(`migraphx_model_cache_dir`) only travels the generic way. The provider's
cache key omits the precision, so f32 and fp16 programs get their own
directories. The probe fingerprint now includes the provider libraries
beside the runtime and the ROCm version, since a distribution's CPU and
ROCm builds are the same file at the same path.

`status().failed` reports only the rungs above the selection, so an AMD
desktop's About line says why MIGraphX won rather than that the NVIDIA
providers are not in the build.

Two examples: `ep_probe` times each provider cold and from cache, and
`ladder` drives `init` as the app does to watch the first-run sequence.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
2026-09-20 19:23:00 +02:00
co-authored by Claude Opus 5
parent 5b4ad11853
commit 39a22875b1
10 changed files with 575 additions and 33 deletions
+75 -15
View File
@@ -75,7 +75,49 @@ TensorRT's job.
worst: nearly five minutes), because it is compiling an engine for this exact GPU. The engine caches to disk and the second load is milliseconds. That
number is what §6 is designed around.
### 1.3 What the numbers say
### 1.3 The desktop — Radeon RX 7900 XT, Threadripper 2920X, 24 threads · 2026-09-20
Arch's `onnxruntime-rocm` 1.29.0 against ROCm 7.2.4 and MIGraphX 7.2.3 (gfx1100), through the
same `ep_probe` harness (`core/dr-inference-engine/examples/ep_probe.rs`). Zero input, three
warm-ups, the median of 15 runs, on a machine doing nothing else. The build columns are the wall
clock of `Session` construction: cold is a MIGraphX compile of the graph for this GPU, cached is the
same session loading the program the cold build wrote.
| Model | ORT CPU f32 | MIGraphX f32 | **MIGraphX fp16** | Compile f32 / fp16 (s) | Cached load (s) |
|---|---|---|---|---|---|
| scrfd_500m (Fast) | 10.4 | 2.8 | **2.4** | 40 / 58 | 0.3 |
| scrfd_2.5g (Balanced) | 20.7 | 3.3 | **2.8** | 37 / 40 | 0.3 |
| scrfd_10g (Thorough) | 57.9 | 4.5 | **3.4** | 40 / 48 | 0.4 |
| arcface_mbf (per face) | 12.8 | 1.8 | 1.6 | 15 / 21 | 0.4 |
| yolo26n-seg | 49.3 | 8.4 | **7.5** | 110 / 136 | 0.9 |
| yolo26s-sem-ade20k | 55.8 | 4.8 | **3.8** | 50 / 60 | 0.5 |
| 2d106det (landmarks) | 9.9 | 1.2 | 1.0 | 16 / 21 | 0.2 |
| ocec_s (eye state) | 2.5 | 0.5 | 0.4 | 17 / 17 | 0.1 |
| xfeat-1024 | 26.8 | 10.5 | 9.9 | 37 / 53 | 0.3 |
| migan-512 (per tile) | 514 | 12.7 | **8.3** | 102 / 132 | 0.8 |
Three things the table settles.
- **The ROCm execution provider does not exist any more.** It was ONNX Runtime's CUDA-provider twin
for AMD, removed in the 1.23 release (AMD's builds dropped it from ROCm 7.1); 1.29's ROCm build ships
`libonnxruntime_providers_migraphx.so` and nothing else for AMD, and asking for `ROCm` answers
"not enabled in this build". So there is no non-compiling AMD rung to sit under MIGraphX the way
the CUDA provider sits under TensorRT: the AMD ladder is MIGraphX, then the CPU.
- **MIGraphX is a compiling provider, and its cache has to be asked for by name.** 15–135 s per
graph cold, under a second from its cache — TensorRT's shape exactly, and §6's design covers it.
Two things the provider does that the code has to know: `ort`'s builder fills the legacy options
struct, which 1.29 reads for the precision flags only, so the cache directory
(`migraphx_model_cache_dir`) reaches it only through the generic key/value registration; and
the cache key is the graph, the GPU and the MIGraphX version *without the precision*, so an fp16
session pointed at the f32 program's directory silently loads the f32 program (the first fp16
row measured here was that, before the directories were split).
- **fp16 is worth 10–35% over f32 on this card, not the 3× it is worth on TensorRT**, because
MIGraphX f32 is already 3–8× the CPU provider and the small graphs are launch-bound. The
detectors at 2.4–3.4 ms sit beside TensorRT fp16's 1.8–3.3 ms on the RTX 3050; the embedder is
1.6–1.8 ms on either precision and stays f32 (§7). The whole face pipeline for one image
(detector + landmarks + eyes + embedder) is under 6 ms.
### 1.4 What the numbers say
- **`tract` is single-threaded.** The tablet's one X4 core and one Raptor Lake core give the same
tract numbers. Replacing it with ONNX Runtime's CPU provider, *no accelerator involved*, is 3–6×
@@ -92,6 +134,8 @@ number is what §6 is designed around.
- **On NVIDIA, TensorRT fp16 ≈ 3× the CUDA provider**, and the CUDA provider ≈ 2× the
multi-threaded CPU; at fp16 the detectors are 1.8–3.3 ms with no quantisation at all. Both leave the twenty cores free for decoding during a batch index, which the table does not
show and which matters more than the ratio.
- **On AMD, MIGraphX fp16 is 4–17× the CPU provider** on the detectors and 60× on the
inpainter, with the same first-run compile cost as TensorRT and no rung between it and the CPU.
---
@@ -105,7 +149,8 @@ winning:
| Android, Qualcomm with a Hexagon the shipped QNN skel covers (V68–V81) | QNN HTP, int8 model | ORT CPU, f32 model | — | tract |
| Android, any other SoC | ORT CPU, f32 | — | — | tract |
| Linux / Windows, NVIDIA GPU | TensorRT, f32 model, fp16 engine | CUDA provider, f32 | ORT CPU, f32 | tract |
| Linux / Windows, no NVIDIA | ORT CPU, f32 | — | — | tract |
| Linux, AMD GPU with ROCm | MIGraphX, f32 model, fp16 program | ORT CPU, f32 | — | tract |
| Linux / Windows, no GPU stack | ORT CPU, f32 | — | — | tract |
| macOS ⁵ | ORT CPU, f32 | — | — | tract |
⁵ CoreML is the obvious rung and is unmeasured; it is listed so its absence is a gap and not an
@@ -113,8 +158,13 @@ oversight.
Deliberately **not** on any ladder, with the measurement that excluded each: NNAPI (no driver),
XNNPACK (slower than CPU, aborts on SCRFD), WebGPU (slower than CPU), the Adreno through QNN (works,
but never where the Hexagon does not also), CUDA int8 (slower than CUDA f32). A rung is added to this
table by a measurement on this page, not by a provider existing.
but never where the Hexagon does not also), CUDA int8 (slower than CUDA f32), the ROCm provider
(gone: §1.3). A rung is added to this table by a measurement on this page, not by a provider
existing.
The AMD ladder has no middle rung. TensorRT falls back to the CUDA provider while its engines
compile; MIGraphX has no such twin, so its fallback is the CPU provider, and the minute or two of
compiling on first run (§6) is spent at the floor's speed rather than at half the GPU's.
Two things the ladder is *not*: it is not a per-model choice — one backend serves every model on a
device, because §7's identity rule needs the detector and embedder on the same runtime for the
@@ -174,9 +224,10 @@ is cheaper than discovering it at packaging time.
| ONNX Runtime | MIT | Yes |
| Qualcomm QNN runtime (`com.qualcomm.qti:qnn-runtime` on Maven) | Qualcomm AI Engine Direct SDK licence — proprietary, redistribution permitted for applications using it | Yes for the APK, with the licence text shipped; not for a source distribution. **To be read in full, not summarised from memory, before the APK gains it.** |
| CUDA runtime, cuDNN, TensorRT | NVIDIA EULAs — redistributable with an application, with the licence text, not modifiable | Yes for a package that bundles them. 600 MB. The alternative is to load them from the user's system install if present and skip the rung otherwise — which is what §4's probe does anyway. |
| ROCm (HIP, MIOpen, rocBLAS), MIGraphX | MIT | Yes, but the HIP SDK MIGraphX needs is ~15 GB installed. Same answer as NVIDIA: the user's system install, or the rung is skipped. |
The position this takes: the **NVIDIA libraries are not bundled**. The desktop package probes for a
system CUDA/TensorRT install and uses it if it is version-compatible; a desktop without one runs on
The position this takes: the **GPU vendors' libraries are not bundled**. The desktop package probes for a
system CUDA/TensorRT or ROCm/MIGraphX install and uses it if it is version-compatible; a desktop without one runs on
ORT CPU, which is still 8–10× today. Bundling 600 MB for a rung that is 2× again is not a trade
worth making unmeasured, and it can be revisited by a measurement on a batch index. The **QNN
runtime is bundled** in the APK, because the Hexagon is the difference between a tablet that
@@ -237,6 +288,7 @@ of which form they load:
| f32 ONNX, shape-fixed, **opset ≥ 13** | `tools/fix-face-model-shapes.sh`, `tools/export-seg-model.sh` | Release time, once | Every rung except Hexagon |
| int8 QDQ ONNX, per-channel, uint8 activations | `tools/quantise-models.sh` (new) | Release time, once, **calibrated on real photographs** | Hexagon |
| TensorRT engine (`.engine`, per GPU architecture and TensorRT version) | The app, from the f32 file | First run on that device, in the background | TensorRT rung |
| MIGraphX program (`.mxr`, per GPU architecture, MIGraphX version and precision) | The app, from the f32 file | First run on that device, in the background | MIGraphX rung |
| QNN context binary | The app, from the int8 file | First run on that device, in the background | Hexagon rung |
Two rules.
@@ -256,9 +308,12 @@ prefers. That is a change to the canonical file and so a change to the shipped m
happens in the same model release as the int8 files.
**Compilation is a device-time step, and it is cached.** A TensorRT engine is specific to the GPU
it was built on and the TensorRT that built it; a QNN context binary is specific to the Hexagon
generation. Neither can ship. Both are built by the app the first time that rung is selected, in
the background (§6), and written beside the probe cache keyed by the same inputs. They are
it was built on and the TensorRT that built it; a MIGraphX program to the GPU and the MIGraphX
that built it; a QNN context binary to the Hexagon generation. None can ship. All are built by
the app the first time that rung is selected, in the background (§6), and written beside the
probe cache keyed by the same inputs. MIGraphX's own key leaves out the precision, so the app
gives its f32 and fp16 programs separate directories — otherwise the embedder's f32 build would
be served the detector's fp16 program, or the reverse. They are
**derived, disposable, and regenerable**: deleting the cache directory costs the next launch a
rebuild and nothing else, and the directory is excluded from anything that syncs (it is a peer of
`thumbs`, not of the catalog).
@@ -267,17 +322,18 @@ rebuild and nothing else, and the directory is excluded from anything that syncs
## 6. First run — building engines without the user waiting for them
The sequence on a device where a compiling rung (TensorRT, Hexagon) is selected:
The sequence on a device where a compiling rung (TensorRT, MIGraphX, Hexagon) is selected:
1. **Launch.** The runtime loads; the probe (§4) starts in the background; the app serves every
model request from the floor. Face indexing, segmentation and scene grading all work, at
today's speed or better (ORT CPU).
2. **Probe reports** — say, TensorRT. The compiling rung is now *selected* but has **no engines**.
Model requests continue on the fallback rung below it (CUDA provider for TensorRT; ORT CPU for
Hexagon), which needs no compilation and is already faster than the floor.
MIGraphX and Hexagon), which needs no compilation and is already faster than the floor.
3. **Engines build**, one model at a time, on a single low-priority background thread, smallest
model first so the detector — the one that runs per image — is ready soonest. On the reference
desktop that is ~1 minute for the first detector and ~10 minutes for all six at fp16; on the tablet the QNN
desktop that is ~1 minute for the first detector and ~10 minutes for all six at fp16; on the AMD
desktop 40 s for the first detector and ~8 minutes for the set; on the tablet the QNN
context binaries take 0.8–1.7 s each and the whole set is ready before the user has opened a
library. Each engine is written to a temporary name and renamed into place, so a request never
sees a half-written file.
@@ -313,7 +369,7 @@ enough to be *offered* at all. f32 on tract, ORT CPU, CUDA and TensorRT-f32 are
same graph, the same arithmetic, differences at the last bit.
**The embedder** is where comparability across devices is the whole point, and it is the one
model that no accelerator helps (§1.3). So: **the embedder runs in f32 on every rung.** On TensorRT
model that no accelerator helps (§1.4). So: **the embedder runs in f32 on every rung.** On TensorRT
that means the embedder's engine is built without fp16 while the detector's is built with it; on
the Hexagon it means the embedder is not on the NPU at all — it runs on the ORT CPU rung at 9 ms,
and the ladder's "one backend per device" is, precisely, one backend *per model role*, with the
@@ -344,7 +400,7 @@ core/dr-inference-engine
src/probe.rs §4 — the ladder per platform, the session-build probe, the cache file
src/engines.rs §6 — background compilation, the cache directory, progress
src/session.rs open(role, bytes) -> ort::Session, applying the rung and the role's precision rule
src/api.rs the one unsafe block: dlopen libonnxruntime, fetch OrtApi, ort::set_api — or ort_tract::api()
src/api.rs the unsafe block that matters: dlopen libonnxruntime, fetch OrtApi, ort::set_api — or ort_tract::api()
```
- `dr-face` and `dr-segment` **delete** their private `install_backend` and their direct
@@ -354,7 +410,11 @@ core/dr-inference-engine
`qnn`: those features add option builders, not linking, under `alternative-backend`. **Verified
for the QNN, CUDA and TensorRT builders on 2026-09-19** — they go through the API table's generic
`SessionOptionsAppendExecutionProvider*`. The NNAPI builder resolves a symbol directly and would
not; it is not needed and is not enabled.
not; it is not needed and is not enabled. MIGraphX uses no `ort` feature at all: `ort`'s builder
fills the legacy options struct, which ONNX Runtime 1.29 reads for the precision flags and
nothing else, and the compiled-program cache directory only travels through the generic
key/value entry point (`migraphx_model_cache_dir`). `session::migraphx` makes that one call
on the API table itself.
- `dr-ui` owns the settings row, the about-screen line and the progress row; it holds one
`Sessions` per process, created at launch, and passes it down. `dr_ui::library` gains
`inference_cache_dir()` beside `shared_face_models_dir()`, on the same account-independent