Add the MIGraphX rung for AMD GPUs
Benchmarks / CPU and I/O (per commit) (push) Failing after 6m20s
Benchmarks / Frame budget (on demand) (push) Skipped
Build and test / Desktop (Linux) (push) Failing after 45s
Build and test / Layer separation (push) Successful in 26s
Traceability / Requirement traces (push) Failing after 46s
🐳 Android image / Build and push (push) Successful in 1s
Build and test / android-image (push) Successful in 1s
🐳 Windows image / Build and push (push) Successful in 1s
Build and test / windows-image (push) Successful in 1s
Build and test / Android (aarch64) (push) Failing after 2m19s
Build and test / Windows (x86_64, cross) (push) Failing after 3m2s

Measured on a Radeon RX 7900 XT against Arch's onnxruntime-rocm 1.29
(docs/inference.md §1.3): MIGraphX fp16 runs the detectors at 2.4–3.4 ms
against 10–58 ms on the CPU provider, the inpainter at 8 ms against 514,
with a 15–135 s compile per graph the first time and under a second from
its cache after. A compiling rung on TensorRT's terms, wired the same way.

The ROCm execution provider is gone (removed in ONNX Runtime 1.23), so the
AMD ladder is MIGraphX then the CPU, with no non-compiling rung between.

MIGraphX is registered through the runtime's generic key/value entry
point rather than ort's builder: 1.29 reads the legacy options struct for
its precision flags only, and the compiled-program cache directory
(`migraphx_model_cache_dir`) only travels the generic way. The provider's
cache key omits the precision, so f32 and fp16 programs get their own
directories. The probe fingerprint now includes the provider libraries
beside the runtime and the ROCm version, since a distribution's CPU and
ROCm builds are the same file at the same path.

`status().failed` reports only the rungs above the selection, so an AMD
desktop's About line says why MIGraphX won rather than that the NVIDIA
providers are not in the build.

Two examples: `ep_probe` times each provider cold and from cache, and
`ladder` drives `init` as the app does to watch the first-run sequence.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
2026-09-20 19:23:00 +02:00
co-authored by Claude Opus 5
parent 5b4ad11853
commit 39a22875b1
10 changed files with 575 additions and 33 deletions
+24
View File
@@ -89,6 +89,30 @@ three `405`s before each `MOVE`. The backend now remembers the collections
it has confirmed (`known_dirs`) for its lifetime, which is one job. When a
per-file operation has a per-batch precondition, satisfy it once.
## Providers: read the runtime's source for the version on disk, not the binding
Two things the MIGraphX rung (2026-09-20) got wrong before it was measured
right, both because `ort`'s builder was trusted to mean what its method
names say.
**A binding's option builder may fill a struct the runtime no longer
reads.** `ep::MIGraphX::with_save_model` sets fields of the legacy
`OrtMIGraphXProviderOptions`; ONNX Runtime 1.29 reads that struct for the
precision flags and ignores the rest, so every session compiled for 40 s
and the cache directory went nowhere. The option that works
(`migraphx_model_cache_dir`) exists only in the generic key/value
registration, which `session::migraphx` calls on the API table directly.
Before wiring a provider option, fetch the provider's source at the
runtime's exact version and find where the option is *read*.
**A provider's cache key may leave out what you are varying.** MIGraphX
keys a compiled program on graph, GPU and its own version — not precision.
The first fp16 measurement built in 0.3 s and matched f32 to the tenth of a
millisecond, because it had loaded the f32 program. A "from cache" build
that is suspiciously fast on the first run of a new configuration is a key
collision, not a fast provider; give each precision its own directory (the
engine does) and check the cache directory gained a file.
## Measuring
`cargo run --release -p dr-ui --example identity_bench -- CATALOG THUMBS`
Generated
+1
View File
@@ -1515,6 +1515,7 @@ dependencies = [
name = "dr-inference-engine"
version = "0.13.5"
dependencies = [
"env_logger",
"libloading",
"log",
"ort",
+9 -1
View File
@@ -28,7 +28,10 @@ ort-sys = { version = "2.0.0-rc.13", default-features = false, features = ["disa
# The NVIDIA rungs exist on the desktop only. These features add `ort`'s
# option builders and nothing else — no linking under `alternative-backend` —
# but an Android binary has no business carrying even the option names, and
# the packaging must never be tempted to (§2, §3.1).
# the packaging must never be tempted to (§2, §3.1). The AMD rung needs no
# feature: MIGraphX is registered through the runtime's generic key/value
# entry point (`session::migraphx`), because `ort`'s own builder cannot
# name the compiled-program cache.
[target.'cfg(not(target_os = "android"))'.dependencies]
ort = { workspace = true, features = ["cuda", "tensorrt"] }
@@ -42,3 +45,8 @@ default = ["tract"]
tract = ["dep:ort-tract"]
# Look for `libonnxruntime` on disk and hand its table to `ort`.
native = ["dep:libloading", "dep:ort-sys"]
[dev-dependencies]
# The `ep_probe` example prints the provider's own diagnostics, which is most
# of what a failed rung tells you.
env_logger.workspace = true
@@ -0,0 +1,207 @@
//! Time each execution provider a runtime offers, on the models this
//! repository ships — the measurement docs/inference.md §1 requires before a
//! rung is added to §2's ladder.
//!
//! DARKROOM_ORT_DIR=/usr/lib \
//! cargo run --release -p dr-inference-engine --features native,tract \
//! --example ep_probe -- models/face/scrfd_500m_640.onnx ...
//!
//! Prints one row per (model, provider): the median of timed runs after
//! warm-ups, and the build time, which for a compiling provider is the
//! number that decides whether it needs an engine cache. MIGraphX is built
//! twice per precision — cold, then again from the cache it just wrote —
//! so both numbers are on the page.
//!
//! The ROCm provider is not in the list: ONNX Runtime removed it in 1.23,
//! and 1.29's `onnxruntime-rocm` ships `libonnxruntime_providers_migraphx.so`
//! and nothing else for AMD.
use std::path::{Path, PathBuf};
use std::time::Instant;
#[derive(Clone, Copy, PartialEq)]
enum Ep {
Cpu,
MiGraphX,
MiGraphXFp16,
}
impl Ep {
fn label(self) -> &'static str {
match self {
Ep::Cpu => "CPU",
Ep::MiGraphX => "MIGraphX f32",
Ep::MiGraphXFp16 => "MIGraphX fp16",
}
}
}
fn build(ep: Ep, bytes: &[u8], threads: usize, cache: &Path) -> ort::Result<ort::session::Session> {
let mut b = ort::session::Session::builder()?.with_intra_threads(threads)?;
match ep {
Ep::Cpu => {}
Ep::MiGraphX => migraphx(&mut b, false, &cache.join("f32"))?,
Ep::MiGraphXFp16 => migraphx(&mut b, true, &cache.join("fp16"))?,
}
b.commit_from_memory(bytes)
}
/// Register MIGraphX through the generic key/value API. `ort`'s own
/// builder fills the legacy `OrtMIGraphXProviderOptions`, which 1.29 reads
/// for its precision flags and nothing else: the model cache directory —
/// the difference between a 40 s load and a 0.3 s one — only travels this
/// way. The cache key is the graph, the GPU and the MIGraphX version, not
/// the precision, so each precision gets its own directory.
fn migraphx(
b: &mut ort::session::builder::SessionBuilder,
fp16: bool,
cache: &Path,
) -> ort::Result<()> {
use ort::AsPointer;
use std::ffi::CString;
std::fs::create_dir_all(cache).map_err(|e| ort::Error::new(e.to_string()))?;
let keys = [c"migraphx_fp16_enable", c"migraphx_model_cache_dir"];
let values = [
CString::new(if fp16 { "1" } else { "0" }).unwrap(),
CString::new(cache.to_string_lossy().as_bytes()).unwrap(),
];
let key_ptrs: Vec<_> = keys.iter().map(|k| k.as_ptr()).collect();
let value_ptrs: Vec<_> = values.iter().map(|v| v.as_ptr()).collect();
// SAFETY: the documented C call, over arrays that outlive it; the
// runtime copies the strings into its own options map.
unsafe {
let status = (ort::api().SessionOptionsAppendExecutionProvider)(
b.ptr_mut(),
c"MIGraphX".as_ptr(),
key_ptrs.as_ptr(),
value_ptrs.as_ptr(),
keys.len(),
);
ort::Error::result_from_status(status)
}
}
/// Median of `runs` timed runs over zeros, in milliseconds, after warm-ups.
fn time(session: &mut ort::session::Session, warmups: usize, runs: usize) -> Result<f64, String> {
let shape: Vec<usize> = session.inputs()[0]
.dtype()
.tensor_shape()
.ok_or("input is not a tensor")?
.iter()
.map(|&d| if d > 0 { d as usize } else { 1 })
.collect();
let zeros = vec![0f32; shape.iter().product()];
let once = |s: &mut ort::session::Session| -> Result<f64, String> {
let input = ort::value::Tensor::from_array((shape.clone(), zeros.clone()))
.map_err(|e| e.to_string())?;
let t = Instant::now();
let out = s.run(ort::inputs![input]).map_err(|e| e.to_string())?;
let _ = out[0]
.try_extract_tensor::<f32>()
.map_err(|e| e.to_string())?;
Ok(t.elapsed().as_secs_f64() * 1e3)
};
for _ in 0..warmups {
once(session)?;
}
let mut times = Vec::with_capacity(runs);
for _ in 0..runs {
times.push(once(session)?);
}
times.sort_by(|a, b| a.partial_cmp(b).unwrap());
Ok(times[times.len() / 2])
}
fn first_line(s: &str) -> String {
s.lines().next().unwrap_or("").chars().take(120).collect()
}
fn main() {
env_logger::Builder::from_env(env_logger::Env::default().default_filter_or("info")).init();
let models: Vec<PathBuf> = std::env::args_os().skip(1).map(PathBuf::from).collect();
if models.is_empty() {
eprintln!("usage: ep_probe MODEL.onnx [MODEL.onnx ...]");
std::process::exit(2);
}
dr_inference_engine::ensure_runtime();
let runtime = dr_inference_engine::status().runtime;
println!("runtime: {}", runtime.label());
if !runtime.is_native() {
println!("(tract: no provider to compare; set DARKROOM_ORT_DIR)");
}
let threads = std::thread::available_parallelism()
.map(|n| n.get().saturating_sub(2).max(1))
.unwrap_or(1);
println!("intra-op threads: {threads}");
let cache = std::env::temp_dir().join("darkroom-ep-probe");
let _ = std::fs::remove_dir_all(&cache);
println!("compiled-program cache: {}\n", cache.display());
println!(
"{:<28} {:<15} {:>10} {:>10}",
"model", "provider", "build s", "median ms"
);
for model in &models {
let bytes = match std::fs::read(model) {
Ok(b) => b,
Err(e) => {
println!("{:<28} read failed: {e}", name(model));
continue;
}
};
// A compiling provider is built twice: the second build reads the
// program the first wrote, and its time is what a launch after the
// first costs.
let plan = [
(Ep::Cpu, false),
(Ep::MiGraphX, false),
(Ep::MiGraphX, true),
(Ep::MiGraphXFp16, false),
(Ep::MiGraphXFp16, true),
];
for (ep, cached) in plan {
let started = Instant::now();
match build(ep, &bytes, threads, &cache) {
Ok(mut session) => {
let built = started.elapsed().as_secs_f64();
match time(&mut session, 3, 15) {
Ok(ms) => println!(
"{:<28} {:<15} {:>10.1} {:>10.1}{}",
name(model),
ep.label(),
built,
ms,
if cached { " (from cache)" } else { "" }
),
Err(e) => println!(
"{:<28} {:<15} {:>10.1} {:>10} {}",
name(model),
ep.label(),
built,
"ran ✗",
first_line(&e)
),
}
}
Err(e) => println!(
"{:<28} {:<15} {:>21} {}",
name(model),
ep.label(),
"build ✗",
first_line(&e.to_string())
),
}
}
println!();
}
}
fn name(p: &Path) -> String {
p.file_name()
.unwrap_or(p.as_os_str())
.to_string_lossy()
.into_owned()
}
+100
View File
@@ -0,0 +1,100 @@
//! Walk the ladder as the app does — probe, engines, then a session — and
//! say what each step chose. The M5 check of docs/inference.md §6 without
//! the app around it.
//!
//! DARKROOM_ORT_DIR=/usr/lib \
//! cargo run --release -p dr-inference-engine --features native,tract \
//! --example ladder -- CACHE_DIR models/face/scrfd_500m_640.onnx [MODEL.onnx ...]
//!
//! Every model named is a `Detector` for the config's purposes, which is
//! enough to see the rung taken, the engines compiled and a session land
//! on it. Delete `CACHE_DIR` to see the first run again; keep it to see the
//! second.
use std::path::PathBuf;
use std::time::{Duration, Instant};
fn main() {
env_logger::Builder::from_env(env_logger::Env::default().default_filter_or("info")).init();
let mut args = std::env::args_os().skip(1).map(PathBuf::from);
let (Some(cache_dir), models) = (args.next(), args.collect::<Vec<_>>()) else {
eprintln!("usage: ladder CACHE_DIR MODEL.onnx [MODEL.onnx ...]");
std::process::exit(2);
};
if models.is_empty() {
eprintln!("usage: ladder CACHE_DIR MODEL.onnx [MODEL.onnx ...]");
std::process::exit(2);
}
let runtime_dirs: Vec<PathBuf> = std::env::var_os("DARKROOM_ORT_DIR")
.map(PathBuf::from)
.into_iter()
.collect();
let started = Instant::now();
dr_inference_engine::init(dr_inference_engine::Config {
runtime_dirs,
cache_dir: cache_dir.clone(),
models: models
.iter()
.map(|p| (dr_inference_engine::Role::Detector, p.clone()))
.collect(),
embedded: Vec::new(),
ceiling: None,
threads: 0,
decay: Duration::ZERO,
});
let mut last = String::new();
loop {
let s = dr_inference_engine::status();
let line = format!(
"{} · {} · engines {}/{}{}",
s.line(),
if s.probing {
"probing"
} else {
s.reason.as_str()
},
s.engines.0,
s.engines.1,
if s.failed.is_empty() {
String::new()
} else {
format!(
" · tried {}",
s.failed
.iter()
.map(|(r, why)| format!("{}: {why}", r.label()))
.collect::<Vec<_>>()
.join(" · ")
)
}
);
if line != last {
println!("{:>6.1} s {line}", started.elapsed().as_secs_f64());
last = line;
}
if !s.probing && s.engines.0 >= s.engines.1 {
break;
}
std::thread::sleep(Duration::from_millis(500));
}
for path in &models {
let bytes = std::fs::read(path).expect("read model");
let t = Instant::now();
let model = dr_inference_engine::open(
dr_inference_engine::Role::Detector,
dr_inference_engine::Form::F32,
&bytes,
)
.expect("open model");
let acquired = model.acquire().expect("acquire session");
println!(
"{} on {} in {:.2} s",
path.file_name().unwrap().to_string_lossy(),
acquired.rung().label(),
t.elapsed().as_secs_f64()
);
}
}
+51 -9
View File
@@ -3,8 +3,8 @@
//!
//! Consumers ask for a session by [`Role`] and get `ort`'s `Session` back;
//! what built it — tract on one core, ONNX Runtime's CPU pool, a TensorRT
//! engine, the Hexagon — is this crate's business and shows up in
//! [`status`] for the settings row and nowhere else.
//! engine, a MIGraphX program, the Hexagon — is this crate's business and
//! shows up in [`status`] for the settings row and nowhere else.
//!
//! The shape follows §3 of the spec: `ort` links nothing (`alternative-backend`),
//! and the first call hands it an API table from either a `libonnxruntime`
@@ -60,7 +60,9 @@ pub enum Form {
/// A rung of the ladder (§2). Ordered: a user override names the highest rung
/// the probe may take, and a compiling rung falls back to the one below it
/// until its engine exists.
/// until its engine exists. The order is within a vendor's ladder — a
/// machine has NVIDIA rungs or an AMD rung, never both — so a ceiling is
/// read as "no higher than this on whichever ladder the device has".
#[derive(Clone, Copy, Debug, PartialEq, Eq, Hash, PartialOrd, Ord, Serialize, Deserialize)]
pub enum Rung {
/// ONNX Runtime's CPU provider, or tract when no runtime file was found.
@@ -69,6 +71,11 @@ pub enum Rung {
Cuda,
/// NVIDIA, through a TensorRT engine compiled on this device. Desktop only.
TensorRt,
/// AMD, through a MIGraphX program compiled on this device. Desktop
/// only. ONNX Runtime's ROCm provider, the CUDA provider's twin, was
/// removed in ONNX Runtime 1.23, so there is no non-compiling AMD rung
/// to fall back to: this one falls back to the CPU.
MiGraphX,
/// Qualcomm's Hexagon NPU through QNN, int8 models only. Android only.
Hexagon,
}
@@ -79,6 +86,7 @@ impl Rung {
Rung::Cpu => "CPU",
Rung::Cuda => "CUDA",
Rung::TensorRt => "TensorRT",
Rung::MiGraphX => "MIGraphX",
Rung::Hexagon => "Hexagon NPU",
}
}
@@ -88,13 +96,13 @@ impl Rung {
fn fallback(self) -> Rung {
match self {
Rung::TensorRt => Rung::Cuda,
Rung::Hexagon | Rung::Cuda | Rung::Cpu => Rung::Cpu,
Rung::MiGraphX | Rung::Hexagon | Rung::Cuda | Rung::Cpu => Rung::Cpu,
}
}
/// Whether a session on this rung needs an engine built first.
fn compiles(self) -> bool {
matches!(self, Rung::TensorRt | Rung::Hexagon)
matches!(self, Rung::TensorRt | Rung::MiGraphX | Rung::Hexagon)
}
/// The model form this rung wants for a role.
@@ -164,7 +172,7 @@ impl Status {
pub fn line(&self) -> String {
let form = match self.rung {
Rung::Hexagon => " · int8",
Rung::TensorRt => " · fp16",
Rung::TensorRt | Rung::MiGraphX => " · fp16",
_ => "",
};
format!("{}{} · {}", self.rung.label(), form, self.runtime.label())
@@ -269,8 +277,8 @@ fn acquire(role: Role, form: Form, bytes: &Arc<[u8]>, hash: u64) -> Result<Acqui
return Ok(Acquired { entry });
}
// Built outside the registry lock: a TensorRT engine load is long enough
// that another role's acquire should not wait on it.
// Built outside the registry lock: a TensorRT or MIGraphX engine load
// is long enough that another role's acquire should not wait on it.
let session = session::build(rung, role, bytes, &cfg)?;
log::debug!("inference: {role:?} loaded on {}", rung.label());
let entry = Arc::new(Loaded {
@@ -401,7 +409,16 @@ pub fn status() -> Status {
runtime: api::runtime(),
rung,
reason: s.cache.reason.clone(),
failed: s.cache.failed.clone(),
// Only what explains the selection: on an AMD machine the NVIDIA
// rungs "not enabled in this build" say nothing about why MIGraphX
// was taken. With the floor selected, everything tried is above it.
failed: s
.cache
.failed
.iter()
.filter(|(r, _)| *r > rung)
.cloned()
.collect(),
probing: s.probing,
engines: if rung.compiles() {
(s.cache.compiled.len(), s.wanted)
@@ -614,8 +631,33 @@ mod tests {
);
}
#[test]
fn the_status_reports_only_the_rungs_above_the_selection() {
let _serial = serial();
let failed = vec![
(Rung::TensorRt, "not enabled".to_string()),
(Rung::Cuda, "not enabled".to_string()),
];
let before = state().lock().unwrap().cache.clone();
state().lock().unwrap().cache = Cache {
rung: Some(Rung::MiGraphX),
failed: failed.clone(),
..Cache::default()
};
// An AMD desktop: the NVIDIA rungs below MIGraphX are not the story.
assert!(status().failed.is_empty());
// An NVIDIA desktop on the CUDA provider: TensorRT's failure is.
state().lock().unwrap().cache.rung = Some(Rung::Cuda);
assert_eq!(status().failed, vec![failed[0].clone()]);
// The floor: everything tried explains it.
state().lock().unwrap().cache.rung = Some(Rung::Cpu);
assert_eq!(status().failed.len(), 2);
state().lock().unwrap().cache = before;
}
#[test]
fn the_status_line_reads_as_the_floor_before_init() {
let _serial = serial();
let s = status();
assert_eq!(s.rung, Rung::Cpu);
assert!(s.line().starts_with("CPU"), "{}", s.line());
+45 -7
View File
@@ -16,8 +16,11 @@ use crate::{api::Runtime, state, Cache, Config, Form, Role, Rung};
fn ladder(ceiling: Option<Rung>) -> Vec<Rung> {
#[cfg(target_os = "android")]
let all = [Rung::Hexagon];
// A desktop has one vendor's GPU; the other vendor's providers are
// "not enabled in this build" or a library that fails to load, and
// either answer arrives in milliseconds.
#[cfg(not(target_os = "android"))]
let all = [Rung::TensorRt, Rung::Cuda];
let all = [Rung::TensorRt, Rung::Cuda, Rung::MiGraphX];
all.into_iter()
.filter(|r| ceiling.is_none_or(|c| *r <= c))
.collect()
@@ -206,15 +209,22 @@ fn first_line(s: &str) -> String {
line[start..].chars().take(200).collect()
}
/// Everything a change of which should re-probe: the runtime and where it
/// came from, this crate, the platform, the driver or SoC, and the models.
/// Everything a change of which should re-probe: the runtime, where it
/// came from and which providers sit beside it, this crate, the platform,
/// the driver or SoC, and the models.
fn fingerprint(runtime: &Runtime, cfg: &Config) -> String {
let mut parts = vec![
format!("engine {}", env!("CARGO_PKG_VERSION")),
format!("{} {}", std::env::consts::OS, std::env::consts::ARCH),
match runtime {
Runtime::Tract => "tract".to_string(),
Runtime::OnnxRuntime { path, version } => format!("ort {version} {}", path.display()),
Runtime::OnnxRuntime { path, version } => {
format!(
"ort {version} {} [{}]",
path.display(),
providers_beside(path)
)
}
},
device_identity(),
];
@@ -237,13 +247,41 @@ fn fingerprint(runtime: &Runtime, cfg: &Config) -> String {
parts.join("\n")
}
/// The `libonnxruntime_providers_*.so` files in the runtime's directory.
/// A distribution's CPU-only and ROCm builds are the same version at the
/// same path; the provider libraries beside them are what differs.
fn providers_beside(runtime: &Path) -> String {
let Some(dir) = runtime.parent() else {
return String::new();
};
let mut names: Vec<String> = std::fs::read_dir(dir)
.into_iter()
.flatten()
.filter_map(|e| e.ok())
.filter_map(|e| e.file_name().into_string().ok())
.filter(|n| {
n.starts_with("libonnxruntime_providers_") || n.starts_with("onnxruntime_providers_")
})
.collect();
names.sort();
names.join(" ")
}
#[cfg(target_os = "linux")]
fn device_identity() -> String {
// The NVIDIA driver's version line; absent means no NVIDIA driver.
std::fs::read_to_string("/proc/driver/nvidia/version")
// The NVIDIA driver's version line, or the ROCm release the AMD stack
// came from (`rocm-core` writes it; the kernel driver has no version
// of its own). Absent means neither.
if let Some(line) = std::fs::read_to_string("/proc/driver/nvidia/version")
.ok()
.and_then(|s| s.lines().next().map(str::to_string))
.unwrap_or_else(|| "no nvidia driver".into())
{
return line;
}
if let Ok(rocm) = std::fs::read_to_string("/opt/rocm/.info/version") {
return format!("rocm {}", rocm.trim());
}
"no nvidia driver, no rocm".into()
}
#[cfg(target_os = "android")]
+58 -1
View File
@@ -83,10 +83,65 @@ fn providers(
ep::CUDA::default().build(),
])?)
}
Rung::MiGraphX => {
// fp16 on the same terms as TensorRT (§7). MIGraphX compiles a
// program per graph — 20–60 s here — and keeps it in the cache
// directory, keyed on the graph, the GPU and its own version
// but not the precision: hence one directory per precision.
// The CPU takes any node it declines.
let fp16 = role != Role::Embedder;
let cache = cfg
.cache_dir
.join("migraphx")
.join(if fp16 { "fp16" } else { "f32" });
let _ = std::fs::create_dir_all(&cache);
let mut b = b;
migraphx(&mut b, fp16, &cache)?;
Ok(b)
}
Rung::Hexagon => unreachable!("the Hexagon rung is not on a desktop ladder"),
}
}
/// Register MIGraphX through ONNX Runtime's generic key/value entry point.
///
/// `ort`'s own builder (`ep::MIGraphX`) fills the legacy
/// `OrtMIGraphXProviderOptions`, and 1.29 reads that struct for its
/// precision flags and nothing else — the compiled-program cache directory
/// is only a key in the generic map (`migraphx_model_cache_dir`), and
/// without it every session is a full compile. Registration through the
/// generic entry point needs no `ort` feature: it is one call on the API
/// table, which is why the crate's `ort` dependency names no AMD feature.
#[cfg(not(target_os = "android"))]
fn migraphx(
b: &mut ort::session::builder::SessionBuilder,
fp16: bool,
cache: &std::path::Path,
) -> ort::Result<()> {
use ort::AsPointer;
use std::ffi::CString;
let keys = [c"migraphx_fp16_enable", c"migraphx_model_cache_dir"];
let values = [
CString::new(if fp16 { "1" } else { "0" }).unwrap(),
CString::new(cache.to_string_lossy().as_bytes())
.map_err(|e| ort::Error::new(e.to_string()))?,
];
let key_ptrs: Vec<_> = keys.iter().map(|k| k.as_ptr()).collect();
let value_ptrs: Vec<_> = values.iter().map(|v| v.as_ptr()).collect();
// SAFETY: the documented C call over arrays that outlive it; the
// runtime copies the strings into its own options map before returning.
unsafe {
let status = (ort::api().SessionOptionsAppendExecutionProvider)(
b.ptr_mut(),
c"MIGraphX".as_ptr(),
key_ptrs.as_ptr(),
value_ptrs.as_ptr(),
keys.len(),
);
ort::Error::result_from_status(status)
}
}
#[cfg(target_os = "android")]
fn providers(
b: ort::session::builder::SessionBuilder,
@@ -120,6 +175,8 @@ fn providers(
.build()
.error_on_failure()])?)
}
Rung::Cuda | Rung::TensorRt => unreachable!("no NVIDIA rung on Android"),
Rung::Cuda | Rung::TensorRt | Rung::MiGraphX => {
unreachable!("no desktop GPU rung on Android")
}
}
}
+75 -15
View File
@@ -75,7 +75,49 @@ TensorRT's job.
worst: nearly five minutes), because it is compiling an engine for this exact GPU. The engine caches to disk and the second load is milliseconds. That
number is what §6 is designed around.
### 1.3 What the numbers say
### 1.3 The desktop — Radeon RX 7900 XT, Threadripper 2920X, 24 threads · 2026-09-20
Arch's `onnxruntime-rocm` 1.29.0 against ROCm 7.2.4 and MIGraphX 7.2.3 (gfx1100), through the
same `ep_probe` harness (`core/dr-inference-engine/examples/ep_probe.rs`). Zero input, three
warm-ups, the median of 15 runs, on a machine doing nothing else. The build columns are the wall
clock of `Session` construction: cold is a MIGraphX compile of the graph for this GPU, cached is the
same session loading the program the cold build wrote.
| Model | ORT CPU f32 | MIGraphX f32 | **MIGraphX fp16** | Compile f32 / fp16 (s) | Cached load (s) |
|---|---|---|---|---|---|
| scrfd_500m (Fast) | 10.4 | 2.8 | **2.4** | 40 / 58 | 0.3 |
| scrfd_2.5g (Balanced) | 20.7 | 3.3 | **2.8** | 37 / 40 | 0.3 |
| scrfd_10g (Thorough) | 57.9 | 4.5 | **3.4** | 40 / 48 | 0.4 |
| arcface_mbf (per face) | 12.8 | 1.8 | 1.6 | 15 / 21 | 0.4 |
| yolo26n-seg | 49.3 | 8.4 | **7.5** | 110 / 136 | 0.9 |
| yolo26s-sem-ade20k | 55.8 | 4.8 | **3.8** | 50 / 60 | 0.5 |
| 2d106det (landmarks) | 9.9 | 1.2 | 1.0 | 16 / 21 | 0.2 |
| ocec_s (eye state) | 2.5 | 0.5 | 0.4 | 17 / 17 | 0.1 |
| xfeat-1024 | 26.8 | 10.5 | 9.9 | 37 / 53 | 0.3 |
| migan-512 (per tile) | 514 | 12.7 | **8.3** | 102 / 132 | 0.8 |
Three things the table settles.
- **The ROCm execution provider does not exist any more.** It was ONNX Runtime's CUDA-provider twin
for AMD, removed in the 1.23 release (AMD's builds dropped it from ROCm 7.1); 1.29's ROCm build ships
`libonnxruntime_providers_migraphx.so` and nothing else for AMD, and asking for `ROCm` answers
"not enabled in this build". So there is no non-compiling AMD rung to sit under MIGraphX the way
the CUDA provider sits under TensorRT: the AMD ladder is MIGraphX, then the CPU.
- **MIGraphX is a compiling provider, and its cache has to be asked for by name.** 15–135 s per
graph cold, under a second from its cache — TensorRT's shape exactly, and §6's design covers it.
Two things the provider does that the code has to know: `ort`'s builder fills the legacy options
struct, which 1.29 reads for the precision flags only, so the cache directory
(`migraphx_model_cache_dir`) reaches it only through the generic key/value registration; and
the cache key is the graph, the GPU and the MIGraphX version *without the precision*, so an fp16
session pointed at the f32 program's directory silently loads the f32 program (the first fp16
row measured here was that, before the directories were split).
- **fp16 is worth 10–35% over f32 on this card, not the 3× it is worth on TensorRT**, because
MIGraphX f32 is already 3–8× the CPU provider and the small graphs are launch-bound. The
detectors at 2.4–3.4 ms sit beside TensorRT fp16's 1.8–3.3 ms on the RTX 3050; the embedder is
1.6–1.8 ms on either precision and stays f32 (§7). The whole face pipeline for one image
(detector + landmarks + eyes + embedder) is under 6 ms.
### 1.4 What the numbers say
- **`tract` is single-threaded.** The tablet's one X4 core and one Raptor Lake core give the same
tract numbers. Replacing it with ONNX Runtime's CPU provider, *no accelerator involved*, is 3–6×
@@ -92,6 +134,8 @@ number is what §6 is designed around.
- **On NVIDIA, TensorRT fp16 ≈ 3× the CUDA provider**, and the CUDA provider ≈ 2× the
multi-threaded CPU; at fp16 the detectors are 1.8–3.3 ms with no quantisation at all. Both leave the twenty cores free for decoding during a batch index, which the table does not
show and which matters more than the ratio.
- **On AMD, MIGraphX fp16 is 4–17× the CPU provider** on the detectors and 60× on the
inpainter, with the same first-run compile cost as TensorRT and no rung between it and the CPU.
---
@@ -105,7 +149,8 @@ winning:
| Android, Qualcomm with a Hexagon the shipped QNN skel covers (V68–V81) | QNN HTP, int8 model | ORT CPU, f32 model | — | tract |
| Android, any other SoC | ORT CPU, f32 | — | — | tract |
| Linux / Windows, NVIDIA GPU | TensorRT, f32 model, fp16 engine | CUDA provider, f32 | ORT CPU, f32 | tract |
| Linux / Windows, no NVIDIA | ORT CPU, f32 | — | — | tract |
| Linux, AMD GPU with ROCm | MIGraphX, f32 model, fp16 program | ORT CPU, f32 | — | tract |
| Linux / Windows, no GPU stack | ORT CPU, f32 | — | — | tract |
| macOS ⁵ | ORT CPU, f32 | — | — | tract |
⁵ CoreML is the obvious rung and is unmeasured; it is listed so its absence is a gap and not an
@@ -113,8 +158,13 @@ oversight.
Deliberately **not** on any ladder, with the measurement that excluded each: NNAPI (no driver),
XNNPACK (slower than CPU, aborts on SCRFD), WebGPU (slower than CPU), the Adreno through QNN (works,
but never where the Hexagon does not also), CUDA int8 (slower than CUDA f32). A rung is added to this
table by a measurement on this page, not by a provider existing.
but never where the Hexagon does not also), CUDA int8 (slower than CUDA f32), the ROCm provider
(gone: §1.3). A rung is added to this table by a measurement on this page, not by a provider
existing.
The AMD ladder has no middle rung. TensorRT falls back to the CUDA provider while its engines
compile; MIGraphX has no such twin, so its fallback is the CPU provider, and the minute or two of
compiling on first run (§6) is spent at the floor's speed rather than at half the GPU's.
Two things the ladder is *not*: it is not a per-model choice — one backend serves every model on a
device, because §7's identity rule needs the detector and embedder on the same runtime for the
@@ -174,9 +224,10 @@ is cheaper than discovering it at packaging time.
| ONNX Runtime | MIT | Yes |
| Qualcomm QNN runtime (`com.qualcomm.qti:qnn-runtime` on Maven) | Qualcomm AI Engine Direct SDK licence — proprietary, redistribution permitted for applications using it | Yes for the APK, with the licence text shipped; not for a source distribution. **To be read in full, not summarised from memory, before the APK gains it.** |
| CUDA runtime, cuDNN, TensorRT | NVIDIA EULAs — redistributable with an application, with the licence text, not modifiable | Yes for a package that bundles them. 600 MB. The alternative is to load them from the user's system install if present and skip the rung otherwise — which is what §4's probe does anyway. |
| ROCm (HIP, MIOpen, rocBLAS), MIGraphX | MIT | Yes, but the HIP SDK MIGraphX needs is ~15 GB installed. Same answer as NVIDIA: the user's system install, or the rung is skipped. |
The position this takes: the **NVIDIA libraries are not bundled**. The desktop package probes for a
system CUDA/TensorRT install and uses it if it is version-compatible; a desktop without one runs on
The position this takes: the **GPU vendors' libraries are not bundled**. The desktop package probes for a
system CUDA/TensorRT or ROCm/MIGraphX install and uses it if it is version-compatible; a desktop without one runs on
ORT CPU, which is still 8–10× today. Bundling 600 MB for a rung that is 2× again is not a trade
worth making unmeasured, and it can be revisited by a measurement on a batch index. The **QNN
runtime is bundled** in the APK, because the Hexagon is the difference between a tablet that
@@ -237,6 +288,7 @@ of which form they load:
| f32 ONNX, shape-fixed, **opset ≥ 13** | `tools/fix-face-model-shapes.sh`, `tools/export-seg-model.sh` | Release time, once | Every rung except Hexagon |
| int8 QDQ ONNX, per-channel, uint8 activations | `tools/quantise-models.sh` (new) | Release time, once, **calibrated on real photographs** | Hexagon |
| TensorRT engine (`.engine`, per GPU architecture and TensorRT version) | The app, from the f32 file | First run on that device, in the background | TensorRT rung |
| MIGraphX program (`.mxr`, per GPU architecture, MIGraphX version and precision) | The app, from the f32 file | First run on that device, in the background | MIGraphX rung |
| QNN context binary | The app, from the int8 file | First run on that device, in the background | Hexagon rung |
Two rules.
@@ -256,9 +308,12 @@ prefers. That is a change to the canonical file and so a change to the shipped m
happens in the same model release as the int8 files.
**Compilation is a device-time step, and it is cached.** A TensorRT engine is specific to the GPU
it was built on and the TensorRT that built it; a QNN context binary is specific to the Hexagon
generation. Neither can ship. Both are built by the app the first time that rung is selected, in
the background (§6), and written beside the probe cache keyed by the same inputs. They are
it was built on and the TensorRT that built it; a MIGraphX program to the GPU and the MIGraphX
that built it; a QNN context binary to the Hexagon generation. None can ship. All are built by
the app the first time that rung is selected, in the background (§6), and written beside the
probe cache keyed by the same inputs. MIGraphX's own key leaves out the precision, so the app
gives its f32 and fp16 programs separate directories — otherwise the embedder's f32 build would
be served the detector's fp16 program, or the reverse. They are
**derived, disposable, and regenerable**: deleting the cache directory costs the next launch a
rebuild and nothing else, and the directory is excluded from anything that syncs (it is a peer of
`thumbs`, not of the catalog).
@@ -267,17 +322,18 @@ rebuild and nothing else, and the directory is excluded from anything that syncs
## 6. First run — building engines without the user waiting for them
The sequence on a device where a compiling rung (TensorRT, Hexagon) is selected:
The sequence on a device where a compiling rung (TensorRT, MIGraphX, Hexagon) is selected:
1. **Launch.** The runtime loads; the probe (§4) starts in the background; the app serves every
model request from the floor. Face indexing, segmentation and scene grading all work, at
today's speed or better (ORT CPU).
2. **Probe reports** — say, TensorRT. The compiling rung is now *selected* but has **no engines**.
Model requests continue on the fallback rung below it (CUDA provider for TensorRT; ORT CPU for
Hexagon), which needs no compilation and is already faster than the floor.
MIGraphX and Hexagon), which needs no compilation and is already faster than the floor.
3. **Engines build**, one model at a time, on a single low-priority background thread, smallest
model first so the detector — the one that runs per image — is ready soonest. On the reference
desktop that is ~1 minute for the first detector and ~10 minutes for all six at fp16; on the tablet the QNN
desktop that is ~1 minute for the first detector and ~10 minutes for all six at fp16; on the AMD
desktop 40 s for the first detector and ~8 minutes for the set; on the tablet the QNN
context binaries take 0.8–1.7 s each and the whole set is ready before the user has opened a
library. Each engine is written to a temporary name and renamed into place, so a request never
sees a half-written file.
@@ -313,7 +369,7 @@ enough to be *offered* at all. f32 on tract, ORT CPU, CUDA and TensorRT-f32 are
same graph, the same arithmetic, differences at the last bit.
**The embedder** is where comparability across devices is the whole point, and it is the one
model that no accelerator helps (§1.3). So: **the embedder runs in f32 on every rung.** On TensorRT
model that no accelerator helps (§1.4). So: **the embedder runs in f32 on every rung.** On TensorRT
that means the embedder's engine is built without fp16 while the detector's is built with it; on
the Hexagon it means the embedder is not on the NPU at all — it runs on the ORT CPU rung at 9 ms,
and the ladder's "one backend per device" is, precisely, one backend *per model role*, with the
@@ -344,7 +400,7 @@ core/dr-inference-engine
src/probe.rs §4 — the ladder per platform, the session-build probe, the cache file
src/engines.rs §6 — background compilation, the cache directory, progress
src/session.rs open(role, bytes) -> ort::Session, applying the rung and the role's precision rule
src/api.rs the one unsafe block: dlopen libonnxruntime, fetch OrtApi, ort::set_api — or ort_tract::api()
src/api.rs the unsafe block that matters: dlopen libonnxruntime, fetch OrtApi, ort::set_api — or ort_tract::api()
```
- `dr-face` and `dr-segment` **delete** their private `install_backend` and their direct
@@ -354,7 +410,11 @@ core/dr-inference-engine
`qnn`: those features add option builders, not linking, under `alternative-backend`. **Verified
for the QNN, CUDA and TensorRT builders on 2026-09-19** — they go through the API table's generic
`SessionOptionsAppendExecutionProvider*`. The NNAPI builder resolves a symbol directly and would
not; it is not needed and is not enabled.
not; it is not needed and is not enabled. MIGraphX uses no `ort` feature at all: `ort`'s builder
fills the legacy options struct, which ONNX Runtime 1.29 reads for the precision flags and
nothing else, and the compiled-program cache directory only travels through the generic
key/value entry point (`migraphx_model_cache_dir`). `session::migraphx` makes that one call
on the API table itself.
- `dr-ui` owns the settings row, the about-screen line and the progress row; it holds one
`Sessions` per process, created at launch, and passes it down. `dr_ui::library` gains
`inference_cache_dir()` beside `shared_face_models_dir()`, on the same account-independent
+5
View File
@@ -13,6 +13,11 @@
# which is what this script exists to fix. Nothing NVIDIA is bundled here:
# the providers load CUDA, cuDNN and TensorRT from the system, and if those
# are missing the probe says so and the app stays on the CPU.
#
# This is the NVIDIA script. On AMD there is nothing to fetch: the
# distribution's ROCm build of ONNX Runtime (Arch's `onnxruntime-rocm`)
# carries the MIGraphX provider, and the app finds it in the system library
# directory (docs/inference.md §1.3).
set -euo pipefail
DEST="${1:-${XDG_DATA_HOME:-${HOME}/.local/share}/darkroom/runtime}"
WORK="$(mktemp -d -p /var/tmp fetch-desktop-runtime.XXXXXX)"