Add the MIGraphX rung for AMD GPUs
Benchmarks / CPU and I/O (per commit) (push) Failing after 6m20s
Benchmarks / Frame budget (on demand) (push) Skipped
Build and test / Desktop (Linux) (push) Failing after 45s
Build and test / Layer separation (push) Successful in 26s
Traceability / Requirement traces (push) Failing after 46s
🐳 Android image / Build and push (push) Successful in 1s
Build and test / android-image (push) Successful in 1s
🐳 Windows image / Build and push (push) Successful in 1s
Build and test / windows-image (push) Successful in 1s
Build and test / Android (aarch64) (push) Failing after 2m19s
Build and test / Windows (x86_64, cross) (push) Failing after 3m2s
Benchmarks / CPU and I/O (per commit) (push) Failing after 6m20s
Benchmarks / Frame budget (on demand) (push) Skipped
Build and test / Desktop (Linux) (push) Failing after 45s
Build and test / Layer separation (push) Successful in 26s
Traceability / Requirement traces (push) Failing after 46s
🐳 Android image / Build and push (push) Successful in 1s
Build and test / android-image (push) Successful in 1s
🐳 Windows image / Build and push (push) Successful in 1s
Build and test / windows-image (push) Successful in 1s
Build and test / Android (aarch64) (push) Failing after 2m19s
Build and test / Windows (x86_64, cross) (push) Failing after 3m2s
Measured on a Radeon RX 7900 XT against Arch's onnxruntime-rocm 1.29 (docs/inference.md §1.3): MIGraphX fp16 runs the detectors at 2.4–3.4 ms against 10–58 ms on the CPU provider, the inpainter at 8 ms against 514, with a 15–135 s compile per graph the first time and under a second from its cache after. A compiling rung on TensorRT's terms, wired the same way. The ROCm execution provider is gone (removed in ONNX Runtime 1.23), so the AMD ladder is MIGraphX then the CPU, with no non-compiling rung between. MIGraphX is registered through the runtime's generic key/value entry point rather than ort's builder: 1.29 reads the legacy options struct for its precision flags only, and the compiled-program cache directory (`migraphx_model_cache_dir`) only travels the generic way. The provider's cache key omits the precision, so f32 and fp16 programs get their own directories. The probe fingerprint now includes the provider libraries beside the runtime and the ROCm version, since a distribution's CPU and ROCm builds are the same file at the same path. `status().failed` reports only the rungs above the selection, so an AMD desktop's About line says why MIGraphX won rather than that the NVIDIA providers are not in the build. Two examples: `ep_probe` times each provider cold and from cache, and `ladder` drives `init` as the app does to watch the first-run sequence. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
@@ -28,7 +28,10 @@ ort-sys = { version = "2.0.0-rc.13", default-features = false, features = ["disa
|
||||
# The NVIDIA rungs exist on the desktop only. These features add `ort`'s
|
||||
# option builders and nothing else — no linking under `alternative-backend` —
|
||||
# but an Android binary has no business carrying even the option names, and
|
||||
# the packaging must never be tempted to (§2, §3.1).
|
||||
# the packaging must never be tempted to (§2, §3.1). The AMD rung needs no
|
||||
# feature: MIGraphX is registered through the runtime's generic key/value
|
||||
# entry point (`session::migraphx`), because `ort`'s own builder cannot
|
||||
# name the compiled-program cache.
|
||||
[target.'cfg(not(target_os = "android"))'.dependencies]
|
||||
ort = { workspace = true, features = ["cuda", "tensorrt"] }
|
||||
|
||||
@@ -42,3 +45,8 @@ default = ["tract"]
|
||||
tract = ["dep:ort-tract"]
|
||||
# Look for `libonnxruntime` on disk and hand its table to `ort`.
|
||||
native = ["dep:libloading", "dep:ort-sys"]
|
||||
|
||||
[dev-dependencies]
|
||||
# The `ep_probe` example prints the provider's own diagnostics, which is most
|
||||
# of what a failed rung tells you.
|
||||
env_logger.workspace = true
|
||||
|
||||
@@ -0,0 +1,207 @@
|
||||
//! Time each execution provider a runtime offers, on the models this
|
||||
//! repository ships — the measurement docs/inference.md §1 requires before a
|
||||
//! rung is added to §2's ladder.
|
||||
//!
|
||||
//! DARKROOM_ORT_DIR=/usr/lib \
|
||||
//! cargo run --release -p dr-inference-engine --features native,tract \
|
||||
//! --example ep_probe -- models/face/scrfd_500m_640.onnx ...
|
||||
//!
|
||||
//! Prints one row per (model, provider): the median of timed runs after
|
||||
//! warm-ups, and the build time, which for a compiling provider is the
|
||||
//! number that decides whether it needs an engine cache. MIGraphX is built
|
||||
//! twice per precision — cold, then again from the cache it just wrote —
|
||||
//! so both numbers are on the page.
|
||||
//!
|
||||
//! The ROCm provider is not in the list: ONNX Runtime removed it in 1.23,
|
||||
//! and 1.29's `onnxruntime-rocm` ships `libonnxruntime_providers_migraphx.so`
|
||||
//! and nothing else for AMD.
|
||||
|
||||
use std::path::{Path, PathBuf};
|
||||
use std::time::Instant;
|
||||
|
||||
#[derive(Clone, Copy, PartialEq)]
|
||||
enum Ep {
|
||||
Cpu,
|
||||
MiGraphX,
|
||||
MiGraphXFp16,
|
||||
}
|
||||
|
||||
impl Ep {
|
||||
fn label(self) -> &'static str {
|
||||
match self {
|
||||
Ep::Cpu => "CPU",
|
||||
Ep::MiGraphX => "MIGraphX f32",
|
||||
Ep::MiGraphXFp16 => "MIGraphX fp16",
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
fn build(ep: Ep, bytes: &[u8], threads: usize, cache: &Path) -> ort::Result<ort::session::Session> {
|
||||
let mut b = ort::session::Session::builder()?.with_intra_threads(threads)?;
|
||||
match ep {
|
||||
Ep::Cpu => {}
|
||||
Ep::MiGraphX => migraphx(&mut b, false, &cache.join("f32"))?,
|
||||
Ep::MiGraphXFp16 => migraphx(&mut b, true, &cache.join("fp16"))?,
|
||||
}
|
||||
b.commit_from_memory(bytes)
|
||||
}
|
||||
|
||||
/// Register MIGraphX through the generic key/value API. `ort`'s own
|
||||
/// builder fills the legacy `OrtMIGraphXProviderOptions`, which 1.29 reads
|
||||
/// for its precision flags and nothing else: the model cache directory —
|
||||
/// the difference between a 40 s load and a 0.3 s one — only travels this
|
||||
/// way. The cache key is the graph, the GPU and the MIGraphX version, not
|
||||
/// the precision, so each precision gets its own directory.
|
||||
fn migraphx(
|
||||
b: &mut ort::session::builder::SessionBuilder,
|
||||
fp16: bool,
|
||||
cache: &Path,
|
||||
) -> ort::Result<()> {
|
||||
use ort::AsPointer;
|
||||
use std::ffi::CString;
|
||||
std::fs::create_dir_all(cache).map_err(|e| ort::Error::new(e.to_string()))?;
|
||||
let keys = [c"migraphx_fp16_enable", c"migraphx_model_cache_dir"];
|
||||
let values = [
|
||||
CString::new(if fp16 { "1" } else { "0" }).unwrap(),
|
||||
CString::new(cache.to_string_lossy().as_bytes()).unwrap(),
|
||||
];
|
||||
let key_ptrs: Vec<_> = keys.iter().map(|k| k.as_ptr()).collect();
|
||||
let value_ptrs: Vec<_> = values.iter().map(|v| v.as_ptr()).collect();
|
||||
// SAFETY: the documented C call, over arrays that outlive it; the
|
||||
// runtime copies the strings into its own options map.
|
||||
unsafe {
|
||||
let status = (ort::api().SessionOptionsAppendExecutionProvider)(
|
||||
b.ptr_mut(),
|
||||
c"MIGraphX".as_ptr(),
|
||||
key_ptrs.as_ptr(),
|
||||
value_ptrs.as_ptr(),
|
||||
keys.len(),
|
||||
);
|
||||
ort::Error::result_from_status(status)
|
||||
}
|
||||
}
|
||||
|
||||
/// Median of `runs` timed runs over zeros, in milliseconds, after warm-ups.
|
||||
fn time(session: &mut ort::session::Session, warmups: usize, runs: usize) -> Result<f64, String> {
|
||||
let shape: Vec<usize> = session.inputs()[0]
|
||||
.dtype()
|
||||
.tensor_shape()
|
||||
.ok_or("input is not a tensor")?
|
||||
.iter()
|
||||
.map(|&d| if d > 0 { d as usize } else { 1 })
|
||||
.collect();
|
||||
let zeros = vec![0f32; shape.iter().product()];
|
||||
let once = |s: &mut ort::session::Session| -> Result<f64, String> {
|
||||
let input = ort::value::Tensor::from_array((shape.clone(), zeros.clone()))
|
||||
.map_err(|e| e.to_string())?;
|
||||
let t = Instant::now();
|
||||
let out = s.run(ort::inputs![input]).map_err(|e| e.to_string())?;
|
||||
let _ = out[0]
|
||||
.try_extract_tensor::<f32>()
|
||||
.map_err(|e| e.to_string())?;
|
||||
Ok(t.elapsed().as_secs_f64() * 1e3)
|
||||
};
|
||||
for _ in 0..warmups {
|
||||
once(session)?;
|
||||
}
|
||||
let mut times = Vec::with_capacity(runs);
|
||||
for _ in 0..runs {
|
||||
times.push(once(session)?);
|
||||
}
|
||||
times.sort_by(|a, b| a.partial_cmp(b).unwrap());
|
||||
Ok(times[times.len() / 2])
|
||||
}
|
||||
|
||||
fn first_line(s: &str) -> String {
|
||||
s.lines().next().unwrap_or("").chars().take(120).collect()
|
||||
}
|
||||
|
||||
fn main() {
|
||||
env_logger::Builder::from_env(env_logger::Env::default().default_filter_or("info")).init();
|
||||
|
||||
let models: Vec<PathBuf> = std::env::args_os().skip(1).map(PathBuf::from).collect();
|
||||
if models.is_empty() {
|
||||
eprintln!("usage: ep_probe MODEL.onnx [MODEL.onnx ...]");
|
||||
std::process::exit(2);
|
||||
}
|
||||
|
||||
dr_inference_engine::ensure_runtime();
|
||||
let runtime = dr_inference_engine::status().runtime;
|
||||
println!("runtime: {}", runtime.label());
|
||||
if !runtime.is_native() {
|
||||
println!("(tract: no provider to compare; set DARKROOM_ORT_DIR)");
|
||||
}
|
||||
|
||||
let threads = std::thread::available_parallelism()
|
||||
.map(|n| n.get().saturating_sub(2).max(1))
|
||||
.unwrap_or(1);
|
||||
println!("intra-op threads: {threads}");
|
||||
let cache = std::env::temp_dir().join("darkroom-ep-probe");
|
||||
let _ = std::fs::remove_dir_all(&cache);
|
||||
println!("compiled-program cache: {}\n", cache.display());
|
||||
|
||||
println!(
|
||||
"{:<28} {:<15} {:>10} {:>10}",
|
||||
"model", "provider", "build s", "median ms"
|
||||
);
|
||||
for model in &models {
|
||||
let bytes = match std::fs::read(model) {
|
||||
Ok(b) => b,
|
||||
Err(e) => {
|
||||
println!("{:<28} read failed: {e}", name(model));
|
||||
continue;
|
||||
}
|
||||
};
|
||||
// A compiling provider is built twice: the second build reads the
|
||||
// program the first wrote, and its time is what a launch after the
|
||||
// first costs.
|
||||
let plan = [
|
||||
(Ep::Cpu, false),
|
||||
(Ep::MiGraphX, false),
|
||||
(Ep::MiGraphX, true),
|
||||
(Ep::MiGraphXFp16, false),
|
||||
(Ep::MiGraphXFp16, true),
|
||||
];
|
||||
for (ep, cached) in plan {
|
||||
let started = Instant::now();
|
||||
match build(ep, &bytes, threads, &cache) {
|
||||
Ok(mut session) => {
|
||||
let built = started.elapsed().as_secs_f64();
|
||||
match time(&mut session, 3, 15) {
|
||||
Ok(ms) => println!(
|
||||
"{:<28} {:<15} {:>10.1} {:>10.1}{}",
|
||||
name(model),
|
||||
ep.label(),
|
||||
built,
|
||||
ms,
|
||||
if cached { " (from cache)" } else { "" }
|
||||
),
|
||||
Err(e) => println!(
|
||||
"{:<28} {:<15} {:>10.1} {:>10} {}",
|
||||
name(model),
|
||||
ep.label(),
|
||||
built,
|
||||
"ran ✗",
|
||||
first_line(&e)
|
||||
),
|
||||
}
|
||||
}
|
||||
Err(e) => println!(
|
||||
"{:<28} {:<15} {:>21} {}",
|
||||
name(model),
|
||||
ep.label(),
|
||||
"build ✗",
|
||||
first_line(&e.to_string())
|
||||
),
|
||||
}
|
||||
}
|
||||
println!();
|
||||
}
|
||||
}
|
||||
|
||||
fn name(p: &Path) -> String {
|
||||
p.file_name()
|
||||
.unwrap_or(p.as_os_str())
|
||||
.to_string_lossy()
|
||||
.into_owned()
|
||||
}
|
||||
@@ -0,0 +1,100 @@
|
||||
//! Walk the ladder as the app does — probe, engines, then a session — and
|
||||
//! say what each step chose. The M5 check of docs/inference.md §6 without
|
||||
//! the app around it.
|
||||
//!
|
||||
//! DARKROOM_ORT_DIR=/usr/lib \
|
||||
//! cargo run --release -p dr-inference-engine --features native,tract \
|
||||
//! --example ladder -- CACHE_DIR models/face/scrfd_500m_640.onnx [MODEL.onnx ...]
|
||||
//!
|
||||
//! Every model named is a `Detector` for the config's purposes, which is
|
||||
//! enough to see the rung taken, the engines compiled and a session land
|
||||
//! on it. Delete `CACHE_DIR` to see the first run again; keep it to see the
|
||||
//! second.
|
||||
|
||||
use std::path::PathBuf;
|
||||
use std::time::{Duration, Instant};
|
||||
|
||||
fn main() {
|
||||
env_logger::Builder::from_env(env_logger::Env::default().default_filter_or("info")).init();
|
||||
let mut args = std::env::args_os().skip(1).map(PathBuf::from);
|
||||
let (Some(cache_dir), models) = (args.next(), args.collect::<Vec<_>>()) else {
|
||||
eprintln!("usage: ladder CACHE_DIR MODEL.onnx [MODEL.onnx ...]");
|
||||
std::process::exit(2);
|
||||
};
|
||||
if models.is_empty() {
|
||||
eprintln!("usage: ladder CACHE_DIR MODEL.onnx [MODEL.onnx ...]");
|
||||
std::process::exit(2);
|
||||
}
|
||||
|
||||
let runtime_dirs: Vec<PathBuf> = std::env::var_os("DARKROOM_ORT_DIR")
|
||||
.map(PathBuf::from)
|
||||
.into_iter()
|
||||
.collect();
|
||||
let started = Instant::now();
|
||||
dr_inference_engine::init(dr_inference_engine::Config {
|
||||
runtime_dirs,
|
||||
cache_dir: cache_dir.clone(),
|
||||
models: models
|
||||
.iter()
|
||||
.map(|p| (dr_inference_engine::Role::Detector, p.clone()))
|
||||
.collect(),
|
||||
embedded: Vec::new(),
|
||||
ceiling: None,
|
||||
threads: 0,
|
||||
decay: Duration::ZERO,
|
||||
});
|
||||
|
||||
let mut last = String::new();
|
||||
loop {
|
||||
let s = dr_inference_engine::status();
|
||||
let line = format!(
|
||||
"{} · {} · engines {}/{}{}",
|
||||
s.line(),
|
||||
if s.probing {
|
||||
"probing"
|
||||
} else {
|
||||
s.reason.as_str()
|
||||
},
|
||||
s.engines.0,
|
||||
s.engines.1,
|
||||
if s.failed.is_empty() {
|
||||
String::new()
|
||||
} else {
|
||||
format!(
|
||||
" · tried {}",
|
||||
s.failed
|
||||
.iter()
|
||||
.map(|(r, why)| format!("{}: {why}", r.label()))
|
||||
.collect::<Vec<_>>()
|
||||
.join(" · ")
|
||||
)
|
||||
}
|
||||
);
|
||||
if line != last {
|
||||
println!("{:>6.1} s {line}", started.elapsed().as_secs_f64());
|
||||
last = line;
|
||||
}
|
||||
if !s.probing && s.engines.0 >= s.engines.1 {
|
||||
break;
|
||||
}
|
||||
std::thread::sleep(Duration::from_millis(500));
|
||||
}
|
||||
|
||||
for path in &models {
|
||||
let bytes = std::fs::read(path).expect("read model");
|
||||
let t = Instant::now();
|
||||
let model = dr_inference_engine::open(
|
||||
dr_inference_engine::Role::Detector,
|
||||
dr_inference_engine::Form::F32,
|
||||
&bytes,
|
||||
)
|
||||
.expect("open model");
|
||||
let acquired = model.acquire().expect("acquire session");
|
||||
println!(
|
||||
"{} on {} in {:.2} s",
|
||||
path.file_name().unwrap().to_string_lossy(),
|
||||
acquired.rung().label(),
|
||||
t.elapsed().as_secs_f64()
|
||||
);
|
||||
}
|
||||
}
|
||||
@@ -3,8 +3,8 @@
|
||||
//!
|
||||
//! Consumers ask for a session by [`Role`] and get `ort`'s `Session` back;
|
||||
//! what built it — tract on one core, ONNX Runtime's CPU pool, a TensorRT
|
||||
//! engine, the Hexagon — is this crate's business and shows up in
|
||||
//! [`status`] for the settings row and nowhere else.
|
||||
//! engine, a MIGraphX program, the Hexagon — is this crate's business and
|
||||
//! shows up in [`status`] for the settings row and nowhere else.
|
||||
//!
|
||||
//! The shape follows §3 of the spec: `ort` links nothing (`alternative-backend`),
|
||||
//! and the first call hands it an API table from either a `libonnxruntime`
|
||||
@@ -60,7 +60,9 @@ pub enum Form {
|
||||
|
||||
/// A rung of the ladder (§2). Ordered: a user override names the highest rung
|
||||
/// the probe may take, and a compiling rung falls back to the one below it
|
||||
/// until its engine exists.
|
||||
/// until its engine exists. The order is within a vendor's ladder — a
|
||||
/// machine has NVIDIA rungs or an AMD rung, never both — so a ceiling is
|
||||
/// read as "no higher than this on whichever ladder the device has".
|
||||
#[derive(Clone, Copy, Debug, PartialEq, Eq, Hash, PartialOrd, Ord, Serialize, Deserialize)]
|
||||
pub enum Rung {
|
||||
/// ONNX Runtime's CPU provider, or tract when no runtime file was found.
|
||||
@@ -69,6 +71,11 @@ pub enum Rung {
|
||||
Cuda,
|
||||
/// NVIDIA, through a TensorRT engine compiled on this device. Desktop only.
|
||||
TensorRt,
|
||||
/// AMD, through a MIGraphX program compiled on this device. Desktop
|
||||
/// only. ONNX Runtime's ROCm provider, the CUDA provider's twin, was
|
||||
/// removed in ONNX Runtime 1.23, so there is no non-compiling AMD rung
|
||||
/// to fall back to: this one falls back to the CPU.
|
||||
MiGraphX,
|
||||
/// Qualcomm's Hexagon NPU through QNN, int8 models only. Android only.
|
||||
Hexagon,
|
||||
}
|
||||
@@ -79,6 +86,7 @@ impl Rung {
|
||||
Rung::Cpu => "CPU",
|
||||
Rung::Cuda => "CUDA",
|
||||
Rung::TensorRt => "TensorRT",
|
||||
Rung::MiGraphX => "MIGraphX",
|
||||
Rung::Hexagon => "Hexagon NPU",
|
||||
}
|
||||
}
|
||||
@@ -88,13 +96,13 @@ impl Rung {
|
||||
fn fallback(self) -> Rung {
|
||||
match self {
|
||||
Rung::TensorRt => Rung::Cuda,
|
||||
Rung::Hexagon | Rung::Cuda | Rung::Cpu => Rung::Cpu,
|
||||
Rung::MiGraphX | Rung::Hexagon | Rung::Cuda | Rung::Cpu => Rung::Cpu,
|
||||
}
|
||||
}
|
||||
|
||||
/// Whether a session on this rung needs an engine built first.
|
||||
fn compiles(self) -> bool {
|
||||
matches!(self, Rung::TensorRt | Rung::Hexagon)
|
||||
matches!(self, Rung::TensorRt | Rung::MiGraphX | Rung::Hexagon)
|
||||
}
|
||||
|
||||
/// The model form this rung wants for a role.
|
||||
@@ -164,7 +172,7 @@ impl Status {
|
||||
pub fn line(&self) -> String {
|
||||
let form = match self.rung {
|
||||
Rung::Hexagon => " · int8",
|
||||
Rung::TensorRt => " · fp16",
|
||||
Rung::TensorRt | Rung::MiGraphX => " · fp16",
|
||||
_ => "",
|
||||
};
|
||||
format!("{}{} · {}", self.rung.label(), form, self.runtime.label())
|
||||
@@ -269,8 +277,8 @@ fn acquire(role: Role, form: Form, bytes: &Arc<[u8]>, hash: u64) -> Result<Acqui
|
||||
return Ok(Acquired { entry });
|
||||
}
|
||||
|
||||
// Built outside the registry lock: a TensorRT engine load is long enough
|
||||
// that another role's acquire should not wait on it.
|
||||
// Built outside the registry lock: a TensorRT or MIGraphX engine load
|
||||
// is long enough that another role's acquire should not wait on it.
|
||||
let session = session::build(rung, role, bytes, &cfg)?;
|
||||
log::debug!("inference: {role:?} loaded on {}", rung.label());
|
||||
let entry = Arc::new(Loaded {
|
||||
@@ -401,7 +409,16 @@ pub fn status() -> Status {
|
||||
runtime: api::runtime(),
|
||||
rung,
|
||||
reason: s.cache.reason.clone(),
|
||||
failed: s.cache.failed.clone(),
|
||||
// Only what explains the selection: on an AMD machine the NVIDIA
|
||||
// rungs "not enabled in this build" say nothing about why MIGraphX
|
||||
// was taken. With the floor selected, everything tried is above it.
|
||||
failed: s
|
||||
.cache
|
||||
.failed
|
||||
.iter()
|
||||
.filter(|(r, _)| *r > rung)
|
||||
.cloned()
|
||||
.collect(),
|
||||
probing: s.probing,
|
||||
engines: if rung.compiles() {
|
||||
(s.cache.compiled.len(), s.wanted)
|
||||
@@ -614,8 +631,33 @@ mod tests {
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn the_status_reports_only_the_rungs_above_the_selection() {
|
||||
let _serial = serial();
|
||||
let failed = vec![
|
||||
(Rung::TensorRt, "not enabled".to_string()),
|
||||
(Rung::Cuda, "not enabled".to_string()),
|
||||
];
|
||||
let before = state().lock().unwrap().cache.clone();
|
||||
state().lock().unwrap().cache = Cache {
|
||||
rung: Some(Rung::MiGraphX),
|
||||
failed: failed.clone(),
|
||||
..Cache::default()
|
||||
};
|
||||
// An AMD desktop: the NVIDIA rungs below MIGraphX are not the story.
|
||||
assert!(status().failed.is_empty());
|
||||
// An NVIDIA desktop on the CUDA provider: TensorRT's failure is.
|
||||
state().lock().unwrap().cache.rung = Some(Rung::Cuda);
|
||||
assert_eq!(status().failed, vec![failed[0].clone()]);
|
||||
// The floor: everything tried explains it.
|
||||
state().lock().unwrap().cache.rung = Some(Rung::Cpu);
|
||||
assert_eq!(status().failed.len(), 2);
|
||||
state().lock().unwrap().cache = before;
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn the_status_line_reads_as_the_floor_before_init() {
|
||||
let _serial = serial();
|
||||
let s = status();
|
||||
assert_eq!(s.rung, Rung::Cpu);
|
||||
assert!(s.line().starts_with("CPU"), "{}", s.line());
|
||||
|
||||
@@ -16,8 +16,11 @@ use crate::{api::Runtime, state, Cache, Config, Form, Role, Rung};
|
||||
fn ladder(ceiling: Option<Rung>) -> Vec<Rung> {
|
||||
#[cfg(target_os = "android")]
|
||||
let all = [Rung::Hexagon];
|
||||
// A desktop has one vendor's GPU; the other vendor's providers are
|
||||
// "not enabled in this build" or a library that fails to load, and
|
||||
// either answer arrives in milliseconds.
|
||||
#[cfg(not(target_os = "android"))]
|
||||
let all = [Rung::TensorRt, Rung::Cuda];
|
||||
let all = [Rung::TensorRt, Rung::Cuda, Rung::MiGraphX];
|
||||
all.into_iter()
|
||||
.filter(|r| ceiling.is_none_or(|c| *r <= c))
|
||||
.collect()
|
||||
@@ -206,15 +209,22 @@ fn first_line(s: &str) -> String {
|
||||
line[start..].chars().take(200).collect()
|
||||
}
|
||||
|
||||
/// Everything a change of which should re-probe: the runtime and where it
|
||||
/// came from, this crate, the platform, the driver or SoC, and the models.
|
||||
/// Everything a change of which should re-probe: the runtime, where it
|
||||
/// came from and which providers sit beside it, this crate, the platform,
|
||||
/// the driver or SoC, and the models.
|
||||
fn fingerprint(runtime: &Runtime, cfg: &Config) -> String {
|
||||
let mut parts = vec![
|
||||
format!("engine {}", env!("CARGO_PKG_VERSION")),
|
||||
format!("{} {}", std::env::consts::OS, std::env::consts::ARCH),
|
||||
match runtime {
|
||||
Runtime::Tract => "tract".to_string(),
|
||||
Runtime::OnnxRuntime { path, version } => format!("ort {version} {}", path.display()),
|
||||
Runtime::OnnxRuntime { path, version } => {
|
||||
format!(
|
||||
"ort {version} {} [{}]",
|
||||
path.display(),
|
||||
providers_beside(path)
|
||||
)
|
||||
}
|
||||
},
|
||||
device_identity(),
|
||||
];
|
||||
@@ -237,13 +247,41 @@ fn fingerprint(runtime: &Runtime, cfg: &Config) -> String {
|
||||
parts.join("\n")
|
||||
}
|
||||
|
||||
/// The `libonnxruntime_providers_*.so` files in the runtime's directory.
|
||||
/// A distribution's CPU-only and ROCm builds are the same version at the
|
||||
/// same path; the provider libraries beside them are what differs.
|
||||
fn providers_beside(runtime: &Path) -> String {
|
||||
let Some(dir) = runtime.parent() else {
|
||||
return String::new();
|
||||
};
|
||||
let mut names: Vec<String> = std::fs::read_dir(dir)
|
||||
.into_iter()
|
||||
.flatten()
|
||||
.filter_map(|e| e.ok())
|
||||
.filter_map(|e| e.file_name().into_string().ok())
|
||||
.filter(|n| {
|
||||
n.starts_with("libonnxruntime_providers_") || n.starts_with("onnxruntime_providers_")
|
||||
})
|
||||
.collect();
|
||||
names.sort();
|
||||
names.join(" ")
|
||||
}
|
||||
|
||||
#[cfg(target_os = "linux")]
|
||||
fn device_identity() -> String {
|
||||
// The NVIDIA driver's version line; absent means no NVIDIA driver.
|
||||
std::fs::read_to_string("/proc/driver/nvidia/version")
|
||||
// The NVIDIA driver's version line, or the ROCm release the AMD stack
|
||||
// came from (`rocm-core` writes it; the kernel driver has no version
|
||||
// of its own). Absent means neither.
|
||||
if let Some(line) = std::fs::read_to_string("/proc/driver/nvidia/version")
|
||||
.ok()
|
||||
.and_then(|s| s.lines().next().map(str::to_string))
|
||||
.unwrap_or_else(|| "no nvidia driver".into())
|
||||
{
|
||||
return line;
|
||||
}
|
||||
if let Ok(rocm) = std::fs::read_to_string("/opt/rocm/.info/version") {
|
||||
return format!("rocm {}", rocm.trim());
|
||||
}
|
||||
"no nvidia driver, no rocm".into()
|
||||
}
|
||||
|
||||
#[cfg(target_os = "android")]
|
||||
|
||||
@@ -83,10 +83,65 @@ fn providers(
|
||||
ep::CUDA::default().build(),
|
||||
])?)
|
||||
}
|
||||
Rung::MiGraphX => {
|
||||
// fp16 on the same terms as TensorRT (§7). MIGraphX compiles a
|
||||
// program per graph — 20–60 s here — and keeps it in the cache
|
||||
// directory, keyed on the graph, the GPU and its own version
|
||||
// but not the precision: hence one directory per precision.
|
||||
// The CPU takes any node it declines.
|
||||
let fp16 = role != Role::Embedder;
|
||||
let cache = cfg
|
||||
.cache_dir
|
||||
.join("migraphx")
|
||||
.join(if fp16 { "fp16" } else { "f32" });
|
||||
let _ = std::fs::create_dir_all(&cache);
|
||||
let mut b = b;
|
||||
migraphx(&mut b, fp16, &cache)?;
|
||||
Ok(b)
|
||||
}
|
||||
Rung::Hexagon => unreachable!("the Hexagon rung is not on a desktop ladder"),
|
||||
}
|
||||
}
|
||||
|
||||
/// Register MIGraphX through ONNX Runtime's generic key/value entry point.
|
||||
///
|
||||
/// `ort`'s own builder (`ep::MIGraphX`) fills the legacy
|
||||
/// `OrtMIGraphXProviderOptions`, and 1.29 reads that struct for its
|
||||
/// precision flags and nothing else — the compiled-program cache directory
|
||||
/// is only a key in the generic map (`migraphx_model_cache_dir`), and
|
||||
/// without it every session is a full compile. Registration through the
|
||||
/// generic entry point needs no `ort` feature: it is one call on the API
|
||||
/// table, which is why the crate's `ort` dependency names no AMD feature.
|
||||
#[cfg(not(target_os = "android"))]
|
||||
fn migraphx(
|
||||
b: &mut ort::session::builder::SessionBuilder,
|
||||
fp16: bool,
|
||||
cache: &std::path::Path,
|
||||
) -> ort::Result<()> {
|
||||
use ort::AsPointer;
|
||||
use std::ffi::CString;
|
||||
let keys = [c"migraphx_fp16_enable", c"migraphx_model_cache_dir"];
|
||||
let values = [
|
||||
CString::new(if fp16 { "1" } else { "0" }).unwrap(),
|
||||
CString::new(cache.to_string_lossy().as_bytes())
|
||||
.map_err(|e| ort::Error::new(e.to_string()))?,
|
||||
];
|
||||
let key_ptrs: Vec<_> = keys.iter().map(|k| k.as_ptr()).collect();
|
||||
let value_ptrs: Vec<_> = values.iter().map(|v| v.as_ptr()).collect();
|
||||
// SAFETY: the documented C call over arrays that outlive it; the
|
||||
// runtime copies the strings into its own options map before returning.
|
||||
unsafe {
|
||||
let status = (ort::api().SessionOptionsAppendExecutionProvider)(
|
||||
b.ptr_mut(),
|
||||
c"MIGraphX".as_ptr(),
|
||||
key_ptrs.as_ptr(),
|
||||
value_ptrs.as_ptr(),
|
||||
keys.len(),
|
||||
);
|
||||
ort::Error::result_from_status(status)
|
||||
}
|
||||
}
|
||||
|
||||
#[cfg(target_os = "android")]
|
||||
fn providers(
|
||||
b: ort::session::builder::SessionBuilder,
|
||||
@@ -120,6 +175,8 @@ fn providers(
|
||||
.build()
|
||||
.error_on_failure()])?)
|
||||
}
|
||||
Rung::Cuda | Rung::TensorRt => unreachable!("no NVIDIA rung on Android"),
|
||||
Rung::Cuda | Rung::TensorRt | Rung::MiGraphX => {
|
||||
unreachable!("no desktop GPU rung on Android")
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
Reference in New Issue
Block a user