Add a CoreML rung on macOS
The macOS ladder was the CPU provider alone, with CoreML listed as a gap. It is now CoreML, then the CPU, then tract — unmeasured, since nobody here has a Mac, and safe to ship unmeasured because the probe's clock rejects a CoreML slower than the CPU and `attempt` refuses one that crashes. - `Rung::CoreMl`, a compiling rung like TensorRT: an ML Program with every compute unit allowed, falling back to the CPU until each model's program is built. The embedder stays on the CPU, as on the Hexagon (§7). - The cache is one directory per model and runtime version. CoreML keys a model committed from memory on its input and node names, not its weights (ONNX Runtime 1.29, coreml_execution_provider.cc), so two exports of one architecture would otherwise share a program. - The fingerprint on macOS is the chip and the OS release, which ships CoreML. - The desktop looks for the runtime in the bundle's Contents/Frameworks and Homebrew's prefixes; fetch-desktop-runtime.sh on a Mac downloads ONNX Runtime 1.29.0 for Apple silicon, which carries CoreML. docs/dev/macos.md says what exists, how to build it, and which log lines to ask a Mac user for.
This commit is contained in:
@@ -38,6 +38,11 @@ ort = { workspace = true, features = ["cuda", "tensorrt"] }
|
||||
[target.'cfg(target_os = "android")'.dependencies]
|
||||
ort = { workspace = true, features = ["qnn"] }
|
||||
|
||||
# The Apple rung: CoreML's option builder, which fills the runtime's generic
|
||||
# key/value map. `ort-sys`'s `coreml` feature is empty; nothing links.
|
||||
[target.'cfg(target_os = "macos")'.dependencies]
|
||||
ort = { workspace = true, features = ["coreml"] }
|
||||
|
||||
[features]
|
||||
# The floor: `tract` supplies the API table when no runtime file is found, or
|
||||
# always, in a build without `native`. Tests want this and nothing else.
|
||||
|
||||
@@ -44,6 +44,20 @@ pub fn context_path(cfg: &Config, bytes: &[u8]) -> PathBuf {
|
||||
.join(format!("{:016x}_ctx.onnx", hash(bytes)))
|
||||
}
|
||||
|
||||
/// Where CoreML compiles `bytes` to: one directory per model, because
|
||||
/// CoreML's own cache key leaves out the weights of a model loaded from
|
||||
/// memory (`session::coreml`), and one per runtime version, which wrote it.
|
||||
pub fn coreml_dir(cfg: &Config, bytes: &[u8]) -> PathBuf {
|
||||
let runtime = match crate::api::runtime() {
|
||||
crate::Runtime::OnnxRuntime { version, .. } => version,
|
||||
crate::Runtime::Tract => "tract".into(),
|
||||
};
|
||||
cfg.cache_dir
|
||||
.join("coreml")
|
||||
.join(runtime)
|
||||
.join(format!("{:016x}", hash(bytes)))
|
||||
}
|
||||
|
||||
/// After the probe: compile every configured model the selected rung can
|
||||
/// take, smallest first, recording each as it lands.
|
||||
pub fn run() {
|
||||
|
||||
@@ -78,6 +78,12 @@ pub enum Rung {
|
||||
MiGraphX,
|
||||
/// Qualcomm's Hexagon NPU through QNN, int8 models only. Android only.
|
||||
Hexagon,
|
||||
/// Apple, through CoreML: the Neural Engine, the GPU or the CPU, as
|
||||
/// CoreML schedules it. macOS only. Compiles an ML Program per model on
|
||||
/// first use, so it is a compiling rung with the CPU below it. The
|
||||
/// embedder stays on the CPU, as on the Hexagon: the Neural Engine
|
||||
/// computes in fp16 (§7).
|
||||
CoreMl,
|
||||
}
|
||||
|
||||
impl Rung {
|
||||
@@ -88,6 +94,7 @@ impl Rung {
|
||||
Rung::TensorRt => "TensorRT",
|
||||
Rung::MiGraphX => "MIGraphX",
|
||||
Rung::Hexagon => "Hexagon NPU",
|
||||
Rung::CoreMl => "CoreML",
|
||||
}
|
||||
}
|
||||
|
||||
@@ -96,13 +103,16 @@ impl Rung {
|
||||
fn fallback(self) -> Rung {
|
||||
match self {
|
||||
Rung::TensorRt => Rung::Cuda,
|
||||
Rung::MiGraphX | Rung::Hexagon | Rung::Cuda | Rung::Cpu => Rung::Cpu,
|
||||
Rung::MiGraphX | Rung::Hexagon | Rung::CoreMl | Rung::Cuda | Rung::Cpu => Rung::Cpu,
|
||||
}
|
||||
}
|
||||
|
||||
/// Whether a session on this rung needs an engine built first.
|
||||
fn compiles(self) -> bool {
|
||||
matches!(self, Rung::TensorRt | Rung::MiGraphX | Rung::Hexagon)
|
||||
matches!(
|
||||
self,
|
||||
Rung::TensorRt | Rung::MiGraphX | Rung::Hexagon | Rung::CoreMl
|
||||
)
|
||||
}
|
||||
|
||||
/// The model form this rung wants for a role.
|
||||
@@ -116,9 +126,11 @@ impl Rung {
|
||||
/// Whether this rung runs `role` at all. The Hexagon takes int8 graphs
|
||||
/// only, and the embedder is never int8 (§7) — it runs on the CPU
|
||||
/// beside a detector on the NPU, so its vectors compare across devices.
|
||||
/// CoreML is kept off the embedder for the same reason: the Neural
|
||||
/// Engine is fp16, and which unit runs a graph is CoreML's choice.
|
||||
fn serves(self, role: Role) -> bool {
|
||||
match self {
|
||||
Rung::Hexagon => role != Role::Embedder,
|
||||
Rung::Hexagon | Rung::CoreMl => role != Role::Embedder,
|
||||
_ => true,
|
||||
}
|
||||
}
|
||||
@@ -637,6 +649,27 @@ mod tests {
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn coreml_takes_a_compiled_detector_and_never_the_embedder() {
|
||||
let hash = engines::hash(b"detector");
|
||||
let mut s = State {
|
||||
config: Config::default(),
|
||||
cache: Cache {
|
||||
rung: Some(Rung::CoreMl),
|
||||
..Cache::default()
|
||||
},
|
||||
probing: false,
|
||||
wanted: 0,
|
||||
};
|
||||
let on = |s: &State, role| effective_rung(s, Rung::CoreMl, role, Form::F32, hash);
|
||||
// Before its program is compiled the detector waits on the CPU.
|
||||
assert_eq!(on(&s, Role::Detector), Rung::Cpu);
|
||||
s.cache.compiled.insert(engines::key_of(Rung::CoreMl, hash));
|
||||
assert_eq!(on(&s, Role::Detector), Rung::CoreMl);
|
||||
// The embedder does not move, compiled or not (§7).
|
||||
assert_eq!(on(&s, Role::Embedder), Rung::Cpu);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn the_status_reports_only_the_rungs_above_the_selection() {
|
||||
let _serial = serial();
|
||||
|
||||
@@ -16,10 +16,15 @@ use crate::{api::Runtime, state, Cache, Config, Form, Role, Rung};
|
||||
fn ladder(ceiling: Option<Rung>) -> Vec<Rung> {
|
||||
#[cfg(target_os = "android")]
|
||||
let all = [Rung::Hexagon];
|
||||
// Unmeasured (§2 ⁵): it is on the ladder because the probe's clock and
|
||||
// `attempt` make a wrong guess cost one slow or failed probe, not a
|
||||
// slow or crashing app.
|
||||
#[cfg(target_os = "macos")]
|
||||
let all = [Rung::CoreMl];
|
||||
// A desktop has one vendor's GPU; the other vendor's providers are
|
||||
// "not enabled in this build" or a library that fails to load, and
|
||||
// either answer arrives in milliseconds.
|
||||
#[cfg(not(target_os = "android"))]
|
||||
#[cfg(not(any(target_os = "android", target_os = "macos")))]
|
||||
let all = [Rung::TensorRt, Rung::Cuda, Rung::MiGraphX];
|
||||
all.into_iter()
|
||||
.filter(|r| ceiling.is_none_or(|c| *r <= c))
|
||||
@@ -360,7 +365,50 @@ fn system_property(name: &str) -> String {
|
||||
String::from_utf8_lossy(&buf[..n.max(0) as usize]).into_owned()
|
||||
}
|
||||
|
||||
#[cfg(not(any(target_os = "linux", target_os = "android")))]
|
||||
#[cfg(target_os = "macos")]
|
||||
fn device_identity() -> String {
|
||||
// The chip, and the OS release: CoreML ships with the OS, so a macOS
|
||||
// update is a new provider as surely as a new driver is on Linux.
|
||||
format!(
|
||||
"{} macOS {}",
|
||||
sysctl("machdep.cpu.brand_string"),
|
||||
sysctl("kern.osproductversion")
|
||||
)
|
||||
}
|
||||
|
||||
#[cfg(target_os = "macos")]
|
||||
fn sysctl(name: &str) -> String {
|
||||
extern "C" {
|
||||
fn sysctlbyname(
|
||||
name: *const std::ffi::c_char,
|
||||
oldp: *mut std::ffi::c_void,
|
||||
oldlenp: *mut usize,
|
||||
newp: *mut std::ffi::c_void,
|
||||
newlen: usize,
|
||||
) -> i32;
|
||||
}
|
||||
let name = std::ffi::CString::new(name).unwrap();
|
||||
let mut buf = [0u8; 256];
|
||||
let mut len = buf.len();
|
||||
// SAFETY: libSystem's documented call; `len` is the buffer's size in and
|
||||
// the string's length, with its terminator, out.
|
||||
let rc = unsafe {
|
||||
sysctlbyname(
|
||||
name.as_ptr(),
|
||||
buf.as_mut_ptr().cast(),
|
||||
&mut len,
|
||||
std::ptr::null_mut(),
|
||||
0,
|
||||
)
|
||||
};
|
||||
if rc != 0 {
|
||||
return String::new();
|
||||
}
|
||||
let s = &buf[..len.min(buf.len())];
|
||||
String::from_utf8_lossy(s.strip_suffix(&[0]).unwrap_or(s)).into_owned()
|
||||
}
|
||||
|
||||
#[cfg(not(any(target_os = "linux", target_os = "android", target_os = "macos")))]
|
||||
fn device_identity() -> String {
|
||||
String::new()
|
||||
}
|
||||
|
||||
@@ -27,13 +27,14 @@ pub fn build(rung: Rung, role: Role, bytes: &[u8], cfg: &Config) -> ort::Result<
|
||||
// what makes the second case rare (§6).
|
||||
let context = (rung == Rung::Hexagon).then(|| crate::engines::context_path(cfg, bytes));
|
||||
let ready = context.as_ref().is_some_and(|p| p.is_file());
|
||||
b = providers(
|
||||
b,
|
||||
rung,
|
||||
role,
|
||||
cfg,
|
||||
if ready { None } else { context.as_deref() },
|
||||
)?;
|
||||
// What the rung keeps for this model: the context the Hexagon is to
|
||||
// write, or the directory CoreML compiles into.
|
||||
let per_model = match rung {
|
||||
Rung::CoreMl => Some(crate::engines::coreml_dir(cfg, bytes)),
|
||||
_ if ready => None,
|
||||
_ => context.clone(),
|
||||
};
|
||||
b = providers(b, rung, role, cfg, per_model.as_deref())?;
|
||||
match (ready, context) {
|
||||
(true, Some(path)) => b.commit_from_file(path),
|
||||
_ => b.commit_from_memory(bytes),
|
||||
@@ -90,11 +91,12 @@ fn providers(
|
||||
rung: Rung,
|
||||
role: Role,
|
||||
cfg: &Config,
|
||||
_generate_context: Option<&std::path::Path>,
|
||||
per_model: Option<&std::path::Path>,
|
||||
) -> ort::Result<ort::session::builder::SessionBuilder> {
|
||||
use ort::ep;
|
||||
match rung {
|
||||
Rung::Cpu => Ok(b),
|
||||
Rung::CoreMl => coreml(b, per_model),
|
||||
Rung::Cuda => {
|
||||
Ok(b.with_execution_providers([ep::CUDA::default().build().error_on_failure()])?)
|
||||
}
|
||||
@@ -139,6 +141,43 @@ fn providers(
|
||||
}
|
||||
}
|
||||
|
||||
/// CoreML, compiling an ML Program — the format with the operators these
|
||||
/// graphs use and the one that reaches the Neural Engine — into `cache`.
|
||||
///
|
||||
/// The option names are those ONNX Runtime 1.29 reads from the generic
|
||||
/// key/value map (`coreml_options.cc`), which is what `ort`'s builder
|
||||
/// fills. The cache is per model because of how CoreML keys it: a model
|
||||
/// committed from memory, as every session here is, has no path, and the
|
||||
/// key falls back to a hash of the graph's input and node names — not its
|
||||
/// weights. Two exports of one architecture would share a program. The
|
||||
/// directory `engines::coreml_dir` names is the hash of the bytes.
|
||||
///
|
||||
/// Every compute unit is allowed, so CoreML may place a graph on the
|
||||
/// Neural Engine, the GPU or the CPU; the probe's clock judges the result.
|
||||
#[cfg(target_os = "macos")]
|
||||
fn coreml(
|
||||
b: ort::session::builder::SessionBuilder,
|
||||
cache: Option<&std::path::Path>,
|
||||
) -> ort::Result<ort::session::builder::SessionBuilder> {
|
||||
use ort::ep::{self, coreml};
|
||||
let mut ep = ep::CoreML::default()
|
||||
.with_model_format(coreml::ModelFormat::MLProgram)
|
||||
.with_compute_units(coreml::ComputeUnits::All);
|
||||
if let Some(dir) = cache {
|
||||
let _ = std::fs::create_dir_all(dir);
|
||||
ep = ep.with_model_cache_dir(dir.to_string_lossy());
|
||||
}
|
||||
Ok(b.with_execution_providers([ep.build().error_on_failure()])?)
|
||||
}
|
||||
|
||||
#[cfg(not(any(target_os = "android", target_os = "macos")))]
|
||||
fn coreml(
|
||||
_b: ort::session::builder::SessionBuilder,
|
||||
_cache: Option<&std::path::Path>,
|
||||
) -> ort::Result<ort::session::builder::SessionBuilder> {
|
||||
unreachable!("the CoreML rung is on the macOS ladder only")
|
||||
}
|
||||
|
||||
/// Register MIGraphX through ONNX Runtime's generic key/value entry point.
|
||||
///
|
||||
/// `ort`'s own builder (`ep::MIGraphX`) fills the legacy
|
||||
@@ -211,8 +250,8 @@ fn providers(
|
||||
.build()
|
||||
.error_on_failure()])?)
|
||||
}
|
||||
Rung::Cuda | Rung::TensorRt | Rung::MiGraphX => {
|
||||
unreachable!("no desktop GPU rung on Android")
|
||||
Rung::Cuda | Rung::TensorRt | Rung::MiGraphX | Rung::CoreMl => {
|
||||
unreachable!("no desktop rung on Android")
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
Reference in New Issue
Block a user