Add a CoreML rung on macOS

The macOS ladder was the CPU provider alone, with CoreML listed as a gap.
It is now CoreML, then the CPU, then tract — unmeasured, since nobody here
has a Mac, and safe to ship unmeasured because the probe's clock rejects a
CoreML slower than the CPU and `attempt` refuses one that crashes.

- `Rung::CoreMl`, a compiling rung like TensorRT: an ML Program with every
  compute unit allowed, falling back to the CPU until each model's program
  is built. The embedder stays on the CPU, as on the Hexagon (§7).
- The cache is one directory per model and runtime version. CoreML keys a
  model committed from memory on its input and node names, not its
  weights (ONNX Runtime 1.29, coreml_execution_provider.cc), so two
  exports of one architecture would otherwise share a program.
- The fingerprint on macOS is the chip and the OS release, which ships
  CoreML.
- The desktop looks for the runtime in the bundle's Contents/Frameworks
  and Homebrew's prefixes; fetch-desktop-runtime.sh on a Mac downloads
  ONNX Runtime 1.29.0 for Apple silicon, which carries CoreML.

docs/dev/macos.md says what exists, how to build it, and which log lines
to ask a Mac user for.
This commit is contained in:
2026-10-03 16:50:37 -04:00
parent 872e35670c
commit c73743394f
10 changed files with 283 additions and 19 deletions
+49 -10
View File
@@ -27,13 +27,14 @@ pub fn build(rung: Rung, role: Role, bytes: &[u8], cfg: &Config) -> ort::Result<
// what makes the second case rare (§6).
let context = (rung == Rung::Hexagon).then(|| crate::engines::context_path(cfg, bytes));
let ready = context.as_ref().is_some_and(|p| p.is_file());
b = providers(
b,
rung,
role,
cfg,
if ready { None } else { context.as_deref() },
)?;
// What the rung keeps for this model: the context the Hexagon is to
// write, or the directory CoreML compiles into.
let per_model = match rung {
Rung::CoreMl => Some(crate::engines::coreml_dir(cfg, bytes)),
_ if ready => None,
_ => context.clone(),
};
b = providers(b, rung, role, cfg, per_model.as_deref())?;
match (ready, context) {
(true, Some(path)) => b.commit_from_file(path),
_ => b.commit_from_memory(bytes),
@@ -90,11 +91,12 @@ fn providers(
rung: Rung,
role: Role,
cfg: &Config,
_generate_context: Option<&std::path::Path>,
per_model: Option<&std::path::Path>,
) -> ort::Result<ort::session::builder::SessionBuilder> {
use ort::ep;
match rung {
Rung::Cpu => Ok(b),
Rung::CoreMl => coreml(b, per_model),
Rung::Cuda => {
Ok(b.with_execution_providers([ep::CUDA::default().build().error_on_failure()])?)
}
@@ -139,6 +141,43 @@ fn providers(
}
}
/// CoreML, compiling an ML Program — the format with the operators these
/// graphs use and the one that reaches the Neural Engine — into `cache`.
///
/// The option names are those ONNX Runtime 1.29 reads from the generic
/// key/value map (`coreml_options.cc`), which is what `ort`'s builder
/// fills. The cache is per model because of how CoreML keys it: a model
/// committed from memory, as every session here is, has no path, and the
/// key falls back to a hash of the graph's input and node names — not its
/// weights. Two exports of one architecture would share a program. The
/// directory `engines::coreml_dir` names is the hash of the bytes.
///
/// Every compute unit is allowed, so CoreML may place a graph on the
/// Neural Engine, the GPU or the CPU; the probe's clock judges the result.
#[cfg(target_os = "macos")]
fn coreml(
b: ort::session::builder::SessionBuilder,
cache: Option<&std::path::Path>,
) -> ort::Result<ort::session::builder::SessionBuilder> {
use ort::ep::{self, coreml};
let mut ep = ep::CoreML::default()
.with_model_format(coreml::ModelFormat::MLProgram)
.with_compute_units(coreml::ComputeUnits::All);
if let Some(dir) = cache {
let _ = std::fs::create_dir_all(dir);
ep = ep.with_model_cache_dir(dir.to_string_lossy());
}
Ok(b.with_execution_providers([ep.build().error_on_failure()])?)
}
#[cfg(not(any(target_os = "android", target_os = "macos")))]
fn coreml(
_b: ort::session::builder::SessionBuilder,
_cache: Option<&std::path::Path>,
) -> ort::Result<ort::session::builder::SessionBuilder> {
unreachable!("the CoreML rung is on the macOS ladder only")
}
/// Register MIGraphX through ONNX Runtime's generic key/value entry point.
///
/// `ort`'s own builder (`ep::MIGraphX`) fills the legacy
@@ -211,8 +250,8 @@ fn providers(
.build()
.error_on_failure()])?)
}
Rung::Cuda | Rung::TensorRt | Rung::MiGraphX => {
unreachable!("no desktop GPU rung on Android")
Rung::Cuda | Rung::TensorRt | Rung::MiGraphX | Rung::CoreMl => {
unreachable!("no desktop rung on Android")
}
}
}