Compare commits

..
Author SHA1 Message Date
dtourolle 064be89a73 Link the inference engine for macOS in a zig container
`docker/macos` builds for aarch64-apple-darwin from Linux with
cargo-zigbuild. Zig carries libSystem and the C headers, so tract's SIMD
kernels compile and the engine's test binaries and examples link as Mach-O
arm64 — the check `cargo check --target` could not do, because tract's
build script needs a macOS C compiler. Crates that link an Apple framework
(dr-plat's keyring, and so the app) still need the Xcode SDK and fail at
the link; macos.md says so.
2026-09-29 22:03:12 -04:00
dtourolle bff12f81a9 Add a CoreML rung on macOS
The macOS ladder was the CPU provider alone, with CoreML listed as a gap.
It is now CoreML, then the CPU, then tract — unmeasured, since nobody here
has a Mac, and safe to ship unmeasured because the probe's clock rejects a
CoreML slower than the CPU and `attempt` refuses one that crashes.

- `Rung::CoreMl`, a compiling rung like TensorRT: an ML Program with every
  compute unit allowed, falling back to the CPU until each model's program
  is built. The embedder stays on the CPU, as on the Hexagon (§7).
- The cache is one directory per model and runtime version. CoreML keys a
  model committed from memory on its input and node names, not its
  weights (ONNX Runtime 1.29, coreml_execution_provider.cc), so two
  exports of one architecture would otherwise share a program.
- The fingerprint on macOS is the chip and the OS release, which ships
  CoreML.
- The desktop looks for the runtime in the bundle's Contents/Frameworks
  and Homebrew's prefixes; fetch-desktop-runtime.sh on a Mac downloads
  ONNX Runtime 1.29.0 for Apple silicon, which carries CoreML.

docs/dev/macos.md says what exists, how to build it, and which log lines
to ask a Mac user for.
2026-09-29 21:33:02 -04:00
dtourolle 07e85cf0b2 Log like a debug build on macOS, where a Mac user can find it
Nobody working on DarkRoom has a Mac, so every macOS build is in the hands
of someone who can send a log and cannot attach a debugger. Three changes
make that log worth sending:

- The desktop's default filter on macOS is `debug` for every `dr_*` crate,
  the desktop crate and `onnxruntime` (the runtime's own session log).
- The state directory — the log and crash records — is `~/Library/Logs`
  on macOS rather than the `~/.local/state` Finder hides; Console.app
  lists it. Config and data keep the Unix rules.
- A `diagnostic` cargo profile: release plus line tables, so a crash
  record's backtrace reads file:line. On macOS the tables are in the
  `.dSYM` beside the executable, which the bundle must keep.
2026-09-29 21:33:02 -04:00
dtourolle b23527d5dc Stop retrying a provider that took the app down
The probe runs in the app's process, and a provider can fail by aborting
rather than by returning an error — XNNPACK did on SCRFD. A rung that does
that once would do it on every launch, before the first photograph is on
screen.

Every session build above the CPU, the probe's and each background
compile's, now writes what it is attempting to `attempt` in the cache
directory first and removes it after. After two launches in a row that
died inside the same attempt it is refused and recorded — a rung in
`failed`, an engine in the new `refused` — until the fingerprint changes.
Two, not one, because quitting during a TensorRT compile leaves the same
file.
2026-09-29 21:33:02 -04:00
dtourolle 3b783d7bd4 Send ONNX Runtime's session log to the app's log
A native session's messages went to ONNX Runtime's stdio logger, which is
nowhere once the app is launched from a menu — and what a provider says
while partitioning a graph (nodes taken, operators declined, a library that
failed to load) is most of what a failed rung tells you. Each session now
forwards them to `log` under the target `onnxruntime`: warnings always,
the runtime's info lines at `debug`, its verbose lines at `trace`.
2026-09-29 21:33:02 -04:00
16 changed files with 615 additions and 31 deletions
+12
View File
@@ -276,6 +276,18 @@ opt-level = 0
lto = "thin"
codegen-units = 1
# A release build that can say where it panicked: line tables, so a crash
# record's backtrace (`dr_plat::crash`) reads `file.rs:123` rather than bare
# addresses. The macOS build uses it (docs/dev/macos.md) — no one here can
# reproduce a Mac bug, so its reports carry what a debugger would have — at
# the price of a larger binary and no slower code. On macOS the tables land
# in a `.dSYM` beside the executable (rustc's default `packed`), and the
# bundle must carry that directory next to the binary for the backtrace to
# find it.
[profile.diagnostic]
inherits = "release"
debug = "line-tables-only"
# Three upstream crates carry a local patch: wgpu-hal and Slint's Skia
# renderer so that the Android build can draw with wgpu on a rotated display
# (technical-debt.md TD-1), and rawler so that a linear DNG wider than 16 700
+35 -4
View File
@@ -15,6 +15,22 @@ use std::path::PathBuf;
use dr_plat::diagnostics::Installed;
/// What the log keeps when `RUST_LOG` does not say.
#[cfg(not(target_os = "macos"))]
const DEFAULT_LOG: &str =
"info,wgpu_core=warn,wgpu_hal=warn,zbus=warn,tracing=warn,calloop=warn,rawler=warn";
/// The same, and `debug` from this application's own crates and from ONNX
/// Runtime, whose `debug` is how many nodes each provider took
/// (docs/dev/macos.md). Nobody here runs a Mac: every macOS build is in
/// the hands of someone who can send us a log and cannot attach a debugger,
/// so the log is written as if for a debug build. `dr_` is a prefix, and
/// `env_logger` matches directives by prefix, so it names every `dr-*`
/// crate — present and future — without naming a dependency.
#[cfg(target_os = "macos")]
const DEFAULT_LOG: &str = "info,dr_=debug,darkroom_desktop=debug,onnxruntime=debug,\
wgpu_core=warn,wgpu_hal=warn,zbus=warn,tracing=warn,calloop=warn,rawler=warn";
fn main() -> anyhow::Result<()> {
// TRACES: FR-PLAT-WIN-3
// Before the logger, the crash hook and everything else: this exists so a
@@ -33,10 +49,9 @@ fn main() -> anyhow::Result<()> {
// (NFR-OPS-1). `filter()` is asked afterwards because the environment may
// have overridden the default below, and the file must not be quieter than
// the terminal.
let console = env_logger::Builder::from_env(env_logger::Env::default().default_filter_or(
"info,wgpu_core=warn,wgpu_hal=warn,zbus=warn,tracing=warn,calloop=warn,rawler=warn",
))
.build();
let console =
env_logger::Builder::from_env(env_logger::Env::default().default_filter_or(DEFAULT_LOG))
.build();
let level = console.filter();
let logging = dr_plat::diagnostics::install(Box::new(console), level);
@@ -109,5 +124,21 @@ fn runtime_dirs() -> Vec<PathBuf> {
PathBuf::from("/usr/lib/darkroom"),
PathBuf::from("/usr/lib"),
]);
// An app bundle keeps its libraries in `Contents/Frameworks`, beside
// the `Contents/MacOS` the executable is in; then Homebrew's
// `onnxruntime`, Apple silicon's prefix before Intel's. Homebrew's build
// may lack CoreML, which the probe finds out for itself.
#[cfg(target_os = "macos")]
{
if let Ok(exe) = std::env::current_exe() {
if let Some(bin) = exe.parent() {
dirs.push(bin.join("../Frameworks"));
}
}
dirs.extend([
PathBuf::from("/opt/homebrew/lib"),
PathBuf::from("/usr/local/lib"),
]);
}
dirs
}
+5
View File
@@ -38,6 +38,11 @@ ort = { workspace = true, features = ["cuda", "tensorrt"] }
[target.'cfg(target_os = "android")'.dependencies]
ort = { workspace = true, features = ["qnn"] }
# The Apple rung: CoreML's option builder, which fills the runtime's generic
# key/value map. `ort-sys`'s `coreml` feature is empty; nothing links.
[target.'cfg(target_os = "macos")'.dependencies]
ort = { workspace = true, features = ["coreml"] }
[features]
# The floor: `tract` supplies the API table when no runtime file is found, or
# always, in a build without `native`. Tests want this and nothing else.
+32 -3
View File
@@ -44,6 +44,20 @@ pub fn context_path(cfg: &Config, bytes: &[u8]) -> PathBuf {
.join(format!("{:016x}_ctx.onnx", hash(bytes)))
}
/// Where CoreML compiles `bytes` to: one directory per model, because
/// CoreML's own cache key leaves out the weights of a model loaded from
/// memory (`session::coreml`), and one per runtime version, which wrote it.
pub fn coreml_dir(cfg: &Config, bytes: &[u8]) -> PathBuf {
let runtime = match crate::api::runtime() {
crate::Runtime::OnnxRuntime { version, .. } => version,
crate::Runtime::Tract => "tract".into(),
};
cfg.cache_dir
.join("coreml")
.join(runtime)
.join(format!("{:016x}", hash(bytes)))
}
/// After the probe: compile every configured model the selected rung can
/// take, smallest first, recording each as it lands.
pub fn run() {
@@ -90,12 +104,27 @@ pub fn run() {
Source::Bytes(b) => (b.to_vec(), format!("embedded {role:?}")),
};
let key = key(rung, &bytes);
if state().lock().unwrap().cache.compiled.contains(&key) {
continue;
{
let s = state().lock().unwrap();
if s.cache.compiled.contains(&key) || s.cache.refused.contains(&key) {
continue;
}
}
log::info!("inference: compiling {name} for {}", rung.label());
let started = std::time::Instant::now();
match crate::session::build(rung, role, &bytes, &cfg) {
let built = match crate::probe::attempt(&cfg, &key, || {
crate::session::build(rung, role, &bytes, &cfg)
}) {
Ok(built) => built,
Err(_) => {
// Refused: the process died inside this compile before.
let mut s = state().lock().unwrap();
s.cache.refused.insert(key);
crate::probe::write_cache(&s.config, &s.cache);
continue;
}
};
match built {
Ok(session) => {
drop(session);
let mut s = state().lock().unwrap();
+42 -3
View File
@@ -78,6 +78,12 @@ pub enum Rung {
MiGraphX,
/// Qualcomm's Hexagon NPU through QNN, int8 models only. Android only.
Hexagon,
/// Apple, through CoreML: the Neural Engine, the GPU or the CPU, as
/// CoreML schedules it. macOS only. Compiles an ML Program per model on
/// first use, so it is a compiling rung with the CPU below it. The
/// embedder stays on the CPU, as on the Hexagon: the Neural Engine
/// computes in fp16 (§7).
CoreMl,
}
impl Rung {
@@ -88,6 +94,7 @@ impl Rung {
Rung::TensorRt => "TensorRT",
Rung::MiGraphX => "MIGraphX",
Rung::Hexagon => "Hexagon NPU",
Rung::CoreMl => "CoreML",
}
}
@@ -96,13 +103,16 @@ impl Rung {
fn fallback(self) -> Rung {
match self {
Rung::TensorRt => Rung::Cuda,
Rung::MiGraphX | Rung::Hexagon | Rung::Cuda | Rung::Cpu => Rung::Cpu,
Rung::MiGraphX | Rung::Hexagon | Rung::CoreMl | Rung::Cuda | Rung::Cpu => Rung::Cpu,
}
}
/// Whether a session on this rung needs an engine built first.
fn compiles(self) -> bool {
matches!(self, Rung::TensorRt | Rung::MiGraphX | Rung::Hexagon)
matches!(
self,
Rung::TensorRt | Rung::MiGraphX | Rung::Hexagon | Rung::CoreMl
)
}
/// The model form this rung wants for a role.
@@ -116,9 +126,11 @@ impl Rung {
/// Whether this rung runs `role` at all. The Hexagon takes int8 graphs
/// only, and the embedder is never int8 (§7) — it runs on the CPU
/// beside a detector on the NPU, so its vectors compare across devices.
/// CoreML is kept off the embedder for the same reason: the Neural
/// Engine is fp16, and which unit runs a graph is CoreML's choice.
fn serves(self, role: Role) -> bool {
match self {
Rung::Hexagon => role != Role::Embedder,
Rung::Hexagon | Rung::CoreMl => role != Role::Embedder,
_ => true,
}
}
@@ -348,6 +360,12 @@ struct Cache {
/// the fingerprint changes: a wedged driver must not cost every launch
/// thirty seconds.
failed: Vec<(Rung, String)>,
/// Engine keys whose compile the process died inside, launch after
/// launch (`probe::attempt`). Left on the fallback until the
/// fingerprint changes. Defaulted, so a cache from before this field
/// still reads.
#[serde(default)]
refused: BTreeSet<String>,
}
struct State {
@@ -631,6 +649,27 @@ mod tests {
);
}
#[test]
fn coreml_takes_a_compiled_detector_and_never_the_embedder() {
let hash = engines::hash(b"detector");
let mut s = State {
config: Config::default(),
cache: Cache {
rung: Some(Rung::CoreMl),
..Cache::default()
},
probing: false,
wanted: 0,
};
let on = |s: &State, role| effective_rung(s, Rung::CoreMl, role, Form::F32, hash);
// Before its program is compiled the detector waits on the CPU.
assert_eq!(on(&s, Role::Detector), Rung::Cpu);
s.cache.compiled.insert(engines::key_of(Rung::CoreMl, hash));
assert_eq!(on(&s, Role::Detector), Rung::CoreMl);
// The embedder does not move, compiled or not (§7).
assert_eq!(on(&s, Role::Embedder), Rung::Cpu);
}
#[test]
fn the_status_reports_only_the_rungs_above_the_selection() {
let _serial = serial();
+147 -3
View File
@@ -16,10 +16,15 @@ use crate::{api::Runtime, state, Cache, Config, Form, Role, Rung};
fn ladder(ceiling: Option<Rung>) -> Vec<Rung> {
#[cfg(target_os = "android")]
let all = [Rung::Hexagon];
// Unmeasured (§2 ⁵): it is on the ladder because the probe's clock and
// `attempt` make a wrong guess cost one slow or failed probe, not a
// slow or crashing app.
#[cfg(target_os = "macos")]
let all = [Rung::CoreMl];
// A desktop has one vendor's GPU; the other vendor's providers are
// "not enabled in this build" or a library that fails to load, and
// either answer arrives in milliseconds.
#[cfg(not(target_os = "android"))]
#[cfg(not(any(target_os = "android", target_os = "macos")))]
let all = [Rung::TensorRt, Rung::Cuda, Rung::MiGraphX];
all.into_iter()
.filter(|r| ceiling.is_none_or(|c| *r <= c))
@@ -81,7 +86,11 @@ pub fn run(runtime: Runtime) {
log::info!("inference: floor {floor:.1} ms on the CPU provider");
for rung in ladder(cfg.ceiling) {
match time_rung(rung, role, &canonical, &cfg) {
let timed = attempt(&cfg, &format!("probe {}", rung.label()), || {
time_rung(rung, role, &canonical, &cfg)
})
.and_then(|timed| timed);
match timed {
Ok((ms, key)) if ms < floor => {
cache.rung = Some(rung);
cache.reason = format!("{ms:.1} ms against {floor:.1} ms on the CPU");
@@ -118,6 +127,52 @@ fn finish(cache: Cache) {
s.probing = false;
}
/// How many launches in a row may die inside one attempt before it is
/// refused. Two, not one: quitting the app while TensorRT spends forty
/// seconds on an engine leaves the same trace as a provider that aborted.
const STRIKES: u32 = 2;
/// Run `f` — a session build on a provider — with `what` written down
/// first, so that if the provider takes the process with it the next launch
/// knows what to stop trying.
///
/// A provider can fail by aborting rather than by returning an error:
/// XNNPACK did on SCRFD (§2), and a C++ exception or a panic across the C
/// API is an abort. The probe runs in the app's own process, so a rung that
/// does this once would do it on every launch, before the first photograph
/// is on screen. The file (`attempt` in the cache directory) holds the
/// attempt and how many launches have started it without finishing;
/// finishing, by success or by error, removes it. After [`STRIKES`] the
/// attempt is refused, and the caller records the refusal in the cache,
/// where it lasts until the fingerprint changes like any other failure.
pub fn attempt<T>(cfg: &Config, what: &str, f: impl FnOnce() -> T) -> Result<T, String> {
if cfg.cache_dir.as_os_str().is_empty() {
return Ok(f());
}
let path = cfg.cache_dir.join("attempt");
let died = std::fs::read_to_string(&path)
.ok()
.and_then(|s| {
let (w, n) = s.split_once('\t')?;
(w == what).then(|| n.trim().parse::<u32>().ok())?
})
.unwrap_or(0);
if died >= STRIKES {
log::error!("inference: the app died during `{what}` on the last {died} launches; not trying it again");
return Err(format!(
"the app died while trying this on {died} launches in a row"
));
}
if died > 0 {
log::warn!("inference: the last launch died during `{what}`; trying it once more");
}
let _ = std::fs::create_dir_all(&cfg.cache_dir);
let _ = std::fs::write(&path, format!("{what}\t{}", died + 1));
let out = f();
let _ = std::fs::remove_file(&path);
Ok(out)
}
/// The smallest detector, or the smallest model of any role if there is
/// none. A ~2 MB detector is the cheapest real test of a provider, and the
/// detector is the role the int8 forms exist for — the eye classifiers are
@@ -310,7 +365,50 @@ fn system_property(name: &str) -> String {
String::from_utf8_lossy(&buf[..n.max(0) as usize]).into_owned()
}
#[cfg(not(any(target_os = "linux", target_os = "android")))]
#[cfg(target_os = "macos")]
fn device_identity() -> String {
// The chip, and the OS release: CoreML ships with the OS, so a macOS
// update is a new provider as surely as a new driver is on Linux.
format!(
"{} macOS {}",
sysctl("machdep.cpu.brand_string"),
sysctl("kern.osproductversion")
)
}
#[cfg(target_os = "macos")]
fn sysctl(name: &str) -> String {
extern "C" {
fn sysctlbyname(
name: *const std::ffi::c_char,
oldp: *mut std::ffi::c_void,
oldlenp: *mut usize,
newp: *mut std::ffi::c_void,
newlen: usize,
) -> i32;
}
let name = std::ffi::CString::new(name).unwrap();
let mut buf = [0u8; 256];
let mut len = buf.len();
// SAFETY: libSystem's documented call; `len` is the buffer's size in and
// the string's length, with its terminator, out.
let rc = unsafe {
sysctlbyname(
name.as_ptr(),
buf.as_mut_ptr().cast(),
&mut len,
std::ptr::null_mut(),
0,
)
};
if rc != 0 {
return String::new();
}
let s = &buf[..len.min(buf.len())];
String::from_utf8_lossy(s.strip_suffix(&[0]).unwrap_or(s)).into_owned()
}
#[cfg(not(any(target_os = "linux", target_os = "android", target_os = "macos")))]
fn device_identity() -> String {
String::new()
}
@@ -338,3 +436,49 @@ pub fn write_cache(cfg: &Config, cache: &Cache) {
}
}
}
#[cfg(test)]
mod tests {
use super::*;
fn a_cache_dir(name: &str) -> Config {
let dir = std::env::temp_dir().join(format!("dr-attempt-{}-{name}", std::process::id()));
let _ = std::fs::remove_dir_all(&dir);
Config {
cache_dir: dir,
..Config::default()
}
}
/// What a launch that died inside `what` leaves behind.
fn died_inside(cfg: &Config, what: &str, launches: u32) {
std::fs::create_dir_all(&cfg.cache_dir).unwrap();
std::fs::write(cfg.cache_dir.join("attempt"), format!("{what}\t{launches}")).unwrap();
}
#[test]
fn a_finished_attempt_leaves_no_trace() {
let cfg = a_cache_dir("finished");
assert_eq!(attempt(&cfg, "probe CoreML", || 7), Ok(7));
assert!(!cfg.cache_dir.join("attempt").exists());
}
#[test]
fn one_death_is_forgiven_and_two_are_not() {
let cfg = a_cache_dir("strikes");
died_inside(&cfg, "probe CoreML", 1);
assert_eq!(attempt(&cfg, "probe CoreML", || 7), Ok(7));
died_inside(&cfg, "probe CoreML", 2);
let mut ran = false;
assert!(attempt(&cfg, "probe CoreML", || ran = true).is_err());
assert!(!ran, "a refused attempt must not run");
}
#[test]
fn another_attempts_deaths_do_not_count() {
let cfg = a_cache_dir("other");
died_inside(&cfg, "probe TensorRT", 2);
assert_eq!(attempt(&cfg, "probe CUDA", || 7), Ok(7));
}
}
+85 -10
View File
@@ -19,24 +19,61 @@ pub fn build(rung: Rung, role: Role, bytes: &[u8], cfg: &Config) -> ort::Result<
// `stack_tensors`) — a panic across the C API, which is an abort. The
// app never asked tract for that and does not start now.
let mut b = Session::builder()?.with_intra_threads(threads(cfg))?;
if crate::api::runtime().is_native() {
b = with_runtime_log(b)?;
}
// A Hexagon session loads the compiled context when there is one and
// compiles it from the model when there is not; the engine thread is
// what makes the second case rare (§6).
let context = (rung == Rung::Hexagon).then(|| crate::engines::context_path(cfg, bytes));
let ready = context.as_ref().is_some_and(|p| p.is_file());
b = providers(
b,
rung,
role,
cfg,
if ready { None } else { context.as_deref() },
)?;
// What the rung keeps for this model: the context the Hexagon is to
// write, or the directory CoreML compiles into.
let per_model = match rung {
Rung::CoreMl => Some(crate::engines::coreml_dir(cfg, bytes)),
_ if ready => None,
_ => context.clone(),
};
b = providers(b, rung, role, cfg, per_model.as_deref())?;
match (ready, context) {
(true, Some(path)) => b.commit_from_file(path),
_ => b.commit_from_memory(bytes),
}
}
/// Send the runtime's own messages for this session to `log`, under the
/// target `onnxruntime`, instead of to ONNX Runtime's stdio logger.
///
/// Its stderr is nowhere once the app is launched from a menu, and what a
/// provider says while it partitions a graph — how many nodes it took, which
/// operator it declined, the library it failed to load — is most of what a
/// failed rung tells you (docs/dev/inference.md §4). The level follows the
/// filter: warnings always, `debug` adds the runtime's info lines (the
/// partition counts), `trace` its verbose ones (every node placement).
fn with_runtime_log(
b: ort::session::builder::SessionBuilder,
) -> ort::Result<ort::session::builder::SessionBuilder> {
use ort::logging::LogLevel;
let level = if log::log_enabled!(target: "onnxruntime", log::Level::Trace) {
LogLevel::Verbose
} else if log::log_enabled!(target: "onnxruntime", log::Level::Debug) {
LogLevel::Info
} else {
LogLevel::Warning
};
let forward = |level: LogLevel, _category: &str, _id: &str, location: &str, message: &str| {
let level = match level {
LogLevel::Verbose => log::Level::Trace,
LogLevel::Info => log::Level::Debug,
LogLevel::Warning => log::Level::Warn,
LogLevel::Error | LogLevel::Fatal => log::Level::Error,
};
log::log!(target: "onnxruntime", level, "{message} ({location})");
};
Ok(b.with_logger(std::sync::Arc::new(forward))?
.with_log_level(level)?)
}
/// The intra-op pool: what the config says, else the cores less two for
/// the compositor and the decoder (§9). tract ignores it.
fn threads(cfg: &Config) -> usize {
@@ -54,11 +91,12 @@ fn providers(
rung: Rung,
role: Role,
cfg: &Config,
_generate_context: Option<&std::path::Path>,
per_model: Option<&std::path::Path>,
) -> ort::Result<ort::session::builder::SessionBuilder> {
use ort::ep;
match rung {
Rung::Cpu => Ok(b),
Rung::CoreMl => coreml(b, per_model),
Rung::Cuda => {
Ok(b.with_execution_providers([ep::CUDA::default().build().error_on_failure()])?)
}
@@ -103,6 +141,43 @@ fn providers(
}
}
/// CoreML, compiling an ML Program — the format with the operators these
/// graphs use and the one that reaches the Neural Engine — into `cache`.
///
/// The option names are those ONNX Runtime 1.29 reads from the generic
/// key/value map (`coreml_options.cc`), which is what `ort`'s builder
/// fills. The cache is per model because of how CoreML keys it: a model
/// committed from memory, as every session here is, has no path, and the
/// key falls back to a hash of the graph's input and node names — not its
/// weights. Two exports of one architecture would share a program. The
/// directory `engines::coreml_dir` names is the hash of the bytes.
///
/// Every compute unit is allowed, so CoreML may place a graph on the
/// Neural Engine, the GPU or the CPU; the probe's clock judges the result.
#[cfg(target_os = "macos")]
fn coreml(
b: ort::session::builder::SessionBuilder,
cache: Option<&std::path::Path>,
) -> ort::Result<ort::session::builder::SessionBuilder> {
use ort::ep::{self, coreml};
let mut ep = ep::CoreML::default()
.with_model_format(coreml::ModelFormat::MLProgram)
.with_compute_units(coreml::ComputeUnits::All);
if let Some(dir) = cache {
let _ = std::fs::create_dir_all(dir);
ep = ep.with_model_cache_dir(dir.to_string_lossy());
}
Ok(b.with_execution_providers([ep.build().error_on_failure()])?)
}
#[cfg(not(any(target_os = "android", target_os = "macos")))]
fn coreml(
_b: ort::session::builder::SessionBuilder,
_cache: Option<&std::path::Path>,
) -> ort::Result<ort::session::builder::SessionBuilder> {
unreachable!("the CoreML rung is on the macOS ladder only")
}
/// Register MIGraphX through ONNX Runtime's generic key/value entry point.
///
/// `ort`'s own builder (`ep::MIGraphX`) fills the legacy
@@ -175,8 +250,8 @@ fn providers(
.build()
.error_on_failure()])?)
}
Rung::Cuda | Rung::TensorRt | Rung::MiGraphX => {
unreachable!("no desktop GPU rung on Android")
Rung::Cuda | Rung::TensorRt | Rung::MiGraphX | Rung::CoreMl => {
unreachable!("no desktop rung on Android")
}
}
}
+48
View File
@@ -0,0 +1,48 @@
# DarkRoom — macOS link check
#
# Compiles and links for macOS from Linux, with zig as the linker
# (cargo-zigbuild). Zig carries macOS's libSystem stubs and C headers, so the
# crates that need only libSystem — the inference engine, dr-plat — build,
# link and produce Mach-O test binaries here. Nothing runs: there is no macOS
# to run them on (docs/dev/macos.md §2). The desktop app needs Apple's
# framework headers (AppKit, Metal, Security), which only the Xcode SDK
# carries, so it does not link here.
#
# Build: docker build -t darkroom-macos:latest docker/macos
# Use: ./docker/macos/build.sh cargo zigbuild --target aarch64-apple-darwin -p dr-inference-engine --all-targets
FROM docker.io/library/debian:trixie-slim
# Pinned, like the Windows and Android images. Rust matches rust-toolchain.toml.
ARG RUST_VERSION=1.92.0
ARG ZIG_VERSION=0.15.2
ARG ZIG_SHA256=02aa270f183da276e5b5920b1dac44a63f1a49e55050ebde3aecc9eb82f93239
ARG CARGO_ZIGBUILD_VERSION=0.23.4
ENV DEBIAN_FRONTEND=noninteractive \
CARGO_HOME=/opt/cargo \
RUSTUP_HOME=/opt/rustup \
PATH=/opt/zig:/opt/cargo/bin:$PATH
RUN apt-get update && apt-get install -y --no-install-recommends \
ca-certificates curl git xz-utils \
# A host C compiler: build scripts and proc-macros are Linux binaries.
gcc libc6-dev \
# `file` says Mach-O; the smoke check in build.sh reads it.
file \
&& rm -rf /var/lib/apt/lists/*
RUN curl -fsSL "https://ziglang.org/download/${ZIG_VERSION}/zig-x86_64-linux-${ZIG_VERSION}.tar.xz" -o /tmp/zig.tar.xz \
&& echo "${ZIG_SHA256} /tmp/zig.tar.xz" | sha256sum -c - \
&& mkdir /opt/zig && tar xJf /tmp/zig.tar.xz -C /opt/zig --strip-components=1 \
&& rm /tmp/zig.tar.xz && zig version
# The components rust-toolchain.toml lists, baked in so rustup does not fetch
# them inside every run.
RUN curl -fsSL https://sh.rustup.rs | sh -s -- -y --profile minimal \
--default-toolchain "${RUST_VERSION}" \
--component rustfmt,clippy,rust-analyzer \
--target aarch64-apple-darwin,x86_64-apple-darwin \
&& cargo install --locked "cargo-zigbuild@${CARGO_ZIGBUILD_VERSION}" \
&& rm -rf /opt/cargo/registry \
&& chmod -R a+rwX /opt/cargo /opt/rustup
+62
View File
@@ -0,0 +1,62 @@
#!/usr/bin/env bash
# Run a command inside the DarkRoom macOS link-check container.
#
# ./docker/macos/build.sh cargo zigbuild --target aarch64-apple-darwin -p dr-inference-engine --all-targets
# ./docker/macos/build.sh # interactive shell
#
# Builds the image on first use; `--rebuild` after editing the Dockerfile.
set -euo pipefail
IMAGE="darkroom-macos:latest"
HERE="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
REPO="$(cd "${HERE}/../.." && pwd)"
if command -v podman >/dev/null 2>&1; then
ENGINE=podman
elif command -v docker >/dev/null 2>&1; then
ENGINE=docker
else
echo "error: neither podman nor docker found" >&2
exit 1
fi
if [[ "${1:-}" == "--rebuild" ]]; then
shift
"${ENGINE}" build -t "${IMAGE}" "${HERE}"
elif ! "${ENGINE}" image inspect "${IMAGE}" >/dev/null 2>&1; then
echo "==> building ${IMAGE} (first run; a few minutes)"
"${ENGINE}" build -t "${IMAGE}" "${HERE}"
fi
# Registry, target and zig's own cache persist across runs.
CACHE="${XDG_CACHE_HOME:-${HOME}/.cache}/darkroom-macos"
mkdir -p "${CACHE}/registry" "${CACHE}/target" "${CACHE}/home"
ARGS=(
--rm
-v "${REPO}:/work:z"
-v "${CACHE}/registry:/opt/cargo/registry:z"
-v "${CACHE}/target:/work/target-macos:z"
-v "${CACHE}/home:/tmp/home:z"
-e HOME=/tmp/home
-e CARGO_TARGET_DIR=/work/target-macos
-w /work
)
# Capped for the same reason as the Windows image: a cross build otherwise
# takes every thread on the host.
JOBS="${DARKROOM_BUILD_JOBS:-8}"
if [[ "${JOBS}" != "0" ]]; then
ARGS+=(--cpus "${JOBS}" -e "CARGO_BUILD_JOBS=${JOBS}")
fi
if [[ "${ENGINE}" == "docker" ]]; then
ARGS+=(--user "$(id -u):$(id -g)")
fi
if [[ $# -eq 0 ]]; then
ARGS+=(-it)
set -- /bin/bash
fi
exec "${ENGINE}" run "${ARGS[@]}" "${IMAGE}" "$@"
+15 -3
View File
@@ -151,10 +151,14 @@ winning:
| Linux / Windows, NVIDIA GPU | TensorRT, f32 model, fp16 engine | CUDA provider, f32 | ORT CPU, f32 | tract |
| Linux, AMD GPU with ROCm | MIGraphX, f32 model, fp16 program | ORT CPU, f32 | — | tract |
| Linux / Windows, no GPU stack | ORT CPU, f32 | — | — | tract |
| macOS ⁵ | ORT CPU, f32 | — | — | tract |
| macOS ⁵ | CoreML, f32 model, ML Program | ORT CPU, f32 | — | tract |
⁵ CoreML is the obvious rung and is unmeasured; it is listed so its absence is a gap and not an
oversight.
⁵ **Unmeasured**, and the one exception to the rule below: nobody here has a Mac. The rung is on
the ladder because the probe makes a wrong guess cheap — a CoreML that is slower than the CPU is
rejected by §4's clock, one that errors is recorded as failed, and one that takes the process
down is refused on the third launch (§4, `attempt`). The embedder stays on the CPU (§7). The first
macOS log that shows a probe line is this row's measurement; [macos.md](macos.md) says what to
ask for.
Deliberately **not** on any ladder, with the measurement that excluded each: NNAPI (no driver),
XNNPACK (slower than CPU, aborts on SCRFD), WebGPU (slower than CPU), the Adreno through QNN (works,
@@ -270,6 +274,14 @@ What the probe may not do:
- **Retry a rung that failed within a session.** A failed probe is cached as a failure with the
same inputs; the rung is tried again when an input changes. Otherwise a wedged driver means a
thirty-second stall on every launch.
- **Crash the app twice for the same reason.** The probe runs in the app's process, and a provider
can fail by aborting rather than by returning an error (XNNPACK on SCRFD, §2). Every session build
on a rung above the CPU — the probe's, and each background compile of §6 — writes what it is
attempting to `attempt` in the cache directory first and removes it after. A launch that finds the
file knows the last one died inside that attempt; after two such launches in a row the attempt is
refused and recorded like any other failure (a rung in `failed`, an engine in `refused`), until
the fingerprint changes. Two, not one, because quitting during a forty-second TensorRT compile
leaves the same file.
- **Choose for the user without saying so.** Settings gains one row, *Inference backend*, showing
what was chosen and why in one line ("Hexagon NPU · int8 · QNN 2.42"; "CPU · ONNX Runtime 1.30 ·
TensorRT probe failed: cuDNN 8 required"), with an override to force any lower rung. The about
+86
View File
@@ -0,0 +1,86 @@
# macOS
macOS is out of scope for v1 ([requirements.md](requirements.md)), and nobody working on
DarkRoom has a Mac. This page records what exists anyway, and how a macOS build is set up so
that someone who does have one can send back enough to fix what they hit.
## 1. What exists
- **Inference** ([inference.md §2](inference.md)). The macOS ladder is CoreML, then ONNX
Runtime's CPU provider, then tract. CoreML is unmeasured. The probe decides whether it is used,
and the crash guard (§4, `attempt`) covers the case where the provider takes the process down.
The device fingerprint is the chip (`machdep.cpu.brand_string`) and the OS release, because
CoreML ships with the OS.
- **Where files go** ([`dr_plat::dirs`](../../platform/dr-plat/src/dirs.rs)). The Unix rules,
except the state directory (the log and crash records), which is `~/Library/Logs/darkroom`.
- **A diagnostic build**, described in §3.
The rest is not built, packaged or run on macOS by anyone here. This covers the window,
Metal through wgpu, the display profile (FR-DSP-8 asks X11 and Wayland), the keyring, the
bundle, and signing. `dr-plat` sends every non-Android Unix to the X11/Wayland dependencies.
## 2. Building
`docker/macos` compiles and links for macOS from Linux, using zig as the linker
(`cargo-zigbuild`). Zig carries libSystem's stubs and the C headers, so tract's SIMD kernels
compile and anything that needs only libSystem links:
./docker/macos/build.sh cargo zigbuild --target aarch64-apple-darwin -p dr-inference-engine --features native --all-targets
./docker/macos/build.sh cargo-zigbuild clippy --target aarch64-apple-darwin -p dr-inference-engine --features native --all-targets -- -D warnings
That produces Mach-O arm64 test binaries and the `ladder` and `ep_probe` examples. Nothing runs
them. Anything that links an Apple framework needs the Xcode SDK, which zig does not carry. That
includes `dr-plat` (through the keyring's Security and CoreFoundation) and so the desktop app, and
its link fails with `unable to find framework`. `cargo check` for those still works in the
container.
Linking the app needs Apple's SDK, which means a Mac. On one:
cargo build --profile diagnostic -p darkroom-desktop
./tools/fetch-desktop-runtime.sh # ONNX Runtime 1.29.0 with CoreML, Apple silicon only
The fetch script puts `libonnxruntime.dylib` in the user's `runtime/` directory, next to the
models. The app also looks in `Contents/Frameworks` of its own bundle, and in Homebrew's
`/opt/homebrew/lib` and `/usr/local/lib`. Homebrew's build may not include CoreML; the probe
reports that as a failed rung and uses the CPU.
**For whoever packages it.** A notarised app runs with the hardened runtime, whose library
validation refuses to `dlopen` a library signed by another team. A bundled
`Contents/Frameworks/libonnxruntime.dylib` must be signed with the app. A runtime the user
fetched needs the `com.apple.security.cs.disable-library-validation` entitlement, or it will not
load, and the app will be the tract build without saying why beyond one log line.
## 3. The diagnostic build
Every macOS build is in the hands of someone who can send a log but cannot attach a debugger,
so it is set up to log like a debug build while running at release speed.
- **The log says more.** With no `RUST_LOG`, the desktop's default filter is `debug` for every
`dr_*` crate, for `darkroom_desktop`, and for `onnxruntime`. That last one is ONNX Runtime's
own session log, which the engine forwards into `log` on every platform (`session.rs`,
`with_runtime_log`). At `debug` it includes how many nodes each provider took. At `trace`
(`RUST_LOG=onnxruntime=trace`) it lists every node's placement, which is long. The log cap is
the same as everywhere (two files of 4 MiB).
- **Backtraces have line numbers.** `--profile diagnostic` is release plus line tables. On
macOS the tables go into a `.dSYM` beside the executable, and the backtrace in a crash record
finds them only if the `.dSYM` stays next to the binary. Keep it in the bundle.
## 4. What to ask a Mac user for
`~/Library/Logs/darkroom/darkroom.log`, plus `darkroom.log.1` if present, after the first launch
and after the first scan with faces. Console.app lists it under *Log Reports*. The Settings
diagnostics bundle collects the same files. The lines that answer the open questions are:
| Line | What it tells us |
|---|---|
| `inference: ONNX Runtime … from …` / `inference: runtime tract` | Whether a runtime was found, and which one |
| `inference: floor … ms on the CPU provider` | The CPU number for §2's table |
| `inference: CoreML session built in … s` | CoreML's first compile of the probe model |
| `inference: CoreML rejected: …` / `failed: …` | Why the CPU was kept |
| `onnxruntime` lines naming `CoreMLExecutionProvider::GetCapability` | How much of the graph CoreML took |
| `inference: the app died during …` | The crash guard fired, and on what |
| `inference: compiling … for CoreML` / `ready on CoreML in … s` | Each model's compile, and any that CoreML refused |
Also ask for the settings row (*Settings › About › Inference*), which is one line and says the
same in short. When one of these logs comes back with CoreML numbers, they go into
[inference.md §1–2](inference.md), and footnote ⁵ becomes a measurement.
+1 -1
View File
@@ -145,7 +145,7 @@ _None._
| FR-PLAT-LIN-2 | [`platform/dr-plat/src/display.rs:1`](../../platform/dr-plat/src/display.rs#L1), [`platform/dr-plat/src/display/wayland.rs:1`](../../platform/dr-plat/src/display/wayland.rs#L1), [`platform/dr-plat/src/display/x11.rs:1`](../../platform/dr-plat/src/display/x11.rs#L1) |
| FR-PLAT-WIN-1 | [`platform/dr-plat/src/dirs.rs:1`](../../platform/dr-plat/src/dirs.rs#L1) |
| FR-PLAT-WIN-2 | [`apps/darkroom-desktop/build.rs:1`](../../apps/darkroom-desktop/build.rs#L1), [`apps/darkroom-desktop/src/main.rs:6`](../../apps/darkroom-desktop/src/main.rs#L6), [`ui/dr-ui/src/launch_ui.rs:937`](../../ui/dr-ui/src/launch_ui.rs#L937) |
| FR-PLAT-WIN-3 | [`apps/darkroom-desktop/src/main.rs:19`](../../apps/darkroom-desktop/src/main.rs#L19) |
| FR-PLAT-WIN-3 | [`apps/darkroom-desktop/src/main.rs:35`](../../apps/darkroom-desktop/src/main.rs#L35) |
| FR-RAW-1 | [`core/dr-decode/src/lib.rs:250`](../../core/dr-decode/src/lib.rs#L250), [`core/dr-types/src/lib.rs:132`](../../core/dr-types/src/lib.rs#L132), [`core/dr-types/src/lib.rs:203`](../../core/dr-types/src/lib.rs#L203) |
| FR-RAW-2 | [`core/dr-decode/src/decoder.rs:1`](../../core/dr-decode/src/decoder.rs#L1), [`core/dr-decode/src/decoder.rs:26`](../../core/dr-decode/src/decoder.rs#L26), [`core/dr-decode/src/decoder.rs:59`](../../core/dr-decode/src/decoder.rs#L59), [`core/dr-decode/src/decoder.rs:95`](../../core/dr-decode/src/decoder.rs#L95), [`ui/dr-ui/src/decoder_seam.rs:183`](../../ui/dr-ui/src/decoder_seam.rs#L183), [`ui/dr-ui/src/decoder_seam.rs:1`](../../ui/dr-ui/src/decoder_seam.rs#L1), [`ui/dr-ui/src/decoder_seam.rs:214`](../../ui/dr-ui/src/decoder_seam.rs#L214), [`ui/dr-ui/src/decoder_seam.rs:255`](../../ui/dr-ui/src/decoder_seam.rs#L255), [`ui/dr-ui/src/export.rs:1096`](../../ui/dr-ui/src/export.rs#L1096), [`ui/dr-ui/src/export.rs:850`](../../ui/dr-ui/src/export.rs#L850), [`ui/dr-ui/src/import.rs:96`](../../ui/dr-ui/src/import.rs#L96), [`ui/dr-ui/src/library/sweep.rs:301`](../../ui/dr-ui/src/library/sweep.rs#L301), [`ui/dr-ui/src/library/sweep.rs:624`](../../ui/dr-ui/src/library/sweep.rs#L624), [`ui/dr-ui/src/library/thumbnails_gen.rs:101`](../../ui/dr-ui/src/library/thumbnails_gen.rs#L101), [`ui/dr-ui/src/merge.rs:113`](../../ui/dr-ui/src/merge.rs#L113), [`ui/dr-ui/src/repairs.rs:243`](../../ui/dr-ui/src/repairs.rs#L243) |
| FR-RAW-3 | [`core/dr-decode/src/lib.rs:146`](../../core/dr-decode/src/lib.rs#L146), [`core/dr-decode/src/lib.rs:537`](../../core/dr-decode/src/lib.rs#L537), [`core/dr-decode/src/locate.rs:1435`](../../core/dr-decode/src/locate.rs#L1435), [`core/dr-gpu/src/demosaic.rs:1182`](../../core/dr-gpu/src/demosaic.rs#L1182), [`core/dr-gpu/src/demosaic.rs:749`](../../core/dr-gpu/src/demosaic.rs#L749), [`core/dr-gpu/tests/hot_pixels.rs:1`](../../core/dr-gpu/tests/hot_pixels.rs#L1) |
+2 -1
View File
@@ -98,7 +98,8 @@ pub fn set_state_dir(dir: PathBuf) {
/// Where this application keeps state that is neither configuration nor cache.
///
/// `$XDG_STATE_HOME/darkroom`, falling back to `~/.local/state/darkroom`.
/// `$XDG_STATE_HOME/darkroom`, falling back to `~/.local/state/darkroom` —
/// `~/Library/Logs/darkroom` on macOS (`dirs`).
/// State rather than cache because a crash record must survive the sweep that
/// a cache directory exists to permit, and rather than config because it is
/// not something the user edits.
+2
View File
@@ -46,6 +46,8 @@
//!
//! * Linux: `$XDG_STATE_HOME/darkroom/darkroom.log`, else
//! `~/.local/state/darkroom/darkroom.log`.
//! * macOS: `$XDG_STATE_HOME/darkroom/darkroom.log`, else
//! `~/Library/Logs/darkroom/darkroom.log`, where Console.app lists it.
//! * Android: `/sdcard/Android/data/paris.tourolle.darkroom/files/darkroom.log`,
//! which `adb pull` reads from an ordinary release build. See
//! [`crate::state`] for why not the internal directory, and
+15 -3
View File
@@ -17,6 +17,11 @@
//! | Unix default | `~/.config` | `~/.local/share` | `~/.local/state` |
//! | Windows | `%APPDATA%` | `%LOCALAPPDATA%` | `%LOCALAPPDATA%`, then `state` |
//! | Windows default | `%USERPROFILE%\AppData\Roaming` | `…\AppData\Local` | `…\AppData\Local` |
//! | macOS default | `~/.config` | `~/.local/share` | `~/Library/Logs` |
//!
//! macOS follows the Unix rules except for the one directory a user is asked
//! to find by hand: the log. Finder hides `~/.local`, and `~/Library/Logs`
//! is where Console.app and a Mac user already look (docs/dev/macos.md).
//!
//! then `darkroom` under each. Config roams on Windows and the rest does not,
//! which is the same split XDG makes between config and everything else, and
@@ -100,11 +105,18 @@ fn resolve(kind: Base, env: impl Fn(&str) -> Option<OsString>) -> PathBuf {
}
}
/// Where state goes under `$HOME` when `XDG_STATE_HOME` does not say.
const STATE_UNDER_HOME: &str = if cfg!(target_os = "macos") {
"Library/Logs"
} else {
".local/state"
};
fn xdg_base(kind: Base, env: &impl Fn(&str) -> Option<OsString>) -> Option<PathBuf> {
let (var, under_home) = match kind {
Base::Config => ("XDG_CONFIG_HOME", ".config"),
Base::Data => ("XDG_DATA_HOME", ".local/share"),
Base::State => ("XDG_STATE_HOME", ".local/state"),
Base::State => ("XDG_STATE_HOME", STATE_UNDER_HOME),
};
absolute(env(var)).or_else(|| absolute(env("HOME")).map(|h| h.join(under_home)))
}
@@ -150,7 +162,7 @@ mod tests {
);
assert_eq!(
xdg_base(Base::State, &e),
Some(PathBuf::from("/home/someone/.local/state"))
Some(PathBuf::from("/home/someone").join(STATE_UNDER_HOME))
);
}
@@ -162,7 +174,7 @@ mod tests {
let e = env(&[("XDG_STATE_HOME", "state"), ("HOME", "/home/someone")]);
assert_eq!(
xdg_base(Base::State, &e),
Some(PathBuf::from("/home/someone/.local/state"))
Some(PathBuf::from("/home/someone").join(STATE_UNDER_HOME))
);
assert_eq!(xdg_base(Base::State, &env(&[("HOME", "")])), None);
}
+26
View File
@@ -20,6 +20,32 @@
# directory (docs/inference.md §1.3).
set -euo pipefail
DEST="${1:-${XDG_DATA_HOME:-${HOME}/.local/share}/darkroom/runtime}"
# macOS: Microsoft's release archive, which carries the CoreML provider in the
# one library. Pinned, because the CoreML options the engine sets were read
# from this version's source (docs/dev/macos.md, CLAUDE.md "Providers").
# Apple silicon only: no Intel archive is published since 1.29; an Intel Mac
# takes Homebrew's `onnxruntime` or stays on tract.
if [[ "$(uname -s)" == Darwin ]]; then
ORT_VERSION=1.29.0
[[ "$(uname -m)" == arm64 ]] || {
echo "error: no ONNX Runtime ${ORT_VERSION} archive for $(uname -m); try: brew install onnxruntime" >&2
exit 1
}
NAME="onnxruntime-osx-arm64-${ORT_VERSION}"
WORK="$(mktemp -d)"
trap 'rm -rf "${WORK}"' EXIT
echo "==> downloading ${NAME}"
curl -fsSL "https://github.com/microsoft/onnxruntime/releases/download/v${ORT_VERSION}/${NAME}.tgz" \
| tar xz -C "${WORK}"
mkdir -p "${DEST}"
cp "${WORK}/${NAME}/lib/libonnxruntime.dylib" "${WORK}/${NAME}/LICENSE" "${DEST}/"
echo "==> runtime in ${DEST}:"
ls -1 "${DEST}" | sed 's/^/ /'
echo " (the app finds it on its next launch; Settings › About › Inference says what it chose)"
exit 0
fi
WORK="$(mktemp -d -p /var/tmp fetch-desktop-runtime.XXXXXX)"
trap 'rm -rf "${WORK}"' EXIT