Add dr-inference-engine and route every model session through it

One crate names the runtime, the providers and the devices; dr-face and
dr-segment ask it for a session by role. It hands ort an API table once
per process — from a libonnxruntime it dlopens when the app names a
directory holding one, otherwise from tract — so the Rust build stays
free of C on every target and a package can install the runtime as a
file (docs/inference.md §3).

Sessions live in a registry behind a Model handle that holds the bytes,
not the session: every use refreshes a timestamp and a reaper unloads
whatever sat idle past the decay. A scan that runs the detector on each
image never lets it go idle; a click in the develop view lets the
segmenter go after thirty seconds; a handle used after that reloads,
and reloads on a higher rung if a compiled engine has landed meanwhile.

The probe walks the platform's ladder by building strict sessions and
timing them against the CPU provider, caches the choice against a
fingerprint of the runtime, driver, hardware and models, and compiles
engines for the selected rung in the background, smallest model first.
Nothing in this commit turns the native path on: the apps still run on
tract until they call init with a runtime directory.
This commit is contained in:
2026-09-19 16:02:37 +02:00
parent caf21bea64
commit d15c41e699
16 changed files with 1368 additions and 90 deletions
+104
View File
@@ -0,0 +1,104 @@
//! Compiled engines: what a rung builds once per device, and the thread that
//! builds them before anyone asks (docs/inference.md §5, §6).
//!
//! TensorRT keeps its own engine cache keyed by graph hash; QNN writes a
//! context model. Both are opaque to this crate, which tracks only *that* a
//! model compiled — by the hash of its bytes — so [`crate::open`] can tell a
//! request whether to expect the rung or its fallback.
use std::path::PathBuf;
use crate::{state, Config, Form, Rung};
/// 64-bit FNV-1a. A cache key, not a checksum: two model files that collide
/// here would have to also be the same size and the same role, and the cost
/// of that is a rebuilt engine.
pub fn hash(bytes: &[u8]) -> u64 {
let mut h = 0xcbf2_9ce4_8422_2325u64;
for &b in bytes {
h ^= b as u64;
h = h.wrapping_mul(0x0000_0100_0000_01b3);
}
h
}
/// The cache entry for `bytes` compiled on `rung`.
pub fn key(rung: Rung, bytes: &[u8]) -> String {
format!("{}:{:016x}", rung.label(), hash(bytes))
}
/// Where QNN's compiled context for `bytes` lives.
pub fn context_path(cfg: &Config, bytes: &[u8]) -> PathBuf {
cfg.cache_dir
.join("qnn")
.join(format!("{:016x}_ctx.onnx", hash(bytes)))
}
/// After the probe: compile every configured model the selected rung can
/// take, smallest first, recording each as it lands.
pub fn run() {
let (rung, cfg) = {
let s = state().lock().unwrap();
(crate::current_rung(&s), s.config.clone())
};
if !rung.compiles() {
return;
}
// Smallest first, so the detector — the one that runs per image — is
// ready soonest (§6 step 3).
let mut jobs: Vec<(crate::Role, PathBuf, u64)> = cfg
.models
.iter()
.filter(|(role, _)| rung.form(*role) != Form::F32 || rung != Rung::Hexagon)
.filter_map(|(role, path)| {
let (path, form) = crate::resolve_model(*role, path);
(form == rung.form(*role)).then(|| {
let size = std::fs::metadata(&path).map(|m| m.len()).unwrap_or(0);
(*role, path, size)
})
})
.collect();
jobs.sort_by_key(|j| j.2);
state().lock().unwrap().wanted = jobs.len();
for (role, path, _) in jobs {
let Ok(bytes) = std::fs::read(&path) else {
continue;
};
let key = key(rung, &bytes);
if state().lock().unwrap().cache.compiled.contains(&key) {
continue;
}
log::info!(
"inference: compiling {} for {}",
path.display(),
rung.label()
);
let started = std::time::Instant::now();
match crate::session::build(rung, role, &bytes, &cfg, false) {
Ok(session) => {
drop(session);
let mut s = state().lock().unwrap();
s.cache.compiled.insert(key);
crate::probe::write_cache(&s.config, &s.cache);
log::info!(
"inference: {} ready on {} in {:.1} s",
path.display(),
rung.label(),
started.elapsed().as_secs_f64()
);
}
Err(e) => {
// This model stays on the fallback; the others still get
// their engine. A corrected model file changes the hash and
// is retried.
log::warn!(
"inference: {} will not compile for {}: {e}",
path.display(),
rung.label()
);
}
}
}
}