Start the inference engine from both apps and show its choice in Settings

The desktop names where a package may have put libonnxruntime — an
override variable, beside the executable, the package's own library
directory, the Flatpak prefix, the system library directory — and
Android points at the APK's native library directory, which is also
what Qualcomm's DSP loader must be told for the Hexagon skel. Android
starts the engine at the end of the model unpack rather than at launch,
because the probe fingerprints the model files and a first launch has
none until then.

The About panel gains an Inference row beside Graphics, re-read every
two seconds while the probe runs and engines land, and faces.model_id
carries the detector's form: an int8 detector finds a different set of
faces and is a different population (docs/inference.md §7). A
low-memory signal drops every idle session with the GPU caches.

The APK assembly bundles ONNX Runtime and the Qualcomm HTP libraries
from Maven, fetched by tools/fetch-android-runtime.sh with their
published checksums; RUNTIME_DIR=none builds the tract-only APK, which
is a slower app and not a broken one. The desktop packages carry no
runtime yet.

Two probe fixes from the first desktop run: the floor must not be
built with CPU fallback disabled, and a versioned libonnxruntime.so is
a runtime too. On the reference desktop the probe now loads ONNX
Runtime 1.30, measures 30 ms on the CPU provider, and selects TensorRT
at 1.5 ms.
This commit is contained in:
2026-09-19 16:02:37 +02:00
parent d15c41e699
commit 05508741af
23 changed files with 465 additions and 57 deletions
+16 -17
View File
@@ -40,16 +40,17 @@ use crate::align::{
};
use crate::eyes::{Eye, EyeReading};
use crate::landmarks::{Landmarker, Landmarks};
use crate::{install_backend, FaceError, Pixels};
use crate::{FaceError, Pixels};
use dr_inference_engine::{Form, Model, Role};
/// A loaded OCEC graph.
pub struct EyeClassifier {
session: ort::session::Session,
session: Model,
}
/// A loaded SGC graph.
pub struct SunglassesClassifier {
session: ort::session::Session,
session: Model,
}
/// Open a single-input, single-output classifier and check it is the shape
@@ -63,12 +64,10 @@ fn open_classifier(
bytes: &[u8],
expected: &'static str,
(h, w): (usize, usize),
) -> Result<ort::session::Session, FaceError> {
install_backend();
let session = ort::session::Session::builder()
.map_err(FaceError::Inference)?
.commit_from_memory(bytes)
.map_err(FaceError::Inference)?;
) -> Result<Model, FaceError> {
let model = dr_inference_engine::open(Role::EyeClassifier, Form::F32, bytes)?;
let acquired = model.acquire()?;
let session = acquired.lock();
let input = session.inputs().first().ok_or(FaceError::WrongModel {
expected,
@@ -93,7 +92,9 @@ fn open_classifier(
detail: format!("{} outputs, expected one", session.outputs().len()),
});
}
Ok(session)
drop(session);
drop(acquired);
Ok(model)
}
/// Lay a `h × w` RGB crop out as the `[1, 3, h, w]` tensor both graphs take.
@@ -110,11 +111,9 @@ fn to_nchw(pixels: &[f32], h: usize, w: usize) -> Array4<f32> {
}
/// Run a one-number classifier and read its sigmoid back, clamped.
fn run_scalar(
session: &mut ort::session::Session,
input: Array4<f32>,
expected: &'static str,
) -> Result<f32, FaceError> {
fn run_scalar(model: &Model, input: Array4<f32>, expected: &'static str) -> Result<f32, FaceError> {
let acquired = model.acquire()?;
let mut session = acquired.lock();
let outputs = session
.run(ort::inputs![
ort::value::Tensor::from_array(input).map_err(FaceError::Inference)?
@@ -149,7 +148,7 @@ impl EyeClassifier {
/// P(open) for one eye.
pub fn classify(&mut self, eye: &EyePatch) -> Result<f32, FaceError> {
let input = to_nchw(eye.pixels(), EYE_PATCH_HEIGHT, EYE_PATCH_WIDTH);
run_scalar(&mut self.session, input, "OCEC")
run_scalar(&self.session, input, "OCEC")
}
}
@@ -170,7 +169,7 @@ impl SunglassesClassifier {
let mut best = 0.0_f32;
for view in head.views() {
let input = to_nchw(view, SUNGLASSES_EDGE, SUNGLASSES_EDGE);
best = best.max(run_scalar(&mut self.session, input, "SGC")?);
best = best.max(run_scalar(&self.session, input, "SGC")?);
}
Ok(best)
}
+12 -10
View File
@@ -32,7 +32,8 @@
use ndarray::Array4;
use crate::align::crop_box;
use crate::{install_backend, FaceError, Pixels};
use crate::{FaceError, Pixels};
use dr_inference_engine::{Form, Model, Role};
/// The graph's input edge, in pixels.
pub const INPUT_EDGE: usize = 192;
@@ -118,7 +119,7 @@ impl Landmarks {
/// A loaded `2d106det` graph.
pub struct Landmarker {
session: ort::session::Session,
session: Model,
}
impl Landmarker {
@@ -128,11 +129,9 @@ impl Landmarker {
}
pub fn from_bytes(bytes: &[u8]) -> Result<Self, FaceError> {
install_backend();
let session = ort::session::Session::builder()
.map_err(FaceError::Inference)?
.commit_from_memory(bytes)
.map_err(FaceError::Inference)?;
let model = dr_inference_engine::open(Role::Landmarks, Form::F32, bytes)?;
let acquired = model.acquire()?;
let session = acquired.lock();
let input = session.inputs().first().ok_or(FaceError::WrongModel {
expected: "2d106det",
@@ -167,7 +166,9 @@ impl Landmarker {
),
});
}
Ok(Self { session })
drop(session);
drop(acquired);
Ok(Self { session: model })
}
/// The landmarks of the face in `bbox` — `(x0, y0, x1, y1)` in source
@@ -206,8 +207,9 @@ impl Landmarker {
}
}
}
let outputs = self
.session
let acquired = self.session.acquire()?;
let mut session = acquired.lock();
let outputs = session
.run(ort::inputs![
ort::value::Tensor::from_array(input).map_err(FaceError::Inference)?
])
+34 -7
View File
@@ -38,10 +38,16 @@ pub fn runtime() -> Runtime {
}
/// Install a table if none is installed yet — tract, since no directories
/// were named. What a test or an example gets.
/// were named. What a test or an example gets, unless `DARKROOM_ORT_DIR`
/// names a runtime: the same variable the desktop honours, so an example
/// can be pointed at the runtime the app uses without learning `init`.
pub fn ensure_installed() {
if RUNTIME.get().is_none() {
install(&[]);
let dirs: Vec<PathBuf> = std::env::var_os("DARKROOM_ORT_DIR")
.map(PathBuf::from)
.into_iter()
.collect();
install(&dirs);
}
}
@@ -93,11 +99,7 @@ fn load_native(dir: &std::path::Path) -> Result<Runtime, String> {
let path = if dir.as_os_str().is_empty() {
PathBuf::from(name)
} else {
let p = dir.join(name);
if !p.is_file() {
return Err("not present".into());
}
p
find_library(dir, name).ok_or("not present")?
};
// SAFETY: the library's initialisers are ONNX Runtime's own; the symbol
@@ -140,3 +142,28 @@ fn load_native(dir: &std::path::Path) -> Result<Runtime, String> {
Ok(Runtime::OnnxRuntime { path, version })
}
}
/// `libonnxruntime.so` in `dir`, or a versioned spelling of it —
/// `libonnxruntime.so.1.30.0` is what the Python wheel ships, and a package
/// that installs only the versioned file is not wrong.
#[cfg(feature = "native")]
fn find_library(dir: &std::path::Path, name: &str) -> Option<PathBuf> {
let exact = dir.join(name);
if exact.is_file() {
return Some(exact);
}
let prefix = format!("{name}.");
let mut versioned: Vec<PathBuf> = std::fs::read_dir(dir)
.ok()?
.filter_map(|e| e.ok())
.map(|e| e.path())
.filter(|p| {
p.is_file()
&& p.file_name()
.and_then(|n| n.to_str())
.is_some_and(|n| n.starts_with(&prefix))
})
.collect();
versioned.sort();
versioned.pop()
}
+26 -15
View File
@@ -10,6 +10,11 @@ use std::path::PathBuf;
use crate::{state, Config, Form, Rung};
enum Source {
File(PathBuf),
Bytes(&'static [u8]),
}
/// 64-bit FNV-1a. A cache key, not a checksum: two model files that collide
/// here would have to also be the same size and the same role, and the cost
/// of that is a rebuilt engine.
@@ -47,34 +52,42 @@ pub fn run() {
// Smallest first, so the detector — the one that runs per image — is
// ready soonest (§6 step 3).
let mut jobs: Vec<(crate::Role, PathBuf, u64)> = cfg
let mut jobs: Vec<(crate::Role, Source, u64)> = cfg
.models
.iter()
.filter(|(role, _)| rung.form(*role) != Form::F32 || rung != Rung::Hexagon)
.filter_map(|(role, path)| {
let (path, form) = crate::resolve_model(*role, path);
(form == rung.form(*role)).then(|| {
let size = std::fs::metadata(&path).map(|m| m.len()).unwrap_or(0);
(*role, path, size)
(*role, Source::File(path), size)
})
})
.chain(cfg.embedded.iter().filter_map(|(role, bytes)| {
// An embedded model has no int8 sibling to offer a rung that
// wants one; it runs on that rung's fallback.
(rung.form(*role) == Form::F32).then_some((
*role,
Source::Bytes(bytes),
bytes.len() as u64,
))
}))
.collect();
jobs.sort_by_key(|j| j.2);
state().lock().unwrap().wanted = jobs.len();
for (role, path, _) in jobs {
let Ok(bytes) = std::fs::read(&path) else {
continue;
for (role, source, _) in jobs {
let (bytes, name) = match &source {
Source::File(path) => match std::fs::read(path) {
Ok(b) => (b, path.display().to_string()),
Err(_) => continue,
},
Source::Bytes(b) => (b.to_vec(), format!("embedded {role:?}")),
};
let key = key(rung, &bytes);
if state().lock().unwrap().cache.compiled.contains(&key) {
continue;
}
log::info!(
"inference: compiling {} for {}",
path.display(),
rung.label()
);
log::info!("inference: compiling {name} for {}", rung.label());
let started = std::time::Instant::now();
match crate::session::build(rung, role, &bytes, &cfg, false) {
Ok(session) => {
@@ -83,8 +96,7 @@ pub fn run() {
s.cache.compiled.insert(key);
crate::probe::write_cache(&s.config, &s.cache);
log::info!(
"inference: {} ready on {} in {:.1} s",
path.display(),
"inference: {name} ready on {} in {:.1} s",
rung.label(),
started.elapsed().as_secs_f64()
);
@@ -94,8 +106,7 @@ pub fn run() {
// their engine. A corrected model file changes the hash and
// is retried.
log::warn!(
"inference: {} will not compile for {}: {e}",
path.display(),
"inference: {name} will not compile for {}: {e}",
rung.label()
);
}
+6
View File
@@ -35,6 +35,10 @@ pub enum Role {
Embedder,
Segmenter,
Scene,
/// The dense landmark model behind the eye reading (docs/faces.md §7c).
Landmarks,
/// The eye-state and sunglasses classifiers, a few hundred kilobytes.
EyeClassifier,
}
/// Which numeric form of a model a session was built from.
@@ -114,6 +118,8 @@ pub struct Config {
/// The canonical model files on this device, so engines can be compiled
/// ahead of the first request for them.
pub models: Vec<(Role, PathBuf)>,
/// Models compiled into the binary, for the same reason.
pub embedded: Vec<(Role, &'static [u8])>,
/// The highest rung the user allows; `None` is "the best that works".
pub ceiling: Option<Rung>,
/// ONNX Runtime's intra-op pool; 0 picks from the core count.
+6
View File
@@ -201,6 +201,12 @@ fn fingerprint(runtime: &Runtime, cfg: &Config) -> String {
},
device_identity(),
];
for (role, bytes) in &cfg.embedded {
parts.push(format!(
"{role:?} embedded {:016x}",
crate::engines::hash(bytes)
));
}
for (role, path) in &cfg.models {
let hash = std::fs::read(path)
.map(|b| crate::engines::hash(&b))
+1 -1
View File
@@ -20,7 +20,7 @@ pub fn build(
let mut b = Session::builder()?
.with_optimization_level(GraphOptimizationLevel::Level3)?
.with_intra_threads(threads(cfg))?;
if strict {
if strict && rung != Rung::Cpu {
b = b.with_config_entry("session.disable_cpu_ep_fallback", "1")?;
}
// A Hexagon session loads the compiled context when there is one and
+2
View File
@@ -71,6 +71,8 @@ pub use refine::{
};
#[cfg(feature = "semantic")]
pub use scene::{Category, Scene, SceneModel};
#[cfg(feature = "embedded-model")]
pub use semantic::embedded_model_bytes;
#[cfg(feature = "semantic")]
pub use semantic::{Instance, SemanticModel, SemanticOptions, Tiling};
+7
View File
@@ -208,6 +208,13 @@ const EMBEDDED_MODEL: &[u8] = include_bytes!("../../../models/segment/yolo26n-se
#[cfg(feature = "embedded-model")]
const EMBEDDED_CLASSES: &str = include_str!("../../../models/segment/yolo26n-seg.classes.json");
/// The bytes of the model that ships with this crate, for whoever compiles
/// engines ahead of the first request (docs/inference.md §6).
#[cfg(feature = "embedded-model")]
pub fn embedded_model_bytes() -> &'static [u8] {
EMBEDDED_MODEL
}
impl SemanticModel {
/// Load the model that ships with this crate.
#[cfg(feature = "embedded-model")]
+14
View File
@@ -260,6 +260,20 @@ impl FaceDetector {
}
}
/// The id when the detector runs in its int8 form (docs/inference.md §7).
///
/// A different detector: it finds a different set of faces, so it is a
/// different population of detections. The embedder half is unchanged,
/// because the embedder never runs in int8, and `embedder_of` keeps the
/// two spellings' vectors in one space.
pub fn model_id_int8(self) -> &'static str {
match self {
FaceDetector::Scrfd500m => "scrfd_500m_i8+w600k_mbf",
FaceDetector::Scrfd2_5g => "scrfd_2.5g_i8+w600k_mbf",
FaceDetector::Scrfd10g => "scrfd_10g_i8+w600k_mbf",
}
}
/// The detector that writes under a pipeline id, if it is one of these.
///
/// The inverse of [`Self::model_id`]. `None` for an id from another