Start the inference engine from both apps and show its choice in Settings
The desktop names where a package may have put libonnxruntime — an override variable, beside the executable, the package's own library directory, the Flatpak prefix, the system library directory — and Android points at the APK's native library directory, which is also what Qualcomm's DSP loader must be told for the Hexagon skel. Android starts the engine at the end of the model unpack rather than at launch, because the probe fingerprints the model files and a first launch has none until then. The About panel gains an Inference row beside Graphics, re-read every two seconds while the probe runs and engines land, and faces.model_id carries the detector's form: an int8 detector finds a different set of faces and is a different population (docs/inference.md §7). A low-memory signal drops every idle session with the GPU caches. The APK assembly bundles ONNX Runtime and the Qualcomm HTP libraries from Maven, fetched by tools/fetch-android-runtime.sh with their published checksums; RUNTIME_DIR=none builds the tract-only APK, which is a slower app and not a broken one. The desktop packages carry no runtime yet. Two probe fixes from the first desktop run: the floor must not be built with CPU fallback disabled, and a versioned libonnxruntime.so is a runtime too. On the reference desktop the probe now loads ONNX Runtime 1.30, measures 30 ms on the CPU provider, and selects TensorRT at 1.5 ms.
This commit is contained in:
@@ -10,6 +10,11 @@ use std::path::PathBuf;
|
||||
|
||||
use crate::{state, Config, Form, Rung};
|
||||
|
||||
enum Source {
|
||||
File(PathBuf),
|
||||
Bytes(&'static [u8]),
|
||||
}
|
||||
|
||||
/// 64-bit FNV-1a. A cache key, not a checksum: two model files that collide
|
||||
/// here would have to also be the same size and the same role, and the cost
|
||||
/// of that is a rebuilt engine.
|
||||
@@ -47,34 +52,42 @@ pub fn run() {
|
||||
|
||||
// Smallest first, so the detector — the one that runs per image — is
|
||||
// ready soonest (§6 step 3).
|
||||
let mut jobs: Vec<(crate::Role, PathBuf, u64)> = cfg
|
||||
let mut jobs: Vec<(crate::Role, Source, u64)> = cfg
|
||||
.models
|
||||
.iter()
|
||||
.filter(|(role, _)| rung.form(*role) != Form::F32 || rung != Rung::Hexagon)
|
||||
.filter_map(|(role, path)| {
|
||||
let (path, form) = crate::resolve_model(*role, path);
|
||||
(form == rung.form(*role)).then(|| {
|
||||
let size = std::fs::metadata(&path).map(|m| m.len()).unwrap_or(0);
|
||||
(*role, path, size)
|
||||
(*role, Source::File(path), size)
|
||||
})
|
||||
})
|
||||
.chain(cfg.embedded.iter().filter_map(|(role, bytes)| {
|
||||
// An embedded model has no int8 sibling to offer a rung that
|
||||
// wants one; it runs on that rung's fallback.
|
||||
(rung.form(*role) == Form::F32).then_some((
|
||||
*role,
|
||||
Source::Bytes(bytes),
|
||||
bytes.len() as u64,
|
||||
))
|
||||
}))
|
||||
.collect();
|
||||
jobs.sort_by_key(|j| j.2);
|
||||
state().lock().unwrap().wanted = jobs.len();
|
||||
|
||||
for (role, path, _) in jobs {
|
||||
let Ok(bytes) = std::fs::read(&path) else {
|
||||
continue;
|
||||
for (role, source, _) in jobs {
|
||||
let (bytes, name) = match &source {
|
||||
Source::File(path) => match std::fs::read(path) {
|
||||
Ok(b) => (b, path.display().to_string()),
|
||||
Err(_) => continue,
|
||||
},
|
||||
Source::Bytes(b) => (b.to_vec(), format!("embedded {role:?}")),
|
||||
};
|
||||
let key = key(rung, &bytes);
|
||||
if state().lock().unwrap().cache.compiled.contains(&key) {
|
||||
continue;
|
||||
}
|
||||
log::info!(
|
||||
"inference: compiling {} for {}",
|
||||
path.display(),
|
||||
rung.label()
|
||||
);
|
||||
log::info!("inference: compiling {name} for {}", rung.label());
|
||||
let started = std::time::Instant::now();
|
||||
match crate::session::build(rung, role, &bytes, &cfg, false) {
|
||||
Ok(session) => {
|
||||
@@ -83,8 +96,7 @@ pub fn run() {
|
||||
s.cache.compiled.insert(key);
|
||||
crate::probe::write_cache(&s.config, &s.cache);
|
||||
log::info!(
|
||||
"inference: {} ready on {} in {:.1} s",
|
||||
path.display(),
|
||||
"inference: {name} ready on {} in {:.1} s",
|
||||
rung.label(),
|
||||
started.elapsed().as_secs_f64()
|
||||
);
|
||||
@@ -94,8 +106,7 @@ pub fn run() {
|
||||
// their engine. A corrected model file changes the hash and
|
||||
// is retried.
|
||||
log::warn!(
|
||||
"inference: {} will not compile for {}: {e}",
|
||||
path.display(),
|
||||
"inference: {name} will not compile for {}: {e}",
|
||||
rung.label()
|
||||
);
|
||||
}
|
||||
|
||||
Reference in New Issue
Block a user