Run each model on the Hexagon in the form measured to hold it

The engine knew f32 and int8, and gave the Hexagon int8 for every role it
served. Measured on the tablet itself (inference.md §1.5), int8 lost
5% of the detector's faces at 40-80 px, moved the landmarks 1.5 px,
emptied the segmenter's scores and cost the denoiser 5-9 dB; fp16 the HTP
refuses outright. `Form` gains A16W8 and A16W16, and `Rung::form` now
names one per role: detectors and landmarks A16W8, the segmenter, scene
model, border filler and denoiser A16W16, XFeat int8. The embedder and
the eye classifiers stay on the CPU.

Each loader resolves its `<stem>.<form>.onnx` sibling; the segmenter and
XFeat, compiled into the binary, embed their quantised forms on Android
only and pick through `choose_embedded`. The probe, the compile step and
the cache fingerprint follow the form instead of assuming int8. Detectors
on the new form write `scrfd_*_a16+w600k_mbf`, and `model_ids` answers
for all three spellings.

On the tablet (ORT 1.29 + QNN 2.42), each shipped file against f32 on the
same inputs, and against the CPU's f32 time:
  SCRFD 500m/2.5g/10g  A16W8   100% of faces in every band   4.2/5.1/9.0 ms vs 17/56/198
  landmarks            A16W8   0.25 px in the 192 crop        0.5 ms vs 2.8
  YOLO26n-seg          A16W16  98.2% found, mask IoU 0.994    12.9 ms vs 90
  scene model          A16W16  98.9% of cells agree           15 ms vs 151
  MI-GAN               A16W16  41 dB from f32 in the fill     87 ms vs 488
  XFeat                int8    pano alignment 0.45 px (f32's own spread 0.41)  6.5 ms vs 58
  denoiser             A16W16  0.00 dB at every ISO            95 ms vs 1510 a tile
Face numbers are over public COCO val2017 photographs, not a library.

The APK carries the siblings (BUNDLED 15 -> 19; the old int8 detectors
removed), about 43 MB more. The Windows installer and its CI count skip
them; the Arch and Flatpak packages list their files and never had them.
The ladder example takes a role per model, which is how the per-role
forms above were seen landing on the NPU from the real probe.
This commit is contained in:
2026-10-04 03:45:46 -04:00
parent 0e6ac09fd5
commit 5a8c3e4c40
39 changed files with 550 additions and 208 deletions
+5 -4
View File
@@ -487,10 +487,11 @@ jobs:
wine "$SETUP" /S 2>/dev/null
INST=$(echo "$HOME"/.wine/drive_c/users/*/AppData/Local/Programs/DarkRoom)
ls "$INST"
# As many files as package.sh stages: everything but the READMEs in
# the directories it copies. A literal here went stale the first
# time a model was added.
WANT=$(find models/face models/scene models/inpaint models/denoise -maxdepth 1 -type f ! -name README.md | wc -l)
# As many files as package.sh stages: everything but the READMEs and
# the Hexagon's quantised siblings in the directories it copies. A
# literal here went stale the first time a model was added.
WANT=$(find models/face models/scene models/inpaint models/denoise -maxdepth 1 -type f ! -name README.md \
! -name '*.int8.onnx' ! -name '*.a16w8.onnx' ! -name '*.a16w16.onnx' | wc -l)
GOT=$(ls "$INST/models" | wc -l)
[ "$GOT" = "$WANT" ] || { echo "FAIL: expected $WANT model files, installed $GOT"; exit 1; }
# The manual, and every picture it shows, counted the same way.
+22 -9
View File
@@ -332,27 +332,38 @@ fn unpack_bundled_models(app: &slint::android::AndroidApp) {
// eyes-open filter has something to read, and a tablet has no other way
// to get them either.
//
// The int8 forms beside the three detectors are what the Hexagon runs
// (docs/dev/inference.md §5); the engine loads the sibling when the probe
// chose that rung and ignores it otherwise.
const BUNDLED: [(&std::ffi::CStr, &str); 15] = [
// The quantised siblings — `.a16w8.onnx`, `.a16w16.onnx` — are what the
// Hexagon runs (docs/dev/inference.md §1.5), each in the narrowest form
// that held that model's accuracy on the tablet; the engine loads the
// sibling when the probe chose that rung and ignores it otherwise. The
// segmenter's and XFeat's forms are compiled into the binary instead,
// beside their f32 graphs.
const BUNDLED: [(&std::ffi::CStr, &str); 19] = [
(c"models/scrfd_500m_640.onnx", "scrfd_500m_640.onnx"),
(
c"models/scrfd_500m_640.int8.onnx",
"scrfd_500m_640.int8.onnx",
c"models/scrfd_500m_640.a16w8.onnx",
"scrfd_500m_640.a16w8.onnx",
),
(c"models/scrfd_2.5g_640.onnx", "scrfd_2.5g_640.onnx"),
(
c"models/scrfd_2.5g_640.int8.onnx",
"scrfd_2.5g_640.int8.onnx",
c"models/scrfd_2.5g_640.a16w8.onnx",
"scrfd_2.5g_640.a16w8.onnx",
),
(c"models/scrfd_10g_640.onnx", "scrfd_10g_640.onnx"),
(c"models/scrfd_10g_640.int8.onnx", "scrfd_10g_640.int8.onnx"),
(
c"models/scrfd_10g_640.a16w8.onnx",
"scrfd_10g_640.a16w8.onnx",
),
(c"models/arcface_mbf_b1.onnx", "arcface_mbf_b1.onnx"),
(c"models/2d106det_b1.onnx", "2d106det_b1.onnx"),
(c"models/2d106det_b1.a16w8.onnx", "2d106det_b1.a16w8.onnx"),
(c"models/ocec_s_b1.onnx", "ocec_s_b1.onnx"),
(c"models/sgc_l_48_b1.onnx", "sgc_l_48_b1.onnx"),
(c"models/yolo26s-sem-ade20k.onnx", "yolo26s-sem-ade20k.onnx"),
(
c"models/yolo26s-sem-ade20k.a16w16.onnx",
"yolo26s-sem-ade20k.a16w16.onnx",
),
(
c"models/yolo26s-sem-ade20k.classes.json",
"yolo26s-sem-ade20k.classes.json",
@@ -360,7 +371,9 @@ fn unpack_bundled_models(app: &slint::android::AndroidApp) {
(c"models/categories.txt", "categories.txt"),
// The panorama border filler (FR-MRG-4); MIT, 28 MB.
(c"models/migan-512.onnx", "migan-512.onnx"),
(c"models/migan-512.a16w16.onnx", "migan-512.a16w16.onnx"),
(c"models/mosaic-1408.onnx", "mosaic-1408.onnx"),
(c"models/mosaic-1408.a16w16.onnx", "mosaic-1408.a16w16.onnx"),
];
let dir = dr_ui::shared_face_models_dir();
+5 -1
View File
@@ -5,7 +5,11 @@
//! returns `rgb`, `1×3×1408×1408` (darkroom-denoise `denoise/export.py`,
//! fixed shape because every model the engine runs is). The engine picks the
//! rung: fp16 on TensorRT and MIGraphX, which measured 0.00 dB from f32; f32
//! on CUDA and the CPU; never the Hexagon, where int8 lost 6–9 dB.
//! on CUDA and the CPU; on the Hexagon the `.a16w16.onnx` sibling, 16-bit
//! activations and weights, 0.00 dB from f32 on the tablet itself where int8
//! lost 5–9 dB (docs/dev/inference.md §1.5). That sibling is the same network
//! with the Bayer packing spelled `SpaceToDepth`, which QNN can hold and the
//! 6-D reshape it replaces it cannot.
use crate::tile::TileNet;
use crate::DenoiseError;
+3 -3
View File
@@ -137,8 +137,8 @@ impl Detection {
/// A loaded SCRFD graph.
pub struct Detector {
session: Model,
/// f32 or int8 — the int8 form finds a different set of faces and is a
/// different detector in `model_id` (docs/dev/inference.md §7).
/// f32 or a quantised form — which finds a different set of faces and is
/// a different detector in `model_id` (docs/dev/inference.md §7).
form: Form,
/// Feature-map count: 3 for strides {8,16,32}, 4 for {8,16,32,64}.
///
@@ -155,7 +155,7 @@ impl Detector {
}
/// Load the canonical f32 file at `path`, or the form the device's
/// backend wants instead — the `.int8.onnx` beside it on a Hexagon —
/// backend wants instead — the `.a16w8.onnx` beside it on a Hexagon —
/// which [`Detector::form`] then reports.
pub fn from_path(path: impl AsRef<std::path::Path>) -> Result<Self, FaceError> {
let (path, form) = dr_inference_engine::resolve_model(Role::Detector, path.as_ref());
+11 -2
View File
@@ -123,13 +123,22 @@ pub struct Landmarker {
}
impl Landmarker {
/// The graph at `path`, or the `.a16w8.onnx` sibling beside it when the
/// device's backend runs that (the Hexagon, inference.md §1.5: 0.25 px
/// from f32 in the 192 crop, where int8 moved the points by 1.5).
pub fn from_path(path: impl AsRef<std::path::Path>) -> Result<Self, FaceError> {
let (path, form) = dr_inference_engine::resolve_model(Role::Landmarks, path.as_ref());
let bytes = std::fs::read(path).map_err(FaceError::ModelRead)?;
Self::from_bytes(&bytes)
Self::from_bytes_in(&bytes, form)
}
pub fn from_bytes(bytes: &[u8]) -> Result<Self, FaceError> {
let model = dr_inference_engine::open(Role::Landmarks, Form::F32, bytes)?;
Self::from_bytes_in(bytes, Form::F32)
}
/// `bytes` in a stated numeric form; the output keeps its meaning.
pub fn from_bytes_in(bytes: &[u8], form: Form) -> Result<Self, FaceError> {
let model = dr_inference_engine::open(Role::Landmarks, form, bytes)?;
let acquired = model.acquire()?;
let session = acquired.lock();
+39 -19
View File
@@ -4,12 +4,14 @@
//!
//! DARKROOM_ORT_DIR=/usr/lib \
//! cargo run --release -p dr-inference-engine --features native,tract \
//! --example ladder -- CACHE_DIR models/face/scrfd_500m_640.onnx [MODEL.onnx ...]
//! --example ladder -- CACHE_DIR models/face/scrfd_500m_640.onnx [ROLE=MODEL.onnx ...]
//!
//! Every model named is a `Detector` for the config's purposes, which is
//! enough to see the rung taken, the engines compiled and a session land
//! on it. Delete `CACHE_DIR` to see the first run again; keep it to see the
//! second.
//! A bare path is a `Detector`; `denoiser=…`, `scene=…`, `inpainter=…`,
//! `landmarks=…` (any `Role`, lower case) says otherwise, so a device can
//! show each role taking its own form (inference.md §1.5). Each is opened
//! through `resolve_model`, as the app opens it, and the line says which
//! form and which rung it landed on. Delete `CACHE_DIR` to see the first
//! run again; keep it to see the second.
use std::path::PathBuf;
use std::time::{Duration, Instant};
@@ -17,7 +19,10 @@ use std::time::{Duration, Instant};
fn main() {
env_logger::Builder::from_env(env_logger::Env::default().default_filter_or("info")).init();
let mut args = std::env::args_os().skip(1).map(PathBuf::from);
let (Some(cache_dir), models) = (args.next(), args.collect::<Vec<_>>()) else {
let (Some(cache_dir), models) = (
args.next(),
args.map(|a| role_and_path(&a)).collect::<Vec<_>>(),
) else {
eprintln!("usage: ladder CACHE_DIR MODEL.onnx [MODEL.onnx ...]");
std::process::exit(2);
};
@@ -34,10 +39,7 @@ fn main() {
dr_inference_engine::init(dr_inference_engine::Config {
runtime_dirs,
cache_dir: cache_dir.clone(),
models: models
.iter()
.map(|p| (dr_inference_engine::Role::Detector, p.clone()))
.collect(),
models: models.clone(),
embedded: Vec::new(),
ceiling: None,
threads: 0,
@@ -80,21 +82,39 @@ fn main() {
std::thread::sleep(Duration::from_millis(500));
}
for path in &models {
let bytes = std::fs::read(path).expect("read model");
for (role, path) in &models {
let (path, form) = dr_inference_engine::resolve_model(*role, path);
let bytes = std::fs::read(&path).expect("read model");
let t = Instant::now();
let model = dr_inference_engine::open(
dr_inference_engine::Role::Detector,
dr_inference_engine::Form::F32,
&bytes,
)
.expect("open model");
let model = dr_inference_engine::open(*role, form, &bytes).expect("open model");
let acquired = model.acquire().expect("acquire session");
println!(
"{} on {} in {:.2} s",
"{role:?}: {} ({form:?}) on {} in {:.2} s",
path.file_name().unwrap().to_string_lossy(),
acquired.rung().label(),
t.elapsed().as_secs_f64()
);
}
}
/// `denoiser=path` → (Denoiser, path); a bare path is a detector.
fn role_and_path(arg: &std::path::Path) -> (dr_inference_engine::Role, PathBuf) {
use dr_inference_engine::Role::*;
let s = arg.to_string_lossy();
let Some((name, path)) = s.split_once('=') else {
return (Detector, arg.to_path_buf());
};
let role = match name {
"detector" => Detector,
"embedder" => Embedder,
"segmenter" => Segmenter,
"scene" => Scene,
"landmarks" => Landmarks,
"eyes" => EyeClassifier,
"keypoints" => Keypoints,
"inpainter" => Inpainter,
"denoiser" => Denoiser,
other => panic!("no role {other:?}"),
};
(role, PathBuf::from(path))
}
+5 -5
View File
@@ -8,7 +8,7 @@
use std::path::PathBuf;
use crate::{state, Config, Form, Rung};
use crate::{state, Config, Rung};
enum Source {
File(PathBuf),
@@ -82,10 +82,10 @@ pub fn run() {
(*role, Source::File(path), size)
})
})
.chain(cfg.embedded.iter().filter_map(|(role, bytes)| {
// An embedded model has no int8 sibling to offer a rung that
// wants one; it runs on that rung's fallback.
(rung.serves(*role) && rung.form(*role) == Form::F32).then_some((
.chain(cfg.embedded.iter().filter_map(|(role, form, bytes)| {
// The embedded form the rung wants, if the build carries it;
// a build without it runs that model on the rung's fallback.
(rung.serves(*role) && rung.form(*role) == *form).then_some((
*role,
Source::Bytes(bytes),
bytes.len() as u64,
+139 -59
View File
@@ -42,25 +42,46 @@ pub enum Role {
/// XFeat, the panorama keypoint detector (docs/dev/panorama.md).
Keypoints,
/// MI-GAN, the panorama border filler (docs/dev/panorama.md §12). Plain
/// convolutions, so any rung serves it; fp16 on TensorRT and int8 on
/// the Hexagon are the point of it.
/// convolutions, so any rung serves it; fp16 on TensorRT and 16-bit
/// activations on the Hexagon (int8 changes the fill, §1.5).
Inpainter,
/// The learned demosaic and denoise on the raw mosaic (docs/dev/denoise.md).
/// fp16 costs it nothing measurable; int8 costs 6–9 dB, because 256
/// levels cannot hold the shadow steps it exists to recover — so the
/// Hexagon does not take it.
/// Hexagon takes it with 16-bit activations and weights (§1.5).
Denoiser,
}
/// Which numeric form of a model a session was built from.
///
/// `Int8` is a different network from `F32` for a detector — it finds a
/// different set of faces — which is why [`form_suffix`] exists and why a
/// The quantised forms are QDQ graphs, per-channel weights, as QNN's HTP
/// takes them (docs/dev/inference.md §1.5): `Int8` is 8-bit activations and
/// weights, `A16W8` 16-bit activations with 8-bit weights, `A16W16` 16-bit
/// both. The Hexagon accepts no float tensor at all, so these are the
/// whole menu; which one a role gets is [`Rung::form`], measured per model.
///
/// A quantised detector is a different network from the f32 one — it finds
/// a different set of faces — which is why [`form_suffix`] exists and why a
/// caller appends it to `model_id`.
#[derive(Clone, Copy, Debug, PartialEq, Eq, Hash, Serialize, Deserialize)]
pub enum Form {
F32,
Int8,
A16W8,
A16W16,
}
impl Form {
/// The infix of the sibling file that holds this form:
/// `scrfd_500m_640.a16w8.onnx` beside `scrfd_500m_640.onnx`.
pub fn file_tag(self) -> Option<&'static str> {
match self {
Form::F32 => None,
Form::Int8 => Some("int8"),
Form::A16W8 => Some("a16w8"),
Form::A16W16 => Some("a16w16"),
}
}
}
/// A rung of the ladder (§2). Ordered: a user override names the highest rung
@@ -81,7 +102,7 @@ pub enum Rung {
/// removed in ONNX Runtime 1.23, so there is no non-compiling AMD rung
/// to fall back to: this one falls back to the CPU.
MiGraphX,
/// Qualcomm's Hexagon NPU through QNN, int8 models only. Android only.
/// Qualcomm's Hexagon NPU through QNN, quantised models only. Android only.
Hexagon,
/// Apple, through CoreML: the Neural Engine, the GPU or the CPU, as
/// CoreML schedules it. macOS only. Compiles an ML Program per model on
@@ -121,23 +142,38 @@ impl Rung {
}
/// The model form this rung wants for a role.
fn form(self, _role: Role) -> Form {
///
/// On the Hexagon, the narrowest form that held each model's accuracy
/// on the tablet itself (§1.5): int8 lost 5% of the detector's faces at
/// 40–80 px, moved the landmarks by 1.5 px and the segmenter's scores
/// to nothing, and the denoiser by 6–9 dB, so those take 16-bit
/// activations; the segmenter, scene model, filler and denoiser also
/// needed 16-bit weights. Only XFeat keeps int8: its panorama alignment
/// moved by no more than f32's own refits do.
pub fn form(self, role: Role) -> Form {
match self {
Rung::Hexagon => Form::Int8,
Rung::Hexagon => match role {
Role::Keypoints => Form::Int8,
Role::Detector | Role::Landmarks => Form::A16W8,
Role::Segmenter | Role::Scene | Role::Inpainter | Role::Denoiser => Form::A16W16,
Role::Embedder | Role::EyeClassifier => Form::F32,
},
_ => Form::F32,
}
}
/// Whether this rung runs `role` at all. The Hexagon takes int8 graphs
/// only, and the embedder is never int8 (§7) — it runs on the CPU
/// beside a detector on the NPU, so its vectors compare across devices.
/// Nor is the denoiser: its int8 form failed the 0.5 dB gate by 6–9 dB
/// (denoise.md §8), so it runs on the CPU there too. CoreML is kept off
/// the embedder for the same reason as the Hexagon: the Neural Engine is
/// fp16, and which unit runs a graph is CoreML's choice.
/// Whether this rung runs `role` at all. The Hexagon takes quantised
/// graphs only, and the embedder is never quantised (§7) — it runs on
/// the CPU beside a detector on the NPU, so its vectors compare across
/// devices; at A16W16 it still missed the 0.999 cosine gate. The eye
/// classifiers stay on the CPU too: a millisecond there, and the two
/// share one role while only one of them held its readings quantised.
/// CoreML is kept off the embedder for the same reason as the Hexagon:
/// the Neural Engine is fp16, and which unit runs a graph is CoreML's
/// choice.
fn serves(self, role: Role) -> bool {
match self {
Rung::Hexagon => !matches!(role, Role::Embedder | Role::Denoiser),
Rung::Hexagon => !matches!(role, Role::Embedder | Role::EyeClassifier),
Rung::CoreMl => role != Role::Embedder,
_ => true,
}
@@ -161,8 +197,10 @@ pub struct Config {
/// The canonical model files on this device, so engines can be compiled
/// ahead of the first request for them.
pub models: Vec<(Role, PathBuf)>,
/// Models compiled into the binary, for the same reason.
pub embedded: Vec<(Role, &'static [u8])>,
/// Models compiled into the binary, for the same reason, each with the
/// form it is. A build that embeds a quantised sibling lists it here
/// beside the f32 graph, and the compile step takes the one the rung wants.
pub embedded: Vec<(Role, Form, &'static [u8])>,
/// The highest rung the user allows; `None` is "the best that works".
pub ceiling: Option<Rung>,
/// ONNX Runtime's intra-op pool; 0 picks from the core count.
@@ -188,10 +226,10 @@ pub struct Status {
}
impl Status {
/// "Hexagon NPU · int8 · ONNX Runtime 1.29" — the settings row's text.
/// "Hexagon NPU · quantised · ONNX Runtime 1.29" — the settings row's text.
pub fn line(&self) -> String {
let form = match self.rung {
Rung::Hexagon => " · int8",
Rung::Hexagon => " · quantised",
Rung::TensorRt | Rung::MiGraphX => " · fp16",
_ => "",
};
@@ -464,26 +502,44 @@ fn current_rung(s: &State) -> Rung {
/// The file to load for `role` under the current selection, and its form.
///
/// A rung that wants int8 gets the `.int8.onnx` sibling of the canonical file
/// if it exists; otherwise the canonical file, on the rung's fallback. A
/// caller adds [`form_suffix`] to the `model_id` it records.
/// A rung that wants a quantised form gets that sibling of the canonical
/// file (`<stem>.a16w8.onnx` and so on, [`Form::file_tag`]) if it exists;
/// otherwise the canonical file, on the rung's fallback. A caller adds
/// [`form_suffix`] to the `model_id` it records.
pub fn resolve_model(role: Role, canonical: &Path) -> (PathBuf, Form) {
let rung = current_rung(&state().lock().unwrap());
if rung.serves(role) && rung.form(role) == Form::Int8 {
let sibling = int8_sibling(canonical);
let want = rung.form(role);
if rung.serves(role) && want != Form::F32 {
let sibling = form_sibling(canonical, want);
if sibling.is_file() {
return (sibling, Form::Int8);
return (sibling, want);
}
}
(canonical.to_path_buf(), Form::F32)
}
fn int8_sibling(canonical: &Path) -> PathBuf {
/// The same choice for a model compiled into the binary: of the forms
/// `offered`, the one the current rung wants for `role`, else the f32 one.
/// `offered` must hold an `F32` entry.
pub fn choose_embedded(role: Role, offered: &[(Form, &'static [u8])]) -> (&'static [u8], Form) {
let rung = current_rung(&state().lock().unwrap());
let want = rung.form(role);
let pick = |form| offered.iter().find(|(f, _)| *f == form);
let (form, bytes) = (rung.serves(role).then(|| pick(want)).flatten())
.or_else(|| pick(Form::F32))
.expect("an embedded model offers its f32 form");
(bytes, *form)
}
fn form_sibling(canonical: &Path, form: Form) -> PathBuf {
let stem = canonical
.file_stem()
.map(|s| s.to_string_lossy().into_owned())
.unwrap_or_default();
canonical.with_file_name(format!("{stem}.int8.onnx"))
match form.file_tag() {
Some(tag) => canonical.with_file_name(format!("{stem}.{tag}.onnx")),
None => canonical.to_path_buf(),
}
}
/// What a form appends to a detector's `model_id` (§7).
@@ -491,6 +547,8 @@ pub fn form_suffix(form: Form) -> &'static str {
match form {
Form::F32 => "",
Form::Int8 => "_i8",
Form::A16W8 => "_a16",
Form::A16W16 => "_a16w16",
}
}
@@ -520,8 +578,8 @@ pub fn open(role: Role, form: Form, bytes: &[u8]) -> Result<Model, Error> {
fn effective_rung(s: &State, selected: Rung, role: Role, form: Form, hash: u64) -> Rung {
let mut rung = selected;
if !rung.serves(role) || rung.form(role) != form {
// The embedder on a Hexagon device, or an f32 detector where the int8
// sibling was missing: neither can go to the NPU.
// The embedder on a Hexagon device, or an f32 detector where the
// quantised sibling was missing: neither can go to the NPU.
rung = rung.fallback();
}
if rung.compiles() && !s.cache.compiled.contains(&engines::key_of(rung, hash)) {
@@ -612,9 +670,10 @@ mod tests {
#[test]
fn the_hexagon_never_takes_the_embedder() {
assert!(!Rung::Hexagon.serves(Role::Embedder));
assert!(!Rung::Hexagon.serves(Role::Denoiser));
assert!(!Rung::Hexagon.serves(Role::EyeClassifier));
assert!(Rung::Hexagon.serves(Role::Detector));
assert_eq!(Rung::Hexagon.form(Role::Detector), Form::Int8);
assert!(Rung::Hexagon.serves(Role::Denoiser));
assert_eq!(Rung::Hexagon.form(Role::Detector), Form::A16W8);
// A detector offered in f32 on a Hexagon device lands on the CPU.
let s = State {
config: Config::default(),
@@ -625,36 +684,57 @@ mod tests {
probing: false,
wanted: 0,
};
let on = |role, form| effective_rung(&s, Rung::Hexagon, role, form, engines::hash(b""));
assert_eq!(on(Role::Embedder, Form::F32), Rung::Cpu);
assert_eq!(on(Role::Detector, Form::F32), Rung::Cpu);
// A form other than the one the role wants is not the NPU's either:
// an int8 detector left over from an older install stays off it.
assert_eq!(on(Role::Detector, Form::Int8), Rung::Cpu);
// The wanted form whose context is not compiled yet: also the CPU.
assert_eq!(on(Role::Detector, Form::A16W8), Rung::Cpu);
}
/// The form each role gets on the Hexagon is the one measured to hold
/// its accuracy there (§1.5); a change to this table is a change to
/// what the tablet computes, and must come with a measurement.
#[test]
fn each_role_has_its_measured_form_on_the_hexagon() {
use Form::*;
for (role, form) in [
(Role::Detector, A16W8),
(Role::Landmarks, A16W8),
(Role::Segmenter, A16W16),
(Role::Scene, A16W16),
(Role::Inpainter, A16W16),
(Role::Denoiser, A16W16),
(Role::Keypoints, Int8),
(Role::Embedder, F32),
(Role::EyeClassifier, F32),
] {
assert_eq!(Rung::Hexagon.form(role), form, "{role:?}");
}
for rung in [
Rung::Cpu,
Rung::Cuda,
Rung::TensorRt,
Rung::MiGraphX,
Rung::CoreMl,
] {
assert_eq!(rung.form(Role::Detector), F32);
}
}
#[test]
fn a_form_lives_in_its_tagged_sibling() {
let canonical = Path::new("/m/scrfd_500m_640.onnx");
assert_eq!(form_sibling(canonical, Form::F32), canonical);
assert_eq!(
effective_rung(
&s,
Rung::Hexagon,
Role::Embedder,
Form::F32,
engines::hash(b"")
),
Rung::Cpu
form_sibling(canonical, Form::A16W8),
Path::new("/m/scrfd_500m_640.a16w8.onnx")
);
assert_eq!(
effective_rung(
&s,
Rung::Hexagon,
Role::Detector,
Form::F32,
engines::hash(b"")
),
Rung::Cpu
);
// An int8 detector whose context is not compiled yet: also the CPU.
assert_eq!(
effective_rung(
&s,
Rung::Hexagon,
Role::Detector,
Form::Int8,
engines::hash(b"")
),
Rung::Cpu
form_sibling(canonical, Form::Int8),
Path::new("/m/scrfd_500m_640.int8.onnx")
);
}
+20 -19
View File
@@ -174,10 +174,9 @@ pub fn attempt<T>(cfg: &Config, what: &str, f: impl FnOnce() -> T) -> Result<T,
}
/// The smallest detector, or the smallest model of any role if there is
/// none. A ~2 MB detector is the cheapest real test of a provider, and the
/// detector is the role the int8 forms exist for — the eye classifiers are
/// smaller still, and a Hexagon probed with one would fail for want of a
/// form nobody ships.
/// none. A ~2 MB detector is the cheapest real test of a provider, and
/// every rung serves it — the eye classifiers are smaller still, but the
/// Hexagon does not take them, and a probe with one would fail it for that.
fn probe_model(cfg: &Config) -> Option<(Role, PathBuf)> {
let smallest = |want: Option<Role>| {
cfg.models
@@ -203,16 +202,14 @@ fn time_rung(
cfg: &Config,
) -> Result<(f64, Option<String>), String> {
let want = rung.form(role);
let path = match want {
Form::Int8 => {
let p = crate::int8_sibling(canonical);
if !p.is_file() {
return Err(format!("no int8 form of {}", canonical.display()));
}
p
}
Form::F32 => canonical.to_path_buf(),
};
let path = crate::form_sibling(canonical, want);
if want != Form::F32 && !path.is_file() {
return Err(format!(
"no {} form of {}",
want.file_tag().unwrap_or("f32"),
canonical.display()
));
}
let bytes = std::fs::read(&path).map_err(|e| e.to_string())?;
let started = Instant::now();
let mut session =
@@ -295,9 +292,9 @@ fn fingerprint(runtime: &Runtime, cfg: &Config) -> String {
},
device_identity(),
];
for (role, bytes) in &cfg.embedded {
for (role, form, bytes) in &cfg.embedded {
parts.push(format!(
"{role:?} embedded {:016x}",
"{role:?} embedded {form:?} {:016x}",
crate::engines::hash(bytes)
));
}
@@ -306,9 +303,13 @@ fn fingerprint(runtime: &Runtime, cfg: &Config) -> String {
.map(|b| crate::engines::hash(&b))
.unwrap_or(0);
parts.push(format!("{role:?} {hash:016x}"));
let int8 = crate::int8_sibling(path);
if let Ok(b) = std::fs::read(&int8) {
parts.push(format!("{role:?} int8 {:016x}", crate::engines::hash(&b)));
for form in [Form::Int8, Form::A16W8, Form::A16W16] {
if let Ok(b) = std::fs::read(crate::form_sibling(path, form)) {
parts.push(format!(
"{role:?} {form:?} {:016x}",
crate::engines::hash(&b)
));
}
}
}
parts.join("\n")
+12 -1
View File
@@ -13,8 +13,13 @@ const MODELS: &[&str] = &[
"../../models/keypoints/xfeat-768.onnx",
];
const QUANTISED: &[&str] = &[
"../../models/keypoints/xfeat-1024.int8.onnx",
"../../models/keypoints/xfeat-768.int8.onnx",
];
fn main() {
for m in MODELS {
for m in MODELS.iter().chain(QUANTISED) {
println!("cargo:rerun-if-changed={m}");
}
println!("cargo:rerun-if-changed=build.rs");
@@ -26,6 +31,12 @@ fn main() {
for model in MODELS.iter().copied() {
check(model);
}
// The Hexagon's int8 forms ride only in an Android build.
if std::env::var("CARGO_CFG_TARGET_OS").as_deref() == Ok("android") {
for model in QUANTISED.iter().copied() {
check(model);
}
}
}
fn check(model: &str) {
+2 -1
View File
@@ -30,7 +30,8 @@ pub struct MiGan {
impl MiGan {
/// From the model file, in whichever form the engine's rung wants
/// (`resolve_model` picks an int8 sibling for the Hexagon).
/// (`resolve_model` picks the `.a16w16.onnx` sibling on the Hexagon:
/// int8 moved the fill 16 dB from f32's, 16-bit about 41).
pub fn from_path(path: &std::path::Path) -> Result<Self, PanoError> {
use dr_inference_engine::{resolve_model, Role};
let (path, form) = resolve_model(Role::Inpainter, path);
+37 -6
View File
@@ -36,18 +36,49 @@ pub struct XFeat {
pub options: DecodeOptions,
}
/// The bytes of both exports compiled into the binary, for whoever compiles
/// engines ahead of the first request (docs/dev/inference.md §6).
/// The Hexagon's forms (docs/dev/inference.md §1.5): int8, from the same
/// network spelled for the HTP (the unfold as SpaceToDepth, the bilinear
/// resizes as matrix products). Only Android has a Hexagon.
#[cfg(all(feature = "embedded-model", target_os = "android"))]
const EMBEDDED_LANDSCAPE_INT8: &[u8] =
include_bytes!("../../../models/keypoints/xfeat-1024.int8.onnx");
#[cfg(all(feature = "embedded-model", target_os = "android"))]
const EMBEDDED_PORTRAIT_INT8: &[u8] =
include_bytes!("../../../models/keypoints/xfeat-768.int8.onnx");
/// Every form of both exports compiled into the binary, landscape then
/// portrait, for whoever compiles engines ahead of the first request
/// (docs/dev/inference.md §6).
#[cfg(feature = "embedded-model")]
pub fn embedded_model_bytes() -> [&'static [u8]; 2] {
[EMBEDDED_LANDSCAPE, EMBEDDED_PORTRAIT]
pub fn embedded_models() -> [Vec<(dr_inference_engine::Form, &'static [u8])>; 2] {
use dr_inference_engine::Form;
#[allow(unused_mut)]
let mut forms = [
vec![(Form::F32, EMBEDDED_LANDSCAPE)],
vec![(Form::F32, EMBEDDED_PORTRAIT)],
];
#[cfg(target_os = "android")]
{
forms[0].push((Form::Int8, EMBEDDED_LANDSCAPE_INT8));
forms[1].push((Form::Int8, EMBEDDED_PORTRAIT_INT8));
}
forms
}
impl XFeat {
/// The weights compiled into the binary.
/// The weights compiled into the binary, in the form the device's
/// backend runs.
#[cfg(feature = "embedded-model")]
pub fn embedded() -> Result<Self, PanoError> {
Self::from_bytes(EMBEDDED_LANDSCAPE, EMBEDDED_PORTRAIT)
use dr_inference_engine::{choose_embedded, open, Role};
let [l, p] = embedded_models();
let (l, lf) = choose_embedded(Role::Keypoints, &l);
let (p, pf) = choose_embedded(Role::Keypoints, &p);
Ok(XFeat {
landscape: open(Role::Keypoints, lf, l)?,
portrait: open(Role::Keypoints, pf, p)?,
options: DecodeOptions::default(),
})
}
/// From the two exports on disk.
+14 -4
View File
@@ -13,9 +13,11 @@
use std::path::Path;
const MODEL: &str = "../../models/segment/yolo26n-seg.onnx";
const QUANTISED: &str = "../../models/segment/yolo26n-seg.a16w16.onnx";
fn main() {
println!("cargo:rerun-if-changed={MODEL}");
println!("cargo:rerun-if-changed={QUANTISED}");
println!("cargo:rerun-if-changed=build.rs");
// Only the embedded path needs the file present; a build without it is
@@ -24,10 +26,18 @@ fn main() {
return;
}
let path = Path::new(MODEL);
check(MODEL);
// The Hexagon's quantised form rides only in an Android build.
if std::env::var("CARGO_CFG_TARGET_OS").as_deref() == Ok("android") {
check(QUANTISED);
}
}
fn check(model: &str) {
let path = Path::new(model);
let Ok(bytes) = std::fs::read(path) else {
panic!(
"\n\n{MODEL} is missing.\n\
"\n\n{model} is missing.\n\
It ships in Git LFS. Run `git lfs install && git lfs pull`, or build \
with `--no-default-features` for a watershed-only build.\n"
);
@@ -40,7 +50,7 @@ fn main() {
// happens in practice.
if bytes.starts_with(b"version https://git-lfs") {
panic!(
"\n\n{MODEL} is a Git LFS pointer, not the model ({} bytes).\n\
"\n\n{model} is a Git LFS pointer, not the model ({} bytes).\n\
Run `git lfs install && git lfs pull` to fetch the real file.\n",
bytes.len()
);
@@ -50,7 +60,7 @@ fn main() {
// export is ~11 MB; anything under a megabyte is a truncated checkout.
if bytes.len() < 1_000_000 {
panic!(
"\n\n{MODEL} is only {} bytes — expected ~11 MB.\n\
"\n\n{model} is only {} bytes — expected several MB.\n\
The checkout looks incomplete; try `git lfs pull`.\n",
bytes.len()
);
+1 -1
View File
@@ -72,7 +72,7 @@ pub use refine::{
#[cfg(feature = "semantic")]
pub use scene::{Category, Scene, SceneModel};
#[cfg(feature = "embedded-model")]
pub use semantic::embedded_model_bytes;
pub use semantic::embedded_models;
#[cfg(feature = "semantic")]
pub use semantic::{Instance, SemanticModel, SemanticOptions, Tiling};
+15 -7
View File
@@ -123,21 +123,29 @@ impl SceneModel {
classes: impl AsRef<std::path::Path>,
categories: impl AsRef<std::path::Path>,
) -> Result<Self, SegmentError> {
// The form the device's backend runs: the `.a16w16.onnx` sibling on
// the Hexagon (attention left in float, inference.md §1.5), else this.
let (model, form) =
dr_inference_engine::resolve_model(dr_inference_engine::Role::Scene, model.as_ref());
let bytes = std::fs::read(model).map_err(SegmentError::ModelRead)?;
let classes = std::fs::read_to_string(classes).map_err(SegmentError::ModelRead)?;
let categories = std::fs::read_to_string(categories).map_err(SegmentError::ModelRead)?;
let classes = crate::semantic::parse_classes(&classes);
let categories = parse_categories(&categories, &classes)?;
Self::from_bytes(&bytes, categories)
Self::from_bytes_in(&bytes, form, categories)
}
pub fn from_bytes(bytes: &[u8], categories: Vec<Category>) -> Result<Self, SegmentError> {
// f32, as for `SemanticModel`; see there.
let session = dr_inference_engine::open(
dr_inference_engine::Role::Scene,
dr_inference_engine::Form::F32,
bytes,
)?;
Self::from_bytes_in(bytes, dr_inference_engine::Form::F32, categories)
}
/// `bytes` in a stated numeric form; the outputs keep their shape.
pub fn from_bytes_in(
bytes: &[u8],
form: dr_inference_engine::Form,
categories: Vec<Category>,
) -> Result<Self, SegmentError> {
let session = dr_inference_engine::open(dr_inference_engine::Role::Scene, form, bytes)?;
Ok(Self {
session,
+31 -13
View File
@@ -208,18 +208,32 @@ const EMBEDDED_MODEL: &[u8] = include_bytes!("../../../models/segment/yolo26n-se
#[cfg(feature = "embedded-model")]
const EMBEDDED_CLASSES: &str = include_str!("../../../models/segment/yolo26n-seg.classes.json");
/// The bytes of the model that ships with this crate, for whoever compiles
/// The Hexagon's form (docs/dev/inference.md §1.5): 16-bit activations and
/// weights, the rows' tail left in float. Only Android has a Hexagon, so only
/// Android carries it.
#[cfg(all(feature = "embedded-model", target_os = "android"))]
const EMBEDDED_A16W16: &[u8] = include_bytes!("../../../models/segment/yolo26n-seg.a16w16.onnx");
/// Every form of the model that ships with this crate, for whoever compiles
/// engines ahead of the first request (docs/dev/inference.md §6).
#[cfg(feature = "embedded-model")]
pub fn embedded_model_bytes() -> &'static [u8] {
EMBEDDED_MODEL
pub fn embedded_models() -> Vec<(dr_inference_engine::Form, &'static [u8])> {
#[allow(unused_mut)]
let mut forms = vec![(dr_inference_engine::Form::F32, EMBEDDED_MODEL)];
#[cfg(target_os = "android")]
forms.push((dr_inference_engine::Form::A16W16, EMBEDDED_A16W16));
forms
}
impl SemanticModel {
/// Load the model that ships with this crate.
/// Load the model that ships with this crate, in the form the device's
/// backend runs.
#[cfg(feature = "embedded-model")]
pub fn embedded() -> Result<Self, SegmentError> {
Self::from_bytes(EMBEDDED_MODEL, parse_classes(EMBEDDED_CLASSES))
let forms = embedded_models();
let (bytes, form) =
dr_inference_engine::choose_embedded(dr_inference_engine::Role::Segmenter, &forms);
Self::from_bytes_in(bytes, form, parse_classes(EMBEDDED_CLASSES))
}
/// Load a model from an ONNX file, with `classes` supplying its vocabulary.
@@ -236,14 +250,18 @@ impl SemanticModel {
}
pub fn from_bytes(bytes: &[u8], classes: Vec<Arc<str>>) -> Result<Self, SegmentError> {
// The f32 graph on whatever the device's backend is. An int8 form
// for the Hexagon waits on docs/dev/inference.md §10 M7 — the mask
// boundary has to be measured before it moves.
let session = dr_inference_engine::open(
dr_inference_engine::Role::Segmenter,
dr_inference_engine::Form::F32,
bytes,
)?;
Self::from_bytes_in(bytes, dr_inference_engine::Form::F32, classes)
}
/// `bytes` in a stated numeric form. The quantised one keeps the same
/// outputs (the rows' tail stays float), so decoding does not change; the
/// masks it draws were measured against f32's (inference.md §1.5).
pub fn from_bytes_in(
bytes: &[u8],
form: dr_inference_engine::Form,
classes: Vec<Arc<str>>,
) -> Result<Self, SegmentError> {
let session = dr_inference_engine::open(dr_inference_engine::Role::Segmenter, form, bytes)?;
Ok(Self { session, classes })
}
+28 -9
View File
@@ -275,13 +275,26 @@ impl FaceDetector {
}
}
/// Both ids this detector writes under — the f32 form and the int8 one —
/// for a question that is about the detector and not about which form
/// of it a device happened to run: "has the chosen detector been over
/// this image", asked by a re-index that must not ping-pong between a
/// desktop that runs it in f32 and a tablet that runs it on the Hexagon.
pub fn model_ids(self) -> [&'static str; 2] {
[self.model_id(), self.model_id_int8()]
/// The id when the detector runs with 16-bit activations and 8-bit
/// weights, the Hexagon's form since the int8 one lost faces at 40–80 px
/// (docs/dev/inference.md §1.5). Different again from both, for the same
/// reason as [`Self::model_id_int8`]; the embedder half is unchanged.
pub fn model_id_a16w8(self) -> &'static str {
match self {
FaceDetector::Scrfd500m => "scrfd_500m_a16+w600k_mbf",
FaceDetector::Scrfd2_5g => "scrfd_2.5g_a16+w600k_mbf",
FaceDetector::Scrfd10g => "scrfd_10g_a16+w600k_mbf",
}
}
/// Every id this detector writes under — the f32 form and each quantised
/// one a device has run — for a question that is about the detector and
/// not about which form of it a device happened to run: "has the chosen
/// detector been over this image", asked by a re-index that must not
/// ping-pong between a desktop that runs it in f32 and a tablet that
/// runs it on the Hexagon.
pub fn model_ids(self) -> [&'static str; 3] {
[self.model_id(), self.model_id_int8(), self.model_id_a16w8()]
}
/// The detector that writes under a pipeline id, if it is one of these.
@@ -1769,11 +1782,16 @@ mod tests {
#[test]
fn a_detectors_two_spellings_share_its_embedder_and_nothing_else() {
for d in FaceDetector::ALL {
let [f32_id, int8_id] = d.model_ids();
let [f32_id, int8_id, a16_id] = d.model_ids();
assert_eq!(f32_id, d.model_id());
assert_eq!(int8_id, d.model_id_int8());
assert_eq!(a16_id, d.model_id_a16w8());
assert_ne!(f32_id, int8_id);
assert_eq!(f32_id.rsplit('+').next(), int8_id.rsplit('+').next());
assert_ne!(f32_id, a16_id);
assert_ne!(int8_id, a16_id);
for q in [int8_id, a16_id] {
assert_eq!(f32_id.rsplit('+').next(), q.rsplit('+').next());
}
}
}
@@ -1797,6 +1815,7 @@ mod tests {
}
assert_eq!(FaceDetector::for_model_id(d.model_id()), Some(d));
assert_eq!(FaceDetector::for_model_id(d.model_id_int8()), Some(d));
assert_eq!(FaceDetector::for_model_id(d.model_id_a16w8()), Some(d));
}
assert_eq!(FaceDetector::for_model_id("scrfd_10g+other"), None);
}
+5 -1
View File
@@ -54,10 +54,14 @@ sed 's/$/\r/' "${REPO}/LICENSE" > "${STAGE}/LICENSE"
# with its two descriptors, the panorama border filler and the denoiser. The installer
# smoke test counts the same directories, so a model added here is expected
# there without a number to update.
#
# Not the quantised siblings (`*.int8.onnx`, `*.a16w8.onnx`, `*.a16w16.onnx`):
# they are the Hexagon's forms (docs/dev/inference.md §1.5), and a Windows
# machine has no Hexagon to load them.
for dir in face scene inpaint denoise; do
for f in "${REPO}/models/${dir}"/*; do
case "$(basename "${f}")" in
README.md) continue ;;
README.md | *.int8.onnx | *.a16w8.onnx | *.a16w16.onnx) continue ;;
esac
if [[ "${f}" == *.onnx && "$(stat -c%s "${f}")" -lt 100000 ]]; then
echo "error: $(basename "${f}") is $(stat -c%s "${f}") bytes — an LFS pointer, not a model." >&2
+11
View File
@@ -331,6 +331,17 @@ If neither holds S's quality within 0.5 dB of fp32 on the real pairs, **v1 is de
tablet shows the classical path. The sidecar still records the intent, so a desktop can render the
learned result for a photograph edited on the tablet.
**Measured 2026-10-04 (inference.md §1.5): the second way holds, without the first.** The shipped
network, with its Bayer packing re-spelled as `SpaceToDepth` so QNN can hold it (the 6-D reshape
it replaces is exact but past the HTP's rank limit), at A16W16 — 16-bit activations and weights —
scores within 0.00 dB of f32 at ISO 400–25600 on the tablet's own HTP, and within 0.09 dB with the
6D's noise model scaled ×0.5, ×2 and ×4 to stand in for other sensors. A16W8 holds the 6D (worst
−0.19 dB at ISO 25600) but not ×4 noise at 25600 (−0.52 dB), so A16W16 is what ships. int8 loses
4.7–9.2 dB and fp16 is refused outright. A 1408 tile takes 95 ms on the Hexagon against 1510 ms on
the tablet's CPU: about 2.3 s for a 20 MP frame. Calibration ranges come from 96 training-day
tiles across every ISO, a third of them with that scaled noise; coverage of other bodies is that
synthetic bracket, not their raws.
## 9. X-Trans
The requirements tie this stage to FR-RAW-5, and the library has no Fuji raws. What we can do
+64 -4
View File
@@ -127,7 +127,8 @@ Three things the table settles.
five of six models. A whole-library face index on the tablet goes from ~100 ms + 39 ms per face
to ~1.4 ms + 12 ms per face, and the "Thorough" detector — 3× the cost of "Fast" today — becomes
free. Its price is that the models must be **quantised to int8**, which is an accuracy question
§5 has to answer before it is believed.
§5 has to answer before it is believed. (§1.5 answered it: int8 lost faces, and every model but
XFeat ships with 16-bit activations, at about three times these timings.)
- **The embedder does not gain from either accelerator.** 112×112 input, per-op overhead
dominates; it is 9 ms on the tablet's CPU and 12 ms on its NPU. It stays float, which §7 turns
from a performance footnote into a correctness rule.
@@ -137,6 +138,60 @@ Three things the table settles.
- **On AMD, MIGraphX fp16 is 4–17× the CPU provider** on the detectors and 60× on the
inpainter, with the same first-run compile cost as TensorRT and no rung between it and the CPU.
### 1.5 The Hexagon at every bit width · 2026-10-04
§1.1's Hexagon column is int8 calibrated on noise: timing only. This is the follow-up — every
model, every bit width the HTP offers, calibrated on real photographs and **scored on the tablet
itself** (ORT 1.29 + QNN 2.42, `htp_arch` 73), against the f32 model on the same inputs. The
"Form shipped" column is the files in `models/`, re-scored on the tablet after
`tools/quantise-models.sh` wrote them. The
calibration and scoring photographs are 800 from the public COCO val2017 set (CC-BY); the face
models' numbers are over the faces in them of at least 32 px. The tools are `tools/quantise-models.sh`
and the scratch harness described with it.
**What the HTP accepts.** fp16: nothing — every fp16 operator fails validation (3110), on QNN
2.42 and 2.50, with `htp_arch` and every `soc_model` tried; the fp16 rung stays off the table until
someone has Qualcomm's own SDK to say why. 4-bit weights (A8W4, A16W4): load, and wreck accuracy
(SCRFD finds 25–35% of f32's faces). What is left: **A8W8 (int8), A16W8 and A16W16**, all running
the whole graph. 16-bit activations cost about 3× int8's time, A16W16 about 4×.
| Model | ORT CPU f32 | Form shipped | Hexagon | On the tablet, against f32 | int8 for comparison |
|---|---|---|---|---|---|
| scrfd_500m / 2.5g / 10g | 17 / 56 / 198 ms | **A16W8** | 4.2 / 5.1 / 9.0 ms | 100% of faces found in every size band; keypoints 0.3–0.6% of the box | 94–95% of faces at 40–80 px |
| 2d106det (landmarks) | 2.8 ms | **A16W8** | 0.5 ms | 0.25 px in the 192 crop (eye points 0.20) | 1.5 px, and 29 partitions at 7.3 ms |
| yolo26n-seg | 90 ms | **A16W16**, tail in float | 12.9 ms | 98.2% of objects, mask IoU 0.994 | 74% (simulated) |
| yolo26s-sem-ade20k | 151 ms | **A16W16**, attention in float | 15 ms | 98.9% of cells agree on the class, TV 0.009 | 67% |
| migan-512 | 488 ms | **A16W16** | 87 ms | 41 dB from f32 in the fill (worst 1%: 30 dB) | 16 dB (simulated) |
| xfeat-1024 / 768 | 58 ms | **int8**, rewritten graph | 6.5 ms | panorama alignment 0.45 px from f32's — f32's own refit on 90% of its matches is 0.41 | — |
| mosaic-1408 (denoiser) | 1510 ms a tile | **A16W16**, rewritten graph | 95 ms a tile | 0.00 dB at every ISO; ≤ 0.09 dB with the noise scaled ×0.5–×4 | −4.7 to −9.2 dB |
| arcface_mbf (embedder) | 8.5 ms | f32, CPU | — | A16W16: cosine 0.9995, p1 0.9967 — misses §7's 0.999 gate | — |
| ocec, sgc (eyes) | 1, 1.7 ms | f32, CPU | — | sgc flips 1.45% of views even at A16W16; not worth a millisecond | — |
Four things the table needed that the f32 graphs did not have, all in `tools/htp_graph.py` and all
checked exact against the f32 graph before they are used:
- **Rank ≤ 5.** QNN's tensors stop at rank 5, and the denoiser packs the mosaic through a 6-D
reshape (6007 at compose). For one channel that reshape is `SpaceToDepth(2)`. XFeat's 8×8 unfold
is 224 Slices and 6-D Concats; it is `SpaceToDepth(8)` (736 nodes to 60).
- **No bilinear Resize at XFeat's sizes** (3110). A half-pixel bilinear resize between fixed sizes
is two constant matrices, so it is two MatMuls.
- **One scale per tensor.** The segmenter's output rows carry boxes in pixels beside scores in
0..1; quantised as one tensor the scores vanish. Everything from the Concat that builds the rows
stays float, on the CPU, where the top-300 selection costs nothing.
- **Float where the HTP's 16-bit arithmetic drifts.** The scene model's one attention block
(two MatMuls and a Softmax over 400 tokens) moved its agreement from 98.7% to 96.7%; it stays
float.
**ORT's CPU simulation of a QDQ graph is not the tablet.** It matched to the hundredth of a dB for
the denoiser and to rounding for the detectors, landmarks and XFeat, and it overstated MI-GAN by
27 dB and the scene model by three points. Every number above is the device's; a new form is not
measured until it has run there.
**XFeat's int8 loses keypoints and not the panorama.** 83% of f32's keypoints come back within
1.5 px; but over the twelve-frame `fixtures/pano/2025-08-05` sweep, the homographies fitted from
int8's matches land 0.45 px from f32's in the overlaps — the spread f32 shows against itself
(0.41).
---
## 2. The shape of the answer
@@ -146,7 +201,7 @@ winning:
| Platform | 1st | 2nd | 3rd | Floor |
|---|---|---|---|---|
| Android, Qualcomm with a Hexagon the shipped QNN skel covers (V68–V81) | QNN HTP, int8 model | ORT CPU, f32 model | — | tract |
| Android, Qualcomm with a Hexagon the shipped QNN skel covers (V68–V81) | QNN HTP, each model's quantised form (§1.5) | ORT CPU, f32 model | — | tract |
| Android, any other SoC | ORT CPU, f32 | — | — | tract |
| Linux / Windows, NVIDIA GPU | TensorRT, f32 model, fp16 engine | CUDA provider, f32 | ORT CPU, f32 | tract |
| Linux, AMD GPU with ROCm | MIGraphX, f32 model, fp16 program | ORT CPU, f32 | — | tract |
@@ -298,10 +353,10 @@ of which form they load:
| Form | Who produces it | When | Needed by |
|---|---|---|---|
| f32 ONNX, shape-fixed, **opset ≥ 13** | `tools/fix-face-model-shapes.sh`, `tools/export-seg-model.sh` | Release time, once | Every rung except Hexagon |
| int8 QDQ ONNX, per-channel, uint8 activations | `tools/quantise-models.sh` (new) | Release time, once, **calibrated on real photographs** | Hexagon |
| QDQ ONNX, per-channel — int8, A16W8 or A16W16 per model (§1.5) | `tools/quantise-models.sh` | Release time, once, **calibrated on real photographs**, scored on the tablet | Hexagon |
| TensorRT engine (`.engine`, per GPU architecture and TensorRT version) | The app, from the f32 file | First run on that device, in the background | TensorRT rung |
| MIGraphX program (`.mxr`, per GPU architecture, MIGraphX version and precision) | The app, from the f32 file | First run on that device, in the background | MIGraphX rung |
| QNN context binary | The app, from the int8 file | First run on that device, in the background | Hexagon rung |
| QNN context binary | The app, from the quantised file | First run on that device, in the background | Hexagon rung |
Two rules.
@@ -380,6 +435,11 @@ already the rule for the detector and because §5 is the gate on whether the int
enough to be *offered* at all. f32 on tract, ORT CPU, CUDA and TensorRT-f32 are one identity: the
same graph, the same arithmetic, differences at the last bit.
After §1.5 the Hexagon runs the detectors in **A16W8**, and that is a third spelling:
`scrfd_500m_a16+w600k_mbf` and its two siblings. Same rule, same reconciliation; a tablet that
indexed under `_i8` keeps those rows, and `FaceDetector::model_ids` answers "has this detector been
over this image" for all three forms.
**The embedder** is where comparability across devices is the whole point, and it is the one
model that no accelerator helps (§1.4). So: **the embedder runs in f32 on every rung.** On TensorRT
that means the embedder's engine is built without fp16 while the detector's is built with it; on
+5 -3
View File
@@ -467,9 +467,11 @@ the tablet. Three ways to make it viable, none built:
1. **Fill at a quarter of the resolution and upsample.** Sky and scree
tolerate it; twenty-odd tiles, about three minutes on the desktop CPU. A
background job with the outbox's patience, not an interactive one.
2. **int8 on the tablet's Hexagon through QNN**, where the plain-conv design
is the point and the whole graph should run in milliseconds. The setup
exists from the eye-state work; MI-GAN is a candidate for the same path.
2. **The tablet's Hexagon through QNN**, where the plain-conv design is the
point. Measured 2026-10-04 (inference.md §1.5): int8 changes the fill
(16 dB from f32's), so it ships with 16-bit activations and weights —
87 ms a tile against 488 ms on the tablet's CPU, the whole graph on the
NPU, 41 dB from f32 in the hole.
3. **A WGSL runtime for those six operators.** A project of its own, and
the only route that would make it interactive on the desktop.
File diff suppressed because one or more lines are too long
+2 -2
View File
@@ -287,7 +287,6 @@ $LOCALAPPDATA\Programs\DarkRoom\
models\
scrfd_500m_640.onnx scrfd_2.5g_640.onnx scrfd_10g_640.onnx arcface_mbf_b1.onnx
2d106det_b1.onnx ocec_s_b1.onnx sgc_l_48_b1.onnx
scrfd_500m_640.int8.onnx scrfd_2.5g_640.int8.onnx scrfd_10g_640.int8.onnx
yolo26s-sem-ade20k.onnx yolo26s-sem-ade20k.classes.json categories.txt
migan-512.onnx
manual\
@@ -297,7 +296,8 @@ $LOCALAPPDATA\Programs\DarkRoom\
```
Plus a Start Menu shortcut, and nothing on the desktop unless the user ticks it. The models are
the same ten files the APK bundles and the PKGBUILD installs; `models\` beside the executable is
the f32 files the APK bundles and the PKGBUILD installs — not the APK's quantised siblings, which
only a Hexagon runs; `models\` beside the executable is
where §3.2's lookup finds them. **No `LICENSE` yet**: the repository has no licence file at its
root (the Arch package points at the system's shared GPL text), so the installer has no licence
page until one is added — a one-file change, and the `.nsi` says where the page then goes. The face weights carry the research-only grant that
+7
View File
@@ -14,6 +14,13 @@ Both come from `https://huggingface.co/Ultralytics/YOLO26`. The face weights in
The keypoint weights in `keypoints/` and the border filler in `inpaint/` are
the other two, and the easiest — see the last two sections.
Every quantised sibling — `*.int8.onnx`, `*.a16w8.onnx`, `*.a16w16.onnx`, the
forms the tablet's Hexagon runs (`tools/quantise-models.sh`) — is the same
weights rounded, and carries exactly the grant of the file it was made from.
What calibration adds is one minimum and maximum per tensor: from public COCO
val2017 photographs (CC-BY 4.0) for the image models, and for the denoiser from
the same training tiles its weights were learned from. No image is in the files.
## The grant
**Ultralytics releases YOLO under AGPL-3.0**, and the weights carry the same
Binary file not shown.
Binary file not shown.
+14 -8
View File
@@ -28,15 +28,21 @@ every reader treats "never read" as unknown, never as closed. They are found in
directory as the pair, so a hand-placed pair does not pick up a package's eye models from a
directory it otherwise outranks.
scrfd_500m_640.int8.onnx 0.8 MB the same three, in the form the Hexagon NPU takes
scrfd_2.5g_640.int8.onnx 0.9 MB (docs/dev/inference.md §5) — opset 17, per-channel int8
scrfd_10g_640.int8.onnx 4.3 MB weights, uint8 activations, calibrated on 96 photographs
scrfd_500m_640.a16w8.onnx the same three, in the form the Hexagon NPU takes
scrfd_2.5g_640.a16w8.onnx (docs/dev/inference.md §1.5) — 16-bit activations,
scrfd_10g_640.a16w8.onnx per-channel 8-bit weights, calibrated on 300 photographs
2d106det_b1.a16w8.onnx the landmarks, likewise
The int8 files are **derived** by `tools/quantise-models.sh` from the f32 ones beside them and
travel with them: the engine loads the `.int8.onnx` sibling when the device's backend wants it and
the canonical file otherwise, and a library indexed on the int8 form records it as a different
detector (`scrfd_500m_i8+w600k_mbf`), because it finds a different set of faces. Every other
platform ignores them. The embedder has no int8 form and never will (§7 of the same document).
These are **derived** by `tools/quantise-models.sh` from the f32 files beside them and travel with
them: the engine loads the `.a16w8.onnx` sibling when the device's backend wants it and the
canonical file otherwise, and a library indexed on that form records it as a different detector
(`scrfd_500m_a16+w600k_mbf`), because it finds a different set of faces. Every other platform
ignores them, and the Windows installer leaves them out. The embedder and the eye classifiers have
no quantised form: the embedder's vectors must compare across devices (inference.md §7), and the
classifiers cost a millisecond on the CPU.
The detectors were int8 until §1.5 measured them on the tablet: int8 found 94–95% of f32's faces
at 40–80 px, A16W8 all of them. A tablet that indexed under the int8 ids keeps those rows.
**A clone without git-lfs gets a ~130-byte pointer where each model should be.** Both packagers check
for exactly that and refuse, rather than shipping the pointer and failing inside tract on the user's
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
+13 -8
View File
@@ -46,12 +46,14 @@ pub fn init(runtime_dirs: Vec<PathBuf>) {
cache_dir: crate::library::inference_cache_dir(),
models,
embedded: {
let [landscape, portrait] = dr_pano::xfeat::embedded_model_bytes();
vec![
(Role::Segmenter, dr_segment::embedded_model_bytes()),
(Role::Keypoints, landscape),
(Role::Keypoints, portrait),
]
let [landscape, portrait] = dr_pano::xfeat::embedded_models();
let tag = |role| move |(form, bytes)| (role, form, bytes);
dr_segment::embedded_models()
.into_iter()
.map(tag(Role::Segmenter))
.chain(landscape.into_iter().map(tag(Role::Keypoints)))
.chain(portrait.into_iter().map(tag(Role::Keypoints)))
.collect()
},
ceiling: None,
threads: 0,
@@ -81,7 +83,7 @@ pub fn user_runtime_dir() -> PathBuf {
///
/// Reads the shared and system directories only. An account-private model
/// directory can override the file `library::face_models` loads, but not
/// which form the backend wants, and the int8 sibling is something a
/// which form the backend wants, and the quantised sibling is something a
/// packager ships, not something a user drops in.
pub fn detector_form(detector: FaceDetector) -> Form {
let canonical = crate::library::shared_model(detector.file_name())
@@ -92,8 +94,11 @@ pub fn detector_form(detector: FaceDetector) -> Form {
/// The `faces.model_id` this device indexes under with `detector`.
pub fn model_id(detector: FaceDetector) -> &'static str {
match detector_form(detector) {
Form::F32 => detector.model_id(),
Form::Int8 => detector.model_id_int8(),
Form::A16W8 => detector.model_id_a16w8(),
// No detector is offered in A16W16 (inference.md §1.5); were one, it
// would be the network f32 is to the last bit that a person can see.
Form::F32 | Form::A16W16 => detector.model_id(),
}
}