Run each model on the Hexagon in the form measured to hold it

The engine knew f32 and int8, and gave the Hexagon int8 for every role it
served. Measured on the tablet itself (inference.md §1.5), int8 lost
5% of the detector's faces at 40-80 px, moved the landmarks 1.5 px,
emptied the segmenter's scores and cost the denoiser 5-9 dB; fp16 the HTP
refuses outright. `Form` gains A16W8 and A16W16, and `Rung::form` now
names one per role: detectors and landmarks A16W8, the segmenter, scene
model, border filler and denoiser A16W16, XFeat int8. The embedder and
the eye classifiers stay on the CPU.

Each loader resolves its `<stem>.<form>.onnx` sibling; the segmenter and
XFeat, compiled into the binary, embed their quantised forms on Android
only and pick through `choose_embedded`. The probe, the compile step and
the cache fingerprint follow the form instead of assuming int8. Detectors
on the new form write `scrfd_*_a16+w600k_mbf`, and `model_ids` answers
for all three spellings.

On the tablet (ORT 1.29 + QNN 2.42), each shipped file against f32 on the
same inputs, and against the CPU's f32 time:
  SCRFD 500m/2.5g/10g  A16W8   100% of faces in every band   4.2/5.1/9.0 ms vs 17/56/198
  landmarks            A16W8   0.25 px in the 192 crop        0.5 ms vs 2.8
  YOLO26n-seg          A16W16  98.2% found, mask IoU 0.994    12.9 ms vs 90
  scene model          A16W16  98.9% of cells agree           15 ms vs 151
  MI-GAN               A16W16  41 dB from f32 in the fill     87 ms vs 488
  XFeat                int8    pano alignment 0.45 px (f32's own spread 0.41)  6.5 ms vs 58
  denoiser             A16W16  0.00 dB at every ISO            95 ms vs 1510 a tile
Face numbers are over public COCO val2017 photographs, not a library.

The APK carries the siblings (BUNDLED 15 -> 19; the old int8 detectors
removed), about 43 MB more. The Windows installer and its CI count skip
them; the Arch and Flatpak packages list their files and never had them.
The ladder example takes a role per model, which is how the per-role
forms above were seen landing on the NPU from the real probe.
This commit is contained in:
2026-10-04 03:45:46 -04:00
parent 0e6ac09fd5
commit 5a8c3e4c40
39 changed files with 550 additions and 208 deletions
+14 -4
View File
@@ -13,9 +13,11 @@
use std::path::Path;
const MODEL: &str = "../../models/segment/yolo26n-seg.onnx";
const QUANTISED: &str = "../../models/segment/yolo26n-seg.a16w16.onnx";
fn main() {
println!("cargo:rerun-if-changed={MODEL}");
println!("cargo:rerun-if-changed={QUANTISED}");
println!("cargo:rerun-if-changed=build.rs");
// Only the embedded path needs the file present; a build without it is
@@ -24,10 +26,18 @@ fn main() {
return;
}
let path = Path::new(MODEL);
check(MODEL);
// The Hexagon's quantised form rides only in an Android build.
if std::env::var("CARGO_CFG_TARGET_OS").as_deref() == Ok("android") {
check(QUANTISED);
}
}
fn check(model: &str) {
let path = Path::new(model);
let Ok(bytes) = std::fs::read(path) else {
panic!(
"\n\n{MODEL} is missing.\n\
"\n\n{model} is missing.\n\
It ships in Git LFS. Run `git lfs install && git lfs pull`, or build \
with `--no-default-features` for a watershed-only build.\n"
);
@@ -40,7 +50,7 @@ fn main() {
// happens in practice.
if bytes.starts_with(b"version https://git-lfs") {
panic!(
"\n\n{MODEL} is a Git LFS pointer, not the model ({} bytes).\n\
"\n\n{model} is a Git LFS pointer, not the model ({} bytes).\n\
Run `git lfs install && git lfs pull` to fetch the real file.\n",
bytes.len()
);
@@ -50,7 +60,7 @@ fn main() {
// export is ~11 MB; anything under a megabyte is a truncated checkout.
if bytes.len() < 1_000_000 {
panic!(
"\n\n{MODEL} is only {} bytes — expected ~11 MB.\n\
"\n\n{model} is only {} bytes — expected several MB.\n\
The checkout looks incomplete; try `git lfs pull`.\n",
bytes.len()
);
+1 -1
View File
@@ -72,7 +72,7 @@ pub use refine::{
#[cfg(feature = "semantic")]
pub use scene::{Category, Scene, SceneModel};
#[cfg(feature = "embedded-model")]
pub use semantic::embedded_model_bytes;
pub use semantic::embedded_models;
#[cfg(feature = "semantic")]
pub use semantic::{Instance, SemanticModel, SemanticOptions, Tiling};
+15 -7
View File
@@ -123,21 +123,29 @@ impl SceneModel {
classes: impl AsRef<std::path::Path>,
categories: impl AsRef<std::path::Path>,
) -> Result<Self, SegmentError> {
// The form the device's backend runs: the `.a16w16.onnx` sibling on
// the Hexagon (attention left in float, inference.md §1.5), else this.
let (model, form) =
dr_inference_engine::resolve_model(dr_inference_engine::Role::Scene, model.as_ref());
let bytes = std::fs::read(model).map_err(SegmentError::ModelRead)?;
let classes = std::fs::read_to_string(classes).map_err(SegmentError::ModelRead)?;
let categories = std::fs::read_to_string(categories).map_err(SegmentError::ModelRead)?;
let classes = crate::semantic::parse_classes(&classes);
let categories = parse_categories(&categories, &classes)?;
Self::from_bytes(&bytes, categories)
Self::from_bytes_in(&bytes, form, categories)
}
pub fn from_bytes(bytes: &[u8], categories: Vec<Category>) -> Result<Self, SegmentError> {
// f32, as for `SemanticModel`; see there.
let session = dr_inference_engine::open(
dr_inference_engine::Role::Scene,
dr_inference_engine::Form::F32,
bytes,
)?;
Self::from_bytes_in(bytes, dr_inference_engine::Form::F32, categories)
}
/// `bytes` in a stated numeric form; the outputs keep their shape.
pub fn from_bytes_in(
bytes: &[u8],
form: dr_inference_engine::Form,
categories: Vec<Category>,
) -> Result<Self, SegmentError> {
let session = dr_inference_engine::open(dr_inference_engine::Role::Scene, form, bytes)?;
Ok(Self {
session,
+31 -13
View File
@@ -208,18 +208,32 @@ const EMBEDDED_MODEL: &[u8] = include_bytes!("../../../models/segment/yolo26n-se
#[cfg(feature = "embedded-model")]
const EMBEDDED_CLASSES: &str = include_str!("../../../models/segment/yolo26n-seg.classes.json");
/// The bytes of the model that ships with this crate, for whoever compiles
/// The Hexagon's form (docs/dev/inference.md §1.5): 16-bit activations and
/// weights, the rows' tail left in float. Only Android has a Hexagon, so only
/// Android carries it.
#[cfg(all(feature = "embedded-model", target_os = "android"))]
const EMBEDDED_A16W16: &[u8] = include_bytes!("../../../models/segment/yolo26n-seg.a16w16.onnx");
/// Every form of the model that ships with this crate, for whoever compiles
/// engines ahead of the first request (docs/dev/inference.md §6).
#[cfg(feature = "embedded-model")]
pub fn embedded_model_bytes() -> &'static [u8] {
EMBEDDED_MODEL
pub fn embedded_models() -> Vec<(dr_inference_engine::Form, &'static [u8])> {
#[allow(unused_mut)]
let mut forms = vec![(dr_inference_engine::Form::F32, EMBEDDED_MODEL)];
#[cfg(target_os = "android")]
forms.push((dr_inference_engine::Form::A16W16, EMBEDDED_A16W16));
forms
}
impl SemanticModel {
/// Load the model that ships with this crate.
/// Load the model that ships with this crate, in the form the device's
/// backend runs.
#[cfg(feature = "embedded-model")]
pub fn embedded() -> Result<Self, SegmentError> {
Self::from_bytes(EMBEDDED_MODEL, parse_classes(EMBEDDED_CLASSES))
let forms = embedded_models();
let (bytes, form) =
dr_inference_engine::choose_embedded(dr_inference_engine::Role::Segmenter, &forms);
Self::from_bytes_in(bytes, form, parse_classes(EMBEDDED_CLASSES))
}
/// Load a model from an ONNX file, with `classes` supplying its vocabulary.
@@ -236,14 +250,18 @@ impl SemanticModel {
}
pub fn from_bytes(bytes: &[u8], classes: Vec<Arc<str>>) -> Result<Self, SegmentError> {
// The f32 graph on whatever the device's backend is. An int8 form
// for the Hexagon waits on docs/dev/inference.md §10 M7 — the mask
// boundary has to be measured before it moves.
let session = dr_inference_engine::open(
dr_inference_engine::Role::Segmenter,
dr_inference_engine::Form::F32,
bytes,
)?;
Self::from_bytes_in(bytes, dr_inference_engine::Form::F32, classes)
}
/// `bytes` in a stated numeric form. The quantised one keeps the same
/// outputs (the rows' tail stays float), so decoding does not change; the
/// masks it draws were measured against f32's (inference.md §1.5).
pub fn from_bytes_in(
bytes: &[u8],
form: dr_inference_engine::Form,
classes: Vec<Arc<str>>,
) -> Result<Self, SegmentError> {
let session = dr_inference_engine::open(dr_inference_engine::Role::Segmenter, form, bytes)?;
Ok(Self { session, classes })
}