Files
dtourolle 5a8c3e4c40 Run each model on the Hexagon in the form measured to hold it
The engine knew f32 and int8, and gave the Hexagon int8 for every role it
served. Measured on the tablet itself (inference.md §1.5), int8 lost
5% of the detector's faces at 40-80 px, moved the landmarks 1.5 px,
emptied the segmenter's scores and cost the denoiser 5-9 dB; fp16 the HTP
refuses outright. `Form` gains A16W8 and A16W16, and `Rung::form` now
names one per role: detectors and landmarks A16W8, the segmenter, scene
model, border filler and denoiser A16W16, XFeat int8. The embedder and
the eye classifiers stay on the CPU.

Each loader resolves its `<stem>.<form>.onnx` sibling; the segmenter and
XFeat, compiled into the binary, embed their quantised forms on Android
only and pick through `choose_embedded`. The probe, the compile step and
the cache fingerprint follow the form instead of assuming int8. Detectors
on the new form write `scrfd_*_a16+w600k_mbf`, and `model_ids` answers
for all three spellings.

On the tablet (ORT 1.29 + QNN 2.42), each shipped file against f32 on the
same inputs, and against the CPU's f32 time:
  SCRFD 500m/2.5g/10g  A16W8   100% of faces in every band   4.2/5.1/9.0 ms vs 17/56/198
  landmarks            A16W8   0.25 px in the 192 crop        0.5 ms vs 2.8
  YOLO26n-seg          A16W16  98.2% found, mask IoU 0.994    12.9 ms vs 90
  scene model          A16W16  98.9% of cells agree           15 ms vs 151
  MI-GAN               A16W16  41 dB from f32 in the fill     87 ms vs 488
  XFeat                int8    pano alignment 0.45 px (f32's own spread 0.41)  6.5 ms vs 58
  denoiser             A16W16  0.00 dB at every ISO            95 ms vs 1510 a tile
Face numbers are over public COCO val2017 photographs, not a library.

The APK carries the siblings (BUNDLED 15 -> 19; the old int8 detectors
removed), about 43 MB more. The Windows installer and its CI count skip
them; the Arch and Flatpak packages list their files and never had them.
The ladder example takes a role per model, which is how the per-role
forms above were seen landing on the NPU from the real probe.
2026-10-04 03:45:46 -04:00

183 lines
7.5 KiB
Rust
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
//! TRACES: FR-MRG-8
//! The XFeat detector — the network under tract, and the decoder after it.
//!
//! Apache-2.0 weights (`models/LICENCE.md`), exported at a fixed shape by
//! `tools/export-xfeat.sh` and loaded through the same `dr-inference-engine`
//! `dr-segment` and `dr-face` use, so this adds no runtime and no C to the
//! tree; what runs it is the device's business (docs/dev/inference.md). ~300 ms
//! per frame on tract on the reference desktop, ~400 ms on the tablet
//! (S15.2, S15.4).
use crate::features::{decode_xfeat, DecodeOptions, Features, XFeatMaps, DESCRIPTOR_LEN};
use crate::image::Gray;
use crate::PanoError;
/// The two input shapes the shipped exports were made for: one landscape,
/// one portrait, the same weights. A frame is fitted into whichever
/// matches its aspect, so a portrait set does not spend half the
/// detector's width on padding — which is what the 6D fixture did before
/// the second export existed (512 × 768 of a 1024 × 768 input). A
/// different size is a different file (`tools/export-xfeat.sh`).
pub const INPUT_LANDSCAPE: (usize, usize) = (1024, 768);
pub const INPUT_PORTRAIT: (usize, usize) = (768, 1024);
/// The long edge of the detector's input, for callers sizing a proxy.
pub const INPUT_LONG_EDGE: usize = 1024;
#[cfg(feature = "embedded-model")]
const EMBEDDED_LANDSCAPE: &[u8] = include_bytes!("../../../models/keypoints/xfeat-1024.onnx");
#[cfg(feature = "embedded-model")]
const EMBEDDED_PORTRAIT: &[u8] = include_bytes!("../../../models/keypoints/xfeat-768.onnx");
/// A loaded detector: the network at both shapes.
pub struct XFeat {
landscape: dr_inference_engine::Model,
portrait: dr_inference_engine::Model,
pub options: DecodeOptions,
}
/// The Hexagon's forms (docs/dev/inference.md §1.5): int8, from the same
/// network spelled for the HTP (the unfold as SpaceToDepth, the bilinear
/// resizes as matrix products). Only Android has a Hexagon.
#[cfg(all(feature = "embedded-model", target_os = "android"))]
const EMBEDDED_LANDSCAPE_INT8: &[u8] =
include_bytes!("../../../models/keypoints/xfeat-1024.int8.onnx");
#[cfg(all(feature = "embedded-model", target_os = "android"))]
const EMBEDDED_PORTRAIT_INT8: &[u8] =
include_bytes!("../../../models/keypoints/xfeat-768.int8.onnx");
/// Every form of both exports compiled into the binary, landscape then
/// portrait, for whoever compiles engines ahead of the first request
/// (docs/dev/inference.md §6).
#[cfg(feature = "embedded-model")]
pub fn embedded_models() -> [Vec<(dr_inference_engine::Form, &'static [u8])>; 2] {
use dr_inference_engine::Form;
#[allow(unused_mut)]
let mut forms = [
vec![(Form::F32, EMBEDDED_LANDSCAPE)],
vec![(Form::F32, EMBEDDED_PORTRAIT)],
];
#[cfg(target_os = "android")]
{
forms[0].push((Form::Int8, EMBEDDED_LANDSCAPE_INT8));
forms[1].push((Form::Int8, EMBEDDED_PORTRAIT_INT8));
}
forms
}
impl XFeat {
/// The weights compiled into the binary, in the form the device's
/// backend runs.
#[cfg(feature = "embedded-model")]
pub fn embedded() -> Result<Self, PanoError> {
use dr_inference_engine::{choose_embedded, open, Role};
let [l, p] = embedded_models();
let (l, lf) = choose_embedded(Role::Keypoints, &l);
let (p, pf) = choose_embedded(Role::Keypoints, &p);
Ok(XFeat {
landscape: open(Role::Keypoints, lf, l)?,
portrait: open(Role::Keypoints, pf, p)?,
options: DecodeOptions::default(),
})
}
/// From the two exports on disk.
pub fn from_paths(
landscape: &std::path::Path,
portrait: &std::path::Path,
) -> Result<Self, PanoError> {
let l = std::fs::read(landscape).map_err(PanoError::ModelRead)?;
let p = std::fs::read(portrait).map_err(PanoError::ModelRead)?;
Self::from_bytes(&l, &p)
}
pub fn from_bytes(landscape: &[u8], portrait: &[u8]) -> Result<Self, PanoError> {
use dr_inference_engine::{Form, Role};
Ok(XFeat {
landscape: dr_inference_engine::open(Role::Keypoints, Form::F32, landscape)?,
portrait: dr_inference_engine::open(Role::Keypoints, Form::F32, portrait)?,
options: DecodeOptions::default(),
})
}
/// Detect keypoints in an upright grayscale image.
///
/// The image is fitted into the network's input of matching aspect —
/// scaled down if larger, never up, and padded to the right and bottom
/// — and the keypoints come back in the coordinates of `image` itself,
/// so a caller that already scaled a frame to a proxy maps them on with
/// the scale it used and nothing else.
pub fn detect(&mut self, image: &Gray) -> Result<Features, PanoError> {
let ((in_w, in_h), model) = if image.height > image.width {
(INPUT_PORTRAIT, &self.portrait)
} else {
(INPUT_LANDSCAPE, &self.landscape)
};
let acquired = model.acquire()?;
let mut session = acquired.lock();
let (fitted, scale) = image.fitted(in_w, in_h);
let padded = fitted.padded(in_w, in_h);
let input =
ndarray::Array::from_shape_vec(ndarray::IxDyn(&[1, 1, in_h, in_w]), padded.data)
.expect("shape matches the buffer by construction");
let tensor = ort::value::Tensor::from_array(input).map_err(PanoError::Inference)?;
let outputs = session
.run(ort::inputs![tensor])
.map_err(PanoError::Inference)?;
let (w8, h8) = (in_w / 8, in_h / 8);
let expect = |i: usize, channels: usize| -> Result<Vec<f32>, PanoError> {
let (shape, data) = outputs[i]
.try_extract_tensor::<f32>()
.map_err(PanoError::Inference)?;
let dims: Vec<i64> = shape.iter().copied().collect();
if dims != [1, channels as i64, h8 as i64, w8 as i64] {
return Err(PanoError::Model(format!(
"output {i} is {dims:?}, expected [1, {channels}, {h8}, {w8}] — \
not the export this decoder was written for"
)));
}
Ok(data.to_vec())
};
let feats = expect(0, DESCRIPTOR_LEN)?;
let keypoints = expect(1, 65)?;
let heatmap = expect(2, 1)?;
let mut features = decode_xfeat(
&XFeatMaps {
feats: &feats,
keypoints: &keypoints,
heatmap: &heatmap,
width: w8,
height: h8,
},
&self.options,
);
// Back to the caller's image: drop anything the padding produced,
// undo the fit.
let border = self.options.border as f32;
let limit_x = fitted.width as f32 - border;
let limit_y = fitted.height as f32 - border;
let mut kept_kp = Vec::with_capacity(features.len());
let mut kept_desc = Vec::with_capacity(features.descriptors.len());
for (i, kp) in features.keypoints.iter().enumerate() {
if kp.x >= limit_x || kp.y >= limit_y {
continue;
}
kept_kp.push(crate::features::Keypoint {
x: (kp.x / scale as f32),
y: (kp.y / scale as f32),
score: kp.score,
});
kept_desc.extend_from_slice(features.descriptor(i));
}
features.keypoints = kept_kp;
features.descriptors = kept_desc;
features.width = image.width;
features.height = image.height;
Ok(features)
}
}