The engine knew f32 and int8, and gave the Hexagon int8 for every role it served. Measured on the tablet itself (inference.md §1.5), int8 lost 5% of the detector's faces at 40-80 px, moved the landmarks 1.5 px, emptied the segmenter's scores and cost the denoiser 5-9 dB; fp16 the HTP refuses outright. `Form` gains A16W8 and A16W16, and `Rung::form` now names one per role: detectors and landmarks A16W8, the segmenter, scene model, border filler and denoiser A16W16, XFeat int8. The embedder and the eye classifiers stay on the CPU. Each loader resolves its `<stem>.<form>.onnx` sibling; the segmenter and XFeat, compiled into the binary, embed their quantised forms on Android only and pick through `choose_embedded`. The probe, the compile step and the cache fingerprint follow the form instead of assuming int8. Detectors on the new form write `scrfd_*_a16+w600k_mbf`, and `model_ids` answers for all three spellings. On the tablet (ORT 1.29 + QNN 2.42), each shipped file against f32 on the same inputs, and against the CPU's f32 time: SCRFD 500m/2.5g/10g A16W8 100% of faces in every band 4.2/5.1/9.0 ms vs 17/56/198 landmarks A16W8 0.25 px in the 192 crop 0.5 ms vs 2.8 YOLO26n-seg A16W16 98.2% found, mask IoU 0.994 12.9 ms vs 90 scene model A16W16 98.9% of cells agree 15 ms vs 151 MI-GAN A16W16 41 dB from f32 in the fill 87 ms vs 488 XFeat int8 pano alignment 0.45 px (f32's own spread 0.41) 6.5 ms vs 58 denoiser A16W16 0.00 dB at every ISO 95 ms vs 1510 a tile Face numbers are over public COCO val2017 photographs, not a library. The APK carries the siblings (BUNDLED 15 -> 19; the old int8 detectors removed), about 43 MB more. The Windows installer and its CI count skip them; the Arch and Flatpak packages list their files and never had them. The ladder example takes a role per model, which is how the per-role forms above were seen landing on the NPU from the real probe.
464 lines
18 KiB
Rust
464 lines
18 KiB
Rust
//! SCRFD face detection (docs/dev/faces.md §4).
|
||
//!
|
||
//! One forward pass produces a box, a confidence and **five landmarks** per
|
||
//! face — the landmarks being the reason for this detector rather than a
|
||
//! general one, since [`crate::align`] cannot work without them.
|
||
//!
|
||
//! # The graph must have fixed input dimensions
|
||
//!
|
||
//! InsightFace ships `det_500m.onnx` with a dynamic H/W input, and **tract
|
||
//! cannot parse it in that form** — it fails at node #0. The same file run
|
||
//! through `tools/fix-face-model-shapes.sh` loads cleanly. Its outputs were
|
||
//! already static at 640, so 640 is not a choice made here: it is the shape
|
||
//! the export was always going to run at.
|
||
|
||
use ndarray::Array4;
|
||
|
||
use crate::FaceError;
|
||
use dr_inference_engine::{Form, Model, Role};
|
||
|
||
/// The graph's input edge, in pixels. See the module note: not configurable.
|
||
pub const INPUT_EDGE: usize = 640;
|
||
|
||
/// Strides, in the order SCRFD emits them.
|
||
const ALL_STRIDES: [usize; 4] = [8, 16, 32, 64];
|
||
|
||
/// Anchors per feature-map location.
|
||
const ANCHORS: usize = 2;
|
||
|
||
/// How detection is tuned.
|
||
#[derive(Debug, Clone, Copy, PartialEq)]
|
||
pub struct DetectOptions {
|
||
/// Minimum detector confidence.
|
||
///
|
||
/// Deliberately *not* the low threshold `dr-segment` chose. There a false
|
||
/// positive costs one spurious row in a list the user is picking from;
|
||
/// here it costs a face in the People view to reject and — worse — a
|
||
/// garbage embedding that can bridge two real clusters into one. A false
|
||
/// negative is recoverable by re-indexing with a better model; a polluted
|
||
/// cluster graph, once the user has confirmed faces inside it, is not.
|
||
pub confidence: f32,
|
||
/// Box IoU above which two detections are judged to be the same face.
|
||
pub nms_iou: f32,
|
||
/// Cheap pre-filter: smallest box to keep, in source pixels on the shorter
|
||
/// edge.
|
||
///
|
||
/// **Not the real size floor** — [`DetectOptions::min_source_px`] is, and
|
||
/// it is measured on the aligned crop rather than the box. This one exists
|
||
/// only to throw away the obviously hopeless before paying for a warp, so
|
||
/// it is deliberately set *below* what the real floor will accept: the
|
||
/// aligned crop spans roughly 1.3x the box's shorter edge, so 24 here
|
||
/// cannot reject a face that would have cleared 32 there.
|
||
pub min_face_px: f32,
|
||
/// Smallest face the embedder may be given, in **source pixels across the
|
||
/// aligned crop** — `crop_px` in the catalog.
|
||
///
|
||
/// The honest statement of "a face must be at least 32x32", because this is
|
||
/// the number of real pixels behind the 112x112 the model actually sees.
|
||
/// The box's own size is not that: the ArcFace template reaches past the
|
||
/// box for forehead and chin, so a 64-pixel box and a 64-pixel crop are
|
||
/// different faces.
|
||
///
|
||
/// Below this the crop was upsampled to reach the embedder, and upsampling
|
||
/// invents no detail — the embedding is of a soft, stretched face and is
|
||
/// correspondingly untrustworthy.
|
||
///
|
||
/// Applied after alignment, so it lives with the sharpness floor rather
|
||
/// than with the detector. See [`DetectOptions::min_sharpness`].
|
||
pub min_source_px: f32,
|
||
/// Least acceptable [`crate::align::Aligned112::sharpness`].
|
||
///
|
||
/// Applied after alignment rather than here, because it is a property of
|
||
/// the warped crop the embedder receives and not of the box. The pipeline
|
||
/// that enforces it is `dr_ui::faces::index_proxy`; it lives on this struct
|
||
/// so that every quality decision about a face is configured in one place
|
||
/// and a caller cannot enable one gate while forgetting the other.
|
||
///
|
||
/// Zero disables it, which is what a measurement run wants.
|
||
///
|
||
/// # It has to move with the size floor
|
||
///
|
||
/// The two are coupled, because an upsampled face scores low here whatever
|
||
/// its original sharpness. Measured over the reference library, with the
|
||
/// size floor at 32 source pixels:
|
||
///
|
||
/// | min sharpness | of what the size floor left, this removes |
|
||
/// |---|---|
|
||
/// | 0.002 | 3% |
|
||
/// | 0.005 | 8% |
|
||
/// | 0.010 | 16% |
|
||
/// | 0.020 | 27% |
|
||
///
|
||
/// At a 64-pixel floor, 0.020 removed 7% — the same *kind* of face, the
|
||
/// large-but-soft one this gate exists for. Holding 0.020 while dropping
|
||
/// the size floor to 32 would have thrown away a quarter of the newly
|
||
/// admitted faces for being small rather than for being blurred, undoing
|
||
/// most of the point of lowering it. 0.005 removes 8% at 32, which is the
|
||
/// same job.
|
||
pub min_sharpness: f32,
|
||
}
|
||
|
||
impl Default for DetectOptions {
|
||
fn default() -> Self {
|
||
Self {
|
||
confidence: 0.5,
|
||
nms_iou: 0.4,
|
||
min_face_px: 24.0,
|
||
min_source_px: 32.0,
|
||
min_sharpness: 0.005,
|
||
}
|
||
}
|
||
}
|
||
|
||
/// One detected face, in **source image pixels**.
|
||
///
|
||
/// Pixels rather than the normalised form the catalog stores, because the
|
||
/// caller still has to crop from this image. Normalisation happens at the
|
||
/// storage boundary, where the long edge is known to be the right divisor.
|
||
#[derive(Debug, Clone, PartialEq)]
|
||
pub struct Detection {
|
||
/// `(x0, y0, x1, y1)`.
|
||
pub bbox: (f32, f32, f32, f32),
|
||
/// Five points in the detector's own order — see [`crate::align`], which
|
||
/// consumes them without reordering.
|
||
pub landmarks: [(f32, f32); 5],
|
||
pub confidence: f32,
|
||
}
|
||
|
||
impl Detection {
|
||
pub fn width(&self) -> f32 {
|
||
self.bbox.2 - self.bbox.0
|
||
}
|
||
pub fn height(&self) -> f32 {
|
||
self.bbox.3 - self.bbox.1
|
||
}
|
||
}
|
||
|
||
/// A loaded SCRFD graph.
|
||
pub struct Detector {
|
||
session: Model,
|
||
/// f32 or a quantised form — which finds a different set of faces and is
|
||
/// a different detector in `model_id` (docs/dev/inference.md §7).
|
||
form: Form,
|
||
/// Feature-map count: 3 for strides {8,16,32}, 4 for {8,16,32,64}.
|
||
///
|
||
/// Discovered from the output count rather than assumed, because both
|
||
/// exports exist and hardcoding 3 silently ignores the largest faces a
|
||
/// four-stride model finds.
|
||
fmc: usize,
|
||
}
|
||
|
||
impl Detector {
|
||
/// Which form this detector was loaded from.
|
||
pub fn form(&self) -> Form {
|
||
self.form
|
||
}
|
||
|
||
/// Load the canonical f32 file at `path`, or the form the device's
|
||
/// backend wants instead — the `.a16w8.onnx` beside it on a Hexagon —
|
||
/// which [`Detector::form`] then reports.
|
||
pub fn from_path(path: impl AsRef<std::path::Path>) -> Result<Self, FaceError> {
|
||
let (path, form) = dr_inference_engine::resolve_model(Role::Detector, path.as_ref());
|
||
let bytes = std::fs::read(path).map_err(FaceError::ModelRead)?;
|
||
Self::from_bytes_in(&bytes, form)
|
||
}
|
||
|
||
/// An f32 graph from memory.
|
||
pub fn from_bytes(bytes: &[u8]) -> Result<Self, FaceError> {
|
||
Self::from_bytes_in(bytes, Form::F32)
|
||
}
|
||
|
||
fn from_bytes_in(bytes: &[u8], form: Form) -> Result<Self, FaceError> {
|
||
let model = dr_inference_engine::open(Role::Detector, form, bytes)?;
|
||
let acquired = model.acquire()?;
|
||
let session = acquired.lock();
|
||
|
||
let n_out = session.outputs().len();
|
||
if n_out % 3 != 0 || !(9..=12).contains(&n_out) {
|
||
return Err(FaceError::WrongModel {
|
||
expected: "InsightFace SCRFD",
|
||
detail: format!("expected 9 or 12 outputs, got {n_out}"),
|
||
});
|
||
}
|
||
let fmc = n_out / 3;
|
||
|
||
// The check that actually distinguishes the models. YuNet also has
|
||
// twelve outputs in three strides, so the count proves nothing — its
|
||
// groups are cls/obj/bbox/kps where SCRFD's are score/bbox/kps, and
|
||
// decoding one as the other yields a page of plausible numbers rather
|
||
// than an error. The last dimension is what separates them.
|
||
for (group, expected_last) in [1_i64, 4, 10].into_iter().enumerate() {
|
||
for s in 0..fmc {
|
||
let idx = group * fmc + s;
|
||
let out = &session.outputs()[idx];
|
||
let last: Option<i64> = out.dtype().tensor_shape().and_then(|d| d.last().copied());
|
||
if last != Some(expected_last) {
|
||
return Err(FaceError::WrongModel {
|
||
expected: "InsightFace SCRFD",
|
||
detail: format!(
|
||
"output '{}' last dim is {:?}, expected {expected_last} \
|
||
(a YuNet export fails exactly here)",
|
||
out.name(),
|
||
last
|
||
),
|
||
});
|
||
}
|
||
}
|
||
}
|
||
|
||
drop(session);
|
||
drop(acquired);
|
||
Ok(Self {
|
||
session: model,
|
||
form,
|
||
fmc,
|
||
})
|
||
}
|
||
|
||
/// Stride levels this graph emits.
|
||
pub fn strides(&self) -> &'static [usize] {
|
||
&ALL_STRIDES[..self.fmc]
|
||
}
|
||
|
||
/// Find the faces in an image.
|
||
///
|
||
/// `rgb` is tightly packed `f32` RGB in `0.0..=1.0`, row-major — the same
|
||
/// convention `dr-segment` and [`crate::align`] use.
|
||
pub fn detect(
|
||
&mut self,
|
||
rgb: &[f32],
|
||
width: usize,
|
||
height: usize,
|
||
options: &DetectOptions,
|
||
) -> Result<Vec<Detection>, FaceError> {
|
||
if width == 0 || height == 0 {
|
||
return Ok(Vec::new());
|
||
}
|
||
if rgb.len() != width * height * 3 {
|
||
return Err(FaceError::ImageShape {
|
||
expected: width * height * 3,
|
||
got: rgb.len(),
|
||
});
|
||
}
|
||
|
||
let lb = Letterbox::fit(width as f32, height as f32);
|
||
let input = lb.sample(rgb, width, height);
|
||
|
||
let acquired = self.session.acquire()?;
|
||
let mut session = acquired.lock();
|
||
let outputs = session
|
||
.run(ort::inputs![
|
||
ort::value::Tensor::from_array(input).map_err(FaceError::Inference)?
|
||
])
|
||
.map_err(FaceError::Inference)?;
|
||
|
||
let mut raw: Vec<Detection> = Vec::new();
|
||
|
||
for (si, &stride) in ALL_STRIDES[..self.fmc].iter().enumerate() {
|
||
let (_, scores) = outputs[si]
|
||
.try_extract_tensor::<f32>()
|
||
.map_err(FaceError::Inference)?;
|
||
let (_, boxes) = outputs[self.fmc + si]
|
||
.try_extract_tensor::<f32>()
|
||
.map_err(FaceError::Inference)?;
|
||
let (_, kps) = outputs[self.fmc * 2 + si]
|
||
.try_extract_tensor::<f32>()
|
||
.map_err(FaceError::Inference)?;
|
||
|
||
let fw = INPUT_EDGE / stride;
|
||
let fh = INPUT_EDGE / stride;
|
||
let s = stride as f32;
|
||
|
||
for r in 0..fh {
|
||
for c in 0..fw {
|
||
for a in 0..ANCHORS {
|
||
let idx = (r * fw + c) * ANCHORS + a;
|
||
let score = scores[idx];
|
||
if score < options.confidence {
|
||
continue;
|
||
}
|
||
|
||
// Anchor centre in input space, then distance-to-box
|
||
// decoding: the four regressed values are distances
|
||
// left/top/right/bottom in units of the stride.
|
||
let (cx, cy) = ((c * stride) as f32, (r * stride) as f32);
|
||
let b = &boxes[idx * 4..idx * 4 + 4];
|
||
let (x0, y0) = lb.into_source(cx - b[0] * s, cy - b[1] * s);
|
||
let (x1, y1) = lb.into_source(cx + b[2] * s, cy + b[3] * s);
|
||
|
||
let k = &kps[idx * 10..idx * 10 + 10];
|
||
let mut landmarks = [(0.0_f32, 0.0_f32); 5];
|
||
for (p, lm) in landmarks.iter_mut().enumerate() {
|
||
*lm = lb.into_source(cx + k[p * 2] * s, cy + k[p * 2 + 1] * s);
|
||
}
|
||
|
||
raw.push(Detection {
|
||
bbox: (x0, y0, x1, y1),
|
||
landmarks,
|
||
confidence: score,
|
||
});
|
||
}
|
||
}
|
||
}
|
||
}
|
||
|
||
let mut kept = non_max_suppress(raw, options.nms_iou);
|
||
|
||
// Size floor last, on the *merged* boxes: a face that only clears the
|
||
// floor once NMS has picked the best of its overlapping detections
|
||
// should be kept.
|
||
kept.retain(|d| d.width().min(d.height()) >= options.min_face_px);
|
||
|
||
// No cap on the count. The reference implementation keeps the ten
|
||
// largest, which is right for a film frame where background extras are
|
||
// noise; it is wrong for a photo library, where a group shot with
|
||
// thirty faces is precisely the picture worth indexing.
|
||
Ok(kept)
|
||
}
|
||
}
|
||
|
||
/// Greedy NMS across all strides together.
|
||
fn non_max_suppress(mut dets: Vec<Detection>, iou_threshold: f32) -> Vec<Detection> {
|
||
dets.sort_by(|a, b| b.confidence.total_cmp(&a.confidence));
|
||
let mut kept: Vec<Detection> = Vec::new();
|
||
for d in dets {
|
||
if kept.iter().all(|k| iou(&k.bbox, &d.bbox) <= iou_threshold) {
|
||
kept.push(d);
|
||
}
|
||
}
|
||
kept
|
||
}
|
||
|
||
fn iou(a: &(f32, f32, f32, f32), b: &(f32, f32, f32, f32)) -> f32 {
|
||
let ix = (a.2.min(b.2) - a.0.max(b.0)).max(0.0);
|
||
let iy = (a.3.min(b.3) - a.1.max(b.1)).max(0.0);
|
||
let inter = ix * iy;
|
||
let area_a = (a.2 - a.0).max(0.0) * (a.3 - a.1).max(0.0);
|
||
let area_b = (b.2 - b.0).max(0.0) * (b.3 - b.1).max(0.0);
|
||
let union = area_a + area_b - inter;
|
||
if union <= 0.0 {
|
||
0.0
|
||
} else {
|
||
inter / union
|
||
}
|
||
}
|
||
|
||
/// How the image is fitted into the graph's fixed square input.
|
||
///
|
||
/// The forward and inverse mappings live in one struct on purpose:
|
||
/// docs/dev/faces.md §4.1 notes that what matters is not *where* the padding goes
|
||
/// but that the two agree. A mismatch offsets every box and landmark by the
|
||
/// padding, producing detections that look plausible and embeddings that
|
||
/// quietly cluster badly three stages later.
|
||
#[derive(Debug, Clone, Copy)]
|
||
struct Letterbox {
|
||
/// Input pixels per source pixel.
|
||
scale: f32,
|
||
pad_x: f32,
|
||
pad_y: f32,
|
||
}
|
||
|
||
impl Letterbox {
|
||
fn fit(w: f32, h: f32) -> Self {
|
||
let scale = (INPUT_EDGE as f32 / w).min(INPUT_EDGE as f32 / h);
|
||
Self {
|
||
scale,
|
||
pad_x: (INPUT_EDGE as f32 - w * scale) * 0.5,
|
||
pad_y: (INPUT_EDGE as f32 - h * scale) * 0.5,
|
||
}
|
||
}
|
||
|
||
/// Resample into `[1, 3, 640, 640]`, normalised as the weights expect.
|
||
///
|
||
/// `(x·255 − 127.5) / 128` — note `/128`, not `/127.5`. The reference
|
||
/// implementation this is ported from uses `/128` for both models, and
|
||
/// every measured number in docs/dev/faces.md §1 came from it.
|
||
///
|
||
/// Padding is grey, matching the reference's `114`: the value the network
|
||
/// reads least as an edge, where black would draw a hard border across the
|
||
/// frame and invite a detection along it.
|
||
fn sample(&self, rgb: &[f32], width: usize, height: usize) -> Array4<f32> {
|
||
const PAD: f32 = 114.0;
|
||
let norm = |v: f32| (v * 255.0 - 127.5) / 128.0;
|
||
|
||
let mut input =
|
||
Array4::<f32>::from_elem((1, 3, INPUT_EDGE, INPUT_EDGE), (PAD - 127.5) / 128.0);
|
||
|
||
for iy in 0..INPUT_EDGE {
|
||
let sy = (iy as f32 + 0.5 - self.pad_y) / self.scale - 0.5;
|
||
if sy < -0.5 || sy > height as f32 - 0.5 {
|
||
continue;
|
||
}
|
||
for ix in 0..INPUT_EDGE {
|
||
let sx = (ix as f32 + 0.5 - self.pad_x) / self.scale - 0.5;
|
||
if sx < -0.5 || sx > width as f32 - 0.5 {
|
||
continue;
|
||
}
|
||
let (x0f, y0f) = (sx.floor(), sy.floor());
|
||
let (fx, fy) = (sx - x0f, sy - y0f);
|
||
let x0 = (x0f as isize).clamp(0, width as isize - 1) as usize;
|
||
let y0 = (y0f as isize).clamp(0, height as isize - 1) as usize;
|
||
let x1 = (x0 + 1).min(width - 1);
|
||
let y1 = (y0 + 1).min(height - 1);
|
||
|
||
for c in 0..3 {
|
||
let at = |x: usize, y: usize| rgb[(y * width + x) * 3 + c];
|
||
let top = at(x0, y0) * (1.0 - fx) + at(x1, y0) * fx;
|
||
let bot = at(x0, y1) * (1.0 - fx) + at(x1, y1) * fx;
|
||
input[[0, c, iy, ix]] = norm(top * (1.0 - fy) + bot * fy);
|
||
}
|
||
}
|
||
}
|
||
|
||
input
|
||
}
|
||
|
||
/// Input-space point back to source pixels.
|
||
fn into_source(self, x: f32, y: f32) -> (f32, f32) {
|
||
((x - self.pad_x) / self.scale, (y - self.pad_y) / self.scale)
|
||
}
|
||
}
|
||
|
||
#[cfg(test)]
|
||
mod tests {
|
||
use super::*;
|
||
|
||
#[test]
|
||
fn letterbox_round_trips_a_point() {
|
||
let lb = Letterbox::fit(1024.0, 683.0);
|
||
for &(x, y) in &[(0.0_f32, 0.0_f32), (512.0, 341.0), (1023.0, 682.0)] {
|
||
let (bx, by) = lb.into_source(x * lb.scale + lb.pad_x, y * lb.scale + lb.pad_y);
|
||
assert!((bx - x).abs() < 1e-2, "{bx} vs {x}");
|
||
assert!((by - y).abs() < 1e-2, "{by} vs {y}");
|
||
}
|
||
}
|
||
|
||
#[test]
|
||
fn letterbox_centres_the_short_axis() {
|
||
let lb = Letterbox::fit(640.0, 320.0);
|
||
assert!((lb.scale - 1.0).abs() < 1e-6);
|
||
assert!(lb.pad_x.abs() < 1e-6);
|
||
assert!((lb.pad_y - 160.0).abs() < 1e-6);
|
||
}
|
||
|
||
#[test]
|
||
fn nms_keeps_the_confident_box_and_drops_its_duplicate() {
|
||
let d = |x: f32, conf: f32| Detection {
|
||
bbox: (x, 0.0, x + 100.0, 100.0),
|
||
landmarks: [(0.0, 0.0); 5],
|
||
confidence: conf,
|
||
};
|
||
let kept = non_max_suppress(vec![d(0.0, 0.8), d(5.0, 0.9), d(500.0, 0.7)], 0.4);
|
||
assert_eq!(kept.len(), 2);
|
||
assert!((kept[0].confidence - 0.9).abs() < 1e-6);
|
||
assert!((kept[1].bbox.0 - 500.0).abs() < 1e-6);
|
||
}
|
||
|
||
#[test]
|
||
fn iou_of_a_box_with_itself_is_one_and_with_a_disjoint_box_is_zero() {
|
||
let a = (0.0, 0.0, 10.0, 10.0);
|
||
assert!((iou(&a, &a) - 1.0).abs() < 1e-6);
|
||
assert!(iou(&a, &(100.0, 100.0, 110.0, 110.0)) < 1e-6);
|
||
}
|
||
}
|