Detect, align and embed faces with SCRFD and MobileFaceNet

Ports the pipeline from the C++ reference in ../scene-actor-extraction
(MIT, same author). End to end on real portraits it separates identities
the way the reference's fitted calibration says it should: 0.596 between
distinct photographs of one person, 0.05 between different people, either
side of MBF's 0.267 boundary.

Three things are structural rather than incidental:

Aligned112 can only be built by align::warp, so Embedder::embed cannot be
handed an unaligned bounding-box crop. That mistake yields 512 plausible
unit-norm numbers and no error, so the type system refuses it instead.

Embedding carries its ModelId and cosine() returns None across models,
because a cross-model similarity is the one mistake that produces
plausible garbage rather than a failure.

The model-free half -- alignment, embedding arithmetic, f16 storage --
sits outside the inference feature and is covered by 11 tests that need
no weights on the machine.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
2026-08-26 19:57:56 +02:00
co-authored by Claude Opus 5
parent 72410f39c6
commit 19981c1033
9 changed files with 1196 additions and 32 deletions
+360
View File
@@ -0,0 +1,360 @@
//! Five-point face alignment (docs/faces.md §5).
//!
//! ArcFace embeddings are trained on faces warped to a canonical 112×112
//! arrangement. Feeding the model a plain bounding-box crop *works* — it
//! produces 512 numbers, they are unit-norm, and cosine similarities between
//! them look entirely reasonable. They are just much worse, and nothing in the
//! system reports it.
//!
//! That is the whole reason this module exists, and the reason [`Aligned112`]
//! is a newtype only [`warp`] can construct: the mistake is not one a reviewer
//! catches, so the type system catches it instead.
//!
//! Model-free, so it builds and tests without the `inference` feature.
/// Canonical landmark positions for a 112×112 ArcFace crop.
///
/// # The naming is a trap; the order is not
///
/// Point 0 sits at x=38 on a 112-wide canvas — left of centre *in the image*,
/// which is the subject's **right** eye. Both namings are in circulation and
/// they are opposite, so the array is written in the detector's order and the
/// comment says whose left is whose:
///
/// ```text
/// 0 subject's right eye (image-left)
/// 1 subject's left eye (image-right)
/// 2 nose tip
/// 3 subject's right mouth corner
/// 4 subject's left mouth corner
/// ```
///
/// SCRFD emits its five points in this same order, so the correct amount of
/// reordering between detector and template is **none**. A detector with a
/// different order carries its own permutation beside its model id rather than
/// this constant growing an assumption.
pub const ARCFACE_TEMPLATE: [(f32, f32); 5] = [
(38.2946, 51.6963),
(73.5318, 51.5014),
(56.0252, 71.7366),
(41.5493, 92.3655),
(70.7299, 92.2041),
];
/// Edge of the aligned crop, in pixels. Fixed by the embedder's input.
pub const ALIGNED_EDGE: usize = 112;
/// A face warped to [`ARCFACE_TEMPLATE`], ready for the embedder.
///
/// Constructible only by [`warp`]. That is the point: an `Embedder` that took
/// a plain `&[f32]` would accept an unaligned bounding-box crop and silently
/// return worse embeddings, which is a failure no test of the embedder itself
/// would catch.
pub struct Aligned112 {
/// `112 × 112 × 3`, row-major RGB in `0.0..=1.0`.
pixels: Vec<f32>,
/// Source pixels across the crop before warping — `crop_px` in the catalog.
///
/// Carried here rather than recomputed later because the scale factor is
/// known exactly at warp time and only approximately from the box
/// afterwards. §7: it is the honest quality signal, and a feature in the
/// calibration.
source_px: f32,
}
impl Aligned112 {
pub fn pixels(&self) -> &[f32] {
&self.pixels
}
/// Source pixels spanned by the 112-pixel crop.
///
/// Below ~112 the face was upsampled to reach the embedder and the
/// embedding is correspondingly weaker; above it, downsampled and healthy.
pub fn source_px(&self) -> f32 {
self.source_px
}
}
/// A similarity transform: rotation, uniform scale, translation.
///
/// Stored as the four independent parameters rather than a 2×3 matrix so that
/// [`Similarity::scale`] is readable without a decomposition.
#[derive(Debug, Clone, Copy, PartialEq)]
pub struct Similarity {
a: f32,
b: f32,
tx: f32,
ty: f32,
}
impl Similarity {
/// `x' = a·x − b·y + tx`, `y' = b·x + a·y + ty`.
pub fn apply(&self, x: f32, y: f32) -> (f32, f32) {
(
self.a * x - self.b * y + self.tx,
self.b * x + self.a * y + self.ty,
)
}
/// Uniform scale factor — destination pixels per source pixel.
pub fn scale(&self) -> f32 {
(self.a * self.a + self.b * self.b).sqrt()
}
fn invert(&self, u: f32, v: f32) -> (f32, f32) {
let det = self.a * self.a + self.b * self.b;
let du = u - self.tx;
let dv = v - self.ty;
((self.a * du + self.b * dv) / det, (-self.b * du + self.a * dv) / det)
}
}
/// Least-squares similarity transform from `src` onto `dst`.
///
/// # Why least squares and not RANSAC
///
/// The reference C++ implementation (docs/faces.md §1.1) fits this with
/// OpenCV's `estimateAffinePartial2D` under RANSAC. RANSAC over five points is
/// a strange fit: the minimal sample for a similarity is two, so it can discard
/// landmarks it judges outliers and solve from a subset — and on a profile face
/// the "outlier" is as likely to be the correct geometry as the wrong one.
/// InsightFace's own pipeline uses plain least squares over all five points,
/// which cannot silently drop anything, and that is what this is.
///
/// # The closed form
///
/// A 2-D similarity is linear in its four parameters:
///
/// ```text
/// x' = a·x − b·y + tx
/// y' = b·x + a·y + ty
/// ```
///
/// so this is an ordinary linear least-squares problem, not an SVD one.
/// Centring both point sets kills `tx`/`ty` from the normal equations and
/// leaves `a` and `b` as two dot products over a common denominator — which is
/// why there is no matrix decomposition anywhere in this function.
///
/// Returns `None` when the source points are degenerate (coincident or
/// collinear to within f32), which does happen: a detector firing on a
/// motion-blurred profile can put all five landmarks on a line.
pub fn fit_similarity(src: &[(f32, f32); 5], dst: &[(f32, f32); 5]) -> Option<Similarity> {
let n = 5.0_f32;
let (mut sx, mut sy, mut dx, mut dy) = (0.0, 0.0, 0.0, 0.0);
for i in 0..5 {
sx += src[i].0;
sy += src[i].1;
dx += dst[i].0;
dy += dst[i].1;
}
let (sx, sy, dx, dy) = (sx / n, sy / n, dx / n, dy / n);
let mut var = 0.0_f32;
let mut num_a = 0.0_f32;
let mut num_b = 0.0_f32;
for i in 0..5 {
let (px, py) = (src[i].0 - sx, src[i].1 - sy);
let (qx, qy) = (dst[i].0 - dx, dst[i].1 - dy);
var += px * px + py * py;
num_a += px * qx + py * qy;
num_b += px * qy - py * qx;
}
// Degenerate: every landmark on one point. Collinear input still solves,
// but with a scale that can be absurd, so the caller's sanity check on
// `scale()` is what catches that case.
if var <= f32::EPSILON {
return None;
}
let a = num_a / var;
let b = num_b / var;
if !a.is_finite() || !b.is_finite() || (a * a + b * b) <= f32::EPSILON {
return None;
}
Some(Similarity {
a,
b,
tx: dx - (a * sx - b * sy),
ty: dy - (b * sx + a * sy),
})
}
/// Warp a face onto the canonical 112×112 arrangement.
///
/// `rgb` is tightly packed `f32` RGB in `0.0..=1.0`, row-major — the same
/// convention `dr-segment` uses, so both read the same proxy.
///
/// Sampling is bilinear **from the source in one step**: never crop-then-warp,
/// which resamples twice and throws away detail the warp could have used.
/// Pixels falling outside the source read as black.
pub fn warp(
rgb: &[f32],
width: usize,
height: usize,
landmarks: &[(f32, f32); 5],
) -> Option<Aligned112> {
if rgb.len() != width * height * 3 {
return None;
}
let m = fit_similarity(landmarks, &ARCFACE_TEMPLATE)?;
let e = ALIGNED_EDGE;
let mut pixels = vec![0.0_f32; e * e * 3];
for v in 0..e {
for u in 0..e {
// Pixel centres, so the transform is not off by half a pixel —
// which is small enough to survive review and large enough to
// matter on a 40-pixel face.
let (x, y) = m.invert(u as f32 + 0.5, v as f32 + 0.5);
let (x, y) = (x - 0.5, y - 0.5);
let out = (v * e + u) * 3;
sample_bilinear(rgb, width, height, x, y, &mut pixels[out..out + 3]);
}
}
Some(Aligned112 {
pixels,
// The warp maps `scale` source pixels to one destination pixel, so the
// crop spans 112/scale of the source.
source_px: ALIGNED_EDGE as f32 / m.scale(),
})
}
fn sample_bilinear(rgb: &[f32], w: usize, h: usize, x: f32, y: f32, out: &mut [f32]) {
let x0 = x.floor();
let y0 = y.floor();
let fx = x - x0;
let fy = y - y0;
let x0 = x0 as isize;
let y0 = y0 as isize;
for (c, o) in out.iter_mut().enumerate() {
let get = |xi: isize, yi: isize| -> f32 {
if xi < 0 || yi < 0 || xi >= w as isize || yi >= h as isize {
0.0
} else {
rgb[(yi as usize * w + xi as usize) * 3 + c]
}
};
let top = get(x0, y0) * (1.0 - fx) + get(x0 + 1, y0) * fx;
let bot = get(x0, y0 + 1) * (1.0 - fx) + get(x0 + 1, y0 + 1) * fx;
*o = top * (1.0 - fy) + bot * fy;
}
}
#[cfg(test)]
mod tests {
use super::*;
fn shifted_scaled(scale: f32, dx: f32, dy: f32, rot: f32) -> [(f32, f32); 5] {
let (s, c) = (rot.sin(), rot.cos());
let mut out = [(0.0, 0.0); 5];
for (i, &(x, y)) in ARCFACE_TEMPLATE.iter().enumerate() {
out[i] = (
scale * (c * x - s * y) + dx,
scale * (s * x + c * y) + dy,
);
}
out
}
#[test]
fn template_onto_itself_is_the_identity() {
let m = fit_similarity(&ARCFACE_TEMPLATE, &ARCFACE_TEMPLATE).unwrap();
for &(x, y) in &ARCFACE_TEMPLATE {
let (u, v) = m.apply(x, y);
assert!((u - x).abs() < 1e-3, "{u} vs {x}");
assert!((v - y).abs() < 1e-3, "{v} vs {y}");
}
assert!((m.scale() - 1.0).abs() < 1e-4);
}
/// The property that matters: whatever similarity the face was seen under,
/// the fit must undo it and land the landmarks back on the template. This
/// is the test that fails if the transform is ever "simplified" into an
/// affine or a bare scale-and-translate.
#[test]
fn any_similarity_of_the_template_maps_back_onto_it() {
for &(scale, dx, dy, rot) in &[
(1.0_f32, 0.0_f32, 0.0_f32, 0.0_f32),
(2.5, 100.0, -40.0, 0.0),
(0.4, -12.0, 300.0, 0.6),
(1.7, 5.0, 5.0, -1.2),
] {
let observed = shifted_scaled(scale, dx, dy, rot);
let m = fit_similarity(&observed, &ARCFACE_TEMPLATE).unwrap();
for (i, &(tx, ty)) in ARCFACE_TEMPLATE.iter().enumerate() {
let (u, v) = m.apply(observed[i].0, observed[i].1);
assert!(
(u - tx).abs() < 1e-2 && (v - ty).abs() < 1e-2,
"scale={scale} rot={rot}: point {i} landed at ({u}, {v}), want ({tx}, {ty})"
);
}
assert!(
(m.scale() - 1.0 / scale).abs() < 1e-3,
"scale {} should invert {scale}",
m.scale()
);
}
}
#[test]
fn coincident_landmarks_are_rejected_rather_than_producing_a_crop() {
let degenerate = [(50.0, 50.0); 5];
assert!(fit_similarity(&degenerate, &ARCFACE_TEMPLATE).is_none());
let rgb = vec![0.5_f32; 64 * 64 * 3];
assert!(warp(&rgb, 64, 64, &degenerate).is_none());
}
#[test]
fn source_px_reports_the_face_size_the_embedder_actually_saw() {
let rgb = vec![0.5_f32; 400 * 400 * 3];
// A face twice the template's size spans 224 source pixels.
let big = shifted_scaled(2.0, 80.0, 80.0, 0.0);
let a = warp(&rgb, 400, 400, &big).unwrap();
assert!((a.source_px() - 224.0).abs() < 0.5, "{}", a.source_px());
// Half-size: 56 source pixels upsampled to 112, which §7 calls the
// degraded bucket.
let small = shifted_scaled(0.5, 10.0, 10.0, 0.0);
let a = warp(&rgb, 400, 400, &small).unwrap();
assert!((a.source_px() - 56.0).abs() < 0.5, "{}", a.source_px());
}
/// A white square on black, warped by a transform that should centre it:
/// checks the sampler's geometry rather than the fit's algebra.
#[test]
fn warp_resamples_the_right_pixels() {
let (w, h) = (224, 224);
let mut rgb = vec![0.0_f32; w * h * 3];
for y in 0..h {
for x in 0..w {
if x >= 56 && x < 168 && y >= 56 && y < 168 {
for c in 0..3 {
rgb[(y * w + x) * 3 + c] = 1.0;
}
}
}
}
// Landmarks placed so the fit is a pure translation of (56, 56):
// the white square maps exactly onto the 112×112 output.
let lm = shifted_scaled(1.0, 56.0, 56.0, 0.0);
let a = warp(&rgb, w, h, &lm).unwrap();
let px = a.pixels();
for (i, v) in px.iter().enumerate() {
assert!((v - 1.0).abs() < 1e-3, "pixel {i} is {v}, expected white");
}
}
#[test]
fn out_of_bounds_samples_read_black_rather_than_wrapping() {
let rgb = vec![1.0_f32; 32 * 32 * 3];
// Face far outside the image: every sample is out of bounds.
let lm = shifted_scaled(1.0, 5000.0, 5000.0, 0.0);
let a = warp(&rgb, 32, 32, &lm).unwrap();
assert!(a.pixels().iter().all(|&v| v == 0.0));
}
}