Detect, align and embed faces with SCRFD and MobileFaceNet
Ports the pipeline from the C++ reference in ../scene-actor-extraction (MIT, same author). End to end on real portraits it separates identities the way the reference's fitted calibration says it should: 0.596 between distinct photographs of one person, 0.05 between different people, either side of MBF's 0.267 boundary. Three things are structural rather than incidental: Aligned112 can only be built by align::warp, so Embedder::embed cannot be handed an unaligned bounding-box crop. That mistake yields 512 plausible unit-norm numbers and no error, so the type system refuses it instead. Embedding carries its ModelId and cosine() returns None across models, because a cross-model similarity is the one mistake that produces plausible garbage rather than a failure. The model-free half -- alignment, embedding arithmetic, f16 storage -- sits outside the inference feature and is covered by 11 tests that need no weights on the machine. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
@@ -0,0 +1,360 @@
|
||||
//! Five-point face alignment (docs/faces.md §5).
|
||||
//!
|
||||
//! ArcFace embeddings are trained on faces warped to a canonical 112×112
|
||||
//! arrangement. Feeding the model a plain bounding-box crop *works* — it
|
||||
//! produces 512 numbers, they are unit-norm, and cosine similarities between
|
||||
//! them look entirely reasonable. They are just much worse, and nothing in the
|
||||
//! system reports it.
|
||||
//!
|
||||
//! That is the whole reason this module exists, and the reason [`Aligned112`]
|
||||
//! is a newtype only [`warp`] can construct: the mistake is not one a reviewer
|
||||
//! catches, so the type system catches it instead.
|
||||
//!
|
||||
//! Model-free, so it builds and tests without the `inference` feature.
|
||||
|
||||
/// Canonical landmark positions for a 112×112 ArcFace crop.
|
||||
///
|
||||
/// # The naming is a trap; the order is not
|
||||
///
|
||||
/// Point 0 sits at x=38 on a 112-wide canvas — left of centre *in the image*,
|
||||
/// which is the subject's **right** eye. Both namings are in circulation and
|
||||
/// they are opposite, so the array is written in the detector's order and the
|
||||
/// comment says whose left is whose:
|
||||
///
|
||||
/// ```text
|
||||
/// 0 subject's right eye (image-left)
|
||||
/// 1 subject's left eye (image-right)
|
||||
/// 2 nose tip
|
||||
/// 3 subject's right mouth corner
|
||||
/// 4 subject's left mouth corner
|
||||
/// ```
|
||||
///
|
||||
/// SCRFD emits its five points in this same order, so the correct amount of
|
||||
/// reordering between detector and template is **none**. A detector with a
|
||||
/// different order carries its own permutation beside its model id rather than
|
||||
/// this constant growing an assumption.
|
||||
pub const ARCFACE_TEMPLATE: [(f32, f32); 5] = [
|
||||
(38.2946, 51.6963),
|
||||
(73.5318, 51.5014),
|
||||
(56.0252, 71.7366),
|
||||
(41.5493, 92.3655),
|
||||
(70.7299, 92.2041),
|
||||
];
|
||||
|
||||
/// Edge of the aligned crop, in pixels. Fixed by the embedder's input.
|
||||
pub const ALIGNED_EDGE: usize = 112;
|
||||
|
||||
/// A face warped to [`ARCFACE_TEMPLATE`], ready for the embedder.
|
||||
///
|
||||
/// Constructible only by [`warp`]. That is the point: an `Embedder` that took
|
||||
/// a plain `&[f32]` would accept an unaligned bounding-box crop and silently
|
||||
/// return worse embeddings, which is a failure no test of the embedder itself
|
||||
/// would catch.
|
||||
pub struct Aligned112 {
|
||||
/// `112 × 112 × 3`, row-major RGB in `0.0..=1.0`.
|
||||
pixels: Vec<f32>,
|
||||
/// Source pixels across the crop before warping — `crop_px` in the catalog.
|
||||
///
|
||||
/// Carried here rather than recomputed later because the scale factor is
|
||||
/// known exactly at warp time and only approximately from the box
|
||||
/// afterwards. §7: it is the honest quality signal, and a feature in the
|
||||
/// calibration.
|
||||
source_px: f32,
|
||||
}
|
||||
|
||||
impl Aligned112 {
|
||||
pub fn pixels(&self) -> &[f32] {
|
||||
&self.pixels
|
||||
}
|
||||
|
||||
/// Source pixels spanned by the 112-pixel crop.
|
||||
///
|
||||
/// Below ~112 the face was upsampled to reach the embedder and the
|
||||
/// embedding is correspondingly weaker; above it, downsampled and healthy.
|
||||
pub fn source_px(&self) -> f32 {
|
||||
self.source_px
|
||||
}
|
||||
}
|
||||
|
||||
/// A similarity transform: rotation, uniform scale, translation.
|
||||
///
|
||||
/// Stored as the four independent parameters rather than a 2×3 matrix so that
|
||||
/// [`Similarity::scale`] is readable without a decomposition.
|
||||
#[derive(Debug, Clone, Copy, PartialEq)]
|
||||
pub struct Similarity {
|
||||
a: f32,
|
||||
b: f32,
|
||||
tx: f32,
|
||||
ty: f32,
|
||||
}
|
||||
|
||||
impl Similarity {
|
||||
/// `x' = a·x − b·y + tx`, `y' = b·x + a·y + ty`.
|
||||
pub fn apply(&self, x: f32, y: f32) -> (f32, f32) {
|
||||
(
|
||||
self.a * x - self.b * y + self.tx,
|
||||
self.b * x + self.a * y + self.ty,
|
||||
)
|
||||
}
|
||||
|
||||
/// Uniform scale factor — destination pixels per source pixel.
|
||||
pub fn scale(&self) -> f32 {
|
||||
(self.a * self.a + self.b * self.b).sqrt()
|
||||
}
|
||||
|
||||
fn invert(&self, u: f32, v: f32) -> (f32, f32) {
|
||||
let det = self.a * self.a + self.b * self.b;
|
||||
let du = u - self.tx;
|
||||
let dv = v - self.ty;
|
||||
((self.a * du + self.b * dv) / det, (-self.b * du + self.a * dv) / det)
|
||||
}
|
||||
}
|
||||
|
||||
/// Least-squares similarity transform from `src` onto `dst`.
|
||||
///
|
||||
/// # Why least squares and not RANSAC
|
||||
///
|
||||
/// The reference C++ implementation (docs/faces.md §1.1) fits this with
|
||||
/// OpenCV's `estimateAffinePartial2D` under RANSAC. RANSAC over five points is
|
||||
/// a strange fit: the minimal sample for a similarity is two, so it can discard
|
||||
/// landmarks it judges outliers and solve from a subset — and on a profile face
|
||||
/// the "outlier" is as likely to be the correct geometry as the wrong one.
|
||||
/// InsightFace's own pipeline uses plain least squares over all five points,
|
||||
/// which cannot silently drop anything, and that is what this is.
|
||||
///
|
||||
/// # The closed form
|
||||
///
|
||||
/// A 2-D similarity is linear in its four parameters:
|
||||
///
|
||||
/// ```text
|
||||
/// x' = a·x − b·y + tx
|
||||
/// y' = b·x + a·y + ty
|
||||
/// ```
|
||||
///
|
||||
/// so this is an ordinary linear least-squares problem, not an SVD one.
|
||||
/// Centring both point sets kills `tx`/`ty` from the normal equations and
|
||||
/// leaves `a` and `b` as two dot products over a common denominator — which is
|
||||
/// why there is no matrix decomposition anywhere in this function.
|
||||
///
|
||||
/// Returns `None` when the source points are degenerate (coincident or
|
||||
/// collinear to within f32), which does happen: a detector firing on a
|
||||
/// motion-blurred profile can put all five landmarks on a line.
|
||||
pub fn fit_similarity(src: &[(f32, f32); 5], dst: &[(f32, f32); 5]) -> Option<Similarity> {
|
||||
let n = 5.0_f32;
|
||||
let (mut sx, mut sy, mut dx, mut dy) = (0.0, 0.0, 0.0, 0.0);
|
||||
for i in 0..5 {
|
||||
sx += src[i].0;
|
||||
sy += src[i].1;
|
||||
dx += dst[i].0;
|
||||
dy += dst[i].1;
|
||||
}
|
||||
let (sx, sy, dx, dy) = (sx / n, sy / n, dx / n, dy / n);
|
||||
|
||||
let mut var = 0.0_f32;
|
||||
let mut num_a = 0.0_f32;
|
||||
let mut num_b = 0.0_f32;
|
||||
for i in 0..5 {
|
||||
let (px, py) = (src[i].0 - sx, src[i].1 - sy);
|
||||
let (qx, qy) = (dst[i].0 - dx, dst[i].1 - dy);
|
||||
var += px * px + py * py;
|
||||
num_a += px * qx + py * qy;
|
||||
num_b += px * qy - py * qx;
|
||||
}
|
||||
|
||||
// Degenerate: every landmark on one point. Collinear input still solves,
|
||||
// but with a scale that can be absurd, so the caller's sanity check on
|
||||
// `scale()` is what catches that case.
|
||||
if var <= f32::EPSILON {
|
||||
return None;
|
||||
}
|
||||
|
||||
let a = num_a / var;
|
||||
let b = num_b / var;
|
||||
if !a.is_finite() || !b.is_finite() || (a * a + b * b) <= f32::EPSILON {
|
||||
return None;
|
||||
}
|
||||
|
||||
Some(Similarity {
|
||||
a,
|
||||
b,
|
||||
tx: dx - (a * sx - b * sy),
|
||||
ty: dy - (b * sx + a * sy),
|
||||
})
|
||||
}
|
||||
|
||||
/// Warp a face onto the canonical 112×112 arrangement.
|
||||
///
|
||||
/// `rgb` is tightly packed `f32` RGB in `0.0..=1.0`, row-major — the same
|
||||
/// convention `dr-segment` uses, so both read the same proxy.
|
||||
///
|
||||
/// Sampling is bilinear **from the source in one step**: never crop-then-warp,
|
||||
/// which resamples twice and throws away detail the warp could have used.
|
||||
/// Pixels falling outside the source read as black.
|
||||
pub fn warp(
|
||||
rgb: &[f32],
|
||||
width: usize,
|
||||
height: usize,
|
||||
landmarks: &[(f32, f32); 5],
|
||||
) -> Option<Aligned112> {
|
||||
if rgb.len() != width * height * 3 {
|
||||
return None;
|
||||
}
|
||||
let m = fit_similarity(landmarks, &ARCFACE_TEMPLATE)?;
|
||||
|
||||
let e = ALIGNED_EDGE;
|
||||
let mut pixels = vec![0.0_f32; e * e * 3];
|
||||
for v in 0..e {
|
||||
for u in 0..e {
|
||||
// Pixel centres, so the transform is not off by half a pixel —
|
||||
// which is small enough to survive review and large enough to
|
||||
// matter on a 40-pixel face.
|
||||
let (x, y) = m.invert(u as f32 + 0.5, v as f32 + 0.5);
|
||||
let (x, y) = (x - 0.5, y - 0.5);
|
||||
let out = (v * e + u) * 3;
|
||||
sample_bilinear(rgb, width, height, x, y, &mut pixels[out..out + 3]);
|
||||
}
|
||||
}
|
||||
|
||||
Some(Aligned112 {
|
||||
pixels,
|
||||
// The warp maps `scale` source pixels to one destination pixel, so the
|
||||
// crop spans 112/scale of the source.
|
||||
source_px: ALIGNED_EDGE as f32 / m.scale(),
|
||||
})
|
||||
}
|
||||
|
||||
fn sample_bilinear(rgb: &[f32], w: usize, h: usize, x: f32, y: f32, out: &mut [f32]) {
|
||||
let x0 = x.floor();
|
||||
let y0 = y.floor();
|
||||
let fx = x - x0;
|
||||
let fy = y - y0;
|
||||
let x0 = x0 as isize;
|
||||
let y0 = y0 as isize;
|
||||
|
||||
for (c, o) in out.iter_mut().enumerate() {
|
||||
let get = |xi: isize, yi: isize| -> f32 {
|
||||
if xi < 0 || yi < 0 || xi >= w as isize || yi >= h as isize {
|
||||
0.0
|
||||
} else {
|
||||
rgb[(yi as usize * w + xi as usize) * 3 + c]
|
||||
}
|
||||
};
|
||||
let top = get(x0, y0) * (1.0 - fx) + get(x0 + 1, y0) * fx;
|
||||
let bot = get(x0, y0 + 1) * (1.0 - fx) + get(x0 + 1, y0 + 1) * fx;
|
||||
*o = top * (1.0 - fy) + bot * fy;
|
||||
}
|
||||
}
|
||||
|
||||
#[cfg(test)]
|
||||
mod tests {
|
||||
use super::*;
|
||||
|
||||
fn shifted_scaled(scale: f32, dx: f32, dy: f32, rot: f32) -> [(f32, f32); 5] {
|
||||
let (s, c) = (rot.sin(), rot.cos());
|
||||
let mut out = [(0.0, 0.0); 5];
|
||||
for (i, &(x, y)) in ARCFACE_TEMPLATE.iter().enumerate() {
|
||||
out[i] = (
|
||||
scale * (c * x - s * y) + dx,
|
||||
scale * (s * x + c * y) + dy,
|
||||
);
|
||||
}
|
||||
out
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn template_onto_itself_is_the_identity() {
|
||||
let m = fit_similarity(&ARCFACE_TEMPLATE, &ARCFACE_TEMPLATE).unwrap();
|
||||
for &(x, y) in &ARCFACE_TEMPLATE {
|
||||
let (u, v) = m.apply(x, y);
|
||||
assert!((u - x).abs() < 1e-3, "{u} vs {x}");
|
||||
assert!((v - y).abs() < 1e-3, "{v} vs {y}");
|
||||
}
|
||||
assert!((m.scale() - 1.0).abs() < 1e-4);
|
||||
}
|
||||
|
||||
/// The property that matters: whatever similarity the face was seen under,
|
||||
/// the fit must undo it and land the landmarks back on the template. This
|
||||
/// is the test that fails if the transform is ever "simplified" into an
|
||||
/// affine or a bare scale-and-translate.
|
||||
#[test]
|
||||
fn any_similarity_of_the_template_maps_back_onto_it() {
|
||||
for &(scale, dx, dy, rot) in &[
|
||||
(1.0_f32, 0.0_f32, 0.0_f32, 0.0_f32),
|
||||
(2.5, 100.0, -40.0, 0.0),
|
||||
(0.4, -12.0, 300.0, 0.6),
|
||||
(1.7, 5.0, 5.0, -1.2),
|
||||
] {
|
||||
let observed = shifted_scaled(scale, dx, dy, rot);
|
||||
let m = fit_similarity(&observed, &ARCFACE_TEMPLATE).unwrap();
|
||||
for (i, &(tx, ty)) in ARCFACE_TEMPLATE.iter().enumerate() {
|
||||
let (u, v) = m.apply(observed[i].0, observed[i].1);
|
||||
assert!(
|
||||
(u - tx).abs() < 1e-2 && (v - ty).abs() < 1e-2,
|
||||
"scale={scale} rot={rot}: point {i} landed at ({u}, {v}), want ({tx}, {ty})"
|
||||
);
|
||||
}
|
||||
assert!(
|
||||
(m.scale() - 1.0 / scale).abs() < 1e-3,
|
||||
"scale {} should invert {scale}",
|
||||
m.scale()
|
||||
);
|
||||
}
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn coincident_landmarks_are_rejected_rather_than_producing_a_crop() {
|
||||
let degenerate = [(50.0, 50.0); 5];
|
||||
assert!(fit_similarity(°enerate, &ARCFACE_TEMPLATE).is_none());
|
||||
let rgb = vec![0.5_f32; 64 * 64 * 3];
|
||||
assert!(warp(&rgb, 64, 64, °enerate).is_none());
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn source_px_reports_the_face_size_the_embedder_actually_saw() {
|
||||
let rgb = vec![0.5_f32; 400 * 400 * 3];
|
||||
// A face twice the template's size spans 224 source pixels.
|
||||
let big = shifted_scaled(2.0, 80.0, 80.0, 0.0);
|
||||
let a = warp(&rgb, 400, 400, &big).unwrap();
|
||||
assert!((a.source_px() - 224.0).abs() < 0.5, "{}", a.source_px());
|
||||
|
||||
// Half-size: 56 source pixels upsampled to 112, which §7 calls the
|
||||
// degraded bucket.
|
||||
let small = shifted_scaled(0.5, 10.0, 10.0, 0.0);
|
||||
let a = warp(&rgb, 400, 400, &small).unwrap();
|
||||
assert!((a.source_px() - 56.0).abs() < 0.5, "{}", a.source_px());
|
||||
}
|
||||
|
||||
/// A white square on black, warped by a transform that should centre it:
|
||||
/// checks the sampler's geometry rather than the fit's algebra.
|
||||
#[test]
|
||||
fn warp_resamples_the_right_pixels() {
|
||||
let (w, h) = (224, 224);
|
||||
let mut rgb = vec![0.0_f32; w * h * 3];
|
||||
for y in 0..h {
|
||||
for x in 0..w {
|
||||
if x >= 56 && x < 168 && y >= 56 && y < 168 {
|
||||
for c in 0..3 {
|
||||
rgb[(y * w + x) * 3 + c] = 1.0;
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
// Landmarks placed so the fit is a pure translation of (56, 56):
|
||||
// the white square maps exactly onto the 112×112 output.
|
||||
let lm = shifted_scaled(1.0, 56.0, 56.0, 0.0);
|
||||
let a = warp(&rgb, w, h, &lm).unwrap();
|
||||
let px = a.pixels();
|
||||
for (i, v) in px.iter().enumerate() {
|
||||
assert!((v - 1.0).abs() < 1e-3, "pixel {i} is {v}, expected white");
|
||||
}
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn out_of_bounds_samples_read_black_rather_than_wrapping() {
|
||||
let rgb = vec![1.0_f32; 32 * 32 * 3];
|
||||
// Face far outside the image: every sample is out of bounds.
|
||||
let lm = shifted_scaled(1.0, 5000.0, 5000.0, 0.0);
|
||||
let a = warp(&rgb, 32, 32, &lm).unwrap();
|
||||
assert!(a.pixels().iter().all(|&v| v == 0.0));
|
||||
}
|
||||
}
|
||||
Reference in New Issue
Block a user