Detect, align and embed faces with SCRFD and MobileFaceNet

Ports the pipeline from the C++ reference in ../scene-actor-extraction
(MIT, same author). End to end on real portraits it separates identities
the way the reference's fitted calibration says it should: 0.596 between
distinct photographs of one person, 0.05 between different people, either
side of MBF's 0.267 boundary.

Three things are structural rather than incidental:

Aligned112 can only be built by align::warp, so Embedder::embed cannot be
handed an unaligned bounding-box crop. That mistake yields 512 plausible
unit-norm numbers and no error, so the type system refuses it instead.

Embedding carries its ModelId and cosine() returns None across models,
because a cross-model similarity is the one mistake that produces
plausible garbage rather than a failure.

The model-free half -- alignment, embedding arithmetic, f16 storage --
sits outside the inference feature and is covered by 11 tests that need
no weights on the machine.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
2026-08-26 19:57:56 +02:00
co-authored by Claude Opus 5
parent 72410f39c6
commit 19981c1033
9 changed files with 1196 additions and 32 deletions
+11
View File
@@ -30,8 +30,19 @@
//! machine with no weights on it — which is what lets CI cover the part most
//! likely to be subtly wrong.
pub mod align;
pub mod embedding;
#[cfg(feature = "inference")]
pub mod detect;
#[cfg(feature = "inference")]
pub mod embed;
pub use align::{warp, Aligned112, Similarity, ALIGNED_EDGE, ARCFACE_TEMPLATE};
pub use embedding::{Embedding, ModelId, EMBEDDING_DIM};
#[cfg(feature = "inference")]
pub use detect::{DetectOptions, Detection, Detector};
#[cfg(feature = "inference")]
pub use embed::Embedder;
/// What can go wrong between an image and a face.
#[derive(Debug, thiserror::Error)]