Detect, align and embed faces with SCRFD and MobileFaceNet

Ports the pipeline from the C++ reference in ../scene-actor-extraction
(MIT, same author). End to end on real portraits it separates identities
the way the reference's fitted calibration says it should: 0.596 between
distinct photographs of one person, 0.05 between different people, either
side of MBF's 0.267 boundary.

Three things are structural rather than incidental:

Aligned112 can only be built by align::warp, so Embedder::embed cannot be
handed an unaligned bounding-box crop. That mistake yields 512 plausible
unit-norm numbers and no error, so the type system refuses it instead.

Embedding carries its ModelId and cosine() returns None across models,
because a cross-model similarity is the one mistake that produces
plausible garbage rather than a failure.

The model-free half -- alignment, embedding arithmetic, f16 storage --
sits outside the inference feature and is covered by 11 tests that need
no weights on the machine.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
2026-08-26 19:57:56 +02:00
co-authored by Claude Opus 5
parent 72410f39c6
commit 19981c1033
9 changed files with 1196 additions and 32 deletions
+29
View File
@@ -706,6 +706,35 @@ three strides — `cls`/`obj`/`bbox`/`kps`, which is a *different layout* from S
licence (§2.3) that makes M9 more interesting than it looked: the permissive detector is also the one
with no shape-fixing step in front of it.
### 12.2 First end-to-end run · 2026-08-26
The Rust port produces the separation it is supposed to. Three distinct portraits of one identity
against two of another, from §1.1's labelled gallery:
| Pair | Cosine |
|---|---|
| Same identity, different photographs | **0.596** |
| Same identity, byte-identical duplicate files | 1.000 |
| Different identities | **0.049 – 0.050** |
Against the reference's fitted MBF boundary of cos 0.267 (§1), 0.596 and 0.05 fall either side with
room to spare — which is the check that the port's pre-processing, letterbox inversion and alignment
are right, since any of them being wrong degrades the same-identity number first.
The duplicate row is not a curiosity: several files in that gallery are byte-identical under
different names, which is exactly the case §8.2's dedup step exists for, and it would otherwise stack
the positive histogram at cos ≈ 1 with pairs that teach the fit nothing.
**M2/M3, provisionally, on the reference desktop:** detection **~1.0–1.4 s** per image, embedding
**~160–280 ms** per face, model load 80–90 ms for the pair. Slower than the extrapolation in §12's M2
row, which guessed SCRFD-500M would come in under YOLO26n-seg's 470 ms — tract is not ORT, and this
is what it costs. A 17k-image library at ~1.3 s plus ~1.5 faces each is on the order of **7 hours** of
background indexing. That is survivable for a resumable, preempted background job (FR-CULL-8) and it
is not survivable on a phone, so NFR-RES-2 needs its own measurement rather than an extrapolation
from this one. Not yet profiled: how much of the detection second is tract and how much is the
scalar letterbox resample in front of it.
---
## 13. Order