FR-CULL-8..12 specify the subsystem in terms of "a 512-dimension embedding from a stated model" and stop there, because D13 was open. This names the models, and grounds them in the measurements and the working C++ pipeline in ../scene-actor-extraction rather than in a literature reading. The licensing half of D13 stays open, but with a route through it: the InsightFace weights are non-commercial and cannot be committed, so the app ships the code and the user fetches the model. faces.model_id already makes that a survivable choice. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
770 lines
46 KiB
Markdown
770 lines
46 KiB
Markdown
# Faces and identity — SCRFD and MobileFaceNet
|
||
|
||
Spec for **S14**, the spike that decides whether §3.9.1 is buildable, and the build that follows it.
|
||
|
||
FR-CULL-8 … FR-CULL-12 specify the subsystem in terms of "a 512-dimension embedding from a stated
|
||
model" and stop there, deliberately: [catalog.md §10.4](catalog.md) records that the model was left
|
||
abstract because D13 was open. This document names the two models, fixes the pre- and
|
||
post-processing they need, specifies the calibration and clustering that sit on top, and states what
|
||
S14 has to measure before any of it is trusted.
|
||
|
||
**It does not close the licensing half of D13.** §2 is the reason, and it comes first because
|
||
[segmentation.md §7](segmentation.md) established the precedent that reading the grant is cheaper
|
||
than discovering it at packaging time.
|
||
|
||
---
|
||
|
||
## 1. The two models, and why these two
|
||
|
||
**Detector — SCRFD.** *Sample and Computation Redistribution for Efficient Face Detection* (Guo et
|
||
al., ICLR 2022). A single-stage anchor-based detector whose contribution is where the FLOPs go, not
|
||
a new architecture: the shallow stages carry more capacity because that is where small faces are
|
||
decided. It emits, per face, a box, a confidence, and **five landmarks in the same forward pass** —
|
||
which is the property that matters here, because FR-CULL-8 requires landmarks and because §4's
|
||
alignment step is not optional. A detector without landmarks would need a second network to supply
|
||
them.
|
||
|
||
The `500M` variant is the one to ship: 500 MFLOPs at VGA, against 10 GFLOPs for `10G`. Face indexing
|
||
is a background sweep over a whole library on a phone as well as a desktop (NFR-RES-2), so the cheap
|
||
variant's accuracy loss on tiny faces is the right trade — a face too small for `500M` to find is
|
||
also too small for §5 to embed usefully.
|
||
|
||
**Embedder — MobileFaceNet trained with ArcFace loss ("MBF").** ~1M parameters, ~0.45 GFLOPs for one
|
||
112×112 crop, 512-d output. ArcFace's additive angular margin is what makes the output space one
|
||
where *cosine similarity means something* — the loss explicitly optimises angular separation between
|
||
identities, which is what §6's calibration then has to convert into a probability.
|
||
|
||
**Why not the larger ResNet50 embedder** (`w600k_r50`, the other half of InsightFace's `buffalo_l`):
|
||
because it is not measurably better here, and it is 13× the file. §1.1's reference project fitted the
|
||
same Platt calibration (§8) to all four candidates over a 2,418-identity gallery, which is a clean
|
||
read on the embedding space alone with no tracking or matching logic in it:
|
||
|
||
| Model | File | Steepness `a` | Boundary at P=0.5 | Held-out macro F1 |
|
||
|---|---|---|---|---|
|
||
| LVFace-B (Glint360K) | 455 MB | 17.7 | sim 0.228 | 67.4% |
|
||
| **ArcFace w600k-MBF** | **13 MB** | **16.2** | **sim 0.267** | **64.4%** |
|
||
| ArcFace w600k-R50 | 174 MB | 15.4 | sim 0.301 | — |
|
||
| ArcFace R18 | 48 MB | 15.3 | sim 0.309 | 63.1% |
|
||
|
||
**MBF's calibration curve is steeper than R50's and R18's**, separating same-identity from
|
||
different-identity pairs more confidently and at a lower boundary, at a fraction of the size. LVFace-B
|
||
is genuinely better and is 35× the file — a defensible choice for a desktop-only feature and not one
|
||
for a subsystem that has to index a library on a phone (NFR-RES-2).
|
||
|
||
Two caveats on that table, both from the source: the R50 gallery was built with ~30% fewer reference
|
||
images per identity, which is why it carries no F1 figure — the calibration row does not depend on the
|
||
image count and is comparable, the benchmark row would not have been. And the F1 column measures a
|
||
*film-cast identification* task, not this one. It ranks the models; it does not predict DarkRoom's
|
||
accuracy.
|
||
|
||
**Both are pure convolutional graphs**, which matters for §3: the runtime is tract, and tract's
|
||
operator coverage is the thing that decides whether a graph runs at all.
|
||
|
||
### 1.1 A working reference implementation exists
|
||
|
||
`../scene-actor-extraction` is a C++ pipeline by the same author that runs **exactly this model
|
||
pair** — SCRFD-500MF then ArcFace over 5-point-aligned 112×112 crops — over feature films, with a
|
||
fitted Platt calibration and a scored benchmark behind it. It is **MIT licensed**, so porting from it
|
||
into GPL-3.0-or-later is straightforward and needs attribution, not permission.
|
||
|
||
That changes what S14 is. The open questions are no longer "does this pair work" and "what are the
|
||
magic numbers" — they are *does tract load these graphs* (§12 M1, and §4.1 says why that is in
|
||
genuine doubt) and *does the accuracy hold on family snapshots rather than film frames*. The
|
||
pre-processing constants, the decode layout, and the calibration algorithm are all readable rather
|
||
than rediscoverable, and several of them are not what the published Python would lead you to write.
|
||
|
||
| What | Reference | Ports to |
|
||
|---|---|---|
|
||
| ArcFace 5-point template and warp | `src/face_utils.hpp` | `align.rs` (§5) |
|
||
| SCRFD pre-process, decode, NMS | `src/backends/ort_backend.cpp` `SCRFDDecoder` | `detect.rs` (§4) |
|
||
| ArcFace pre-process and L2 normalise | same file, `ArcFaceEmbedder` | `embed.rs` (§6) |
|
||
| Platt fit over a similarity histogram | `src/gallery/gallery_calibration.hpp` | `calibrate.rs` (§8) |
|
||
| Model-free unit tests for both | `tests/test_face_utils.cpp`, `tests/test_calibration.cpp` | the §3 feature split |
|
||
|
||
That last row is worth noticing: the reference already separates the geometry and arithmetic from the
|
||
inference well enough to unit-test them with no model on the machine. §3's `inference` feature flag is
|
||
the same boundary, enforced by Cargo rather than by discipline.
|
||
|
||
**What does not port.** The reference matches faces against a *known gallery* of named identities;
|
||
DarkRoom clusters an unlabelled library (§9). And it is a video pipeline, so it has a tracker, a
|
||
`max_faces` cap, and frame-to-frame identity annealing that have no analogue here.
|
||
|
||
---
|
||
|
||
## 2. Licensing — the open half of D13
|
||
|
||
The architectures are published research. **The weights everyone actually uses are not
|
||
redistributable by this project.**
|
||
|
||
InsightFace's code is MIT. Its *pretrained models* — `buffalo_l`, `buffalo_s`, `buffalo_sc`, and
|
||
every `det_*` and `w600k_*` checkpoint inside them — carry a **non-commercial research-only** grant,
|
||
stated in the model-zoo README and applying to both the auto-downloaded and the manually downloaded
|
||
artefacts. `w600k_mbf` is trained on WebFace600K, whose own terms are research-only as well, so the
|
||
restriction has two independent sources rather than one that might be renegotiated.
|
||
|
||
These are the exact artefacts in question — §1.1's reference project fetches them from InsightFace's
|
||
own release page, which is where the grant attaches:
|
||
|
||
| File | Source |
|
||
|---|---|
|
||
| `det_500m.onnx` (2.5 MB) | `insightface/releases/download/v0.7/buffalo_sc.zip` |
|
||
| `w600k_mbf.onnx` (13 MB) | `insightface/releases/download/v0.7/buffalo_s.zip` |
|
||
|
||
DarkRoom is GPL-3.0-or-later. A non-commercial field-of-use restriction is not a GPL-compatible
|
||
grant, so:
|
||
|
||
- these weights **cannot be committed to this repository**, the way `yolo26n-seg.onnx` is (D14);
|
||
- they cannot ship inside an APK, a Flatpak, or an F-Droid build;
|
||
- and the restriction binds the *user*, not only the project — a professional photographer using
|
||
DarkRoom commercially is outside the grant even if they fetched the file themselves.
|
||
|
||
That last point is the one that is easy to miss and dishonest to leave implicit. It is why §2.2's
|
||
recommended route puts the licence text in front of the user rather than in a footnote.
|
||
|
||
### 2.1 The routes, priced
|
||
|
||
| Route | What ships | Cost |
|
||
|---|---|---|
|
||
| **A — commit the weights** | Everything works out of the box, one `git lfs pull` | **Not available.** Licence-incompatible with GPL-3.0 and with all three distribution channels (NFR-COMPAT-2). Listed only so it is visibly rejected rather than quietly assumed. |
|
||
| **B — permissively licensed weights** | Same, legally | Requires weights that do not exist off the shelf in this architecture pair. Either found (§2.3) or trained, and training an ArcFace embedder needs a face dataset whose own terms permit redistribution — the harder half of the problem. |
|
||
| **C — the user fetches them** | The app ships the *code*, not the weights; face indexing is off until the user obtains a model | Works today, keeps the repository clean, and is honest. Costs a first-run flow, an F-Droid anti-feature declaration, and a licence notice the user has to read. |
|
||
|
||
**Recommendation: C now, B when it becomes possible.** §1.1's project already works this way — a
|
||
`scripts/download_models.sh` that fetches the weights rather than a repository that carries them —
|
||
so route C is a practised pattern here rather than a theory. Two things DarkRoom must do that the
|
||
reference does not: pin a checksum for **every** model (the reference verifies only one of the four),
|
||
and put the licence in front of the user, because a photo editor's users are not all researchers.
|
||
|
||
The schema already forces this to be a survivable choice — `faces.model_id` (catalog.md §10.1) exists precisely so that a model change is
|
||
detectable and re-indexable rather than silently poisoning every similarity in the library. Under C,
|
||
swapping in a permissive model later is a new `model_id` and a re-index, not a migration.
|
||
|
||
### 2.2 What route C requires of the code
|
||
|
||
- **The weights are never a build input.** `dr-face` (§3) takes bytes; it has no `embedded-model`
|
||
feature and no `models/` directory. This is the one structural difference from `dr-segment`, and
|
||
it is deliberate — a feature flag that *could* embed weights is a feature flag someone eventually
|
||
turns on in a packaging script.
|
||
- **The fetch does not live in `dr-face`.** It lives in the app layer, so the inference crate keeps
|
||
no network dependency at all. §11 makes that a checkable property rather than a convention.
|
||
- **The licence is shown, not linked.** Before the first download, the app states in plain language
|
||
that the weights are research-only, that commercial use is outside the grant, and who the grantor
|
||
is. NFR-SEC-5's last bullet already requires the models to be named, versioned and licensed in the
|
||
about screen; this is the same obligation moved to the moment where it can still change a
|
||
decision.
|
||
- **A checksum is pinned.** The app verifies the digest of what it fetched against a value compiled
|
||
in, and refuses a mismatch. An unverified model file is an arbitrary graph executed over the
|
||
user's family photographs.
|
||
- **Indexing stays off until a model is present and verified**, and disabling it deletes nothing —
|
||
that is a separate control (NFR-SEC-5, §11).
|
||
|
||
### 2.3 Before S14 writes any code
|
||
|
||
Three questions to answer by reading, in this order, and to record with the date they were checked —
|
||
the same discipline [segmentation.md §7](segmentation.md) applied:
|
||
|
||
1. **Is there a permissively licensed SCRFD export?** Third-party ONNX exports of SCRFD are
|
||
plentiful; a re-export inherits the original weights' grant, so the question is whether anyone has
|
||
*trained* SCRFD on a redistributable dataset, not whether they have converted the InsightFace one.
|
||
2. **Is there a permissively licensed 512-d MobileFaceNet/ArcFace checkpoint?** Candidates worth
|
||
reading the terms on rather than assuming: EdgeFace (Idiap), the ONNX Model Zoo's ArcFace entry
|
||
(MIT, but a ResNet100 at ~250 MB — licence-clean and the wrong size), and any MobileFaceNet
|
||
trained on a dataset with distribution terms.
|
||
3. **What does a permissive *detector* alternative cost in accuracy?** YuNet (OpenCV Zoo) is 232 KB,
|
||
emits five landmarks, and comes from a permissively licensed repository. If it is close enough,
|
||
the detector half of the licence problem disappears and only the embedder remains — a materially
|
||
better position than either half alone.
|
||
|
||
Question 3 is nearly free to answer: `face_detection_yunet_2023mar.onnx` is **already sitting in
|
||
§1.1's `models/` directory**, fetched from `opencv/opencv_zoo`. It needs its own decode path (its
|
||
output layout is different, which is why the reference's `SCRFDDecoder` explicitly rejects it at load
|
||
rather than silently misreading it), and then it is one more row in §12's table.
|
||
|
||
---
|
||
|
||
## 3. Crate shape — `core/dr-face`
|
||
|
||
A new workspace member, modelled on `dr-segment` and for the same reason: everything that reasons
|
||
about faces is testable with no GPU adapter and no model present (ARCH §6.5a).
|
||
|
||
```
|
||
core/dr-face/
|
||
src/lib.rs FaceError, re-exports
|
||
src/detect.rs SCRFD: preprocess, decode, NMS
|
||
src/align.rs 5-point similarity transform → 112×112 crop
|
||
src/embed.rs MBF: preprocess, forward, L2 normalise
|
||
src/calibrate.rs cosine → P(same person) (FR-CULL-9)
|
||
src/cluster.rs constrained agglomeration (FR-CULL-10)
|
||
```
|
||
|
||
```toml
|
||
[features]
|
||
# No `embedded-model`. §2.2 — the weights are not a build input, ever.
|
||
default = []
|
||
# The ONNX runtime. Off by default so `calibrate` and `cluster` — which are
|
||
# arithmetic over embeddings and have no model in them — stay testable in a
|
||
# build that carries no inference engine at all.
|
||
inference = ["dep:ort", "dep:ort-tract", "dep:ndarray"]
|
||
```
|
||
|
||
The `ort` + `ort-tract` pairing is settled by D13's 2026-08-21 update and already in the workspace
|
||
manifest: `ort`'s API, tract's pure-Rust engine, no C under the NDK.
|
||
|
||
**The split between `inference` and the rest is load-bearing.** Calibration and clustering are where
|
||
the subsystem's accuracy actually lives, they are pure functions of embeddings and labels, and they
|
||
must be testable against synthetic embedding sets with no weights on the machine. A test suite that
|
||
needs a research-licensed download to run is a test suite that does not run in CI.
|
||
|
||
### 3.1 The API
|
||
|
||
```rust
|
||
/// A loaded detector. Fixed input shape — see §4.
|
||
pub struct Detector { /* session, input edge */ }
|
||
|
||
impl Detector {
|
||
pub fn from_bytes(onnx: &[u8], id: ModelId) -> Result<Self, FaceError>;
|
||
/// `rgb` is f32 0..=1, row-major, three per pixel.
|
||
pub fn detect(&self, rgb: &[f32], w: usize, h: usize, opts: &DetectOptions)
|
||
-> Result<Vec<Detection>, FaceError>;
|
||
}
|
||
|
||
pub struct Detection {
|
||
/// Normalised to the image's long edge (catalog.md §10.1).
|
||
pub bbox: (f32, f32, f32, f32),
|
||
/// Five points, same normalisation, in the model's own order (§5).
|
||
pub landmarks: [(f32, f32); 5],
|
||
pub confidence: f32,
|
||
}
|
||
|
||
pub struct Embedder { /* session */ }
|
||
|
||
impl Embedder {
|
||
pub fn from_bytes(onnx: &[u8], id: ModelId) -> Result<Self, FaceError>;
|
||
/// Takes an *aligned* 112×112 crop. Alignment is `align::warp`; passing a
|
||
/// raw bbox crop here is the silent-accuracy-loss bug §5 exists to prevent.
|
||
pub fn embed(&self, aligned: &Aligned112) -> Result<Embedding, FaceError>;
|
||
}
|
||
|
||
/// L2-normalised, 512-d. Carries its model id so a comparison across models
|
||
/// is a type error rather than a plausible-looking number (catalog.md §10.1).
|
||
pub struct Embedding { pub model: ModelId, pub v: [f32; 512] }
|
||
```
|
||
|
||
`Aligned112` is a newtype over the pixel buffer that only `align::warp` can construct. That is the
|
||
whole defence against §5's failure mode, and it costs nothing.
|
||
|
||
---
|
||
|
||
## 4. Detection
|
||
|
||
### 4.1 Fixed input, and how the image is fitted to it
|
||
|
||
**`det_500m.onnx` has a dynamic H/W input, and this is the largest single risk in the document.**
|
||
§1.1's reference had to use ONNX Runtime rather than OpenCV's `dnn` module precisely because OpenCV
|
||
could not load SCRFD's dynamic `Shape` nodes. tract is in the same family of problem: `dr-segment`
|
||
exists in its current shape because tract failed shape inference on YOLO's dynamic export, which is
|
||
why `tools/export-seg-model.sh` passes `dynamic=False` and why `semantic.rs` has a fixed
|
||
`INPUT_EDGE`.
|
||
|
||
So the graph almost certainly needs its input dimensions frozen before tract will touch it. That is a
|
||
known, cheap operation — `onnxruntime.tools.make_dynamic_shape_fixed` rewrites the declared dims
|
||
without retraining or re-exporting from PyTorch — but it is a step, it produces an artefact that must
|
||
itself be reproducible (a `tools/fix-face-model-shapes.sh`, in the spirit of the existing export
|
||
script), and **it must be tried before anything else in this document is scheduled.** §12's M1.
|
||
|
||
Fixed at **640**, with 320 available as a faster, blinder option to measure.
|
||
|
||
**Letterbox.** Scale by `min(640/w, 640/h)` preserving aspect, then paste into a 640×640 canvas. The
|
||
reference centres the image and fills the margin with grey `114`, inverting with
|
||
`x_src = (x_model − pad_x) / scale`. What matters is not *where* the padding goes but that the
|
||
forward and inverse agree and that the fill value is treated as part of the contract: a mismatch
|
||
between them offsets every box and landmark the model returns by the padding, which yields detections
|
||
that look plausible, landmarks that align badly, and an embedding-quality loss that only shows up as
|
||
poor clustering three stages later. Port the reference's convention rather than inventing one, and
|
||
assert it with a round-trip test.
|
||
|
||
Normalisation is `(x·255 − 127.5) / 128`, RGB, NCHW — note `/128`, not `/127.5`; see §6.
|
||
|
||
### 4.2 Outputs, and decoding them
|
||
|
||
**Nine tensors for three strides `{8, 16, 32}`, twelve for four `{8, 16, 32, 64}`.** Which one a
|
||
given export produces is discovered at load — `strides = output_count / 3` — not assumed, because
|
||
both variants exist and hardcoding three silently ignores the largest faces in a twelve-output model.
|
||
With two anchors per location and `N_s = (640/s)²` locations:
|
||
|
||
| Tensor | Shape | Meaning |
|
||
|---|---|---|
|
||
| `score_s` | `[N_s·2, 1]` | sigmoid already applied inside the graph |
|
||
| `bbox_s` | `[N_s·2, 4]` | distances left, top, right, bottom **in units of the stride** |
|
||
| `kps_s` | `[N_s·2, 10]` | five `(dx, dy)` offsets, same units |
|
||
|
||
Anchor centres are `(x·s, y·s)` for each grid cell, repeated once per anchor. Decoding is therefore
|
||
|
||
```
|
||
x1 = cx − l·s x2 = cx + r·s kx_i = cx + dx_i·s
|
||
y1 = cy − t·s y2 = cy + b·s ky_i = cy + dy_i·s
|
||
```
|
||
|
||
then divide by the letterbox scale to return to source pixels, then normalise by the long edge before
|
||
storage.
|
||
|
||
Flat index within a stride is `(row · fw + col) · 2 + anchor`.
|
||
|
||
**A shape assertion at load time, not a decode-time surprise.** `dr-segment` already learned this —
|
||
`SegmentError::OutputShape` exists because a different export of the same model produces confidently
|
||
wrong results otherwise. The reference implements exactly this and it is worth porting verbatim:
|
||
output count divisible by three and between 9 and 12, then the **last dimension of each group checked
|
||
against `{1, 4, 10}`**. That single check is what catches a YuNet file passed where an SCRFD one was
|
||
meant — YuNet also has twelve outputs, so the count alone does not distinguish them, and the failure
|
||
without the check is not an error but a page of plausible garbage.
|
||
|
||
### 4.3 Thresholds
|
||
|
||
Score ≥ **0.5**, NMS IoU **0.4**, plain greedy NMS per image across all strides together.
|
||
|
||
Deliberately *not* the low threshold `dr-segment` chose. There, a false positive costs one spurious
|
||
entry in a list the user is picking from. Here it costs an entry in the People view that the user has
|
||
to reject, in a library with thousands of images — and worse, a garbage embedding that participates
|
||
in clustering and can bridge two real clusters into one. False negatives are recoverable by a later
|
||
re-index with a better model; a polluted cluster graph is not, once the user has confirmed faces
|
||
inside it.
|
||
|
||
A minimum box size of **40 px on the source proxy** is applied on top — the reference's figure,
|
||
chosen there for the same reason it holds here: below it there is not enough face left to align
|
||
reliably, and §7's table shows what the embedder actually receives at that size. §12's M4 is what
|
||
moves this number, if measurement says it should move.
|
||
|
||
**One thing not to port:** the reference caps detections at ten per frame, largest first. That is
|
||
right for a film frame, where the extras in the background are noise. It is wrong for a photo
|
||
library, where a group shot with thirty faces in it is precisely the picture the user wants indexed.
|
||
No cap; the min-size floor is the only filter.
|
||
|
||
---
|
||
|
||
## 5. Alignment — the step that is silently wrong when skipped
|
||
|
||
ArcFace embeddings are trained on faces warped to a canonical 112×112 arrangement. Feeding the model
|
||
a plain bounding-box crop *works* — it produces 512 numbers, they are unit-norm, and cosine
|
||
similarities between them look entirely reasonable. They are just much worse, and nothing in the
|
||
system reports it.
|
||
|
||
The transform is a **similarity transform** — rotation, uniform scale, translation, four degrees of
|
||
freedom — from the five detected landmarks to this template, which is the arrangement the weights were
|
||
trained against:
|
||
|
||
```
|
||
(38.2946, 51.6963) subject's right eye ─┐ image-left of centre
|
||
(73.5318, 51.5014) subject's left eye ─┘
|
||
(56.0252, 71.7366) nose tip
|
||
(41.5493, 92.3655) subject's right mouth corner
|
||
(70.7299, 92.2041) subject's left mouth corner
|
||
```
|
||
|
||
Similarity, not affine: a full affine fit to five noisy points will happily shear a face, and shear is
|
||
exactly the deformation the embedder was never shown.
|
||
|
||
**The naming is a trap and the order is not.** Point 0 sits at x=38 on a 112-wide canvas — left of
|
||
centre *in the image*, which is the subject's **right** eye. Both namings are in circulation and they
|
||
are opposite. What matters is that SCRFD emits its five points in this same order, so the correct
|
||
amount of reordering between detector and template is **none**; the reference states this explicitly
|
||
in `types.hpp` for the benefit of whoever next reads it and doubts it. A future detector with a
|
||
different order carries its own permutation next to its `model_id`, rather than this file growing an
|
||
assumption.
|
||
|
||
**One divergence from the reference to settle by measurement.** It fits the transform with OpenCV's
|
||
`estimateAffinePartial2D` under **RANSAC** at a 3-pixel threshold, and drops the detection when the
|
||
fit fails. RANSAC over five points is a strange fit — the minimal sample for a similarity is two, so
|
||
it can discard landmarks it judges outliers and solve from a subset, which on a profile face is as
|
||
likely to be the correct geometry as the wrong one. InsightFace's own pipeline uses a plain
|
||
least-squares (Umeyama) similarity over all five points, which cannot silently drop anything and is
|
||
the one to write first. §12's M5 measures whether the choice matters; the degenerate-input path still
|
||
has to exist either way, because collinear landmarks do occur.
|
||
|
||
**A trick worth keeping.** The reference retries detection on a frame where nothing was found, after
|
||
replicate-padding by 25% and applying CLAHE to the L channel. Padding gives a face at the very edge a
|
||
context the detector needs; CLAHE rescues underexposed frames. Both are plausible on scanned negatives
|
||
and backlit photographs — worth trying at low priority, since a retry doubles the cost of exactly the
|
||
images that yielded nothing.
|
||
|
||
Sampling is bilinear from the *source proxy*, in one step — never crop-then-warp, which resamples
|
||
twice and throws away detail the warp could have used. Pixels falling outside the source are black.
|
||
|
||
**The landmark order must match the template order.** The template above is written in the detector's
|
||
own output order; if a future detector emits them differently, the template is reordered with it and
|
||
that mapping belongs next to the model id, not compiled in as an assumption.
|
||
|
||
*Test:* warp a synthetic image with a known rotation and scale, and assert the five points land on
|
||
the template within a fraction of a pixel. This is testable with no model present, which is why
|
||
`align` sits outside the `inference` feature.
|
||
|
||
---
|
||
|
||
## 6. Embedding
|
||
|
||
112×112 **RGB** — the crop comes out of the warp in whatever order the source was in, and the
|
||
reference converts BGR→RGB explicitly before the blob rather than relying on a flag, which is the
|
||
readable way to do it.
|
||
|
||
Normalisation is `(x·255 − 127.5) / 128`. **Note `/128`, not `/127.5`:** InsightFace's published
|
||
Python uses `1.0/127.5` for the recognition model, the reference uses `1.0/128` for both models, and
|
||
every measured number in §1's table was produced with `/128`. The difference is 0.4% of scale and
|
||
almost certainly immaterial, but "almost certainly" is not a reason to pick silently — write `/128` to
|
||
match the numbers we have, and settle it with one back-to-back run in §12.
|
||
|
||
Output is 512 floats; **L2-normalise before storing**, so every downstream comparison is a dot product
|
||
and no code path has to remember to normalise. The reference clamps the norm at 1e-6 before dividing,
|
||
which costs nothing and removes a NaN path.
|
||
|
||
Some exports of these graphs are fp16 in, fp16 out. The reference detects this from the graph's
|
||
declared element type rather than from the filename; worth porting, because the alternative failure is
|
||
a silent garbage tensor.
|
||
|
||
Storage is `512 × f16` (catalog.md §10.1) — 1 KB per face, 25 MB for a 25,000-face library. The f16
|
||
round-trip perturbs a unit vector by ~1e-3 in cosine, three orders of magnitude below the separation
|
||
between a match and a non-match, and the halving matters because these rows are the ones NFR-SEC-5
|
||
contemplates optionally syncing.
|
||
|
||
Re-normalise on load after the f16 widen. It is one pass over 512 floats and it removes a class of
|
||
drift that is otherwise invisible.
|
||
|
||
---
|
||
|
||
## 7. Which pixels the pipeline actually sees
|
||
|
||
FR-CULL-8 pins indexing to the FR-CULL-2 ladder: **the proxy tier, never a full decode.** The
|
||
relevant tier is `ThumbSize::Large` — 1024 px on the long edge (`dr-thumbs`).
|
||
|
||
That has a consequence worth stating in numbers rather than discovering in a clustering report. On a
|
||
1024 px proxy:
|
||
|
||
| Face size in frame | Pixels across | What §6 receives |
|
||
|---|---|---|
|
||
| A portrait, face fills a third of the frame | ~340 | Downsampled to 112. Ideal. |
|
||
| Two people, half-length | ~120 | Roughly native. Good. |
|
||
| A group of eight | ~50 | **Upsampled** to 112. Degraded, and usable. |
|
||
| A figure in a landscape | ~20 | At or below §4.3's floor. Rejected. |
|
||
|
||
So `faces` records one column beyond catalog.md §10.1's schema:
|
||
|
||
```sql
|
||
ALTER TABLE faces ADD COLUMN crop_px INTEGER NOT NULL; -- source pixels across the aligned crop
|
||
```
|
||
|
||
`crop_px` earns its place three times over. It is the honest quality signal for the UI; it is a
|
||
**feature in §8's calibration**, which FR-CULL-9 explicitly demands ("a raw cosine means something
|
||
different for every model, every population, and *every face size*"); and it is what a future
|
||
higher-resolution re-embedding pass would select on, so that pass becomes a query rather than a
|
||
re-index of everything.
|
||
|
||
**No RAW decode is added.** Where the Large proxy is missing, the job enqueues a `Thumbnail` job at
|
||
background priority and re-queues itself, exactly as FR-CULL-8 requires.
|
||
|
||
---
|
||
|
||
## 8. Calibration — cosine to probability
|
||
|
||
FR-CULL-9 makes this a hard requirement: no code path may threshold a bare cosine, every threshold
|
||
in the subsystem is stated as a probability, and the fit is per library and reports its own validity.
|
||
This section is how that is obtained, and the interesting part is where the training pairs come from
|
||
when the user has labelled nothing yet.
|
||
|
||
### 8.1 Where the pairs come from
|
||
|
||
**Negatives are free and abundant.** Two faces detected in *the same photograph* are almost never the
|
||
same person. That gives every multi-face image in the library a full set of negative pairs at no
|
||
labelling cost — and they are *hard* negatives, drawn from the same camera, lighting, and processing,
|
||
which is exactly the population where a threshold tuned on easy negatives fails. The known exceptions
|
||
— mirrors, photographs of photographs, collages, a spliced panorama — are rare enough to be noise at
|
||
this scale and worth naming so the exception is not mistaken for a bug later.
|
||
|
||
**Positives, in order of trustworthiness:**
|
||
|
||
1. **User confirmations** (FR-CULL-10). Every pair of faces confirmed to the same person. The gold
|
||
standard, and empty on day one.
|
||
2. **Burst siblings.** FR-CULL-5 already groups bursts. Two faces in adjacent frames of one burst,
|
||
in nearly the same position, are near-certainly the same person. Free, requires no labelling, and
|
||
available immediately on any library with continuous-shooting frames in it. **Their purity is an
|
||
S14 measurement** (§12), not an assumption — if bursts turn out to be dirtier than expected, this
|
||
source is dropped and the calibration simply stays invalid for longer.
|
||
3. Nothing else. Bootstrapping positives from high cosine similarity is circular — it fits the
|
||
calibration to the belief it was supposed to test.
|
||
|
||
### 8.2 The fit
|
||
|
||
§1.1's reference already implements this and its shape should be ported rather than reinvented:
|
||
|
||
```
|
||
P(same | cos) = σ(a·cos + b + log_prior_odds)
|
||
```
|
||
|
||
Four details in it are the difference between working and nearly working.
|
||
|
||
**Fit against a histogram, not against pairs.** A 25,000-face library has 3×10⁸ pairs; no gradient
|
||
descent is running over that. The reference buckets every pair into **200 bins over cos ∈ [−1, 1]**,
|
||
carrying a positive and a negative count per bin, and fits the two parameters against the per-bin
|
||
counts. The cost becomes the similarity matrix plus 200 numbers, and the matrix is a single GEMM
|
||
(§9). This is the trick that makes a per-library fit affordable at all, and it is not obvious from
|
||
the outside.
|
||
|
||
**Balance the classes explicitly.** `w_pos = total/(2·n_pos)`, `w_neg = total/(2·n_neg)`. Negatives
|
||
outnumber positives by orders of magnitude; an unweighted fit produces a well-shaped curve sitting at
|
||
the wrong height — precisely the "plausible number all the way to the user interface" failure
|
||
FR-CULL-9 describes.
|
||
|
||
**The base rate is a runtime argument, not part of the fit.** `log_prior_odds` is added at evaluation
|
||
time, so the balanced fit is stored once and the prior varies per query — the odds that two faces
|
||
drawn from a 40-image holiday album match are not the odds for a 40,000-image archive. Its inverse,
|
||
`boundary_at(p)`, gives the cosine at which the probability crosses `p`, which is what turns §9's
|
||
"merge above 0.9" into an actual comparison. Baking a prior into `a` and `b` would need a refit per
|
||
context and would make the stored parameters mean something different depending on where they came
|
||
from.
|
||
|
||
**Deduplicate before pairing.** Near-identical embeddings (cos > 1 − 1e-7) are the same photograph
|
||
counted twice; the reference drops them per identity first. In a photo library the equivalent is
|
||
duplicates and virtual copies (FR-CAT-11), and leaving them in stacks the positive histogram at
|
||
cos ≈ 1 with pairs that teach the fit nothing about hard cases.
|
||
|
||
**A third feature this design adds.** The reference fits on cosine alone; DarkRoom should carry
|
||
`crop_px` too:
|
||
|
||
```
|
||
logit P(same) = w₀ + w₁·cos + w₂·log₂(min(crop_px_a, crop_px_b))
|
||
```
|
||
|
||
FR-CULL-9 explicitly names face size as an axis along which an uncalibrated similarity misbehaves,
|
||
§7 shows a real library spans 40 px to 340 px of face, and `crop_px` is already in hand. The
|
||
*minimum* of the pair, because a comparison is only as good as its worse crop. It generalises the
|
||
histogram to a small 2-D grid of bins, which changes nothing structural. If M7 says the extra feature
|
||
buys nothing, drop it and match the reference exactly — but a film pipeline working from broadcast
|
||
frames had far less size variation to explain than this one does.
|
||
|
||
### 8.3 Validity, and saying so
|
||
|
||
Stored per catalog.md §10.3: parameters, a validity flag, and a hash of the face set fitted from.
|
||
|
||
Invalid when fewer than **200 positive pairs** or **2,000 negative pairs** are available, or when the
|
||
reliability check fails.
|
||
|
||
Deliberately far stricter than the reference's floor of two positives and one negative. That floor is
|
||
reasonable there: its pairs come from a curated gallery of labelled reference portraits, where a
|
||
positive pair is trustworthy by construction. Here the positives are bootstrapped from bursts and
|
||
early confirmations (§8.1) and the whole risk is fitting confidently to a handful of them. The
|
||
reference also refuses to draw positives from an identity with fewer than five distinct embeddings,
|
||
letting it contribute negatives only — the same asymmetry applies to a thinly-confirmed person and is
|
||
worth keeping. In that state the UI says confidence is unavailable and the People view
|
||
still works — clustering falls back to a documented default operating point, labelled in the
|
||
interface as an untuned default, and no probability is displayed. That is FR-CULL-9's requirement
|
||
read literally: not presenting an untuned default *as though it were measured*.
|
||
|
||
Refit is triggered by the same debounce as clustering (§9), and when the confirmed-pair count grows
|
||
materially.
|
||
|
||
*Acceptance (FR-CULL-9):* a reliability diagram over ten probability bins, on a held-out labelled
|
||
split, with observed match rate within a stated tolerance of the predicted probability in each
|
||
populated bin. A single accuracy figure is not an answer to this requirement.
|
||
|
||
---
|
||
|
||
## 9. Clustering
|
||
|
||
Not a job kind, for the reason catalog.md §10.2 gives: it is a whole-library operation with no
|
||
natural `subject_id`. It runs as a debounced library pass when detection has been idle and the face
|
||
count has moved materially.
|
||
|
||
**The graph.** For each face, its `k = 20` nearest neighbours by cosine, then each candidate edge
|
||
scored through §8 to a probability.
|
||
|
||
Brute force is honest arithmetic here, and §1.1's project is the evidence rather than the estimate: it
|
||
computes the full similarity matrix as **one GEMM** (`E · Eᵀ` over L2-normalised rows) at gallery
|
||
scale, on CPU, and instruments it well enough to have compared CPU against OpenCL and found the
|
||
question worth asking rather than urgent. For DarkRoom that is a 25,000 × 512 matrix at 25 MB in and
|
||
~2.5 GB out — so the matrix is computed **in row blocks**, with each block reduced to its top-k and
|
||
its histogram contribution before the next is started, and never materialised whole. Once, in the
|
||
background, after an indexing sweep that took an hour. An approximate index is an optimisation to
|
||
reach for when §12 says it is needed, not before.
|
||
|
||
**Constraints, not just thresholds:**
|
||
|
||
- **Cannot-link on co-occurrence.** Two faces in the same image are never merged. This is the same
|
||
observation §8.1 mines for negatives, used here as a hard constraint, and it is the single
|
||
cheapest defence against the over-merging FR-CULL-10 warns about.
|
||
- **Confirmed faces are anchors.** A confirmation is user data (FR-CULL-12) and clustering never
|
||
moves it. Two clusters each containing confirmations of *different* people cannot merge; a cluster
|
||
containing confirmations of one person absorbs suggestions but never reassigns the confirmed.
|
||
|
||
**The algorithm.** Constrained average-link agglomeration over the probability graph, merging while
|
||
the average pairwise probability exceeds **0.9** and no cannot-link is violated. Average-link rather
|
||
than single-link because single-link chains — one bad edge welds two identities together, which is
|
||
the documented way face clustering fails on families.
|
||
|
||
**Incremental by default.** A new face joins the existing cluster whose average probability against
|
||
it is highest, if that clears the threshold; otherwise it starts an unnamed group. A full
|
||
re-agglomeration happens on model change, on recalibration, and on explicit user request. Suggestions
|
||
are recomputed freely; confirmations survive all of it (FR-CULL-10).
|
||
|
||
**Splitting** re-agglomerates within one person at a raised threshold and offers the resulting groups
|
||
as the split. FR-CULL-10 requires splitting to be as easy as merging, and a split that hands the user
|
||
a pile of loose faces to re-sort is not that.
|
||
|
||
---
|
||
|
||
## 10. Catalog and jobs
|
||
|
||
**Schema** is catalog.md §10.1's v5 migration, plus `faces.crop_px` (§7) and the calibration table
|
||
(§8.3). Nothing else changes.
|
||
|
||
**One new job kind:**
|
||
|
||
```rust
|
||
/// Detect and embed faces in one image, from its proxy (FR-CULL-8).
|
||
DetectFaces = 8,
|
||
```
|
||
|
||
Background priority, coalesced per `image_id`, interruptible, resumable — it inherits FR-CAT-3's
|
||
properties with no exceptions, which is what FR-CULL-8's acceptance criterion is about. It is not a
|
||
network job.
|
||
|
||
One job does detection *and* embedding for every face in the image, rather than splitting them.
|
||
Splitting would double the queue's row count for no benefit: the proxy is already decoded and in
|
||
memory, and the natural unit of resumable work is one photograph.
|
||
|
||
**Selector term** (FR-CULL-11):
|
||
|
||
```rust
|
||
Person { id: PersonId, include_suggested: bool }, // defaults to false
|
||
```
|
||
|
||
Confirmed-only by default, so a saved smart collection does not silently change membership when a
|
||
later indexing pass revises a guess.
|
||
|
||
---
|
||
|
||
## 11. Privacy obligations that constrain the code's shape
|
||
|
||
NFR-SEC-5 is not a policy to remember; it is a set of properties the code either has structurally or
|
||
does not have at all. The checkable ones:
|
||
|
||
- **`dr-face` has no network dependency.** No `reqwest`, no `ureq`, transitively. Worth a CI check
|
||
over `cargo tree`, alongside the existing lints — the crate that holds the embeddings should be
|
||
provably unable to send them anywhere.
|
||
- **The model fetch (§2.2) is in the app layer**, which is why the API in §3.1 takes bytes.
|
||
- **The diagnostics bundle is an allowlist** (NFR-OPS-1), so `faces`, `face_person` and the
|
||
calibration table are excluded by not being named, and a future table cannot become uploadable by
|
||
existing.
|
||
- **Delete-all is one transaction and one control**: faces, links, people, calibration, and the
|
||
cached crops if any. Separately, a switch that stops indexing so no such data is produced. The two
|
||
are different actions and are not collapsed into one.
|
||
- **The about screen reads the model metadata from the loader** — id, version, licence — rather than
|
||
from a hardcoded string that will drift from what is actually running.
|
||
|
||
---
|
||
|
||
## 12. What S14 measures
|
||
|
||
The corpus is a real personal library — the ~2,000-image sample S14 already specifies — with a
|
||
hand-labelled identity subset. Fixed in advance, so the result is a measurement rather than an
|
||
argument:
|
||
|
||
| # | Measure | Why it decides something |
|
||
|---|---|---|
|
||
| **M1** | **Does tract load both graphs?** As shipped, then with the input dims frozen | Go/no-go, and it is *first*. `det_500m.onnx` has a **dynamic H/W input** (§4.1), which is the exact thing tract failed on for YOLO and which OpenCV's `dnn` could not load either — so the as-shipped answer is expected to be no, and the real question is whether `make_dynamic_shape_fixed` is enough or whether a PyTorch re-export is needed. Costs an afternoon; everything below is void without it. |
|
||
| **M2** | Detection latency per image at 640, and at 320, on the reference desktop and one Android device | Whether a 17k library indexes in a background sweep or a weekend. The one measured datapoint we have — YOLO26n-seg, ~470 ms at 640×640 in tract — suggests SCRFD-500M lands well under it, but tract is not ORT and an extrapolation is not a measurement. |
|
||
| **M3** | Embedding latency per face, and faces per image in a real library | The multiplier on M2. A library averaging 1.5 faces per image at 100 ms per face is an hour for 25k faces; at 400 ms it is four. |
|
||
| **M4** | Detection recall against hand-labelled faces, bucketed by `crop_px` | Where §4.3's floor should actually sit, and what proportion of a real library's faces are in the degraded bucket §7 predicts. |
|
||
| **M5** | **Aligned versus unaligned embeddings**, same corpus, same clustering. Then two A/Bs that cost one run each: least-squares versus RANSAC alignment (§5), and `/128` versus `/127.5` normalisation (§6) | Quantifies §5. If the gap is small the alignment code is still correct; if it is large, this is the measurement that stops someone "simplifying" it later. The two A/Bs close divergences between the reference and the published Python that are currently settled by assertion. |
|
||
| **M6** | Cluster purity and completeness against the labelled subset | Whether the subsystem is worth building at all. §1's F1 figures rank the models on *film frames against a known cast*; nothing yet says how MBF behaves on family snapshots it must cluster blind. |
|
||
| **M7** | Calibration reliability, and **the sample size at which the fit becomes valid** | FR-CULL-9's acceptance criterion, plus the practical question of whether a 2,000-image library ever gets a valid fit or whether §8.3's "unavailable" is the normal state. |
|
||
| **M8** | Burst-derived positive-pair purity (§8.1) | Whether the free positives are usable or the fit waits for user confirmations. |
|
||
| **M9** | YuNet as a drop-in detector: M4 and M6, re-run | §2.3's question 3, and the model file is already on disk. If the answer is "close enough", half the licence problem disappears. Needs its own decode path — its outputs are not SCRFD's. |
|
||
| **M10** | Peak RSS during an indexing sweep | NFR-RES-2. Two loaded graphs plus a proxy plus a batch of crops, on a phone. |
|
||
|
||
---
|
||
|
||
## 13. Order
|
||
|
||
1. **M1** — load both graphs in tract, freezing the input dims if needed (§4.1). Go/no-go, and it
|
||
now comes first: it is an afternoon, and §2.3's reading is wasted effort if the graphs will not
|
||
run at all. This is a change from the S14 brief's ordering, made because §1.1 removed the
|
||
uncertainty that justified reading first — the models are known to work, so the open question is
|
||
the runtime, not the choice.
|
||
2. **Licence reading** (§2.3), in parallel with 3. Before anything *ships*, per D13.
|
||
3. `dr-face` skeleton, `detect` + `align` + `embed`, ported from §1.1's table, with an example binary
|
||
that draws boxes and landmarks on a JPEG — the same shape as `dr-segment`'s `examples/detect.rs`,
|
||
and for the same reason: the thing worth looking at is whether the landmarks land on a real
|
||
photograph. Port `tests/test_face_utils.cpp`'s cases first; they are model-free and they fail
|
||
loudly on exactly the mistakes §5 describes.
|
||
4. M2–M5 on the corpus. Alignment is validated here, before anything depends on embedding quality.
|
||
5. `calibrate` + `cluster` against the labelled subset. M6–M8. The Platt fit ports from
|
||
`gallery_calibration.hpp`; the pair *sourcing* (§8.1) is new and is the part to get wrong.
|
||
6. The v5 migration, the `DetectFaces` job, the debounced clustering pass.
|
||
7. UI: People view, confirm and reject, merge and split, the `Person` selector term.
|
||
8. The route-C first-run flow, the about-screen entry, and the delete-all control — with the feature
|
||
shipped disabled until they exist.
|
||
|
||
Steps 1–5 are the spike. Steps 6–8 are the build, and they are only justified if §12 says so.
|
||
|
||
---
|
||
|
||
## 14. What this does not settle
|
||
|
||
- **Whether a permissive model pair exists.** §2.3. If it does, route B replaces route C and step 8's
|
||
first-run flow shrinks to nothing.
|
||
- **Whether tract runs these graphs at all.** §12 M1, and §4.1 says why the answer is in real doubt.
|
||
Every other line of this document is conditional on it.
|
||
- **Faces in trashed images.** catalog.md §10.4's open question, unchanged: probably excluded from
|
||
suggestions but not deleted, so a restore does not re-index.
|
||
- **Whether embeddings sync.** NFR-SEC-5 permits it, opt-in, separately consented. Nothing here
|
||
depends on the answer.
|
||
- **Re-embedding at higher resolution.** §7's `crop_px` makes it a query rather than a full re-index,
|
||
but whether it is worth doing is an M4 question.
|
||
- **Approximate nearest neighbours.** §9 says brute force until measured otherwise. A 100k-face
|
||
library is where this stops being true.
|
||
|
||
---
|
||
|
||
## 15. Register entries
|
||
|
||
**D13** — *face inference runtime and model licensing* · the runtime half stays answered (`ort` +
|
||
`ort-tract`, unchanged since 2026-08-21). **The licensing half is answered conditionally by §2:**
|
||
route C — ship the code, not the weights — unblocks the build without breaking GPL-3.0 or any
|
||
distribution channel, and §2.3's reading may yet convert it to route B. Closing D13 outright waits on
|
||
that reading.
|
||
|
||
**D17** — *face model pair and distribution route* · **PROPOSED**. SCRFD-500M for detection,
|
||
MobileFaceNet/ArcFace for embedding, weights obtained by the user rather than redistributed, with a
|
||
pinned checksum and a licence notice shown before the first fetch. The alternative considered and
|
||
rejected is committing the InsightFace weights (§2.1 route A), which is not available at any price.
|
||
The model *pair* is better evidenced than a proposal usually is — §1's table is a measured comparison
|
||
over a 2,418-identity gallery, not a literature reading — so the open half of this decision is the
|
||
route, not the pair. Relates to: D13, D14, NFR-COMPAT-2, FR-CULL-8.
|
||
|
||
**D18** — *porting from `scene-actor-extraction`* · **PROPOSED, and the easy half of a decision**.
|
||
That project is MIT and by this project's author, so MIT-into-GPL-3.0-or-later is a compatible
|
||
one-way combination requiring attribution, not permission. Ported files carry a header naming the
|
||
origin. Worth recording because "we already have this working in another language" is exactly the
|
||
provenance that goes undocumented and then cannot be answered three years later.
|
||
|
||
**S14** — *face pipeline in Rust* · scope sharpened, and **substantially de-risked**, by this
|
||
document. §12 replaces "measure per-image latency, cluster purity and calibration convergence" with
|
||
ten specific measures. Two changes to the brief itself: its instruction to resolve the licence
|
||
question *before writing any of it* is relaxed to "before shipping any of it", because M1 is an
|
||
afternoon and is worth knowing first (§13); and its central question — whether the pipeline works at
|
||
all — is largely answered in advance by §1.1, leaving the runtime and the domain shift from film
|
||
frames to family snapshots as the real unknowns.
|
||
|
||
---
|
||
|
||
## 16. Requirements touched
|
||
|
||
| ID | How this document addresses it |
|
||
|---|---|
|
||
| FR-CULL-8 | §4 detection, §6 embedding, §7 the proxy tier and its consequences, §10 the `DetectFaces` job |
|
||
| FR-CULL-9 | §8 — pairs, fit, validity, and the reliability-diagram acceptance test |
|
||
| FR-CULL-10 | §9 constrained agglomeration, confirmations as anchors, split by re-agglomeration |
|
||
| FR-CULL-11 | §10 the `Person` selector term, confirmed-only by default |
|
||
| FR-CULL-12 | §10 schema unchanged from catalog.md §10.1: embeddings derived, names to the sidecar |
|
||
| NFR-SEC-5 | §11 — the obligations restated as structural properties, one of them CI-checkable |
|
||
| NFR-COMPAT-2 | §2 — why the obvious weights cannot ship, and what does instead |
|
||
| NFR-RES-2 | §1 the cheap model pair, §12 M2/M3/M10 on a phone |
|
||
| NFR-ARCH-2 | §10 background priority, preempted by visible work |
|