FR-CULL-8..12 specify the subsystem in terms of "a 512-dimension embedding from a stated model" and stop there, because D13 was open. This names the models, and grounds them in the measurements and the working C++ pipeline in ../scene-actor-extraction rather than in a literature reading. The licensing half of D13 stays open, but with a route through it: the InsightFace weights are non-commercial and cannot be committed, so the app ships the code and the user fetches the model. faces.model_id already makes that a survivable choice. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
46 KiB
Faces and identity — SCRFD and MobileFaceNet
Spec for S14, the spike that decides whether §3.9.1 is buildable, and the build that follows it.
FR-CULL-8 … FR-CULL-12 specify the subsystem in terms of "a 512-dimension embedding from a stated model" and stop there, deliberately: catalog.md §10.4 records that the model was left abstract because D13 was open. This document names the two models, fixes the pre- and post-processing they need, specifies the calibration and clustering that sit on top, and states what S14 has to measure before any of it is trusted.
It does not close the licensing half of D13. §2 is the reason, and it comes first because segmentation.md §7 established the precedent that reading the grant is cheaper than discovering it at packaging time.
1. The two models, and why these two
Detector — SCRFD. Sample and Computation Redistribution for Efficient Face Detection (Guo et al., ICLR 2022). A single-stage anchor-based detector whose contribution is where the FLOPs go, not a new architecture: the shallow stages carry more capacity because that is where small faces are decided. It emits, per face, a box, a confidence, and five landmarks in the same forward pass — which is the property that matters here, because FR-CULL-8 requires landmarks and because §4's alignment step is not optional. A detector without landmarks would need a second network to supply them.
The 500M variant is the one to ship: 500 MFLOPs at VGA, against 10 GFLOPs for 10G. Face indexing
is a background sweep over a whole library on a phone as well as a desktop (NFR-RES-2), so the cheap
variant's accuracy loss on tiny faces is the right trade — a face too small for 500M to find is
also too small for §5 to embed usefully.
Embedder — MobileFaceNet trained with ArcFace loss ("MBF"). ~1M parameters, ~0.45 GFLOPs for one 112×112 crop, 512-d output. ArcFace's additive angular margin is what makes the output space one where cosine similarity means something — the loss explicitly optimises angular separation between identities, which is what §6's calibration then has to convert into a probability.
Why not the larger ResNet50 embedder (w600k_r50, the other half of InsightFace's buffalo_l):
because it is not measurably better here, and it is 13× the file. §1.1's reference project fitted the
same Platt calibration (§8) to all four candidates over a 2,418-identity gallery, which is a clean
read on the embedding space alone with no tracking or matching logic in it:
| Model | File | Steepness a |
Boundary at P=0.5 | Held-out macro F1 |
|---|---|---|---|---|
| LVFace-B (Glint360K) | 455 MB | 17.7 | sim 0.228 | 67.4% |
| ArcFace w600k-MBF | 13 MB | 16.2 | sim 0.267 | 64.4% |
| ArcFace w600k-R50 | 174 MB | 15.4 | sim 0.301 | — |
| ArcFace R18 | 48 MB | 15.3 | sim 0.309 | 63.1% |
MBF's calibration curve is steeper than R50's and R18's, separating same-identity from different-identity pairs more confidently and at a lower boundary, at a fraction of the size. LVFace-B is genuinely better and is 35× the file — a defensible choice for a desktop-only feature and not one for a subsystem that has to index a library on a phone (NFR-RES-2).
Two caveats on that table, both from the source: the R50 gallery was built with ~30% fewer reference images per identity, which is why it carries no F1 figure — the calibration row does not depend on the image count and is comparable, the benchmark row would not have been. And the F1 column measures a film-cast identification task, not this one. It ranks the models; it does not predict DarkRoom's accuracy.
Both are pure convolutional graphs, which matters for §3: the runtime is tract, and tract's operator coverage is the thing that decides whether a graph runs at all.
1.1 A working reference implementation exists
../scene-actor-extraction is a C++ pipeline by the same author that runs exactly this model
pair — SCRFD-500MF then ArcFace over 5-point-aligned 112×112 crops — over feature films, with a
fitted Platt calibration and a scored benchmark behind it. It is MIT licensed, so porting from it
into GPL-3.0-or-later is straightforward and needs attribution, not permission.
That changes what S14 is. The open questions are no longer "does this pair work" and "what are the magic numbers" — they are does tract load these graphs (§12 M1, and §4.1 says why that is in genuine doubt) and does the accuracy hold on family snapshots rather than film frames. The pre-processing constants, the decode layout, and the calibration algorithm are all readable rather than rediscoverable, and several of them are not what the published Python would lead you to write.
| What | Reference | Ports to |
|---|---|---|
| ArcFace 5-point template and warp | src/face_utils.hpp |
align.rs (§5) |
| SCRFD pre-process, decode, NMS | src/backends/ort_backend.cpp SCRFDDecoder |
detect.rs (§4) |
| ArcFace pre-process and L2 normalise | same file, ArcFaceEmbedder |
embed.rs (§6) |
| Platt fit over a similarity histogram | src/gallery/gallery_calibration.hpp |
calibrate.rs (§8) |
| Model-free unit tests for both | tests/test_face_utils.cpp, tests/test_calibration.cpp |
the §3 feature split |
That last row is worth noticing: the reference already separates the geometry and arithmetic from the
inference well enough to unit-test them with no model on the machine. §3's inference feature flag is
the same boundary, enforced by Cargo rather than by discipline.
What does not port. The reference matches faces against a known gallery of named identities;
DarkRoom clusters an unlabelled library (§9). And it is a video pipeline, so it has a tracker, a
max_faces cap, and frame-to-frame identity annealing that have no analogue here.
2. Licensing — the open half of D13
The architectures are published research. The weights everyone actually uses are not redistributable by this project.
InsightFace's code is MIT. Its pretrained models — buffalo_l, buffalo_s, buffalo_sc, and
every det_* and w600k_* checkpoint inside them — carry a non-commercial research-only grant,
stated in the model-zoo README and applying to both the auto-downloaded and the manually downloaded
artefacts. w600k_mbf is trained on WebFace600K, whose own terms are research-only as well, so the
restriction has two independent sources rather than one that might be renegotiated.
These are the exact artefacts in question — §1.1's reference project fetches them from InsightFace's own release page, which is where the grant attaches:
| File | Source |
|---|---|
det_500m.onnx (2.5 MB) |
insightface/releases/download/v0.7/buffalo_sc.zip |
w600k_mbf.onnx (13 MB) |
insightface/releases/download/v0.7/buffalo_s.zip |
DarkRoom is GPL-3.0-or-later. A non-commercial field-of-use restriction is not a GPL-compatible grant, so:
- these weights cannot be committed to this repository, the way
yolo26n-seg.onnxis (D14); - they cannot ship inside an APK, a Flatpak, or an F-Droid build;
- and the restriction binds the user, not only the project — a professional photographer using DarkRoom commercially is outside the grant even if they fetched the file themselves.
That last point is the one that is easy to miss and dishonest to leave implicit. It is why §2.2's recommended route puts the licence text in front of the user rather than in a footnote.
2.1 The routes, priced
| Route | What ships | Cost |
|---|---|---|
| A — commit the weights | Everything works out of the box, one git lfs pull |
Not available. Licence-incompatible with GPL-3.0 and with all three distribution channels (NFR-COMPAT-2). Listed only so it is visibly rejected rather than quietly assumed. |
| B — permissively licensed weights | Same, legally | Requires weights that do not exist off the shelf in this architecture pair. Either found (§2.3) or trained, and training an ArcFace embedder needs a face dataset whose own terms permit redistribution — the harder half of the problem. |
| C — the user fetches them | The app ships the code, not the weights; face indexing is off until the user obtains a model | Works today, keeps the repository clean, and is honest. Costs a first-run flow, an F-Droid anti-feature declaration, and a licence notice the user has to read. |
Recommendation: C now, B when it becomes possible. §1.1's project already works this way — a
scripts/download_models.sh that fetches the weights rather than a repository that carries them —
so route C is a practised pattern here rather than a theory. Two things DarkRoom must do that the
reference does not: pin a checksum for every model (the reference verifies only one of the four),
and put the licence in front of the user, because a photo editor's users are not all researchers.
The schema already forces this to be a survivable choice — faces.model_id (catalog.md §10.1) exists precisely so that a model change is
detectable and re-indexable rather than silently poisoning every similarity in the library. Under C,
swapping in a permissive model later is a new model_id and a re-index, not a migration.
2.2 What route C requires of the code
- The weights are never a build input.
dr-face(§3) takes bytes; it has noembedded-modelfeature and nomodels/directory. This is the one structural difference fromdr-segment, and it is deliberate — a feature flag that could embed weights is a feature flag someone eventually turns on in a packaging script. - The fetch does not live in
dr-face. It lives in the app layer, so the inference crate keeps no network dependency at all. §11 makes that a checkable property rather than a convention. - The licence is shown, not linked. Before the first download, the app states in plain language that the weights are research-only, that commercial use is outside the grant, and who the grantor is. NFR-SEC-5's last bullet already requires the models to be named, versioned and licensed in the about screen; this is the same obligation moved to the moment where it can still change a decision.
- A checksum is pinned. The app verifies the digest of what it fetched against a value compiled in, and refuses a mismatch. An unverified model file is an arbitrary graph executed over the user's family photographs.
- Indexing stays off until a model is present and verified, and disabling it deletes nothing — that is a separate control (NFR-SEC-5, §11).
2.3 Before S14 writes any code
Three questions to answer by reading, in this order, and to record with the date they were checked — the same discipline segmentation.md §7 applied:
- Is there a permissively licensed SCRFD export? Third-party ONNX exports of SCRFD are plentiful; a re-export inherits the original weights' grant, so the question is whether anyone has trained SCRFD on a redistributable dataset, not whether they have converted the InsightFace one.
- Is there a permissively licensed 512-d MobileFaceNet/ArcFace checkpoint? Candidates worth reading the terms on rather than assuming: EdgeFace (Idiap), the ONNX Model Zoo's ArcFace entry (MIT, but a ResNet100 at ~250 MB — licence-clean and the wrong size), and any MobileFaceNet trained on a dataset with distribution terms.
- What does a permissive detector alternative cost in accuracy? YuNet (OpenCV Zoo) is 232 KB, emits five landmarks, and comes from a permissively licensed repository. If it is close enough, the detector half of the licence problem disappears and only the embedder remains — a materially better position than either half alone.
Question 3 is nearly free to answer: face_detection_yunet_2023mar.onnx is already sitting in
§1.1's models/ directory, fetched from opencv/opencv_zoo. It needs its own decode path (its
output layout is different, which is why the reference's SCRFDDecoder explicitly rejects it at load
rather than silently misreading it), and then it is one more row in §12's table.
3. Crate shape — core/dr-face
A new workspace member, modelled on dr-segment and for the same reason: everything that reasons
about faces is testable with no GPU adapter and no model present (ARCH §6.5a).
core/dr-face/
src/lib.rs FaceError, re-exports
src/detect.rs SCRFD: preprocess, decode, NMS
src/align.rs 5-point similarity transform → 112×112 crop
src/embed.rs MBF: preprocess, forward, L2 normalise
src/calibrate.rs cosine → P(same person) (FR-CULL-9)
src/cluster.rs constrained agglomeration (FR-CULL-10)
[features]
# No `embedded-model`. §2.2 — the weights are not a build input, ever.
default = []
# The ONNX runtime. Off by default so `calibrate` and `cluster` — which are
# arithmetic over embeddings and have no model in them — stay testable in a
# build that carries no inference engine at all.
inference = ["dep:ort", "dep:ort-tract", "dep:ndarray"]
The ort + ort-tract pairing is settled by D13's 2026-08-21 update and already in the workspace
manifest: ort's API, tract's pure-Rust engine, no C under the NDK.
The split between inference and the rest is load-bearing. Calibration and clustering are where
the subsystem's accuracy actually lives, they are pure functions of embeddings and labels, and they
must be testable against synthetic embedding sets with no weights on the machine. A test suite that
needs a research-licensed download to run is a test suite that does not run in CI.
3.1 The API
/// A loaded detector. Fixed input shape — see §4.
pub struct Detector { /* session, input edge */ }
impl Detector {
pub fn from_bytes(onnx: &[u8], id: ModelId) -> Result<Self, FaceError>;
/// `rgb` is f32 0..=1, row-major, three per pixel.
pub fn detect(&self, rgb: &[f32], w: usize, h: usize, opts: &DetectOptions)
-> Result<Vec<Detection>, FaceError>;
}
pub struct Detection {
/// Normalised to the image's long edge (catalog.md §10.1).
pub bbox: (f32, f32, f32, f32),
/// Five points, same normalisation, in the model's own order (§5).
pub landmarks: [(f32, f32); 5],
pub confidence: f32,
}
pub struct Embedder { /* session */ }
impl Embedder {
pub fn from_bytes(onnx: &[u8], id: ModelId) -> Result<Self, FaceError>;
/// Takes an *aligned* 112×112 crop. Alignment is `align::warp`; passing a
/// raw bbox crop here is the silent-accuracy-loss bug §5 exists to prevent.
pub fn embed(&self, aligned: &Aligned112) -> Result<Embedding, FaceError>;
}
/// L2-normalised, 512-d. Carries its model id so a comparison across models
/// is a type error rather than a plausible-looking number (catalog.md §10.1).
pub struct Embedding { pub model: ModelId, pub v: [f32; 512] }
Aligned112 is a newtype over the pixel buffer that only align::warp can construct. That is the
whole defence against §5's failure mode, and it costs nothing.
4. Detection
4.1 Fixed input, and how the image is fitted to it
det_500m.onnx has a dynamic H/W input, and this is the largest single risk in the document.
§1.1's reference had to use ONNX Runtime rather than OpenCV's dnn module precisely because OpenCV
could not load SCRFD's dynamic Shape nodes. tract is in the same family of problem: dr-segment
exists in its current shape because tract failed shape inference on YOLO's dynamic export, which is
why tools/export-seg-model.sh passes dynamic=False and why semantic.rs has a fixed
INPUT_EDGE.
So the graph almost certainly needs its input dimensions frozen before tract will touch it. That is a
known, cheap operation — onnxruntime.tools.make_dynamic_shape_fixed rewrites the declared dims
without retraining or re-exporting from PyTorch — but it is a step, it produces an artefact that must
itself be reproducible (a tools/fix-face-model-shapes.sh, in the spirit of the existing export
script), and it must be tried before anything else in this document is scheduled. §12's M1.
Fixed at 640, with 320 available as a faster, blinder option to measure.
Letterbox. Scale by min(640/w, 640/h) preserving aspect, then paste into a 640×640 canvas. The
reference centres the image and fills the margin with grey 114, inverting with
x_src = (x_model − pad_x) / scale. What matters is not where the padding goes but that the
forward and inverse agree and that the fill value is treated as part of the contract: a mismatch
between them offsets every box and landmark the model returns by the padding, which yields detections
that look plausible, landmarks that align badly, and an embedding-quality loss that only shows up as
poor clustering three stages later. Port the reference's convention rather than inventing one, and
assert it with a round-trip test.
Normalisation is (x·255 − 127.5) / 128, RGB, NCHW — note /128, not /127.5; see §6.
4.2 Outputs, and decoding them
Nine tensors for three strides {8, 16, 32}, twelve for four {8, 16, 32, 64}. Which one a
given export produces is discovered at load — strides = output_count / 3 — not assumed, because
both variants exist and hardcoding three silently ignores the largest faces in a twelve-output model.
With two anchors per location and N_s = (640/s)² locations:
| Tensor | Shape | Meaning |
|---|---|---|
score_s |
[N_s·2, 1] |
sigmoid already applied inside the graph |
bbox_s |
[N_s·2, 4] |
distances left, top, right, bottom in units of the stride |
kps_s |
[N_s·2, 10] |
five (dx, dy) offsets, same units |
Anchor centres are (x·s, y·s) for each grid cell, repeated once per anchor. Decoding is therefore
x1 = cx − l·s x2 = cx + r·s kx_i = cx + dx_i·s
y1 = cy − t·s y2 = cy + b·s ky_i = cy + dy_i·s
then divide by the letterbox scale to return to source pixels, then normalise by the long edge before storage.
Flat index within a stride is (row · fw + col) · 2 + anchor.
A shape assertion at load time, not a decode-time surprise. dr-segment already learned this —
SegmentError::OutputShape exists because a different export of the same model produces confidently
wrong results otherwise. The reference implements exactly this and it is worth porting verbatim:
output count divisible by three and between 9 and 12, then the last dimension of each group checked
against {1, 4, 10}. That single check is what catches a YuNet file passed where an SCRFD one was
meant — YuNet also has twelve outputs, so the count alone does not distinguish them, and the failure
without the check is not an error but a page of plausible garbage.
4.3 Thresholds
Score ≥ 0.5, NMS IoU 0.4, plain greedy NMS per image across all strides together.
Deliberately not the low threshold dr-segment chose. There, a false positive costs one spurious
entry in a list the user is picking from. Here it costs an entry in the People view that the user has
to reject, in a library with thousands of images — and worse, a garbage embedding that participates
in clustering and can bridge two real clusters into one. False negatives are recoverable by a later
re-index with a better model; a polluted cluster graph is not, once the user has confirmed faces
inside it.
A minimum box size of 40 px on the source proxy is applied on top — the reference's figure, chosen there for the same reason it holds here: below it there is not enough face left to align reliably, and §7's table shows what the embedder actually receives at that size. §12's M4 is what moves this number, if measurement says it should move.
One thing not to port: the reference caps detections at ten per frame, largest first. That is right for a film frame, where the extras in the background are noise. It is wrong for a photo library, where a group shot with thirty faces in it is precisely the picture the user wants indexed. No cap; the min-size floor is the only filter.
5. Alignment — the step that is silently wrong when skipped
ArcFace embeddings are trained on faces warped to a canonical 112×112 arrangement. Feeding the model a plain bounding-box crop works — it produces 512 numbers, they are unit-norm, and cosine similarities between them look entirely reasonable. They are just much worse, and nothing in the system reports it.
The transform is a similarity transform — rotation, uniform scale, translation, four degrees of freedom — from the five detected landmarks to this template, which is the arrangement the weights were trained against:
(38.2946, 51.6963) subject's right eye ─┐ image-left of centre
(73.5318, 51.5014) subject's left eye ─┘
(56.0252, 71.7366) nose tip
(41.5493, 92.3655) subject's right mouth corner
(70.7299, 92.2041) subject's left mouth corner
Similarity, not affine: a full affine fit to five noisy points will happily shear a face, and shear is exactly the deformation the embedder was never shown.
The naming is a trap and the order is not. Point 0 sits at x=38 on a 112-wide canvas — left of
centre in the image, which is the subject's right eye. Both namings are in circulation and they
are opposite. What matters is that SCRFD emits its five points in this same order, so the correct
amount of reordering between detector and template is none; the reference states this explicitly
in types.hpp for the benefit of whoever next reads it and doubts it. A future detector with a
different order carries its own permutation next to its model_id, rather than this file growing an
assumption.
One divergence from the reference to settle by measurement. It fits the transform with OpenCV's
estimateAffinePartial2D under RANSAC at a 3-pixel threshold, and drops the detection when the
fit fails. RANSAC over five points is a strange fit — the minimal sample for a similarity is two, so
it can discard landmarks it judges outliers and solve from a subset, which on a profile face is as
likely to be the correct geometry as the wrong one. InsightFace's own pipeline uses a plain
least-squares (Umeyama) similarity over all five points, which cannot silently drop anything and is
the one to write first. §12's M5 measures whether the choice matters; the degenerate-input path still
has to exist either way, because collinear landmarks do occur.
A trick worth keeping. The reference retries detection on a frame where nothing was found, after replicate-padding by 25% and applying CLAHE to the L channel. Padding gives a face at the very edge a context the detector needs; CLAHE rescues underexposed frames. Both are plausible on scanned negatives and backlit photographs — worth trying at low priority, since a retry doubles the cost of exactly the images that yielded nothing.
Sampling is bilinear from the source proxy, in one step — never crop-then-warp, which resamples twice and throws away detail the warp could have used. Pixels falling outside the source are black.
The landmark order must match the template order. The template above is written in the detector's own output order; if a future detector emits them differently, the template is reordered with it and that mapping belongs next to the model id, not compiled in as an assumption.
Test: warp a synthetic image with a known rotation and scale, and assert the five points land on
the template within a fraction of a pixel. This is testable with no model present, which is why
align sits outside the inference feature.
6. Embedding
112×112 RGB — the crop comes out of the warp in whatever order the source was in, and the reference converts BGR→RGB explicitly before the blob rather than relying on a flag, which is the readable way to do it.
Normalisation is (x·255 − 127.5) / 128. Note /128, not /127.5: InsightFace's published
Python uses 1.0/127.5 for the recognition model, the reference uses 1.0/128 for both models, and
every measured number in §1's table was produced with /128. The difference is 0.4% of scale and
almost certainly immaterial, but "almost certainly" is not a reason to pick silently — write /128 to
match the numbers we have, and settle it with one back-to-back run in §12.
Output is 512 floats; L2-normalise before storing, so every downstream comparison is a dot product and no code path has to remember to normalise. The reference clamps the norm at 1e-6 before dividing, which costs nothing and removes a NaN path.
Some exports of these graphs are fp16 in, fp16 out. The reference detects this from the graph's declared element type rather than from the filename; worth porting, because the alternative failure is a silent garbage tensor.
Storage is 512 × f16 (catalog.md §10.1) — 1 KB per face, 25 MB for a 25,000-face library. The f16
round-trip perturbs a unit vector by ~1e-3 in cosine, three orders of magnitude below the separation
between a match and a non-match, and the halving matters because these rows are the ones NFR-SEC-5
contemplates optionally syncing.
Re-normalise on load after the f16 widen. It is one pass over 512 floats and it removes a class of drift that is otherwise invisible.
7. Which pixels the pipeline actually sees
FR-CULL-8 pins indexing to the FR-CULL-2 ladder: the proxy tier, never a full decode. The
relevant tier is ThumbSize::Large — 1024 px on the long edge (dr-thumbs).
That has a consequence worth stating in numbers rather than discovering in a clustering report. On a 1024 px proxy:
| Face size in frame | Pixels across | What §6 receives |
|---|---|---|
| A portrait, face fills a third of the frame | ~340 | Downsampled to 112. Ideal. |
| Two people, half-length | ~120 | Roughly native. Good. |
| A group of eight | ~50 | Upsampled to 112. Degraded, and usable. |
| A figure in a landscape | ~20 | At or below §4.3's floor. Rejected. |
So faces records one column beyond catalog.md §10.1's schema:
ALTER TABLE faces ADD COLUMN crop_px INTEGER NOT NULL; -- source pixels across the aligned crop
crop_px earns its place three times over. It is the honest quality signal for the UI; it is a
feature in §8's calibration, which FR-CULL-9 explicitly demands ("a raw cosine means something
different for every model, every population, and every face size"); and it is what a future
higher-resolution re-embedding pass would select on, so that pass becomes a query rather than a
re-index of everything.
No RAW decode is added. Where the Large proxy is missing, the job enqueues a Thumbnail job at
background priority and re-queues itself, exactly as FR-CULL-8 requires.
8. Calibration — cosine to probability
FR-CULL-9 makes this a hard requirement: no code path may threshold a bare cosine, every threshold in the subsystem is stated as a probability, and the fit is per library and reports its own validity. This section is how that is obtained, and the interesting part is where the training pairs come from when the user has labelled nothing yet.
8.1 Where the pairs come from
Negatives are free and abundant. Two faces detected in the same photograph are almost never the same person. That gives every multi-face image in the library a full set of negative pairs at no labelling cost — and they are hard negatives, drawn from the same camera, lighting, and processing, which is exactly the population where a threshold tuned on easy negatives fails. The known exceptions — mirrors, photographs of photographs, collages, a spliced panorama — are rare enough to be noise at this scale and worth naming so the exception is not mistaken for a bug later.
Positives, in order of trustworthiness:
- User confirmations (FR-CULL-10). Every pair of faces confirmed to the same person. The gold standard, and empty on day one.
- Burst siblings. FR-CULL-5 already groups bursts. Two faces in adjacent frames of one burst, in nearly the same position, are near-certainly the same person. Free, requires no labelling, and available immediately on any library with continuous-shooting frames in it. Their purity is an S14 measurement (§12), not an assumption — if bursts turn out to be dirtier than expected, this source is dropped and the calibration simply stays invalid for longer.
- Nothing else. Bootstrapping positives from high cosine similarity is circular — it fits the calibration to the belief it was supposed to test.
8.2 The fit
§1.1's reference already implements this and its shape should be ported rather than reinvented:
P(same | cos) = σ(a·cos + b + log_prior_odds)
Four details in it are the difference between working and nearly working.
Fit against a histogram, not against pairs. A 25,000-face library has 3×10⁸ pairs; no gradient descent is running over that. The reference buckets every pair into 200 bins over cos ∈ [−1, 1], carrying a positive and a negative count per bin, and fits the two parameters against the per-bin counts. The cost becomes the similarity matrix plus 200 numbers, and the matrix is a single GEMM (§9). This is the trick that makes a per-library fit affordable at all, and it is not obvious from the outside.
Balance the classes explicitly. w_pos = total/(2·n_pos), w_neg = total/(2·n_neg). Negatives
outnumber positives by orders of magnitude; an unweighted fit produces a well-shaped curve sitting at
the wrong height — precisely the "plausible number all the way to the user interface" failure
FR-CULL-9 describes.
The base rate is a runtime argument, not part of the fit. log_prior_odds is added at evaluation
time, so the balanced fit is stored once and the prior varies per query — the odds that two faces
drawn from a 40-image holiday album match are not the odds for a 40,000-image archive. Its inverse,
boundary_at(p), gives the cosine at which the probability crosses p, which is what turns §9's
"merge above 0.9" into an actual comparison. Baking a prior into a and b would need a refit per
context and would make the stored parameters mean something different depending on where they came
from.
Deduplicate before pairing. Near-identical embeddings (cos > 1 − 1e-7) are the same photograph counted twice; the reference drops them per identity first. In a photo library the equivalent is duplicates and virtual copies (FR-CAT-11), and leaving them in stacks the positive histogram at cos ≈ 1 with pairs that teach the fit nothing about hard cases.
A third feature this design adds. The reference fits on cosine alone; DarkRoom should carry
crop_px too:
logit P(same) = w₀ + w₁·cos + w₂·log₂(min(crop_px_a, crop_px_b))
FR-CULL-9 explicitly names face size as an axis along which an uncalibrated similarity misbehaves,
§7 shows a real library spans 40 px to 340 px of face, and crop_px is already in hand. The
minimum of the pair, because a comparison is only as good as its worse crop. It generalises the
histogram to a small 2-D grid of bins, which changes nothing structural. If M7 says the extra feature
buys nothing, drop it and match the reference exactly — but a film pipeline working from broadcast
frames had far less size variation to explain than this one does.
8.3 Validity, and saying so
Stored per catalog.md §10.3: parameters, a validity flag, and a hash of the face set fitted from.
Invalid when fewer than 200 positive pairs or 2,000 negative pairs are available, or when the reliability check fails.
Deliberately far stricter than the reference's floor of two positives and one negative. That floor is reasonable there: its pairs come from a curated gallery of labelled reference portraits, where a positive pair is trustworthy by construction. Here the positives are bootstrapped from bursts and early confirmations (§8.1) and the whole risk is fitting confidently to a handful of them. The reference also refuses to draw positives from an identity with fewer than five distinct embeddings, letting it contribute negatives only — the same asymmetry applies to a thinly-confirmed person and is worth keeping. In that state the UI says confidence is unavailable and the People view still works — clustering falls back to a documented default operating point, labelled in the interface as an untuned default, and no probability is displayed. That is FR-CULL-9's requirement read literally: not presenting an untuned default as though it were measured.
Refit is triggered by the same debounce as clustering (§9), and when the confirmed-pair count grows materially.
Acceptance (FR-CULL-9): a reliability diagram over ten probability bins, on a held-out labelled split, with observed match rate within a stated tolerance of the predicted probability in each populated bin. A single accuracy figure is not an answer to this requirement.
9. Clustering
Not a job kind, for the reason catalog.md §10.2 gives: it is a whole-library operation with no
natural subject_id. It runs as a debounced library pass when detection has been idle and the face
count has moved materially.
The graph. For each face, its k = 20 nearest neighbours by cosine, then each candidate edge
scored through §8 to a probability.
Brute force is honest arithmetic here, and §1.1's project is the evidence rather than the estimate: it
computes the full similarity matrix as one GEMM (E · Eᵀ over L2-normalised rows) at gallery
scale, on CPU, and instruments it well enough to have compared CPU against OpenCL and found the
question worth asking rather than urgent. For DarkRoom that is a 25,000 × 512 matrix at 25 MB in and
~2.5 GB out — so the matrix is computed in row blocks, with each block reduced to its top-k and
its histogram contribution before the next is started, and never materialised whole. Once, in the
background, after an indexing sweep that took an hour. An approximate index is an optimisation to
reach for when §12 says it is needed, not before.
Constraints, not just thresholds:
- Cannot-link on co-occurrence. Two faces in the same image are never merged. This is the same observation §8.1 mines for negatives, used here as a hard constraint, and it is the single cheapest defence against the over-merging FR-CULL-10 warns about.
- Confirmed faces are anchors. A confirmation is user data (FR-CULL-12) and clustering never moves it. Two clusters each containing confirmations of different people cannot merge; a cluster containing confirmations of one person absorbs suggestions but never reassigns the confirmed.
The algorithm. Constrained average-link agglomeration over the probability graph, merging while the average pairwise probability exceeds 0.9 and no cannot-link is violated. Average-link rather than single-link because single-link chains — one bad edge welds two identities together, which is the documented way face clustering fails on families.
Incremental by default. A new face joins the existing cluster whose average probability against it is highest, if that clears the threshold; otherwise it starts an unnamed group. A full re-agglomeration happens on model change, on recalibration, and on explicit user request. Suggestions are recomputed freely; confirmations survive all of it (FR-CULL-10).
Splitting re-agglomerates within one person at a raised threshold and offers the resulting groups as the split. FR-CULL-10 requires splitting to be as easy as merging, and a split that hands the user a pile of loose faces to re-sort is not that.
10. Catalog and jobs
Schema is catalog.md §10.1's v5 migration, plus faces.crop_px (§7) and the calibration table
(§8.3). Nothing else changes.
One new job kind:
/// Detect and embed faces in one image, from its proxy (FR-CULL-8).
DetectFaces = 8,
Background priority, coalesced per image_id, interruptible, resumable — it inherits FR-CAT-3's
properties with no exceptions, which is what FR-CULL-8's acceptance criterion is about. It is not a
network job.
One job does detection and embedding for every face in the image, rather than splitting them. Splitting would double the queue's row count for no benefit: the proxy is already decoded and in memory, and the natural unit of resumable work is one photograph.
Selector term (FR-CULL-11):
Person { id: PersonId, include_suggested: bool }, // defaults to false
Confirmed-only by default, so a saved smart collection does not silently change membership when a later indexing pass revises a guess.
11. Privacy obligations that constrain the code's shape
NFR-SEC-5 is not a policy to remember; it is a set of properties the code either has structurally or does not have at all. The checkable ones:
dr-facehas no network dependency. Noreqwest, noureq, transitively. Worth a CI check overcargo tree, alongside the existing lints — the crate that holds the embeddings should be provably unable to send them anywhere.- The model fetch (§2.2) is in the app layer, which is why the API in §3.1 takes bytes.
- The diagnostics bundle is an allowlist (NFR-OPS-1), so
faces,face_personand the calibration table are excluded by not being named, and a future table cannot become uploadable by existing. - Delete-all is one transaction and one control: faces, links, people, calibration, and the cached crops if any. Separately, a switch that stops indexing so no such data is produced. The two are different actions and are not collapsed into one.
- The about screen reads the model metadata from the loader — id, version, licence — rather than from a hardcoded string that will drift from what is actually running.
12. What S14 measures
The corpus is a real personal library — the ~2,000-image sample S14 already specifies — with a hand-labelled identity subset. Fixed in advance, so the result is a measurement rather than an argument:
| # | Measure | Why it decides something |
|---|---|---|
| M1 | Does tract load both graphs? As shipped, then with the input dims frozen | Go/no-go, and it is first. det_500m.onnx has a dynamic H/W input (§4.1), which is the exact thing tract failed on for YOLO and which OpenCV's dnn could not load either — so the as-shipped answer is expected to be no, and the real question is whether make_dynamic_shape_fixed is enough or whether a PyTorch re-export is needed. Costs an afternoon; everything below is void without it. |
| M2 | Detection latency per image at 640, and at 320, on the reference desktop and one Android device | Whether a 17k library indexes in a background sweep or a weekend. The one measured datapoint we have — YOLO26n-seg, ~470 ms at 640×640 in tract — suggests SCRFD-500M lands well under it, but tract is not ORT and an extrapolation is not a measurement. |
| M3 | Embedding latency per face, and faces per image in a real library | The multiplier on M2. A library averaging 1.5 faces per image at 100 ms per face is an hour for 25k faces; at 400 ms it is four. |
| M4 | Detection recall against hand-labelled faces, bucketed by crop_px |
Where §4.3's floor should actually sit, and what proportion of a real library's faces are in the degraded bucket §7 predicts. |
| M5 | Aligned versus unaligned embeddings, same corpus, same clustering. Then two A/Bs that cost one run each: least-squares versus RANSAC alignment (§5), and /128 versus /127.5 normalisation (§6) |
Quantifies §5. If the gap is small the alignment code is still correct; if it is large, this is the measurement that stops someone "simplifying" it later. The two A/Bs close divergences between the reference and the published Python that are currently settled by assertion. |
| M6 | Cluster purity and completeness against the labelled subset | Whether the subsystem is worth building at all. §1's F1 figures rank the models on film frames against a known cast; nothing yet says how MBF behaves on family snapshots it must cluster blind. |
| M7 | Calibration reliability, and the sample size at which the fit becomes valid | FR-CULL-9's acceptance criterion, plus the practical question of whether a 2,000-image library ever gets a valid fit or whether §8.3's "unavailable" is the normal state. |
| M8 | Burst-derived positive-pair purity (§8.1) | Whether the free positives are usable or the fit waits for user confirmations. |
| M9 | YuNet as a drop-in detector: M4 and M6, re-run | §2.3's question 3, and the model file is already on disk. If the answer is "close enough", half the licence problem disappears. Needs its own decode path — its outputs are not SCRFD's. |
| M10 | Peak RSS during an indexing sweep | NFR-RES-2. Two loaded graphs plus a proxy plus a batch of crops, on a phone. |
13. Order
- M1 — load both graphs in tract, freezing the input dims if needed (§4.1). Go/no-go, and it now comes first: it is an afternoon, and §2.3's reading is wasted effort if the graphs will not run at all. This is a change from the S14 brief's ordering, made because §1.1 removed the uncertainty that justified reading first — the models are known to work, so the open question is the runtime, not the choice.
- Licence reading (§2.3), in parallel with 3. Before anything ships, per D13.
dr-faceskeleton,detect+align+embed, ported from §1.1's table, with an example binary that draws boxes and landmarks on a JPEG — the same shape asdr-segment'sexamples/detect.rs, and for the same reason: the thing worth looking at is whether the landmarks land on a real photograph. Porttests/test_face_utils.cpp's cases first; they are model-free and they fail loudly on exactly the mistakes §5 describes.- M2–M5 on the corpus. Alignment is validated here, before anything depends on embedding quality.
calibrate+clusteragainst the labelled subset. M6–M8. The Platt fit ports fromgallery_calibration.hpp; the pair sourcing (§8.1) is new and is the part to get wrong.- The v5 migration, the
DetectFacesjob, the debounced clustering pass. - UI: People view, confirm and reject, merge and split, the
Personselector term. - The route-C first-run flow, the about-screen entry, and the delete-all control — with the feature shipped disabled until they exist.
Steps 1–5 are the spike. Steps 6–8 are the build, and they are only justified if §12 says so.
14. What this does not settle
- Whether a permissive model pair exists. §2.3. If it does, route B replaces route C and step 8's first-run flow shrinks to nothing.
- Whether tract runs these graphs at all. §12 M1, and §4.1 says why the answer is in real doubt. Every other line of this document is conditional on it.
- Faces in trashed images. catalog.md §10.4's open question, unchanged: probably excluded from suggestions but not deleted, so a restore does not re-index.
- Whether embeddings sync. NFR-SEC-5 permits it, opt-in, separately consented. Nothing here depends on the answer.
- Re-embedding at higher resolution. §7's
crop_pxmakes it a query rather than a full re-index, but whether it is worth doing is an M4 question. - Approximate nearest neighbours. §9 says brute force until measured otherwise. A 100k-face library is where this stops being true.
15. Register entries
D13 — face inference runtime and model licensing · the runtime half stays answered (ort +
ort-tract, unchanged since 2026-08-21). The licensing half is answered conditionally by §2:
route C — ship the code, not the weights — unblocks the build without breaking GPL-3.0 or any
distribution channel, and §2.3's reading may yet convert it to route B. Closing D13 outright waits on
that reading.
D17 — face model pair and distribution route · PROPOSED. SCRFD-500M for detection, MobileFaceNet/ArcFace for embedding, weights obtained by the user rather than redistributed, with a pinned checksum and a licence notice shown before the first fetch. The alternative considered and rejected is committing the InsightFace weights (§2.1 route A), which is not available at any price. The model pair is better evidenced than a proposal usually is — §1's table is a measured comparison over a 2,418-identity gallery, not a literature reading — so the open half of this decision is the route, not the pair. Relates to: D13, D14, NFR-COMPAT-2, FR-CULL-8.
D18 — porting from scene-actor-extraction · PROPOSED, and the easy half of a decision.
That project is MIT and by this project's author, so MIT-into-GPL-3.0-or-later is a compatible
one-way combination requiring attribution, not permission. Ported files carry a header naming the
origin. Worth recording because "we already have this working in another language" is exactly the
provenance that goes undocumented and then cannot be answered three years later.
S14 — face pipeline in Rust · scope sharpened, and substantially de-risked, by this document. §12 replaces "measure per-image latency, cluster purity and calibration convergence" with ten specific measures. Two changes to the brief itself: its instruction to resolve the licence question before writing any of it is relaxed to "before shipping any of it", because M1 is an afternoon and is worth knowing first (§13); and its central question — whether the pipeline works at all — is largely answered in advance by §1.1, leaving the runtime and the domain shift from film frames to family snapshots as the real unknowns.
16. Requirements touched
| ID | How this document addresses it |
|---|---|
| FR-CULL-8 | §4 detection, §6 embedding, §7 the proxy tier and its consequences, §10 the DetectFaces job |
| FR-CULL-9 | §8 — pairs, fit, validity, and the reliability-diagram acceptance test |
| FR-CULL-10 | §9 constrained agglomeration, confirmations as anchors, split by re-agglomeration |
| FR-CULL-11 | §10 the Person selector term, confirmed-only by default |
| FR-CULL-12 | §10 schema unchanged from catalog.md §10.1: embeddings derived, names to the sidecar |
| NFR-SEC-5 | §11 — the obligations restated as structural properties, one of them CI-checkable |
| NFR-COMPAT-2 | §2 — why the obvious weights cannot ship, and what does instead |
| NFR-RES-2 | §1 the cheap model pair, §12 M2/M3/M10 on a phone |
| NFR-ARCH-2 | §10 background priority, preempted by visible work |